Chapter 33Lesson 02~280 minutes

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Guided Hands-On Workflow

Simulate a partial deployment locally, preserve its evidence, make reconciliation idempotent, compare rerun and new-run semantics, and practice guarded compensation without production infrastructure.

Hands-onPartial failureReconciliationCompensationSafe retry

Learning objectives

  • Build a disposable local target where a durable side effect survives a failed command.
  • Use a receipt and idempotency key to make the second execution reconcile instead of duplicate.
  • Preserve state and logs before repeating a failed operation.
  • Map the local simulation to GitHub run/attempt semantics and explain why hosted-runner files are not durable across reruns.
  • Choose between a targeted rerun and a new run based on whether event/source identity must change.

1. Scenario: a deployment activates, then verification fails

We will model an external deployment target as a local directory that deliberately survives multiple script invocations. The “artifact” is a tiny text file, the “release store” is a directory, and the “deployment receipt” records the idempotency key. The first invocation activates the release and then fails. The second invocation reads the receipt, proves the release already exists, and continues without creating another logical deployment.

Mandatory path: this entire exercise is local and free. Use a disposable directory. No cloud account, package registry, environment secret, self-hosted runner or write token is required.

2. Preflight and assumptions

  • Use Bash with sha256sum, awk, grep and ordinary filesystem permissions.
  • Create the lab in a temporary repository/directory; do not point ROOT at production or personal data.
  • The fake source SHA is synthetic. When mapped to Actions, replace it with the exact GITHUB_SHA.
  • The directory itself is the durable external target for this local exercise. A GitHub-hosted runner workspace would not survive a rerun.
mkdir -p gha-recovery-lab/scripts
cd gha-recovery-lab
printf '%s
' '# disposable recovery lab' > README.md
# Save the script from the next section as scripts/fake-deploy.sh
chmod +x scripts/fake-deploy.sh
rm -rf .incident-target

3. Build the smallest idempotent fake deployment

The script computes an artifact digest and a logical key from environment + source SHA + artifact digest. Before creating a release it checks for an existing receipt. That read-before-write step is the idempotency boundary. The fault marker is also durable, so the intentionally injected failure happens once for that logical deployment.

#!/usr/bin/env bash
set -euo pipefail

ROOT=${1:-.incident-target}
SOURCE_SHA=${SOURCE_SHA:-1111111111111111111111111111111111111111}
ENVIRONMENT=${ENVIRONMENT:-staging}
FAIL_ONCE=${FAIL_ONCE:-1}
mkdir -p "$ROOT/releases" "$ROOT/receipts" "$ROOT/evidence"

artifact="$ROOT/artifact.txt"
printf 'source=%s\nenvironment=%s\n' "$SOURCE_SHA" "$ENVIRONMENT" > "$artifact"
digest="sha256:$(sha256sum "$artifact" | awk '{print $1}')"
key="$(printf '%s|%s|%s' "$ENVIRONMENT" "$SOURCE_SHA" "$digest" | sha256sum | awk '{print $1}')"
receipt="$ROOT/receipts/$key.txt"
release="$ROOT/releases/${digest#sha256:}.txt"
current="$ROOT/current.txt"

printf 'source_sha=%s\ndigest=%s\nidempotency_key=%s\n' "$SOURCE_SHA" "$digest" "$key" | tee "$ROOT/evidence/identity.txt"

if [[ -f "$receipt" ]]; then
  echo "receipt exists; reconcile instead of creating another side effect"
  grep -F "digest=$digest" "$receipt" >/dev/null
  test -f "$release"
  sha256sum "$release"
else
  cp "$artifact" "$release"
  printf 'key=%s\nsource_sha=%s\ndigest=%s\nstatus=activated\n' "$key" "$SOURCE_SHA" "$digest" > "$receipt"
  printf '%s\n' "$digest" > "$current"
  echo "created one durable fake release"
fi

# A durable marker models a fault that should fire exactly once for this logical deployment.
fault="$ROOT/receipts/$key.fault-injected"
if [[ "$FAIL_ONCE" == 1 && ! -f "$fault" ]]; then
  : > "$fault"
  echo "simulated failure after activation" | tee "$ROOT/evidence/first-failure.txt"
  exit 23
fi

test "$(cat "$current")" = "$digest"
grep -F 'status=activated' "$receipt" >/dev/null
printf 'health=healthy\n' >> "$receipt"
echo "reconciled: target is healthy and no duplicate release was created"

4. First execution: preserve the partial state

set +e
SOURCE_SHA=1111111111111111111111111111111111111111   FAIL_ONCE=1 ./scripts/fake-deploy.sh .incident-target   >attempt-1.log 2>&1
rc=$?
set -e
printf 'attempt-1 exit=%s
' "$rc"
cat attempt-1.log
find .incident-target -maxdepth 3 -type f -print | sort
cat .incident-target/current.txt
cat .incident-target/receipts/*.txt

Expected observation: exit code 23, one release file, one deployment receipt, one current pointer and one fault marker. The red process result does not mean the target is unchanged—the target is already activated. Copy attempt-1.log and the directory listing before any repair.

5. Classify before rerun: what is safe to repeat?

Operation Current state after attempt 1 Recovery decision
Artifact creation Deterministic from same source SHA Safe to recompute and compare digest
Release copy Already exists at digest-addressed path Do not create another logical release; verify existing bytes
Current pointer update Already points to desired digest Reconcile; repeating identical write is harmless but unnecessary
Fault injection Durable marker proves it already occurred Must not repeat
Health verification Never completed Safe and necessary to retry

This table is the heart of incident recovery. Recovery is not “run the whole script again and hope.” It is a per-side-effect classification. The script can be rerun because its create operation has been converted to read/verify/reconcile semantics.

6. Second execution: reconcile idempotently

SOURCE_SHA=1111111111111111111111111111111111111111   FAIL_ONCE=1 ./scripts/fake-deploy.sh .incident-target   | tee attempt-2.log

# Prove there is still exactly one release and one logical receipt.
find .incident-target/releases -type f -maxdepth 1 -print | wc -l
find .incident-target/receipts -type f -name '*.txt' -maxdepth 1 -print | wc -l
cat .incident-target/receipts/*.txt

Expected observation: the script says the receipt exists, verifies the exact stored release, skips duplicate activation, and completes the health reconciliation. The second execution is safe because it converges on the same logical target state.

7. Compare with the broken “create blindly” pattern

# INTENTIONALLY BROKEN — demonstration only.
release_id="release-$(date +%s%N)"
cp artifact.txt ".incident-target/releases/$release_id.txt"
# A retry always invents a new ID, so the second attempt creates a second release.

Time/random/run-attempt-based resource names often make a mutation non-idempotent by design. They are useful for disposable test instances, but not when “rerun the deployment” is supposed to converge on the same desired release. Production APIs frequently provide create-or-update, conditional requests, resource version checks or explicit idempotency keys; use those mechanisms rather than inventing a duplicate-prone workflow convention.

8. Add a guarded local rollback

Rollback needs a known previous pointer. Save it before activation, then require the operator to name the exact current digest and exact rollback digest. The guard prevents “rollback whatever is current” from racing with another deployment.

CURRENT=$(cat .incident-target/current.txt)
EXPECTED_CURRENT="$CURRENT"        # capture from preserved incident evidence
ROLLBACK_TO='sha256:KNOWN_GOOD_DIGEST_FROM_LEDGER'

# Guard example: refuse if target has changed since evidence capture.
test "$(cat .incident-target/current.txt)" = "$EXPECTED_CURRENT" || {
  echo 'target changed; stop and re-inspect' >&2
  exit 1
}
# In this local lesson, do not execute a fake/unknown rollback digest.
printf 'Would rollback exact current %s to verified %s
' "$EXPECTED_CURRENT" "$ROLLBACK_TO"

The example intentionally stops short of writing an unknown target. A recovery command is safe only when the rollback target has an independently verified artifact/state identity.

9. Map the local model to GitHub Actions attempts

Local concept GitHub Actions mapping
Invocation ID GITHUB_RUN_ID + GITHUB_RUN_ATTEMPT
Fake source SHA GITHUB_SHA
Persistent target directory Real deployment system / registry / state store / dedicated lab ledger
Receipt file External deployment ID, release record, API idempotency receipt
attempt-1.log Attempt-specific GitHub logs + bounded evidence artifact
Second invocation Targeted rerun if same event/source is still desired

A critical difference is runner lifetime. GitHub-hosted reruns start on fresh compute, so .incident-target would disappear if it lived only in the workspace. Lesson 5 therefore uses a dedicated disposable state branch as a durable fake target so the side effect genuinely survives attempt 1.

10. Preserve an Actions attempt before rerunning it

RUN_ID=123456789
# Exact run identity and first-attempt result.
gh run view "$RUN_ID" --attempt 1 \
  --json databaseId,attempt,headSha,event,status,conclusion,url

# Save first-attempt logs outside the workflow before recovery.
gh run view "$RUN_ID" --attempt 1 --log > "run-$RUN_ID-attempt-1.log"

# If failure is isolated and repetition is safe after inspection:
gh run rerun "$RUN_ID" --failed

Do not run gh run rerun until the side-effect table says repeating the failed region is safe. If the fix requires editing source/workflow code, commit that change and create a new run instead; a rerun cannot adopt the new SHA.

11. Challenge: choose the recovery layer

Suppose attempt 1 created release sha256:aaaa, updated staging to that digest, and failed because an external health endpoint timed out. You confirm the endpoint is healthy manually and the target still runs sha256:aaaa. Which layer should change? The best answer is recovery/evidence: retry or rerun only the verification/reconciliation scope, using the existing deployment receipt. Rebuilding the artifact or creating a new release would change state unnecessarily.

12. Cleanup

cd ..
rm -rf gha-recovery-lab

Cleanup is safe here because the target is explicitly local and disposable. In a real incident, do not delete failed target state merely to make a dashboard green; retention requirements and forensic needs may require preserving it.

13. Lesson summary

You made a partial failure visible, preserved it, and converted retry from duplicate creation into reconciliation. The next lesson generalizes that choice: when should a pipeline retry, rerun, roll back, compensate, serialize with GitHub concurrency, or coordinate through an external lock?

Next lesson

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Configuration, Design Patterns, and Trade-Offs

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why does the local target directory deliberately persist between script invocations?

What evidence proves the second invocation did not duplicate the release?

Why can a GitHub-hosted workspace not be the cross-attempt side-effect ledger?

The fix requires changing the deployment script. Should you use “rerun failed jobs” and expect the change?

A health check timed out after activation, but target identity is correct. Which operation should you avoid?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.