Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Guided Hands-On Workflow
Simulate a partial deployment locally, preserve its evidence, make reconciliation idempotent, compare rerun and new-run semantics, and practice guarded compensation without production infrastructure.
Learning objectives
- Build a disposable local target where a durable side effect survives a failed command.
- Use a receipt and idempotency key to make the second execution reconcile instead of duplicate.
- Preserve state and logs before repeating a failed operation.
- Map the local simulation to GitHub run/attempt semantics and explain why hosted-runner files are not durable across reruns.
- Choose between a targeted rerun and a new run based on whether event/source identity must change.
1. Scenario: a deployment activates, then verification fails
We will model an external deployment target as a local directory that deliberately survives multiple script invocations. The “artifact” is a tiny text file, the “release store” is a directory, and the “deployment receipt” records the idempotency key. The first invocation activates the release and then fails. The second invocation reads the receipt, proves the release already exists, and continues without creating another logical deployment.
Mandatory path: this entire exercise is local and free. Use a disposable directory. No cloud account, package registry, environment secret, self-hosted runner or write token is required.
2. Preflight and assumptions
-
Use Bash with
sha256sum,awk,grepand ordinary filesystem permissions. -
Create the lab in a temporary repository/directory; do not point
ROOTat production or personal data. -
The fake source SHA is synthetic. When mapped to Actions, replace
it with the exact
GITHUB_SHA. - The directory itself is the durable external target for this local exercise. A GitHub-hosted runner workspace would not survive a rerun.
mkdir -p gha-recovery-lab/scripts
cd gha-recovery-lab
printf '%s
' '# disposable recovery lab' > README.md
# Save the script from the next section as scripts/fake-deploy.sh
chmod +x scripts/fake-deploy.sh
rm -rf .incident-target
3. Build the smallest idempotent fake deployment
The script computes an artifact digest and a logical key from environment + source SHA + artifact digest. Before creating a release it checks for an existing receipt. That read-before-write step is the idempotency boundary. The fault marker is also durable, so the intentionally injected failure happens once for that logical deployment.
#!/usr/bin/env bash
set -euo pipefail
ROOT=${1:-.incident-target}
SOURCE_SHA=${SOURCE_SHA:-1111111111111111111111111111111111111111}
ENVIRONMENT=${ENVIRONMENT:-staging}
FAIL_ONCE=${FAIL_ONCE:-1}
mkdir -p "$ROOT/releases" "$ROOT/receipts" "$ROOT/evidence"
artifact="$ROOT/artifact.txt"
printf 'source=%s\nenvironment=%s\n' "$SOURCE_SHA" "$ENVIRONMENT" > "$artifact"
digest="sha256:$(sha256sum "$artifact" | awk '{print $1}')"
key="$(printf '%s|%s|%s' "$ENVIRONMENT" "$SOURCE_SHA" "$digest" | sha256sum | awk '{print $1}')"
receipt="$ROOT/receipts/$key.txt"
release="$ROOT/releases/${digest#sha256:}.txt"
current="$ROOT/current.txt"
printf 'source_sha=%s\ndigest=%s\nidempotency_key=%s\n' "$SOURCE_SHA" "$digest" "$key" | tee "$ROOT/evidence/identity.txt"
if [[ -f "$receipt" ]]; then
echo "receipt exists; reconcile instead of creating another side effect"
grep -F "digest=$digest" "$receipt" >/dev/null
test -f "$release"
sha256sum "$release"
else
cp "$artifact" "$release"
printf 'key=%s\nsource_sha=%s\ndigest=%s\nstatus=activated\n' "$key" "$SOURCE_SHA" "$digest" > "$receipt"
printf '%s\n' "$digest" > "$current"
echo "created one durable fake release"
fi
# A durable marker models a fault that should fire exactly once for this logical deployment.
fault="$ROOT/receipts/$key.fault-injected"
if [[ "$FAIL_ONCE" == 1 && ! -f "$fault" ]]; then
: > "$fault"
echo "simulated failure after activation" | tee "$ROOT/evidence/first-failure.txt"
exit 23
fi
test "$(cat "$current")" = "$digest"
grep -F 'status=activated' "$receipt" >/dev/null
printf 'health=healthy\n' >> "$receipt"
echo "reconciled: target is healthy and no duplicate release was created"
4. First execution: preserve the partial state
set +e
SOURCE_SHA=1111111111111111111111111111111111111111 FAIL_ONCE=1 ./scripts/fake-deploy.sh .incident-target >attempt-1.log 2>&1
rc=$?
set -e
printf 'attempt-1 exit=%s
' "$rc"
cat attempt-1.log
find .incident-target -maxdepth 3 -type f -print | sort
cat .incident-target/current.txt
cat .incident-target/receipts/*.txt
Expected observation: exit code 23, one release file, one deployment
receipt, one current pointer and one fault marker. The red process
result does not mean the target is unchanged—the target is already
activated. Copy attempt-1.log and the directory listing
before any repair.
5. Classify before rerun: what is safe to repeat?
| Operation | Current state after attempt 1 | Recovery decision |
|---|---|---|
| Artifact creation | Deterministic from same source SHA | Safe to recompute and compare digest |
| Release copy | Already exists at digest-addressed path | Do not create another logical release; verify existing bytes |
| Current pointer update | Already points to desired digest | Reconcile; repeating identical write is harmless but unnecessary |
| Fault injection | Durable marker proves it already occurred | Must not repeat |
| Health verification | Never completed | Safe and necessary to retry |
This table is the heart of incident recovery. Recovery is not “run the whole script again and hope.” It is a per-side-effect classification. The script can be rerun because its create operation has been converted to read/verify/reconcile semantics.
6. Second execution: reconcile idempotently
SOURCE_SHA=1111111111111111111111111111111111111111 FAIL_ONCE=1 ./scripts/fake-deploy.sh .incident-target | tee attempt-2.log
# Prove there is still exactly one release and one logical receipt.
find .incident-target/releases -type f -maxdepth 1 -print | wc -l
find .incident-target/receipts -type f -name '*.txt' -maxdepth 1 -print | wc -l
cat .incident-target/receipts/*.txt
Expected observation: the script says the receipt exists, verifies the exact stored release, skips duplicate activation, and completes the health reconciliation. The second execution is safe because it converges on the same logical target state.
7. Compare with the broken “create blindly” pattern
# INTENTIONALLY BROKEN — demonstration only.
release_id="release-$(date +%s%N)"
cp artifact.txt ".incident-target/releases/$release_id.txt"
# A retry always invents a new ID, so the second attempt creates a second release.
Time/random/run-attempt-based resource names often make a mutation non-idempotent by design. They are useful for disposable test instances, but not when “rerun the deployment” is supposed to converge on the same desired release. Production APIs frequently provide create-or-update, conditional requests, resource version checks or explicit idempotency keys; use those mechanisms rather than inventing a duplicate-prone workflow convention.
8. Add a guarded local rollback
Rollback needs a known previous pointer. Save it before activation, then require the operator to name the exact current digest and exact rollback digest. The guard prevents “rollback whatever is current” from racing with another deployment.
CURRENT=$(cat .incident-target/current.txt)
EXPECTED_CURRENT="$CURRENT" # capture from preserved incident evidence
ROLLBACK_TO='sha256:KNOWN_GOOD_DIGEST_FROM_LEDGER'
# Guard example: refuse if target has changed since evidence capture.
test "$(cat .incident-target/current.txt)" = "$EXPECTED_CURRENT" || {
echo 'target changed; stop and re-inspect' >&2
exit 1
}
# In this local lesson, do not execute a fake/unknown rollback digest.
printf 'Would rollback exact current %s to verified %s
' "$EXPECTED_CURRENT" "$ROLLBACK_TO"
The example intentionally stops short of writing an unknown target. A recovery command is safe only when the rollback target has an independently verified artifact/state identity.
9. Map the local model to GitHub Actions attempts
| Local concept | GitHub Actions mapping |
|---|---|
| Invocation ID |
GITHUB_RUN_ID + GITHUB_RUN_ATTEMPT
|
| Fake source SHA | GITHUB_SHA |
| Persistent target directory | Real deployment system / registry / state store / dedicated lab ledger |
| Receipt file | External deployment ID, release record, API idempotency receipt |
| attempt-1.log | Attempt-specific GitHub logs + bounded evidence artifact |
| Second invocation | Targeted rerun if same event/source is still desired |
A critical difference is runner lifetime. GitHub-hosted reruns start
on fresh compute, so .incident-target would disappear
if it lived only in the workspace. Lesson 5 therefore uses a
dedicated disposable state branch as a durable fake target so the
side effect genuinely survives attempt 1.
10. Preserve an Actions attempt before rerunning it
RUN_ID=123456789
# Exact run identity and first-attempt result.
gh run view "$RUN_ID" --attempt 1 \
--json databaseId,attempt,headSha,event,status,conclusion,url
# Save first-attempt logs outside the workflow before recovery.
gh run view "$RUN_ID" --attempt 1 --log > "run-$RUN_ID-attempt-1.log"
# If failure is isolated and repetition is safe after inspection:
gh run rerun "$RUN_ID" --failed
Do not run gh run rerun until the side-effect table
says repeating the failed region is safe. If the fix requires
editing source/workflow code, commit that change and create a new
run instead; a rerun cannot adopt the new SHA.
11. Challenge: choose the recovery layer
Suppose attempt 1 created release sha256:aaaa, updated
staging to that digest, and failed because an external health
endpoint timed out. You confirm the endpoint is healthy manually and
the target still runs sha256:aaaa. Which layer should
change? The best answer is recovery/evidence: retry
or rerun only the verification/reconciliation scope, using the
existing deployment receipt. Rebuilding the artifact or creating a
new release would change state unnecessarily.
12. Cleanup
cd ..
rm -rf gha-recovery-lab
Cleanup is safe here because the target is explicitly local and disposable. In a real incident, do not delete failed target state merely to make a dashboard green; retention requirements and forensic needs may require preserving it.
13. Lesson summary
You made a partial failure visible, preserved it, and converted retry from duplicate creation into reconciliation. The next lesson generalizes that choice: when should a pipeline retry, rerun, roll back, compensate, serialize with GitHub concurrency, or coordinate through an external lock?
Knowledge check
Why does the local target directory deliberately persist between script invocations?
It models a durable external system whose side effects survive a failed workflow attempt. Without persistence, you could not test idempotent reconciliation.
What evidence proves the second invocation did not duplicate the release?
The same idempotency receipt/digest is reused and the release/receipt counts remain one while health status advances.
Why can a GitHub-hosted workspace not be the cross-attempt side-effect ledger?
Hosted reruns/jobs execute on fresh runner compute; workspace files are runner-local unless transferred to durable storage.
The fix requires changing the deployment script. Should you use “rerun failed jobs” and expect the change?
No. A rerun keeps the original source/workflow ref and SHA. Commit the repair and create a new run, linking it to the incident.
A health check timed out after activation, but target identity is correct. Which operation should you avoid?
Avoid rebuilding/re-publishing/redeploying blindly. Preserve evidence and retry/reconcile the verification layer first.
Official references and version notes
- Re-running workflows and jobs — Current rerun window, attempt behavior, actor privileges, same ref/SHA semantics and targeted rerun options.
- Variables reference — Definitions of GITHUB_RUN_ID, GITHUB_RUN_ATTEMPT, GITHUB_RUN_NUMBER, GITHUB_ACTOR and GITHUB_TRIGGERING_ACTOR.
- Contexts reference — Current github context fields and rerun identity semantics.
- REST API: workflow runs — Attempt-specific logs, rerun/cancel endpoints and exact workflow-run resource state.
- GitHub CLI: gh run rerun — Current --failed, --job and --debug rerun controls; job reruns require the database ID.
- GitHub CLI: gh run view — Current --attempt, --log, --log-failed and JSON run/job evidence inspection.
- Concurrency — Current serialization, cancel-in-progress and queueing behavior.
- Workflow syntax: concurrency — Current workflow/job concurrency syntax including queue:max.
- Deployment environments — Environment gates, secret timing and deployment target separation.
- Deployments and environments — Protection rules and deployment governance boundaries.
- actions/checkout v7.0.1 — Full commit SHA used by executable examples.
- actions/upload-artifact v7.0.1 — Full commit SHA used for bounded attempt-specific evidence artifacts.
- upload-artifact behavior — Current immutable-artifact behavior, unique names, outputs and overwrite semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.