Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Diagnostics, Failure Modes, and Production Practices
Diagnose duplicate publication, destructive cleanup, revision drift, unknown rollback targets and unsafe always() recovery by preserving evidence before mutation.
Learning objectives
- Use a fixed evidence-first sequence before retrying or rolling back any failed workflow.
- Diagnose blind publication retries, destructive cleanup, source-revision drift and unknown rollback targets causally.
- Explain why privileged always() cleanup and force-cancel can compound incidents.
- Use attempt-specific logs, job IDs, artifacts and external resource IDs to preserve first failure.
- Apply the least destructive recovery and rerun only an equivalent, safe scope.
1. Diagnostic rule: a red job is not an external-state description
The highest-risk recovery failures begin with the same mistaken inference: “the job failed, therefore the side effect failed.” A package registry, cloud API or deployment controller can commit the mutation and then lose the response; the next verification step can fail after a successful deployment; cancellation can interrupt the workflow after the target already changed. Diagnose the external state independently.
Do not start by deleting the partial target or rerunning. Start by freezing the incident identity: run ID, attempt, source/workflow SHA, actor, target and every resource/receipt returned before the failure.
2. Evidence-first diagnostic sequence
- Preserve run/attempt and first-failure evidence. Save attempt-specific logs and current target metadata.
- Confirm event/ref/SHA and workflow revision. Reruns use the original ref/SHA; do not assume a later commit is present.
- Confirm evaluated conditions, permissions and inputs. Check original actor privilege and rerun initiator separately.
- Inspect job graph and queue. Know which dependencies a targeted rerun would repeat.
- Confirm runner/image/toolchain. Separate runner drift from target-state failures.
- Inspect failing step/action/runtime/network. Find the first causal error, not the last cleanup error.
- Inspect outputs/artifacts/caches. Validate IDs/digests and preserve attempt evidence.
- Inspect environment/deployment/external target. Query exact target/resource ID directly.
- Apply the least destructive correction. Prefer verify/reconcile over delete/recreate.
- Rerun the smallest equivalent scope. Or create a new run when source/workflow/input must change.
3. Broken example: publish, fail verification, delete and republish
name: BROKEN — blind publish recovery
on:
workflow_dispatch:
permissions:
contents: write
jobs:
publish-and-clean:
runs-on: ubuntu-24.04
steps:
- run: ./publish-package.sh # may succeed remotely
- run: ./verify-package.sh # times out, job becomes red
- if: ${{ always() }}
run: ./delete-and-republish-latest.sh # DANGEROUS: ambiguous + destructive
This workflow contains three causal hazards. First,
publish-package.sh may have succeeded remotely even if
the next step fails. Second, “latest” is not an exact resource
identity. Third, always() schedules privileged cleanup
even after cancellation/failure, turning a diagnostic failure into a
destructive mutation. The red conclusion hides whether zero, one or
multiple versions now exist.
4. Repair blind package publication without erasing evidence
Preserve the publish response/version/digest immediately. On failure, query that exact version. If it exists and matches the expected digest, treat publication as completed and retry only downstream verification. If it does not exist and the API guarantees the request was not committed, an idempotent publish may be retried. If state is ambiguous, stop and escalate—do not mint another version or delete the namespace.
Evidence packet before repair:
run_id: 123456789
attempt: 1
source_sha: 012345...
intended_version: 1.2.3
expected_digest: sha256:...
publish_request_id: req-4821
provider_response: timeout after request body sent
registry_query(1.2.3): EXISTS / digest matches
Decision: publication completed -> do not republish -> rerun verification only
5. Failure mode: deleting a partial target before understanding it
“Delete and recreate” destroys the strongest evidence you have: the exact state produced by the failed attempt. It may also remove healthy subresources, release a lock, invalidate a rollback target or race with another actor. Prefer read-only target snapshots, deployment history and exact resource versions first.
Cleanup should occur only after the target is reconciled and
evidence retention requirements are met. In the checkpoint,
if: always() is allowed only for read-only evidence
collection/upload—never for privileged rollback or branch mutation.
6. Failure mode: rerunning with a different revision and calling it the same attempt
GitHub prevents one common form of this confusion: an actual rerun
keeps the original GITHUB_SHA/GITHUB_REF.
The confusion appears when teams commit a fix, trigger a new run and
describe it as “attempt 2.” That loses the audit link between the
failed revision and the repaired revision.
Use precise language: run 123 attempt 1 failed; run 123 attempt 2 reproduced/reconciled the same SHA; run 124 at SHA B tested the code change. If the fix changes workflow behavior, preserve both run URLs and SHAs.
7. Failure mode: “rollback” without an exact artifact/state target
A command such as deploy previous is ambiguous when
multiple deployments occurred or another actor changed the target.
The rollback input should be an immutable digest, revision ID or
state generation captured before the failed mutation. Then compare
the current target to the expected failed deployment before writing.
EXPECTED_CURRENT='sha256:failed-but-known-digest'
ROLLBACK_TO='sha256:verified-previous-digest'
ACTUAL_CURRENT=$(fake-provider get staging --output digest)
if [[ "$ACTUAL_CURRENT" != "$EXPECTED_CURRENT" ]]; then
echo 'Target changed since incident capture; refuse rollback.' >&2
exit 1
fi
fake-provider promote --environment staging --digest "$ROLLBACK_TO"
The guard is more important than the rollback command. It prevents a delayed recovery job from overwriting a newer legitimate deployment.
8. Failure mode: privileged always() cleanup
always() is useful for summaries, logs and evidence
artifacts. It is hazardous for high-impact cleanup because it runs
in failure/cancellation paths where assumptions may be invalid.
GitHub even provides a force-cancel endpoint that bypasses
conditions such as always() for stuck runs, reinforcing
that always-running logic is not a transaction manager.
Move recovery mutations into an explicit job/workflow with exact inputs and guards. Require environment approval or incident authorization when appropriate. Keep the failure handler itself read-only whenever possible.
9. Diagnose targeted rerun dependencies before clicking
RUN_ID=123456789
# Preserve and list exact job database IDs.
gh run view "$RUN_ID" --attempt 1 --json jobs --jq '.jobs[] | {name,databaseId,status,conclusion}'
# Review first failed logs before mutation.
gh run view "$RUN_ID" --attempt 1 --log-failed > "incident-$RUN_ID-a1-failed.log"
# Only after classifying dependencies as safe:
gh run rerun "$RUN_ID" --failed
# or: gh run rerun --job JOB_DATABASE_ID
The CLI documents that rerunning a job includes dependencies. A dependency that publishes or deploys must therefore be idempotent/reconciling even if it was green in the original attempt.
10. Cancellation is not rollback
A normal cancel request returns asynchronously, and a force-cancel exists for workflows that ignore normal cancellation. Neither operation restores an external target. After cancellation, inspect deployment/resource IDs, partial artifacts and locks. If the target changed, enter the same reconcile/rollback decision process as any other partial failure.
11. Separate causal failure layers
| Observed symptom | Likely layer to inspect first | Do not jump directly to |
|---|---|---|
| Workflow never ran | event/filter/workflow selection/policy | deployment rollback |
| Job queued unusually long | runner capacity/concurrency | re-publish artifact |
| 403 on recovery API | token/OIDC/environment authorization | overly broad write token |
| Artifact digest mismatch | build/artifact identity | deploy “latest” |
| Deploy API success, health failure | external target/runtime/health | assume deploy did not happen |
| Rollback guard says target changed | concurrent/external actor state | force overwrite |
| Evidence upload failed | artifact service/evidence path | delete target to “clean up” |
12. Production incident-safe practices
- Create explicit idempotency keys and persist provider/resource receipts.
- Name evidence by run ID + attempt; retain first failure separately.
- Serialize side-effecting deployments where ordering matters, but still design idempotency.
- Do not rebuild release artifacts during recovery; promote/rollback exact verified digests.
- Use narrow recovery identities and environment/incident approvals for privileged mutations.
- Record rollback compatibility and last-known-good target before deployment.
- Test fault injection in disposable environments so the recovery path is exercised before production.
13. Lesson summary
A safe incident response preserves evidence, proves external state, and changes only the causal layer. Blind reruns, delete-and-recreate, ambiguous “previous” targets and privileged always() handlers are shortcuts that trade a visible failure for hidden state corruption. The checkpoint now combines run attempts, a durable fake target, idempotency, fault injection and guarded recovery in one reproducible drill.
Knowledge check
Why is “the publish step was green” stronger evidence than “the whole job was red” for external state?
The job can fail later. Step/provider receipts and direct target queries tell you whether publication actually committed.
What should happen if a rollback guard finds the current target no longer equals the failed deployment you inspected?
Stop and re-inspect. Another actor/change occurred, so the old rollback decision is stale.
Why should privileged recovery not live in an unconditional always() cleanup step?
Failure/cancellation paths can invalidate assumptions, and always() can cause another mutation before the incident is understood.
Does force-cancel restore external state?
No. It stops workflow execution; any external side effects must still be inspected and reconciled.
A targeted job rerun includes dependencies. What must be true of a deployment dependency?
It must be safe to repeat or detect/reuse its existing side effect; otherwise the targeted rerun can duplicate deployment state.
Official references and version notes
- Re-running workflows and jobs — Current rerun window, attempt behavior, actor privileges, same ref/SHA semantics and targeted rerun options.
- Variables reference — Definitions of GITHUB_RUN_ID, GITHUB_RUN_ATTEMPT, GITHUB_RUN_NUMBER, GITHUB_ACTOR and GITHUB_TRIGGERING_ACTOR.
- Contexts reference — Current github context fields and rerun identity semantics.
- REST API: workflow runs — Attempt-specific logs, rerun/cancel endpoints and exact workflow-run resource state.
- GitHub CLI: gh run rerun — Current --failed, --job and --debug rerun controls; job reruns require the database ID.
- GitHub CLI: gh run view — Current --attempt, --log, --log-failed and JSON run/job evidence inspection.
- Concurrency — Current serialization, cancel-in-progress and queueing behavior.
- Workflow syntax: concurrency — Current workflow/job concurrency syntax including queue:max.
- Deployment environments — Environment gates, secret timing and deployment target separation.
- Deployments and environments — Protection rules and deployment governance boundaries.
- actions/checkout v7.0.1 — Full commit SHA used by executable examples.
- actions/upload-artifact v7.0.1 — Full commit SHA used for bounded attempt-specific evidence artifacts.
- upload-artifact behavior — Current immutable-artifact behavior, unique names, outputs and overwrite semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.