Chapter 33Lesson 04~245 minutes

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Diagnostics, Failure Modes, and Production Practices

Diagnose duplicate publication, destructive cleanup, revision drift, unknown rollback targets and unsafe always() recovery by preserving evidence before mutation.

DiagnosticsFirst failureDuplicate side effectsDriftSafe repair

Learning objectives

  • Use a fixed evidence-first sequence before retrying or rolling back any failed workflow.
  • Diagnose blind publication retries, destructive cleanup, source-revision drift and unknown rollback targets causally.
  • Explain why privileged always() cleanup and force-cancel can compound incidents.
  • Use attempt-specific logs, job IDs, artifacts and external resource IDs to preserve first failure.
  • Apply the least destructive recovery and rerun only an equivalent, safe scope.

1. Diagnostic rule: a red job is not an external-state description

The highest-risk recovery failures begin with the same mistaken inference: “the job failed, therefore the side effect failed.” A package registry, cloud API or deployment controller can commit the mutation and then lose the response; the next verification step can fail after a successful deployment; cancellation can interrupt the workflow after the target already changed. Diagnose the external state independently.

Do not start by deleting the partial target or rerunning. Start by freezing the incident identity: run ID, attempt, source/workflow SHA, actor, target and every resource/receipt returned before the failure.

2. Evidence-first diagnostic sequence

  1. Preserve run/attempt and first-failure evidence. Save attempt-specific logs and current target metadata.
  2. Confirm event/ref/SHA and workflow revision. Reruns use the original ref/SHA; do not assume a later commit is present.
  3. Confirm evaluated conditions, permissions and inputs. Check original actor privilege and rerun initiator separately.
  4. Inspect job graph and queue. Know which dependencies a targeted rerun would repeat.
  5. Confirm runner/image/toolchain. Separate runner drift from target-state failures.
  6. Inspect failing step/action/runtime/network. Find the first causal error, not the last cleanup error.
  7. Inspect outputs/artifacts/caches. Validate IDs/digests and preserve attempt evidence.
  8. Inspect environment/deployment/external target. Query exact target/resource ID directly.
  9. Apply the least destructive correction. Prefer verify/reconcile over delete/recreate.
  10. Rerun the smallest equivalent scope. Or create a new run when source/workflow/input must change.

3. Broken example: publish, fail verification, delete and republish

name: BROKEN — blind publish recovery
on:
  workflow_dispatch:

permissions:
  contents: write

jobs:
  publish-and-clean:
    runs-on: ubuntu-24.04
    steps:
      - run: ./publish-package.sh               # may succeed remotely
      - run: ./verify-package.sh                # times out, job becomes red
      - if: ${{ always() }}
        run: ./delete-and-republish-latest.sh   # DANGEROUS: ambiguous + destructive

This workflow contains three causal hazards. First, publish-package.sh may have succeeded remotely even if the next step fails. Second, “latest” is not an exact resource identity. Third, always() schedules privileged cleanup even after cancellation/failure, turning a diagnostic failure into a destructive mutation. The red conclusion hides whether zero, one or multiple versions now exist.

4. Repair blind package publication without erasing evidence

Preserve the publish response/version/digest immediately. On failure, query that exact version. If it exists and matches the expected digest, treat publication as completed and retry only downstream verification. If it does not exist and the API guarantees the request was not committed, an idempotent publish may be retried. If state is ambiguous, stop and escalate—do not mint another version or delete the namespace.

Evidence packet before repair:
  run_id: 123456789
  attempt: 1
  source_sha: 012345...
  intended_version: 1.2.3
  expected_digest: sha256:...
  publish_request_id: req-4821
  provider_response: timeout after request body sent
  registry_query(1.2.3): EXISTS / digest matches

Decision: publication completed -> do not republish -> rerun verification only

5. Failure mode: deleting a partial target before understanding it

“Delete and recreate” destroys the strongest evidence you have: the exact state produced by the failed attempt. It may also remove healthy subresources, release a lock, invalidate a rollback target or race with another actor. Prefer read-only target snapshots, deployment history and exact resource versions first.

Cleanup should occur only after the target is reconciled and evidence retention requirements are met. In the checkpoint, if: always() is allowed only for read-only evidence collection/upload—never for privileged rollback or branch mutation.

6. Failure mode: rerunning with a different revision and calling it the same attempt

GitHub prevents one common form of this confusion: an actual rerun keeps the original GITHUB_SHA/GITHUB_REF. The confusion appears when teams commit a fix, trigger a new run and describe it as “attempt 2.” That loses the audit link between the failed revision and the repaired revision.

Use precise language: run 123 attempt 1 failed; run 123 attempt 2 reproduced/reconciled the same SHA; run 124 at SHA B tested the code change. If the fix changes workflow behavior, preserve both run URLs and SHAs.

7. Failure mode: “rollback” without an exact artifact/state target

A command such as deploy previous is ambiguous when multiple deployments occurred or another actor changed the target. The rollback input should be an immutable digest, revision ID or state generation captured before the failed mutation. Then compare the current target to the expected failed deployment before writing.

EXPECTED_CURRENT='sha256:failed-but-known-digest'
ROLLBACK_TO='sha256:verified-previous-digest'
ACTUAL_CURRENT=$(fake-provider get staging --output digest)

if [[ "$ACTUAL_CURRENT" != "$EXPECTED_CURRENT" ]]; then
  echo 'Target changed since incident capture; refuse rollback.' >&2
  exit 1
fi
fake-provider promote --environment staging --digest "$ROLLBACK_TO"

The guard is more important than the rollback command. It prevents a delayed recovery job from overwriting a newer legitimate deployment.

8. Failure mode: privileged always() cleanup

always() is useful for summaries, logs and evidence artifacts. It is hazardous for high-impact cleanup because it runs in failure/cancellation paths where assumptions may be invalid. GitHub even provides a force-cancel endpoint that bypasses conditions such as always() for stuck runs, reinforcing that always-running logic is not a transaction manager.

Move recovery mutations into an explicit job/workflow with exact inputs and guards. Require environment approval or incident authorization when appropriate. Keep the failure handler itself read-only whenever possible.

9. Diagnose targeted rerun dependencies before clicking

RUN_ID=123456789
# Preserve and list exact job database IDs.
gh run view "$RUN_ID" --attempt 1 --json jobs   --jq '.jobs[] | {name,databaseId,status,conclusion}'

# Review first failed logs before mutation.
gh run view "$RUN_ID" --attempt 1 --log-failed > "incident-$RUN_ID-a1-failed.log"

# Only after classifying dependencies as safe:
gh run rerun "$RUN_ID" --failed
# or: gh run rerun --job JOB_DATABASE_ID

The CLI documents that rerunning a job includes dependencies. A dependency that publishes or deploys must therefore be idempotent/reconciling even if it was green in the original attempt.

10. Cancellation is not rollback

A normal cancel request returns asynchronously, and a force-cancel exists for workflows that ignore normal cancellation. Neither operation restores an external target. After cancellation, inspect deployment/resource IDs, partial artifacts and locks. If the target changed, enter the same reconcile/rollback decision process as any other partial failure.

11. Separate causal failure layers

Observed symptom Likely layer to inspect first Do not jump directly to
Workflow never ran event/filter/workflow selection/policy deployment rollback
Job queued unusually long runner capacity/concurrency re-publish artifact
403 on recovery API token/OIDC/environment authorization overly broad write token
Artifact digest mismatch build/artifact identity deploy “latest”
Deploy API success, health failure external target/runtime/health assume deploy did not happen
Rollback guard says target changed concurrent/external actor state force overwrite
Evidence upload failed artifact service/evidence path delete target to “clean up”

12. Production incident-safe practices

  • Create explicit idempotency keys and persist provider/resource receipts.
  • Name evidence by run ID + attempt; retain first failure separately.
  • Serialize side-effecting deployments where ordering matters, but still design idempotency.
  • Do not rebuild release artifacts during recovery; promote/rollback exact verified digests.
  • Use narrow recovery identities and environment/incident approvals for privileged mutations.
  • Record rollback compatibility and last-known-good target before deployment.
  • Test fault injection in disposable environments so the recovery path is exercised before production.

13. Lesson summary

A safe incident response preserves evidence, proves external state, and changes only the causal layer. Blind reruns, delete-and-recreate, ambiguous “previous” targets and privileged always() handlers are shortcuts that trade a visible failure for hidden state corruption. The checkpoint now combines run attempts, a durable fake target, idempotency, fault injection and guarded recovery in one reproducible drill.

Next lesson

Checkpoint Lab — Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why is “the publish step was green” stronger evidence than “the whole job was red” for external state?

What should happen if a rollback guard finds the current target no longer equals the failed deployment you inspected?

Why should privileged recovery not live in an unconditional always() cleanup step?

Does force-cancel restore external state?

A targeted job rerun includes dependencies. What must be true of a deployment dependency?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.