Chapter 33Lesson 01~235 minutes

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Core Concepts and Mental Model

Design GitHub Actions recovery so reruns repeat safe work, durable side effects are idempotent, rollback targets are known, and first-failure evidence survives the incident.

RecoveryRerunsIdempotencyRollbackIncident evidence

Learning objectives

  • Explain run, attempt, retry, rerun, new run, compensation and rollback as different recovery operations.
  • Use run/SHA/attempt plus an idempotency key to reason about durable side effects independently of job success.
  • Classify a failure as safe to retry, safe to reconcile, requiring compensation, or requiring operator-approved rollback.
  • Preserve first-attempt logs and external-state evidence before any rerun or cleanup.
  • Explain why concurrency, artifacts, environments and external locks solve different parts of incident safety.

1. The practical problem: the second attempt can be more dangerous than the first

A GitHub Actions failure is not necessarily a clean “nothing happened” state. A job can upload a package, create a release, change a deployment target, write a database migration marker, rotate a credential, update a DNS record or push infrastructure state and then fail on the next step. If an operator responds by pressing Re-run jobs, the workflow may execute the side effect again. The resulting incident is no longer “the deployment failed”; it is “the recovery path duplicated or corrupted state.”

Chapter 30 taught you to preserve observability evidence. This chapter adds the operational half: every mutation must have an identity, a durable checkpoint and a recovery rule. A green rerun is useful only when you can prove what happened in the first attempt, what the rerun repeated, and what external state exists now.

2. Mental model: attempt → side effect → evidence → recovery decision

Start with the workflow run, identified by GITHUB_RUN_ID. A rerun creates another attempt of that same run, identified by the incrementing GITHUB_RUN_ATTEMPT. The workflow executes a side-effecting operation against some external target. Before or immediately after that operation, it records an idempotency key or checkpoint that identifies the logical change. If a later step fails, recovery starts by preserving the run/attempt and target state, not by repeating the mutation.

Incident-safe recovery causality
flowchart TD
  A[Run ID + source SHA] --> B[Attempt N]
  B --> C[Idempotency key / side-effect checkpoint]
  C --> D[External mutation]
  D --> E{Later step fails?}
  E -->|No| F[Verify final state]
  E -->|Yes| G[Preserve logs + target evidence]
  G --> H{Recovery classification}
  H -->|Safe retry/reconcile| I[Targeted rerun same run/SHA]
  H -->|Compensate / rollback| J[Guarded recovery to known target]
  I --> K[Verify no duplicate side effect]
  J --> K
  K --> L[Final reconciliation record]

The arrows are causal. The side effect exists before the later failure. Evidence is preserved before recovery. The recovery choice depends on what already changed. Final success therefore means more than “attempt 2 is green”: the external target, side-effect ledger, source SHA and recovery record must agree.

3. Run identity: what changes on a rerun and what does not

Field/state Rerun of same workflow run New workflow run
GITHUB_RUN_ID Stays the same New unique run ID
GITHUB_RUN_NUMBER Stays the same Advances for a new run of the workflow
GITHUB_RUN_ATTEMPT Increments from 1 Starts at 1
GITHUB_SHA / GITHUB_REF Same as original event Derived from the new event/ref
Privileges Use the actor that triggered the original run Use the new run's triggering identity
Event inputs Same original event payload May be different
External side effects May already exist from an earlier attempt Must be inspected independently

As verified on 2026-09-10, GitHub allows reruns for up to 30 days after the initial run and caps a workflow run at 50 reruns. Those are platform limits, not a recovery policy. Your policy should normally permit far fewer attempts and should stop automatically when state identity is ambiguous.

4. The recovery state ledger

Before a side-effecting job runs, write down the states you need to reconcile. This prevents the common incident mistake of treating a failed step conclusion as proof that the external mutation never happened.

Layer Identity/evidence Recovery question
Event / source event, ref, source SHA, workflow SHA Are we still operating on the exact intended revision?
Run / attempt run ID, run number, attempt, actor, triggering actor Is this a rerun of the same event or a new execution?
Runner / job job database ID, runner label/image, step conclusions What actually executed and where?
Side effect target ID, deployment/release ID, resource version Did the mutation complete before the visible failure?
Idempotency logical key, receipt/checkpoint, expected digest Would repeating the request create a duplicate?
Evidence attempt logs, artifact IDs/digests, API responses Can we reconstruct the first failure after recovery?
Rollback known-good target/digest/state generation What exact state can we safely restore?
Final reconciliation desired SHA/digest versus target state Does the external system match the recorded outcome?

5. Idempotency is a contract, not a GitHub Actions keyword

GitHub Actions has no magic idempotent: true switch. Idempotency comes from the command or external API. A safe pattern derives a stable key from correctness-relevant identity—such as repository ID + environment + artifact digest—and asks the external system to create-or-update or to return the existing receipt for that key. The second execution then converges on the same state instead of creating another release/deployment/payment/message.

Logical operation: deploy repository 1234 / staging / artifact sha256:abcd...
Idempotency key: sha256("1234|staging|sha256:abcd...")

Attempt 1:
  key absent -> create deployment receipt -> activate target -> later health probe fails

Attempt 2:
  same key present -> verify receipt + digest -> do NOT create second deployment -> retry probe/reconcile

The key should not be based only on GITHUB_RUN_ATTEMPT, because that would intentionally make every rerun a different logical mutation. Run/attempt belong in the evidence record; source/artifact/target identity usually belongs in the idempotency key.

6. Checkpoints: prove what completed before the failure

A side-effect checkpoint is a durable receipt written by the external system or by a controlled ledger. Good receipts contain the logical key, target, source SHA, artifact digest, external resource ID, timestamp and originating run ID. A local file on a GitHub-hosted runner is not durable across jobs or reruns; if the real target is remote, the receipt must live with the target or in another durable control plane.

Write the checkpoint at the boundary that matters. For example, if an API returns deployment ID dep-4821 after accepting a mutation, store that ID before starting a health check. If the health check fails, you now know exactly which deployment to inspect or roll back.

7. Retry, targeted rerun, full rerun and new run

Operation Use when Main hazard
In-step retry Transient, side-effect-free or idempotent call with bounded backoff Hiding persistent failures or multiplying mutations
Rerun one job Failure is isolated and dependencies can be safely repeated Dependent jobs may rerun too; side effects must remain idempotent
Rerun failed jobs Failed region of the DAG is safe to repeat Earlier successful side effects may still constrain recovery
Rerun all jobs Whole run can be repeated safely against same SHA/event Broad repetition increases cost and side-effect surface
New run Workflow/source/input must change It is not the same incident attempt; correlate explicitly

The current CLI provides gh run rerun RUN_ID --failed and gh run rerun --job JOB_DATABASE_ID. The browser URL number for a job is not necessarily the database ID expected by --job; inspect it with gh run view RUN_ID --json jobs.

8. Rollback is a new mutation to a known state

Rollback is not time travel and it is not “delete whatever the failed workflow created.” A rollback is an explicit state transition to a previously verified target: artifact digest, deployment revision, database-compatible release, infrastructure state generation or configuration version. You need that target identity before the incident or must derive it from preserved evidence.

Some changes cannot be safely rolled back: irreversible schema migrations, data deletion, external notifications or security credential rotations may require compensation or forward repair. The workflow should encode that distinction rather than pretending every failure has a symmetric undo command.

9. Concurrency prevents overlap; it does not make mutations idempotent

A deployment concurrency group can serialize runs so two workflows do not update the same environment simultaneously. Current GitHub.com also supports queue: max to retain multiple pending entries rather than replacing the previous pending run. This protects ordering, but the same job can still be rerun after a partial side effect. Concurrency and idempotency solve different failure modes.

concurrency:
  group: deploy-staging
  queue: max
# No cancel-in-progress: a deployment already mutating the target should normally
# finish or enter an explicit recovery state instead of being killed casually.

10. First-failure evidence is part of the recovery state

Before rerunning, preserve attempt 1. Current GitHub APIs expose attempt-specific logs, and gh run view RUN_ID --attempt 1 --log lets an operator inspect a previous attempt. Download or copy the relevant evidence before destructive cleanup. Never overwrite a first-failure evidence artifact with a generic name and overwrite: true just to make a rerun pass.

RUN_ID=123456789
# Read the exact failed attempt before any recovery mutation.
gh run view "$RUN_ID" --attempt 1 --json attempt,headSha,status,conclusion,url

gh api \
  -H 'Accept: application/vnd.github+json' \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  repos/{owner}/{repo}/actions/runs/$RUN_ID/attempts/1/logs \
  --include >/tmp/attempt-1-log-response.txt

The REST log endpoint redirects to a short-lived archive URL, so scripts must follow or record the redirect appropriately. The important point is exact run + exact attempt identity, not “download the latest logs.”

11. Rerun identity is subtle: triggering actor and effective privileges can differ

On a rerun, github.triggering_actor identifies who initiated the rerun, while github.actor remains the actor associated with the original workflow run. GitHub documents that the rerun uses the original actor’s privileges. Do not treat “Alice clicked rerun” as proof that the rerun now has Alice’s authorization boundary.

For sensitive recovery, independently re-check environment approvals, external authorization and target guards. A rerun is an execution mechanism, not a substitute for an incident/change-management decision.

12. Lesson summary

Incident-safe GitHub Actions recovery keeps three identities separate: the immutable event/source identity, the evolving run-attempt identity, and the durable external side-effect identity. Preserve the first attempt, inspect what already changed, then choose retry, reconcile, compensate, rollback or a new run. The next lesson builds this model first with local files so the failure mechanics are visible without a cloud account.

Next lesson

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Guided Hands-On Workflow

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Which value changes when you rerun the same workflow run: GITHUB_RUN_ID or GITHUB_RUN_ATTEMPT?

Why is run attempt a poor idempotency key for a deployment?

A deploy step succeeded, the health-check step failed, and the job is red. Can you infer that nothing reached the target?

When should a source-code fix be treated as a new run rather than another attempt?

Does a concurrency group remove the need for idempotency?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.