Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Core Concepts and Mental Model
Design GitHub Actions recovery so reruns repeat safe work, durable side effects are idempotent, rollback targets are known, and first-failure evidence survives the incident.
Learning objectives
- Explain run, attempt, retry, rerun, new run, compensation and rollback as different recovery operations.
- Use run/SHA/attempt plus an idempotency key to reason about durable side effects independently of job success.
- Classify a failure as safe to retry, safe to reconcile, requiring compensation, or requiring operator-approved rollback.
- Preserve first-attempt logs and external-state evidence before any rerun or cleanup.
- Explain why concurrency, artifacts, environments and external locks solve different parts of incident safety.
1. The practical problem: the second attempt can be more dangerous than the first
A GitHub Actions failure is not necessarily a clean “nothing happened” state. A job can upload a package, create a release, change a deployment target, write a database migration marker, rotate a credential, update a DNS record or push infrastructure state and then fail on the next step. If an operator responds by pressing Re-run jobs, the workflow may execute the side effect again. The resulting incident is no longer “the deployment failed”; it is “the recovery path duplicated or corrupted state.”
Chapter 30 taught you to preserve observability evidence. This chapter adds the operational half: every mutation must have an identity, a durable checkpoint and a recovery rule. A green rerun is useful only when you can prove what happened in the first attempt, what the rerun repeated, and what external state exists now.
2. Mental model: attempt → side effect → evidence → recovery decision
Start with the workflow run, identified by
GITHUB_RUN_ID. A rerun creates another
attempt of that same run, identified by the
incrementing GITHUB_RUN_ATTEMPT. The workflow executes
a side-effecting operation against some external target. Before or
immediately after that operation, it records an
idempotency key or checkpoint that identifies the
logical change. If a later step fails, recovery starts by preserving
the run/attempt and target state, not by repeating the mutation.
flowchart TD
A[Run ID + source SHA] --> B[Attempt N]
B --> C[Idempotency key / side-effect checkpoint]
C --> D[External mutation]
D --> E{Later step fails?}
E -->|No| F[Verify final state]
E -->|Yes| G[Preserve logs + target evidence]
G --> H{Recovery classification}
H -->|Safe retry/reconcile| I[Targeted rerun same run/SHA]
H -->|Compensate / rollback| J[Guarded recovery to known target]
I --> K[Verify no duplicate side effect]
J --> K
K --> L[Final reconciliation record]
The arrows are causal. The side effect exists before the later failure. Evidence is preserved before recovery. The recovery choice depends on what already changed. Final success therefore means more than “attempt 2 is green”: the external target, side-effect ledger, source SHA and recovery record must agree.
3. Run identity: what changes on a rerun and what does not
| Field/state | Rerun of same workflow run | New workflow run |
|---|---|---|
GITHUB_RUN_ID |
Stays the same | New unique run ID |
GITHUB_RUN_NUMBER |
Stays the same | Advances for a new run of the workflow |
GITHUB_RUN_ATTEMPT |
Increments from 1 | Starts at 1 |
GITHUB_SHA / GITHUB_REF |
Same as original event | Derived from the new event/ref |
| Privileges | Use the actor that triggered the original run | Use the new run's triggering identity |
| Event inputs | Same original event payload | May be different |
| External side effects | May already exist from an earlier attempt | Must be inspected independently |
As verified on 2026-09-10, GitHub allows reruns for up to 30 days after the initial run and caps a workflow run at 50 reruns. Those are platform limits, not a recovery policy. Your policy should normally permit far fewer attempts and should stop automatically when state identity is ambiguous.
4. The recovery state ledger
Before a side-effecting job runs, write down the states you need to reconcile. This prevents the common incident mistake of treating a failed step conclusion as proof that the external mutation never happened.
| Layer | Identity/evidence | Recovery question |
|---|---|---|
| Event / source | event, ref, source SHA, workflow SHA | Are we still operating on the exact intended revision? |
| Run / attempt | run ID, run number, attempt, actor, triggering actor | Is this a rerun of the same event or a new execution? |
| Runner / job | job database ID, runner label/image, step conclusions | What actually executed and where? |
| Side effect | target ID, deployment/release ID, resource version | Did the mutation complete before the visible failure? |
| Idempotency | logical key, receipt/checkpoint, expected digest | Would repeating the request create a duplicate? |
| Evidence | attempt logs, artifact IDs/digests, API responses | Can we reconstruct the first failure after recovery? |
| Rollback | known-good target/digest/state generation | What exact state can we safely restore? |
| Final reconciliation | desired SHA/digest versus target state | Does the external system match the recorded outcome? |
5. Idempotency is a contract, not a GitHub Actions keyword
GitHub Actions has no magic idempotent: true switch.
Idempotency comes from the command or external API. A safe pattern
derives a stable key from correctness-relevant identity—such as
repository ID + environment + artifact digest—and asks the external
system to create-or-update or to return the existing receipt for
that key. The second execution then converges on the same state
instead of creating another release/deployment/payment/message.
Logical operation: deploy repository 1234 / staging / artifact sha256:abcd...
Idempotency key: sha256("1234|staging|sha256:abcd...")
Attempt 1:
key absent -> create deployment receipt -> activate target -> later health probe fails
Attempt 2:
same key present -> verify receipt + digest -> do NOT create second deployment -> retry probe/reconcile
The key should not be based only on GITHUB_RUN_ATTEMPT,
because that would intentionally make every rerun a different
logical mutation. Run/attempt belong in the evidence record;
source/artifact/target identity usually belongs in the idempotency
key.
6. Checkpoints: prove what completed before the failure
A side-effect checkpoint is a durable receipt written by the external system or by a controlled ledger. Good receipts contain the logical key, target, source SHA, artifact digest, external resource ID, timestamp and originating run ID. A local file on a GitHub-hosted runner is not durable across jobs or reruns; if the real target is remote, the receipt must live with the target or in another durable control plane.
Write the checkpoint at the boundary that matters. For example, if
an API returns deployment ID dep-4821 after accepting a
mutation, store that ID before starting a health check. If the
health check fails, you now know exactly which deployment to inspect
or roll back.
7. Retry, targeted rerun, full rerun and new run
| Operation | Use when | Main hazard |
|---|---|---|
| In-step retry | Transient, side-effect-free or idempotent call with bounded backoff | Hiding persistent failures or multiplying mutations |
| Rerun one job | Failure is isolated and dependencies can be safely repeated | Dependent jobs may rerun too; side effects must remain idempotent |
| Rerun failed jobs | Failed region of the DAG is safe to repeat | Earlier successful side effects may still constrain recovery |
| Rerun all jobs | Whole run can be repeated safely against same SHA/event | Broad repetition increases cost and side-effect surface |
| New run | Workflow/source/input must change | It is not the same incident attempt; correlate explicitly |
The current CLI provides
gh run rerun RUN_ID --failed and
gh run rerun --job JOB_DATABASE_ID. The browser URL
number for a job is not necessarily the database ID expected by
--job; inspect it with
gh run view RUN_ID --json jobs.
8. Rollback is a new mutation to a known state
Rollback is not time travel and it is not “delete whatever the failed workflow created.” A rollback is an explicit state transition to a previously verified target: artifact digest, deployment revision, database-compatible release, infrastructure state generation or configuration version. You need that target identity before the incident or must derive it from preserved evidence.
Some changes cannot be safely rolled back: irreversible schema migrations, data deletion, external notifications or security credential rotations may require compensation or forward repair. The workflow should encode that distinction rather than pretending every failure has a symmetric undo command.
9. Concurrency prevents overlap; it does not make mutations idempotent
A deployment concurrency group can serialize runs so two workflows
do not update the same environment simultaneously. Current
GitHub.com also supports queue: max to retain multiple
pending entries rather than replacing the previous pending run. This
protects ordering, but the same job can still be rerun after a
partial side effect. Concurrency and idempotency solve different
failure modes.
concurrency:
group: deploy-staging
queue: max
# No cancel-in-progress: a deployment already mutating the target should normally
# finish or enter an explicit recovery state instead of being killed casually.
10. First-failure evidence is part of the recovery state
Before rerunning, preserve attempt 1. Current GitHub APIs expose
attempt-specific logs, and
gh run view RUN_ID --attempt 1 --log lets an operator
inspect a previous attempt. Download or copy the relevant evidence
before destructive cleanup. Never overwrite a first-failure evidence
artifact with a generic name and overwrite: true just
to make a rerun pass.
RUN_ID=123456789
# Read the exact failed attempt before any recovery mutation.
gh run view "$RUN_ID" --attempt 1 --json attempt,headSha,status,conclusion,url
gh api \
-H 'Accept: application/vnd.github+json' \
-H 'X-GitHub-Api-Version: 2026-03-10' \
repos/{owner}/{repo}/actions/runs/$RUN_ID/attempts/1/logs \
--include >/tmp/attempt-1-log-response.txt
The REST log endpoint redirects to a short-lived archive URL, so scripts must follow or record the redirect appropriately. The important point is exact run + exact attempt identity, not “download the latest logs.”
11. Rerun identity is subtle: triggering actor and effective privileges can differ
On a rerun, github.triggering_actor identifies who
initiated the rerun, while github.actor remains the
actor associated with the original workflow run. GitHub documents
that the rerun uses the original actor’s privileges. Do not treat
“Alice clicked rerun” as proof that the rerun now has Alice’s
authorization boundary.
For sensitive recovery, independently re-check environment approvals, external authorization and target guards. A rerun is an execution mechanism, not a substitute for an incident/change-management decision.
12. Lesson summary
Incident-safe GitHub Actions recovery keeps three identities separate: the immutable event/source identity, the evolving run-attempt identity, and the durable external side-effect identity. Preserve the first attempt, inspect what already changed, then choose retry, reconcile, compensate, rollback or a new run. The next lesson builds this model first with local files so the failure mechanics are visible without a cloud account.
Knowledge check
Which value changes when you rerun the same workflow run: GITHUB_RUN_ID or GITHUB_RUN_ATTEMPT?
GITHUB_RUN_ATTEMPT increments. GITHUB_RUN_ID and GITHUB_RUN_NUMBER stay the same for reruns of that run.
Why is run attempt a poor idempotency key for a deployment?
Because every rerun would intentionally produce a new key and could create another side effect. The logical desired state—target plus artifact/source identity—should drive the key.
A deploy step succeeded, the health-check step failed, and the job is red. Can you infer that nothing reached the target?
No. The durable mutation may already exist. Inspect the target/deployment receipt before deciding whether a rerun is safe.
When should a source-code fix be treated as a new run rather than another attempt?
When the workflow or source revision changes. A rerun keeps the original event ref/SHA, so new source requires a new event/run.
Does a concurrency group remove the need for idempotency?
No. Concurrency prevents overlapping executions; a later rerun can still repeat a completed side effect unless the operation itself is idempotent/reconciled.
Official references and version notes
- Re-running workflows and jobs — Current rerun window, attempt behavior, actor privileges, same ref/SHA semantics and targeted rerun options.
- Variables reference — Definitions of GITHUB_RUN_ID, GITHUB_RUN_ATTEMPT, GITHUB_RUN_NUMBER, GITHUB_ACTOR and GITHUB_TRIGGERING_ACTOR.
- Contexts reference — Current github context fields and rerun identity semantics.
- REST API: workflow runs — Attempt-specific logs, rerun/cancel endpoints and exact workflow-run resource state.
- GitHub CLI: gh run rerun — Current --failed, --job and --debug rerun controls; job reruns require the database ID.
- GitHub CLI: gh run view — Current --attempt, --log, --log-failed and JSON run/job evidence inspection.
- Concurrency — Current serialization, cancel-in-progress and queueing behavior.
- Workflow syntax: concurrency — Current workflow/job concurrency syntax including queue:max.
- Deployment environments — Environment gates, secret timing and deployment target separation.
- Deployments and environments — Protection rules and deployment governance boundaries.
- actions/checkout v7.0.1 — Full commit SHA used by executable examples.
- actions/upload-artifact v7.0.1 — Full commit SHA used for bounded attempt-specific evidence artifacts.
- upload-artifact behavior — Current immutable-artifact behavior, unique names, outputs and overwrite semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.