Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Configuration, Design Patterns, and Trade-Offs
Choose retry, rerun, compensation, rollback, concurrency and fail-open/fail-closed behavior from the side-effect model rather than from convenience.
Learning objectives
- Choose retry, rerun, rollback or compensation from side-effect semantics and evidence.
- Compare create-or-update with blind create, and understand conditional/idempotent APIs.
- Separate GitHub concurrency from external distributed locks and resource-version guards.
- Choose automatic versus operator-approved rollback from reversibility and blast radius.
- Use fail-open/fail-closed decisions explicitly instead of accidental error handling.
1. Recovery architecture starts from side effects, not from YAML syntax
The same gh run rerun --failed command can be perfectly
safe for a pure test workflow and disastrous for package
publication. The difference is not GitHub syntax; it is the external
state transition. Design the recovery path while you design the
mutation, before the first production incident.
For every side effect, answer four questions: How do I detect that it already happened? Can repeating it converge safely? What exact state can I compensate/roll back to? What evidence must survive? The answers determine the workflow shape.
2. Retry versus rerun versus rollback
| Choice | Correct fit | Evidence prerequisite | Typical failure if misused |
|---|---|---|---|
| Bounded retry | Transient read or idempotent write within one step | request/operation ID + bounded attempts | duplicates or hidden persistent failure |
| Targeted rerun | Same event/SHA; failed job region safe to repeat | run ID, attempt, job IDs, side-effect receipt | repeats a completed mutation |
| Full rerun | All jobs are repeatable and same event/SHA still desired | whole-run side-effect inventory | broad duplicate/cost surface |
| Rollback | Known previous state is compatible and reversible | exact current + rollback target IDs/digests | roll back wrong/newer target |
| Compensation | Original mutation cannot be undone exactly | original transaction/resource ID | fake “undo” loses data or creates more damage |
| New run / forward fix | Source/workflow/input must change | new SHA + linkage to failed run | calling changed code “same attempt” |
3. Prefer create-or-update/reconcile over create blindly
A workflow that executes POST /releases with a random
name on every attempt has no safe rerun story. Better external APIs
support a stable natural key, an idempotency key, a conditional
update token (ETag/resourceVersion), or declarative reconciliation.
When such primitives exist, make them part of the contract and store
the returned resource ID.
If the provider has no idempotency primitive, add a durable ledger under your control and guard it transactionally. Do not pretend a runner-local environment variable can serialize two independent workflow runs.
4. Automatic rollback versus operator-approved rollback
| Automatic rollback can fit when… | Operator approval is safer when… |
|---|---|
| The failure signal is high-confidence and target is single-purpose | Signal is ambiguous or depends on external business state |
| Rollback artifact/state is precomputed and compatible | Schema/data compatibility is uncertain |
| Rollback itself is idempotent and tested | Rollback invokes destructive or irreversible operations |
| Blast radius is small and bounded | Production/regulated/high-value systems are affected |
| Evidence is captured before rollback starts | First-failure evidence could be overwritten or lost |
“Automatic” must not mean if: failure() followed by a
privileged deletion. It means an explicitly modeled recovery
transition with exact resource guards, known rollback identity,
bounded permissions and independent post-rollback verification.
5. GitHub concurrency versus an external lock
GitHub concurrency serializes jobs/runs that share a
concurrency group inside the repository. It is excellent
for preventing two repository deployments from overlapping. It does
not coordinate another repository, a manual operator, a GitOps
controller or a cloud-native deployment process that can mutate the
same target.
| Mechanism | Scope | Best use |
|---|---|---|
| GitHub concurrency group | Actions jobs/runs sharing the repository-level group | Serialize one repository's staging/production deployment jobs |
| External distributed lock/lease | All actors that honor the same external lock | Coordinate cross-repository/controllers/operators |
| Resource version / compare-and-swap | One target object/version | Reject stale writes even if scheduling overlaps |
| Idempotency key | One logical operation | Deduplicate repeated requests/attempts |
A robust system often combines them: Actions concurrency reduces accidental overlap; a target resource version prevents stale writes; an idempotency key deduplicates retries.
6. cancel-in-progress is safe only for work whose interruption is understood
For lint/test on obsolete pull-request commits,
cancel-in-progress: true can be an efficient choice.
For a deployment that may be halfway through an external mutation,
cancellation can create exactly the partial state this chapter is
trying to control. Prefer serialization with
queue: max for side-effecting deployment streams unless
the provider itself guarantees transactional cancellation/rollback.
Current GitHub force-cancel can bypass conditions such as
always(). Treat force cancellation as an incident
control for stuck runs, not a normal rollback mechanism. External
target state must still be inspected afterwards.
7. Fail closed versus fail open
| Control | Fail closed | Fail open |
|---|---|---|
| Artifact identity cannot be verified | Stop deployment | Deploy anyway — generally unsafe |
| Idempotency receipt is unreadable | Stop mutation and escalate | Create a new resource — duplicate risk |
| Health telemetry unavailable | Hold/review or use alternate verified signal | Assume healthy — availability/risk trade-off |
| Evidence upload fails after target mutation | Preserve target receipt elsewhere; flag incident | Ignore evidence loss — weak auditability |
| Cleanup fails | Leave known residue and report it | Delete more aggressively — dangerous |
Fail-closed defaults are usually correct for identity, authorization and destructive state transitions. Availability-sensitive systems may intentionally fail open for some telemetry/readiness layers, but the policy must be explicit, bounded and observable.
8. Attempt evidence: immutable names, no overwrite shortcut
Use attempt-qualified evidence names such as
incident-${{ github.run_id }}-${{ github.run_attempt }}. Current artifact actions create immutable artifacts and fail on
duplicate names unless overwrite is requested. In recovery
workflows, overwrite is usually the wrong choice because each
attempt is forensic evidence. Preserve IDs/digests and retention
instead of presenting one mutable “latest incident” artifact.
- name: Preserve attempt evidence
if: ${{ always() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
with:
name: incident-${{ github.run_id }}-${{ github.run_attempt }}
path: evidence/
retention-days: 7
if-no-files-found: error
9. Smallest equivalent rerun scope
Rerun only the scope that can reproduce the failure without adding unrelated side effects. If a failed health-check job depends on a deploy job, GitHub may include dependencies when rerunning a job. That means the deploy dependency must also be safe to repeat or must detect/reuse its existing deployment receipt. The job graph is therefore part of the recovery contract.
If you cannot make the upstream mutation repeatable, separate it from verification into jobs/workflows whose rerun semantics are explicit. A “verify deployment” workflow keyed by exact deployment ID is often safer than making operators rerun an entire publish-and-deploy DAG.
10. Recovery permissions should be narrower than normal operator access
Separate read/evidence jobs from mutation/recovery jobs. A
diagnostic job can normally use permissions: {} or
read-only permissions. The rollback job should receive only the
permission or federated identity needed for the exact target and
should be protected by environment/approval rules where appropriate.
Do not give every job a broad token so any failure handler can “fix”
anything.
11. Worked decision scenarios
| Scenario | Decision | Why |
|---|---|---|
| Package publish returned success; metadata fetch timed out | Read exact package/version; do not republish blindly | Publish may already be durable and version is usually immutable |
| Unit test runner lost network before any external write | Targeted rerun | Side-effect inventory is empty |
| Deployment activated exact digest; health probe timed out | Reconcile/verify same deployment first | Target identity is already correct |
| Migration changed schema incompatibly; app health failed | Operator-approved recovery/forward fix | Rollback compatibility is uncertain |
| Two repos deploy same cluster namespace | External lease/resourceVersion in addition to repository concurrency | GitHub concurrency cannot coordinate independent repos |
| Workflow bug fixed in a new commit | New run | Rerun would execute original workflow/source identity |
12. Production recovery contract
- Identify: exact event, run, attempt, SHA, workflow SHA and effective authorization.
- Inventory: list every external side effect that may have completed.
- Checkpoint: record operation/resource IDs and stable idempotency key.
- Preserve: keep attempt logs/artifacts/API/target evidence before mutation.
- Classify: retry, reconcile, compensate, rollback, forward-fix or stop.
- Guard: verify exact current target/version before recovery write.
- Act: apply the smallest reversible correction with narrow permissions.
- Verify: prove final external state and record reconciliation outcome.
13. Lesson summary
Recovery design is a transaction/state-machine problem wrapped in CI/CD orchestration. GitHub reruns, concurrency and artifacts provide useful primitives, but safe recovery depends on the external operation being identifiable, repeatable or compensatable. Lesson 4 now turns those principles into an evidence-first diagnostic procedure for the failures that cause the most damage.
Knowledge check
When is an in-step retry preferable to rerunning the job?
For a bounded transient operation that is side-effect-free or explicitly idempotent, where keeping the same job context is useful.
Why can automatic rollback from an always() block be unsafe?
always() says only when the step runs, not whether rollback identity, authorization, compatibility or target guards are valid. A privileged failure handler can compound damage.
Two different repositories deploy the same namespace. Is a matching GitHub concurrency group enough?
No. Concurrency groups are repository-scoped; use an external lock/lease or target resource-version guard shared by all actors.
A provider supports ETags/resource versions. What recovery property does that add?
Compare-and-swap protection: a stale workflow can be rejected if the target changed since it was inspected.
Why should attempt evidence artifacts use run ID plus attempt?
It preserves each attempt as a distinct forensic object rather than overwriting the first failure with a later rerun.
Official references and version notes
- Re-running workflows and jobs — Current rerun window, attempt behavior, actor privileges, same ref/SHA semantics and targeted rerun options.
- Variables reference — Definitions of GITHUB_RUN_ID, GITHUB_RUN_ATTEMPT, GITHUB_RUN_NUMBER, GITHUB_ACTOR and GITHUB_TRIGGERING_ACTOR.
- Contexts reference — Current github context fields and rerun identity semantics.
- REST API: workflow runs — Attempt-specific logs, rerun/cancel endpoints and exact workflow-run resource state.
- GitHub CLI: gh run rerun — Current --failed, --job and --debug rerun controls; job reruns require the database ID.
- GitHub CLI: gh run view — Current --attempt, --log, --log-failed and JSON run/job evidence inspection.
- Concurrency — Current serialization, cancel-in-progress and queueing behavior.
- Workflow syntax: concurrency — Current workflow/job concurrency syntax including queue:max.
- Deployment environments — Environment gates, secret timing and deployment target separation.
- Deployments and environments — Protection rules and deployment governance boundaries.
- actions/checkout v7.0.1 — Full commit SHA used by executable examples.
- actions/upload-artifact v7.0.1 — Full commit SHA used for bounded attempt-specific evidence artifacts.
- upload-artifact behavior — Current immutable-artifact behavior, unique names, outputs and overwrite semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.