Chapter 33Lesson 03~235 minutes

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Configuration, Design Patterns, and Trade-Offs

Choose retry, rerun, compensation, rollback, concurrency and fail-open/fail-closed behavior from the side-effect model rather than from convenience.

Trade-offsConcurrencyExternal locksFail closedRollback policy

Learning objectives

  • Choose retry, rerun, rollback or compensation from side-effect semantics and evidence.
  • Compare create-or-update with blind create, and understand conditional/idempotent APIs.
  • Separate GitHub concurrency from external distributed locks and resource-version guards.
  • Choose automatic versus operator-approved rollback from reversibility and blast radius.
  • Use fail-open/fail-closed decisions explicitly instead of accidental error handling.

1. Recovery architecture starts from side effects, not from YAML syntax

The same gh run rerun --failed command can be perfectly safe for a pure test workflow and disastrous for package publication. The difference is not GitHub syntax; it is the external state transition. Design the recovery path while you design the mutation, before the first production incident.

For every side effect, answer four questions: How do I detect that it already happened? Can repeating it converge safely? What exact state can I compensate/roll back to? What evidence must survive? The answers determine the workflow shape.

2. Retry versus rerun versus rollback

Choice Correct fit Evidence prerequisite Typical failure if misused
Bounded retry Transient read or idempotent write within one step request/operation ID + bounded attempts duplicates or hidden persistent failure
Targeted rerun Same event/SHA; failed job region safe to repeat run ID, attempt, job IDs, side-effect receipt repeats a completed mutation
Full rerun All jobs are repeatable and same event/SHA still desired whole-run side-effect inventory broad duplicate/cost surface
Rollback Known previous state is compatible and reversible exact current + rollback target IDs/digests roll back wrong/newer target
Compensation Original mutation cannot be undone exactly original transaction/resource ID fake “undo” loses data or creates more damage
New run / forward fix Source/workflow/input must change new SHA + linkage to failed run calling changed code “same attempt”

3. Prefer create-or-update/reconcile over create blindly

A workflow that executes POST /releases with a random name on every attempt has no safe rerun story. Better external APIs support a stable natural key, an idempotency key, a conditional update token (ETag/resourceVersion), or declarative reconciliation. When such primitives exist, make them part of the contract and store the returned resource ID.

If the provider has no idempotency primitive, add a durable ledger under your control and guard it transactionally. Do not pretend a runner-local environment variable can serialize two independent workflow runs.

4. Automatic rollback versus operator-approved rollback

Automatic rollback can fit when… Operator approval is safer when…
The failure signal is high-confidence and target is single-purpose Signal is ambiguous or depends on external business state
Rollback artifact/state is precomputed and compatible Schema/data compatibility is uncertain
Rollback itself is idempotent and tested Rollback invokes destructive or irreversible operations
Blast radius is small and bounded Production/regulated/high-value systems are affected
Evidence is captured before rollback starts First-failure evidence could be overwritten or lost

“Automatic” must not mean if: failure() followed by a privileged deletion. It means an explicitly modeled recovery transition with exact resource guards, known rollback identity, bounded permissions and independent post-rollback verification.

5. GitHub concurrency versus an external lock

GitHub concurrency serializes jobs/runs that share a concurrency group inside the repository. It is excellent for preventing two repository deployments from overlapping. It does not coordinate another repository, a manual operator, a GitOps controller or a cloud-native deployment process that can mutate the same target.

Mechanism Scope Best use
GitHub concurrency group Actions jobs/runs sharing the repository-level group Serialize one repository's staging/production deployment jobs
External distributed lock/lease All actors that honor the same external lock Coordinate cross-repository/controllers/operators
Resource version / compare-and-swap One target object/version Reject stale writes even if scheduling overlaps
Idempotency key One logical operation Deduplicate repeated requests/attempts

A robust system often combines them: Actions concurrency reduces accidental overlap; a target resource version prevents stale writes; an idempotency key deduplicates retries.

6. cancel-in-progress is safe only for work whose interruption is understood

For lint/test on obsolete pull-request commits, cancel-in-progress: true can be an efficient choice. For a deployment that may be halfway through an external mutation, cancellation can create exactly the partial state this chapter is trying to control. Prefer serialization with queue: max for side-effecting deployment streams unless the provider itself guarantees transactional cancellation/rollback.

Current GitHub force-cancel can bypass conditions such as always(). Treat force cancellation as an incident control for stuck runs, not a normal rollback mechanism. External target state must still be inspected afterwards.

7. Fail closed versus fail open

Control Fail closed Fail open
Artifact identity cannot be verified Stop deployment Deploy anyway — generally unsafe
Idempotency receipt is unreadable Stop mutation and escalate Create a new resource — duplicate risk
Health telemetry unavailable Hold/review or use alternate verified signal Assume healthy — availability/risk trade-off
Evidence upload fails after target mutation Preserve target receipt elsewhere; flag incident Ignore evidence loss — weak auditability
Cleanup fails Leave known residue and report it Delete more aggressively — dangerous

Fail-closed defaults are usually correct for identity, authorization and destructive state transitions. Availability-sensitive systems may intentionally fail open for some telemetry/readiness layers, but the policy must be explicit, bounded and observable.

8. Attempt evidence: immutable names, no overwrite shortcut

Use attempt-qualified evidence names such as incident-${{ github.run_id }}-${{ github.run_attempt }}. Current artifact actions create immutable artifacts and fail on duplicate names unless overwrite is requested. In recovery workflows, overwrite is usually the wrong choice because each attempt is forensic evidence. Preserve IDs/digests and retention instead of presenting one mutable “latest incident” artifact.

- name: Preserve attempt evidence
  if: ${{ always() }}
  uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
  with:
    name: incident-${{ github.run_id }}-${{ github.run_attempt }}
    path: evidence/
    retention-days: 7
    if-no-files-found: error

9. Smallest equivalent rerun scope

Rerun only the scope that can reproduce the failure without adding unrelated side effects. If a failed health-check job depends on a deploy job, GitHub may include dependencies when rerunning a job. That means the deploy dependency must also be safe to repeat or must detect/reuse its existing deployment receipt. The job graph is therefore part of the recovery contract.

If you cannot make the upstream mutation repeatable, separate it from verification into jobs/workflows whose rerun semantics are explicit. A “verify deployment” workflow keyed by exact deployment ID is often safer than making operators rerun an entire publish-and-deploy DAG.

10. Recovery permissions should be narrower than normal operator access

Separate read/evidence jobs from mutation/recovery jobs. A diagnostic job can normally use permissions: {} or read-only permissions. The rollback job should receive only the permission or federated identity needed for the exact target and should be protected by environment/approval rules where appropriate. Do not give every job a broad token so any failure handler can “fix” anything.

11. Worked decision scenarios

Scenario Decision Why
Package publish returned success; metadata fetch timed out Read exact package/version; do not republish blindly Publish may already be durable and version is usually immutable
Unit test runner lost network before any external write Targeted rerun Side-effect inventory is empty
Deployment activated exact digest; health probe timed out Reconcile/verify same deployment first Target identity is already correct
Migration changed schema incompatibly; app health failed Operator-approved recovery/forward fix Rollback compatibility is uncertain
Two repos deploy same cluster namespace External lease/resourceVersion in addition to repository concurrency GitHub concurrency cannot coordinate independent repos
Workflow bug fixed in a new commit New run Rerun would execute original workflow/source identity

12. Production recovery contract

  1. Identify: exact event, run, attempt, SHA, workflow SHA and effective authorization.
  2. Inventory: list every external side effect that may have completed.
  3. Checkpoint: record operation/resource IDs and stable idempotency key.
  4. Preserve: keep attempt logs/artifacts/API/target evidence before mutation.
  5. Classify: retry, reconcile, compensate, rollback, forward-fix or stop.
  6. Guard: verify exact current target/version before recovery write.
  7. Act: apply the smallest reversible correction with narrow permissions.
  8. Verify: prove final external state and record reconciliation outcome.

13. Lesson summary

Recovery design is a transaction/state-machine problem wrapped in CI/CD orchestration. GitHub reruns, concurrency and artifacts provide useful primitives, but safe recovery depends on the external operation being identifiable, repeatable or compensatable. Lesson 4 now turns those principles into an evidence-first diagnostic procedure for the failures that cause the most damage.

Next lesson

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Diagnostics, Failure Modes, and Production Practices

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

When is an in-step retry preferable to rerunning the job?

Why can automatic rollback from an always() block be unsafe?

Two different repositories deploy the same namespace. Is a matching GitHub concurrency group enough?

A provider supports ETags/resource versions. What recovery property does that add?

Why should attempt evidence artifacts use run ID plus attempt?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.