Chapter 32Lesson 04~255 minutes

Production Capstone: Govern a GitHub Organization and Build a Secure End-to-End Delivery System: Failure Injection, Troubleshooting, and Recovery Drill

Inject failures across interacting controls, credentials, provenance, access, automation, and cost pressure; diagnose from preserved evidence, contain risk, recover safely, and record preventive actions.

Failure injectionIncident responseRecoveryResidual riskRunbooks

Learning objectives

  • Use an evidence-first diagnostic sequence that preserves the original cause before containment and repair.
  • Diagnose an interaction failure between a required-check ruleset and over-broad path filtering without bypassing policy.
  • Contain runner/workflow authority, credential leaks, provenance mismatches, residual access, and incomplete API inventory safely.
  • Treat performance and cost pressure as a control-design input instead of a reason to disable governance.
  • Record recovery time, evidence, residual risk, preventive control, and owner for each injected incident.

1. The recovery loop

Use the same sequence for every capstone incident: preserve evidence → identify actor/repository/workflow/ref/artifact → inspect rules/permissions/logs/audit/API state → contain risk → choose the least disruptive correction → verify → capture lessons learned. This ordering prevents a frantic “fix” from destroying evidence or widening the incident.

Record for every drill Why it matters
Detection time / recovery time Shows whether the team can meet an operational objective, not merely describe a procedure.
Original evidence Preserves the cause instead of replacing it with the post-fix state.
Containment action Stops additional harm before optimization or cleanup.
Verification Proves the correction worked independently.
Residual risk States what remains uncertain or intentionally accepted.
Preventive action + owner Turns the drill into platform improvement.

2. Interaction failure: ruleset + required check + path filter

This failure is intentionally caused by two controls interacting. The ruleset requires the status check named verify. A cost-optimization change then filters the entire workflow so docs-only PRs do not create that check. The PR is safe but cannot merge because the required check never exists.

# BROKEN experiment — only on a disposable branch.
# The entire workflow is filtered, so docs-only changes may never create the required "verify" check.
on:
  pull_request:
    paths:
      - "src/**"
  1. Preserve: record PR number/head SHA, ruleset JSON, Actions run list, and the absence/pending state of the required check.
  2. Scope: the repository policy is working as configured; the workflow trigger is the incompatible control.
  3. Contain: do not bypass the ruleset just to merge the docs change.
  4. Repair: keep pull_request: broad enough that the stable aggregate verify check is always created. Put selective/expensive execution inside jobs/steps while the aggregate check reports a deterministic result.
  5. Verify: push the repaired workflow on the branch, observe verify complete, then confirm the same ruleset now permits the normal merge path.
Lesson: path filtering is not merely a cost feature. When a check is required by policy, trigger design becomes part of availability and governance.

3. Compromised workflow or runner has excessive authority

Do not create a privileged self-hosted runner for this drill. Analyze the following fixture as if it appeared during review:

runs-on: [self-hosted, production-network]
permissions:
  contents: write
  packages: write
  id-token: write
# Runner also has network reachability to internal deployment systems.

The problem is not one permission. A compromised job could alter repository content, publish packages, request OIDC identity, and reach a sensitive network. Containment must therefore span GitHub credentials, external provider trust, and runner/network isolation.

  1. Stop or disable the affected workflow/runner path and preserve job/runner logs.
  2. Revoke or invalidate reachable durable credentials and provider sessions; inspect OIDC/provider audit events if external federation exists.
  3. Remove unnecessary GitHub token permissions; move release/deploy authority into a narrowly triggered job/environment.
  4. Quarantine/rebuild the runner from trusted image if self-hosted; review egress and ambient credentials.
  5. Re-run on a trusted executor and compare source/artifact identities before restoring promotion.

4. Credential leak response starts with Git cleanup instead of revocation

The intentionally broken response is: “remove the file, rewrite history, force push, then decide whether to rotate.” That reverses containment priority. If an attacker copied the credential before cleanup, history rewriting does nothing to invalidate it.

Correct order for a real credential: revoke/rotate immediately → preserve incident evidence → determine where/how it was used → update dependent systems → remove exposed content from current Git/non-Git surfaces → assess whether history rewrite is justified → coordinate rewrite if approved → verify old credential is unusable.

The capstone never creates a real credential, so the drill uses a redacted synthetic incident record. Any history rewrite or force update is destructive; do not run it merely to make a training repository look clean.

5. Release, package, and attestation identities disagree

Copy the legitimate artifact, change one byte, and prove that provenance verification rejects the modified subject. This is a safe, local failure injection.

cp release-download/capstone.tar.gz /tmp/capstone-tampered.tar.gz
printf 'training-tamper' >> /tmp/capstone-tampered.tar.gz
sha256sum release-download/capstone.tar.gz /tmp/capstone-tampered.tar.gz

set +e
gh attestation verify /tmp/capstone-tampered.tar.gz   --repo "$FULL"   --signer-workflow "$FULL/.github/workflows/capstone.yml"   --source-ref refs/heads/main   --deny-self-hosted-runners
VERIFY_RC=$?
set -e
printf 'tampered_verify_exit=%s\n' "$VERIFY_RC"
test "$VERIFY_RC" -ne 0

Now imagine a package tag points to different bytes than the GitHub Release asset while the attestation belongs to only one digest. Freeze promotion. Compare the source tag SHA, release asset digest, registry digest, and attestation subject. Repair by producing a new, coherent version—not by moving the old tag or pretending the attestation covers different bytes.

6. Offboarding leaves residual direct access

GitHub access is additive. Removing one team does not prove a person lost access if a direct grant, second team, base permission, GitHub App installation, deploy key, or token still authorizes access. The mandatory lab models this without using another person’s account.

cat > /tmp/access-fixture.json <<'JSON'
{
  "user": "former-maintainer",
  "team_roles": [],
  "direct_repository_role": "write",
  "app_installations": [],
  "deploy_keys_owned": [],
  "expected_effective_role": "none"
}
JSON

python - <<'PY'
import json, sys
x=json.load(open('/tmp/access-fixture.json'))
actual=x['direct_repository_role'] if x['direct_repository_role'] != 'none' else 'none'
print({'expected':x['expected_effective_role'],'actual':actual})
sys.exit(0 if actual==x['expected_effective_role'] else 1)
PY

The checker intentionally fails. Repair the fixture to direct_repository_role: "none", rerun, and record the before/after evidence. In a real organization, enumerate every effective access path rather than assuming the membership operation was sufficient.

7. Automation treats the first API page as complete

Create two disposable issues, then run an intentionally incomplete inventory query that asks for one item and does not paginate. The failure is semantic: the HTTP request succeeds, but the client silently undercounts state.

gh issue create -R "$FULL" --title "automation-fixture-a" --body "Disposable"
gh issue create -R "$FULL" --title "automation-fixture-b" --body "Disposable"

# BROKEN inventory: only first page, one item.
gh api -H "X-GitHub-Api-Version: 2026-03-10"   "repos/$FULL/issues?state=open&per_page=1"   --jq 'map({number,title})'

# Independent inventory shows the discrepancy.
gh issue list -R "$FULL" --state open --limit 100 --json number,title

# REPAIR: traverse all pages.
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10"   "repos/$FULL/issues?state=open&per_page=100"   --jq '.[] | {number,title}'

A 200 response is not evidence of completeness. Production clients must understand pagination, filtering, rate limits, and event consistency; webhooks can reduce polling but introduce delivery deduplication and ordering concerns.

8. Cost/performance regression creates pressure to bypass controls

A platform can become insecure through economics. If every tiny docs change runs a 45-minute matrix, engineers eventually seek bypasses. The answer is not to remove required verification; it is to separate a fast, always-created gate from selective expensive work, measure cache/artifact retention, cancel superseded work safely, and budget the workflow.

Signal Bad reaction System repair
Long PR queue Disable required checks. Measure queue/run duration, trim matrix, cache carefully, parallelize only where useful, preserve stable gate.
Storage growth Delete evidence indiscriminately. Classify artifact purpose; shorten disposable retention; preserve release/compliance evidence according to policy.
Frequent reruns Add broad write authority so automation “fixes” state. Diagnose flake/root cause; bound retries; avoid retrying ambiguous non-idempotent mutations.
Search/API latency Treat search results as authoritative inventory. Use correct API/list endpoints with pagination and rate-limit handling; record query scope/time.

9. Failure-drill record template

incident_id: CAPSTONE-DRILL-01
failure_class: ruleset-required-check interaction
detected_at: 2026-08-20T00:00:00Z
contained_at: 2026-08-20T00:06:00Z
recovered_at: 2026-08-20T00:18:00Z
source_sha: <exact SHA>
evidence:
  - ruleset-evidence.json
  - pr-checks.json
containment: "No bypass; freeze merge while trigger is repaired"
correction: "Always create verify; move selectivity inside the job"
verification: "Required verify completed for docs-only PR"
residual_risk: "Future workflow rename could break required-check identity"
preventive_control: "Ruleset/workflow contract test"
owner: "platform-maintainer"

Use actual timestamps when you run the drill. Recovery time without a preserved cause is a vanity metric; preserved evidence without restoration is an incomplete incident response.

Knowledge checks

A required check never appears on a docs-only PR after path optimization. Should an admin bypass the ruleset?

Why can a self-hosted runner be high privilege even with contents: read?

What must happen before history rewriting after a real secret leak?

A tag and release asset look correct, but attestation verification fails after one byte changes. Which identity wins?

Why is a successful first-page API request dangerous in governance automation?

Summary

The recovery drills show why GitHub controls must be operated as a system. A ruleset can conflict with a cost optimization; a runner can carry authority outside GitHub tokens; Git cleanup can distract from credential containment; and valid-looking release metadata can disagree with artifact evidence. Lesson 5 repeats the architecture as an operational handoff and requires independent proof of critical invariants.

Next lesson

Production Capstone: Govern a GitHub Organization and Build a Secure End-to-End Delivery System: Final Operational Review and Handoff

Official references

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.