Production Capstone: Govern a GitHub Organization and Build a Secure End-to-End Delivery System: Failure Injection, Troubleshooting, and Recovery Drill
Inject failures across interacting controls, credentials, provenance, access, automation, and cost pressure; diagnose from preserved evidence, contain risk, recover safely, and record preventive actions.
Learning objectives
- Use an evidence-first diagnostic sequence that preserves the original cause before containment and repair.
- Diagnose an interaction failure between a required-check ruleset and over-broad path filtering without bypassing policy.
- Contain runner/workflow authority, credential leaks, provenance mismatches, residual access, and incomplete API inventory safely.
- Treat performance and cost pressure as a control-design input instead of a reason to disable governance.
- Record recovery time, evidence, residual risk, preventive control, and owner for each injected incident.
1. The recovery loop
Use the same sequence for every capstone incident: preserve evidence → identify actor/repository/workflow/ref/artifact → inspect rules/permissions/logs/audit/API state → contain risk → choose the least disruptive correction → verify → capture lessons learned. This ordering prevents a frantic “fix” from destroying evidence or widening the incident.
| Record for every drill | Why it matters |
|---|---|
| Detection time / recovery time | Shows whether the team can meet an operational objective, not merely describe a procedure. |
| Original evidence | Preserves the cause instead of replacing it with the post-fix state. |
| Containment action | Stops additional harm before optimization or cleanup. |
| Verification | Proves the correction worked independently. |
| Residual risk | States what remains uncertain or intentionally accepted. |
| Preventive action + owner | Turns the drill into platform improvement. |
2. Interaction failure: ruleset + required check + path filter
This failure is intentionally caused by
two controls interacting. The ruleset requires the status
check named verify. A cost-optimization change then
filters the entire workflow so docs-only PRs do not create that
check. The PR is safe but cannot merge because the required check
never exists.
# BROKEN experiment — only on a disposable branch.
# The entire workflow is filtered, so docs-only changes may never create the required "verify" check.
on:
pull_request:
paths:
- "src/**"
- Preserve: record PR number/head SHA, ruleset JSON, Actions run list, and the absence/pending state of the required check.
- Scope: the repository policy is working as configured; the workflow trigger is the incompatible control.
- Contain: do not bypass the ruleset just to merge the docs change.
-
Repair: keep
pull_request:broad enough that the stable aggregateverifycheck is always created. Put selective/expensive execution inside jobs/steps while the aggregate check reports a deterministic result. -
Verify: push the repaired workflow on the branch,
observe
verifycomplete, then confirm the same ruleset now permits the normal merge path.
3. Compromised workflow or runner has excessive authority
Do not create a privileged self-hosted runner for this drill. Analyze the following fixture as if it appeared during review:
runs-on: [self-hosted, production-network]
permissions:
contents: write
packages: write
id-token: write
# Runner also has network reachability to internal deployment systems.
The problem is not one permission. A compromised job could alter repository content, publish packages, request OIDC identity, and reach a sensitive network. Containment must therefore span GitHub credentials, external provider trust, and runner/network isolation.
- Stop or disable the affected workflow/runner path and preserve job/runner logs.
- Revoke or invalidate reachable durable credentials and provider sessions; inspect OIDC/provider audit events if external federation exists.
- Remove unnecessary GitHub token permissions; move release/deploy authority into a narrowly triggered job/environment.
- Quarantine/rebuild the runner from trusted image if self-hosted; review egress and ambient credentials.
- Re-run on a trusted executor and compare source/artifact identities before restoring promotion.
4. Credential leak response starts with Git cleanup instead of revocation
The intentionally broken response is: “remove the file, rewrite history, force push, then decide whether to rotate.” That reverses containment priority. If an attacker copied the credential before cleanup, history rewriting does nothing to invalidate it.
The capstone never creates a real credential, so the drill uses a redacted synthetic incident record. Any history rewrite or force update is destructive; do not run it merely to make a training repository look clean.
5. Release, package, and attestation identities disagree
Copy the legitimate artifact, change one byte, and prove that provenance verification rejects the modified subject. This is a safe, local failure injection.
cp release-download/capstone.tar.gz /tmp/capstone-tampered.tar.gz
printf 'training-tamper' >> /tmp/capstone-tampered.tar.gz
sha256sum release-download/capstone.tar.gz /tmp/capstone-tampered.tar.gz
set +e
gh attestation verify /tmp/capstone-tampered.tar.gz --repo "$FULL" --signer-workflow "$FULL/.github/workflows/capstone.yml" --source-ref refs/heads/main --deny-self-hosted-runners
VERIFY_RC=$?
set -e
printf 'tampered_verify_exit=%s\n' "$VERIFY_RC"
test "$VERIFY_RC" -ne 0
Now imagine a package tag points to different bytes than the GitHub Release asset while the attestation belongs to only one digest. Freeze promotion. Compare the source tag SHA, release asset digest, registry digest, and attestation subject. Repair by producing a new, coherent version—not by moving the old tag or pretending the attestation covers different bytes.
6. Offboarding leaves residual direct access
GitHub access is additive. Removing one team does not prove a person lost access if a direct grant, second team, base permission, GitHub App installation, deploy key, or token still authorizes access. The mandatory lab models this without using another person’s account.
cat > /tmp/access-fixture.json <<'JSON'
{
"user": "former-maintainer",
"team_roles": [],
"direct_repository_role": "write",
"app_installations": [],
"deploy_keys_owned": [],
"expected_effective_role": "none"
}
JSON
python - <<'PY'
import json, sys
x=json.load(open('/tmp/access-fixture.json'))
actual=x['direct_repository_role'] if x['direct_repository_role'] != 'none' else 'none'
print({'expected':x['expected_effective_role'],'actual':actual})
sys.exit(0 if actual==x['expected_effective_role'] else 1)
PY
The checker intentionally fails. Repair the fixture to
direct_repository_role: "none", rerun, and record the
before/after evidence. In a real organization, enumerate every
effective access path rather than assuming the membership operation
was sufficient.
7. Automation treats the first API page as complete
Create two disposable issues, then run an intentionally incomplete inventory query that asks for one item and does not paginate. The failure is semantic: the HTTP request succeeds, but the client silently undercounts state.
gh issue create -R "$FULL" --title "automation-fixture-a" --body "Disposable"
gh issue create -R "$FULL" --title "automation-fixture-b" --body "Disposable"
# BROKEN inventory: only first page, one item.
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$FULL/issues?state=open&per_page=1" --jq 'map({number,title})'
# Independent inventory shows the discrepancy.
gh issue list -R "$FULL" --state open --limit 100 --json number,title
# REPAIR: traverse all pages.
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" "repos/$FULL/issues?state=open&per_page=100" --jq '.[] | {number,title}'
A 200 response is not evidence of completeness. Production clients must understand pagination, filtering, rate limits, and event consistency; webhooks can reduce polling but introduce delivery deduplication and ordering concerns.
8. Cost/performance regression creates pressure to bypass controls
A platform can become insecure through economics. If every tiny docs change runs a 45-minute matrix, engineers eventually seek bypasses. The answer is not to remove required verification; it is to separate a fast, always-created gate from selective expensive work, measure cache/artifact retention, cancel superseded work safely, and budget the workflow.
| Signal | Bad reaction | System repair |
|---|---|---|
| Long PR queue | Disable required checks. | Measure queue/run duration, trim matrix, cache carefully, parallelize only where useful, preserve stable gate. |
| Storage growth | Delete evidence indiscriminately. | Classify artifact purpose; shorten disposable retention; preserve release/compliance evidence according to policy. |
| Frequent reruns | Add broad write authority so automation “fixes” state. | Diagnose flake/root cause; bound retries; avoid retrying ambiguous non-idempotent mutations. |
| Search/API latency | Treat search results as authoritative inventory. | Use correct API/list endpoints with pagination and rate-limit handling; record query scope/time. |
9. Failure-drill record template
incident_id: CAPSTONE-DRILL-01
failure_class: ruleset-required-check interaction
detected_at: 2026-08-20T00:00:00Z
contained_at: 2026-08-20T00:06:00Z
recovered_at: 2026-08-20T00:18:00Z
source_sha: <exact SHA>
evidence:
- ruleset-evidence.json
- pr-checks.json
containment: "No bypass; freeze merge while trigger is repaired"
correction: "Always create verify; move selectivity inside the job"
verification: "Required verify completed for docs-only PR"
residual_risk: "Future workflow rename could break required-check identity"
preventive_control: "Ruleset/workflow contract test"
owner: "platform-maintainer"
Use actual timestamps when you run the drill. Recovery time without a preserved cause is a vanity metric; preserved evidence without restoration is an incomplete incident response.
Knowledge checks
A required check never appears on a docs-only PR after path optimization. Should an admin bypass the ruleset?
No as the default repair. Preserve the state, confirm the workflow was filtered out, restore an always-created stable required check, and keep expensive selectivity inside the workflow.
Why can a self-hosted runner be high privilege even with
contents: read?
Because ambient credentials, filesystem persistence, network reachability, and external systems are authority too. GitHub token scope is only one part of the trust boundary.
What must happen before history rewriting after a real secret leak?
Revoke/rotate the credential, preserve evidence, assess use, and migrate dependent systems. Rewrite only after containment and coordination if it adds security/compliance value.
A tag and release asset look correct, but attestation verification fails after one byte changes. Which identity wins?
The digest/attestation mismatch is decisive evidence that the bytes are not the attested subject. Stop promotion and produce a coherent new artifact/version rather than overriding the verifier.
Why is a successful first-page API request dangerous in governance automation?
Because it can silently create a partial inventory. HTTP success does not prove pagination was exhausted or the query covered all relevant resources.
Summary
The recovery drills show why GitHub controls must be operated as a system. A ruleset can conflict with a cost optimization; a runner can carry authority outside GitHub tokens; Git cleanup can distract from credential containment; and valid-looking release metadata can disagree with artifact evidence. Lesson 5 repeats the architecture as an operational handoff and requires independent proof of critical invariants.
Official references
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.