Production Capstone: Design, Automate, Secure, Scale, Upgrade, and Recover an Enterprise Jenkins Platform: Failure Injection, Troubleshooting, and Recovery Drill
Production readiness is proven under failure, not only on the happy path. This drill injects bounded faults across configuration, queue/agents, artifact identity, upgrade and recovery layers, preserving evidence before every intervention.
Learning objectives
- Inject safe, bounded failures without using production systems.
- Preserve build/queue/source/agent/controller/external evidence before intervention.
- Classify failure by causal layer rather than symptom.
- Choose restart, retry, rollback, reconciliation or restore only when evidence supports it.
- Produce an incident handoff and runbook improvement.
1. Failure injection has a guardrail
The goal is not chaos for its own sake. Every injection must name the expected state change, blast radius, stop condition, evidence path and rollback. If you cannot identify the exact disposable resource, do not run the fault.
2. Failure matrix
| Injected failure | Observable symptom | Primary layer | First recovery direction |
|---|---|---|---|
| Configuration drift | JCasC source differs from runtime/export | Controller/config ownership | Capture source/runtime diff; restore code-owned state |
| Plugin incompatibility | Startup/step failure after change | Controller/plugin dependency graph | Preserve plugin/core/Java inventory and logs; rollback/forward-fix from tested point |
| Privileged crossover | Untrusted build reaches privileged agent/secret | SCM trust + identity + agent | Abort/contain; preserve build/source/credential-binding evidence; correct trust policy |
| Secret exposure | Sensitive value appears in log/artifact | Credential/process/report | Restrict evidence, rotate/revoke real secret, fix binding/output path |
| Mutable artifact | Same version resolves to different digest | Repository/release | Block promotion; reconcile immutable digest/source/build identity |
| Queue exhaustion | Long waits, healthy controller | Queue/capacity/labels | Preserve queue reasons; adjust bounded capacity or demand |
| Controller restart/agent loss | Durable/nonresumable steps differ | Pipeline/CPS/agent | Classify resumable vs retryable vs manual reconcile |
| Failed upgrade | Core/plugin/Java mismatch | Upgrade/recovery | Stop change; use tested rollback/restore gate |
| Unusable backup | Restore misses key/plugin/external dependency | Recovery | Keep original; repair backup scope/key/external inventory, retest |
3. Evidence-first diagnostic sequence
- Freeze incident time window and exact job/build/queue/source/agent IDs.
- Capture controller core/Java/plugin baseline.
- Confirm item/job/Jenkinsfile/library/source/cause.
- Inspect queue reason, labels and executor eligibility.
- Inspect agent/Remoting/workspace/tool identity.
- Inspect Pipeline/CPS/step state and durable-task evidence.
- Inspect credentials authorization, plugin/network/external service.
- Inspect report/artifact/repository/deployment side effects.
- Apply the least destructive correction.
- Retry/rerun only the smallest safe scope and prove recovery.
INCIDENT=capstone-drill-001
mkdir -p "evidence/incidents/$INCIDENT"
D="evidence/incidents/$INCIDENT"
date -u +%FT%TZ > "$D/start.txt"
curl -fsS "$JENKINS_URL/queue/api/json" > "$D/queue-before.json"
curl -fsS "$JENKINS_URL/computer/api/json" > "$D/nodes-before.json"
cp plugins.txt "$D/plugins-intended.txt"
cp platform-contract.yaml "$D/platform-contract.txt"
printf 'Do not restart, reconnect, delete, retry, or republish until first-failure evidence is captured.\n' > "$D/guard.txt"
4. Injection 1: configuration drift
Do not mutate the running controller for this mandatory exercise. Instead create a simulated runtime copy and prove the platform can detect divergence.
# Safe local failure injection: change a COPY, not the running controller.
cp casc/jenkins.yaml evidence/jcasc-runtime-sim.yaml
printf '\n# DRIFT: lab-only marker\n' >> evidence/jcasc-runtime-sim.yaml
diff -u casc/jenkins.yaml evidence/jcasc-runtime-sim.yaml > evidence/config-drift.diff || true
test -s evidence/config-drift.diff
The diagnosis is configuration ownership. A restart alone would not explain or prevent the drift. The corrective action is to restore the code-managed source of truth and determine which actor/process created the unauthorized delta.
5. Injection 2: agent/queue failure
Take only the disposable capstone-lab agent offline
while two bounded builds request that label. Predict: one running
build may fail/wait according to the interrupted step; queued work
should show a label/capacity reason; controller JVM health should
remain normal.
Capture queue JSON, node offline cause, build URLs, agent log and timestamps before reconnect/recreate. If the queue says no matching online node, adding controller heap is unrelated.
6. Injection 3: mutable/tampered artifact
Never change the immutable repository-sim object. Tamper with a separate copy and prove digest mismatch.
set -eu
cp dist/app.txt evidence/tampered-app.txt
printf 'tampered\n' >> evidence/tampered-app.txt
expected=$(cut -d' ' -f1 evidence/artifact.sha256)
actual=$(sha256sum evidence/tampered-app.txt | cut -d' ' -f1)
printf 'expected=%s\nactual=%s\n' "$expected" "$actual"
test "$expected" != "$actual"
The repair is not “update the expected digest.” Re-establish the chain from source/build to the original immutable bytes, reject the tampered copy and investigate how mutability entered the path.
7. Injection 4: plugin compatibility failure — faithful simulation
Instead of intentionally installing a vulnerable/incompatible
plugin, edit a copy of plugins.txt to an
advisory-affected version and run the policy validator. Expected
result: validation fails before controller mutation. This
demonstrates a preferred production control: catch unsafe changes in
staging/policy before installation.
8. Injection 5: failed upgrade decision
Create a hypothetical canary result where the target core starts but one representative Pipeline fails and one agent is on an unsupported Java runtime. The correct gate is do not cut over. Preserve the canary logs/inventory, fix agent/runtime/plugin compatibility in staging, and rerun the exact acceptance matrix. Production cutover is not a debugging environment.
9. Injection 6: unusable backup
Restore a disposable copy with one deliberately missing non-secret marker file or external-dependency manifest—not with destroyed real key material. The restore gate should fail. Document the gap, rebuild the backup correctly, then repeat the clean-room restore. This demonstrates why “backup job succeeded” and “recovery is possible” are different claims.
10. Intentionally broken response
BAD RESPONSE
1. Restart controller immediately.
2. Reconnect/delete agent without saving logs.
3. Retry deployment three times.
4. Replace the artifact under the same release version.
5. Delete the failed build because it is noisy.
WHY IT FAILS
- destroys first-failure evidence
- can duplicate external side effects
- hides queue/agent/plugin causal layers
- breaks artifact identity
- removes audit/recovery evidence
The repaired response captures identity and evidence, reconciles external state, fixes one causal layer, then validates the smallest safe recovery path.
11. Incident handoff template
| Field | Required content |
|---|---|
| Impact/window | Who/what was affected; UTC start/end |
| Identity | Controller, job/build/queue, source SHA, agent, artifact/external IDs |
| First evidence | Logs/metrics/thread/queue/node/plugin/config snapshots and checksums |
| Hypothesis | Causal layer and why evidence supports it |
| Containment | Exact bounded action and actor/time |
| Recovery | What changed and what did not |
| Validation | Canary, queue/agent state, artifact/external verification |
| Follow-up | Runbook/code/policy change, owner, due date |
12. Cleanup
Return the lab agent online, restore the clean simulated JCasC copy, delete only the tampered artifact copy and temporary compatibility manifest, run a normal canary, and retain the incident packet until the review is complete. Never delete evidence merely to make the lab look green.
Knowledge check
1. A build is queued because no capstone-lab node
is online. Which layer is primary?
Queue/agent capacity and label eligibility, not Pipeline CPS or controller heap.
2. Why must external side effects be reconciled before retry?
The request may have succeeded remotely even if Jenkins saw a timeout; retry can duplicate or corrupt state.
3. Why simulate an incompatible plugin in the manifest instead of installing it?
It tests the preventive validation path without intentionally weakening or destabilizing even the lab controller.
4. What does a failed restore test change?
The backup can no longer be claimed usable; recovery scope/runbook/key/dependency gaps must be repaired and retested.
5. Why preserve the failed build and queue IDs?
They anchor logs, metrics, agent events and external evidence to the exact incident execution.
13. Summary
You have proven that the platform can detect and recover from failures without erasing evidence or repeating unsafe side effects. Lesson 5 turns all chapter evidence into a final operational handoff.
Official references and version notes
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.