Checkpoint Lab — Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks
The checkpoint is a controlled production-style incident drill: collect evidence before intervention, separate primary cause from downstream symptoms, recover minimally, and leave a handoff another operator can continue.
Learning objectives
- Run bounded work on a disposable non-controller agent.
- Create queue pressure and one agent/network symptom.
- Preserve exact build/queue/agent/source IDs before intervention.
- Capture and interpret thread/support evidence.
- Recover only the failed layer and prove it.
- Produce an incident handoff and runbook improvement.
1. Setup and assumptions
Use disposable controller jenkins-incident-drill and
agent incident-lab-agent with one executor and matching
label. Built-in node executors remain zero. No production external
endpoint or credential is permitted.
| Component | Assumption | Evidence |
|---|---|---|
| Controller | 2.568.3 / Java 21 | System Information/About + uptime |
| Support Core | 1863.vdddd6f8d12a_8 if installed | Plugin version + reviewed bundle manifest |
| Agent | Java 21 / one executor / disposable | Node JSON + launcher log + workspace |
| Security | 16 Sep 2026 advisory reviewed | Plugin inventory/notes |
| External | Local simulation only | Explicit statement of no production side effect |
2. Predict state changes first
- When the only eligible executor is occupied, the second build will remain queued with a recorded queue reason while the controller remains healthy.
- When the disposable agent channel is lost, controller node/offline evidence and build console state will change; the controller process need not fail.
- Thread-dump capture will preserve build/queue identities rather than mutate them.
- Agent reconnect creates new connection events, so pre-reconnect logs must be preserved.
3. Preflight
- Record core/Java/plugin inventory and controller uptime.
- Confirm built-in executors are zero.
- Confirm lab agent identity/Java/label/executor.
- Record disk space and clock.
- Create restricted local evidence directory.
- Verify job has no real external endpoint/credential.
mkdir -m 700 -p incident-evidence/INC-0040
cd incident-evidence/INC-0040
date -u +%FT%TZ > preflight-time.txt
printf '%s
' 'controller=jenkins-incident-drill' 'scope=disposable-local-only' > scope.txt
4. Exact checkpoint Pipeline
pipeline {
agent { label 'incident-lab-agent' }
options { timestamps(); disableConcurrentBuilds() }
parameters { booleanParam(name: 'INJECT_STALL', defaultValue: false, description: 'Synthetic long step') }
stages {
stage('Evidence seed') {
steps {
sh '''set -eu
mkdir -p evidence
printf 'job=%s\nbuild=%s\nnode=%s\nworkspace=%s\n' \
"$JOB_NAME" "$BUILD_NUMBER" "$NODE_NAME" "$WORKSPACE" > evidence/identity.txt
date -u +%FT%TZ > evidence/start-utc.txt
'''
}
}
stage('Synthetic work') {
steps {
script {
if (params.INJECT_STALL) { sh 'echo stall-begin; sleep 120; echo stall-end' }
else { sh 'sleep 3; echo healthy > evidence/result.txt' }
}
}
}
}
post { always { archiveArtifacts artifacts: 'evidence/**', allowEmptyArchive: true, fingerprint: true } }
}
Run one normal control build. Then start
INJECT_STALL=true. Once that build occupies the only
executor, start a second build so the queue/capacity state is
observable.
5. Evidence packet
| Evidence | Required fields | Purpose |
|---|---|---|
| Controller baseline | Core, Java, plugin manifest, uptime | Platform identity |
| Build | Job, build URL/number, cause, source/Jenkinsfile revision | Exact execution identity |
| Queue | Queue item ID, why, time, label | Scheduling evidence |
| Agent | Node, executor, Java/Remoting, workspace, offline cause | Execution-environment evidence |
| Thread dump | Timestamp + SHA-256 | Controller/JVM state before intervention |
| Support review | Manifest + sensitivity note | Cross-layer evidence with disclosure control |
| External note | No production target; simulation only | Side-effect boundary |
| Actions | UTC time, actor, action, reason | Audit/replay |
6. Induce the incident
- Record queued item and reason while stall build occupies executor.
- Capture build/node/queue state and console tail.
- Capture controller thread dump and hash it.
- Stop/disconnect only the disposable agent process/network.
- Preserve agent-side log and controller-side offline/channel evidence before reconnect.
- If Support Core is installed, create/review a restricted bundle.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,why,blocked,stuck,inQueueSince,task[name,url]]" > queue-before.json
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,numExecutors,busyExecutors,offlineCauseReason]" > nodes-before.json
curl -fsS "$JENKINS_URL/threadDump" > controller-threadDump-before.txt
sha256sum controller-threadDump-before.txt > SHA256SUMS
7. Classify the evidence
| Observation | Layer | Interpretation |
|---|---|---|
| Second build waited before agent loss | Queue/capacity | Expected single-executor pressure |
| Agent offline during running build | Agent/Remoting/network | Execution channel changed |
| Controller UI/API/thread dump responsive | Controller/JVM | No controller failure evidence |
| Running step stops after channel loss | Pipeline/agent boundary | Inspect channel/durable-step behavior |
| No external transaction exists | External state | Lab retry risk is bounded by design |
Root cause should identify the injected agent/channel loss. Queue starvation after that is a downstream consequence, not the primary cause.
8. Recover minimally
Reconnect only the disposable agent after preserving logs. Verify Java/label/executor. Preserve the interrupted build result; if needed, run a new bounded control canary rather than rewriting history. Do not restart the controller because the evidence does not identify it as causal.
9. Verification checklist
- Pre-failure queue item/reason retained.
- Pre-reconnect agent log/offline cause retained.
- Thread-dump SHA-256 verifies.
- Controller availability history documented.
- Agent reconnects with expected Java/labels/executor.
- New normal canary succeeds on that agent.
- No build ran on built-in node.
- No production external side effect occurred.
- Every intervention has timestamp/reason.
10. Incident handoff
# INC-0040 handoff
## Impact
Disposable Jenkins lab only. Build #17 interrupted; build #18 queued.
## Identity
- controller: jenkins-incident-drill / 2.568.3 / Java 21
- job: incident-lab
- affected build: #17
- queue item: #481 for build #18
- agent: incident-lab-agent / executor 0
## Timeline
- 16:00:00Z #17 began synthetic stall
- 16:00:08Z #18 queued; capacity reason captured
- 16:00:25Z controller thread dump captured and hashed
- 16:00:40Z agent channel intentionally stopped
- 16:00:43Z agent/controller channel logs preserved
- 16:02:00Z agent restarted after evidence capture
- 16:03:10Z normal canary succeeded
## Root cause
Injected loss of the only eligible agent channel; queue starvation followed.
## Recovery
Reconnect same disposable agent after evidence capture; verify Java/label; run bounded canary.
## Runbook update
Export ephemeral-agent logs before teardown and correlate job/build/node/timestamp automatically.
11. Mandatory runbook improvement
Choose one evidence-derived improvement: automatic ephemeral-agent log export, a standard queue/node JSON capture command, documented thread-dump capture, support-bundle review ownership, or external transaction-ID reconciliation before retries. Give it an owner and review date.
12. Cleanup
- Confirm canary and evidence packet.
- Remove only synthetic job and disposable agent.
- Delete extracted support working copies after approved evidence is secured.
- Keep the failed build until exercise review completes.
- Remove the controller only after confirming no other lab depends on it.
Knowledge check
1. Primary root cause when the only eligible agent is disconnected?
Agent/channel loss; queue starvation is a downstream capacity consequence.
2. Why preserve the failed build after canary recovery?
It contains the original console/timing/build-state evidence needed for review.
3. Controller remains responsive after agent loss. Restart anyway?
No. That mutates an unrelated layer without evidence that it is causal.
4. Which evidence is especially sensitive?
Support bundles and heap dumps can contain broad sensitive operational data and need restricted review.
5. What makes recovery minimal?
Only the proven failed agent/channel layer is restored; unrelated controller/plugins/job/external state remains unchanged.
6. What does the canary prove?
That the recovered path runs a bounded job on the expected agent—not that every production integration is healthy.
13. What Chapter 40 adds
You now have an incident method that preserves state, correlates controller/job/build/queue/agent/external evidence, protects diagnostic data, chooses intervention by causal layer, validates recovery and feeds learning back into runbooks. Chapter 41 brings these practices together in the final enterprise Jenkins platform capstone.
Official references and version notes
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.