Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Guided Hands-On Workflow and Core Operations
A useful incident lab bounds the failure and preserves the evidence. This workflow creates queue pressure, a long-running step and an agent-channel loss without touching production infrastructure.
Learning objectives
- Create a reproducible disposable controller/project/agent incident scenario.
- Diagnose queue pressure, a running long step and an agent disconnect as different layers.
- Capture thread/support evidence before intervention.
- Use stop/reconnect/restart options only after evidence capture.
- Produce a compact incident timeline and evidence packet.
1. Scenario and lab boundary
Use local controller jenkins-incident-lab and one
disposable agent labeled incident-lab-agent with one
executor. Built-in controller executors stay at zero. All “external”
latency is simulated with sleep; no real cloud, SCM organization,
artifact repository or production credential is needed.
2. Preflight
| Check | Expected evidence | Stop condition |
|---|---|---|
| Controller | 2.568.3 / Java 21 / uptime | Undocumented baseline difference |
| Support Core | 1863.vdddd6f8d12a_8 if bundle exercise is used | Unknown plugin/security baseline |
| Controller node | 0 executors | Routine work can land on controller |
| Agent | Online, Java 21, one executor, expected label | Unknown execution identity |
| Storage | Space for logs/dumps/bundle | Existing disk pressure |
| External targets | None; local simulation only | Any real endpoint configured |
3. Create the synthetic Pipeline
Every build records identity first, then performs one bounded mode. The control mode proves the harness. The other modes produce deterministic delay without external side effects.
pipeline {
agent { label 'incident-lab-agent' }
options { timestamps(); disableConcurrentBuilds() }
parameters {
choice(name: 'MODE', choices: ['normal', 'slow-external', 'local-stall'], description: 'Synthetic incident mode')
}
stages {
stage('Identity') {
steps {
sh '''set -eu
mkdir -p evidence
printf 'job=%s\nbuild=%s\nnode=%s\nworkspace=%s\n' \
"$JOB_NAME" "$BUILD_NUMBER" "$NODE_NAME" "$WORKSPACE" > evidence/identity.txt
git rev-parse HEAD 2>/dev/null > evidence/source-sha.txt || printf 'synthetic-no-scm\n' > evidence/source-sha.txt
'''
}
}
stage('Bounded work') {
steps {
script {
if (params.MODE == 'normal') {
sh 'sleep 5; echo ok > evidence/work-result.txt'
} else if (params.MODE == 'slow-external') {
sh 'echo synthetic-external-start; sleep 45; echo synthetic-external-done'
} else {
sh 'echo synthetic-stall-marker; sleep 120'
}
}
}
}
}
post { always { archiveArtifacts artifacts: 'evidence/**', allowEmptyArchive: true, fingerprint: true } }
}
4. Run a normal control build
Run MODE=normal and record build number, queue item if
visible, assigned agent, workspace and source identity. Confirm the
archived evidence exists. Do not inject failures until the control
case is stable.
5. Case A — queue pressure
Occupy the agent's only executor with MODE=local-stall,
then enqueue another build. Inspect queue/node JSON.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,why,blocked,stuck,inQueueSince,task[name,url]]" | tee evidence/queue-pressure.json
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,numExecutors,busyExecutors,offlineCauseReason]" | tee evidence/nodes-pressure.json
The expected diagnosis is capacity/eligibility. No second Pipeline executor thread exists yet because the second build has not started.
6. Case B — a running build appears stuck
With the long step active, record the exact build URL/executor, console tail and node state. Capture a controller thread dump before changing the build.
curl -fsS "$JENKINS_URL/job/incident-lab/$BUILD_NUMBER/api/json?tree=number,url,building,result,timestamp,duration" > evidence/build-before-stop.json
curl -fsS "$JENKINS_URL/threadDump" > evidence/controller-threadDump.txt
sha256sum evidence/controller-threadDump.txt >> evidence/SHA256SUMS
If you need to stop the exact disposable build, use normal Pipeline
/stop first. Jenkins documents /term and
then /kill as progressively more forceful options; hard
kill is a last resort, not a routine diagnostic.
7. Case C — lose the disposable agent channel
While MODE=slow-external runs, stop only the lab agent
process/network. Preserve the controller-side offline cause and
agent-side launcher/service log before reconnecting. The controller
can remain healthy while the running step loses its channel.
curl -fsS "$JENKINS_URL/computer/incident-lab-agent/api/json?tree=displayName,offline,temporarilyOffline,offlineCauseReason,executors[currentExecutable[url]]" > evidence/agent-loss.json
java -version 2> evidence/agent-java.txt
# Preserve the disposable agent service/container log before restart.
8. Review a support bundle before sharing
If Support Core is installed, generate a lab-only bundle through the Jenkins Support action or documented CLI. List contents, extract only in a restricted directory, and search likely sensitive terms. This is a review aid, not proof of complete redaction.
mkdir -m 700 -p evidence/support-review
unzip -l support-bundle.zip > evidence/support-review/member-list.txt
unzip -q support-bundle.zip -d evidence/support-review/extracted
grep -RniE 'token|password|secret|authorization|cookie|private|internal\.' evidence/support-review/extracted | head -n 100 || true
9. Broken incident response: restart first
BAD: immediately hard-kill build, restart controller, delete workspace, rerun deployment.
LOST: live thread state, queue/agent evidence, original workspace/process state, external side-effect certainty.
The repair is to capture IDs and read-only state first, reconcile external effects, and record every intervention as a timestamped event.
10. Write the incident timeline
15:20:04Z build #17 queued; item 381 recorded
15:21:10Z #17 assigned incident-lab-agent executor 0
15:21:42Z agent channel lost; node JSON + agent log preserved
15:22:05Z controller thread dump captured; SHA-256 recorded
15:23:12Z agent restarted after evidence capture
15:24:08Z bounded canary healthy; no real external side effect existed
11. Layer-selection challenge
A PR build waits 20 minutes. Controller CPU/heap are normal. Three
agents are online but none has label browser-linux. The
next action belongs to queue/node-label capacity, not Pipeline CPS,
controller heap or support-bundle generation. State the evidence you
would preserve before adding/changing capacity.
12. Verification and cleanup
- Verify built-in executors remain zero.
- Reconnect the lab agent and run one normal canary.
- Preserve the original failed/queued build evidence until review.
- Hash the final timeline/manifest.
- Delete only the synthetic job/agent/evidence copies after the exercise is closed.
Knowledge check
1. Why run a control build before failure injection?
It proves the harness and evidence path work normally, so injected failures have a stable comparison.
2. A queue item exists but no executor is assigned. Is a running shell process relevant?
No. Execution has not started; inspect queue reasons, labels and capacity.
3. When is /kill appropriate?
As a last resort after evidence is preserved and less destructive stop/term approaches fail.
4. Why preserve agent logs before reconnecting?
Reconnect/restart can rotate or overwrite first-failure evidence and alter the offline cause.
5. What do you do if a support bundle contains internal URLs?
Keep it restricted; review/redact or create a safe subset before sharing.
13. Summary
Queue pressure, long-running work and agent loss can all look like “stuck Jenkins,” but they are different causal layers. Lesson 3 now turns these operations into explicit design tradeoffs.
Official references and version notes
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.