Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Concepts, Architecture, and Mental Model
Production troubleshooting is not “restart Jenkins and see.” It is a controlled reduction of uncertainty: preserve the incident, localize the causal layer, intervene minimally, and prove recovery.
Learning objectives
- Translate vague symptoms into exact controller, job, build, queue and agent identities.
- Explain what thread dumps, support bundles, logs and metrics can and cannot prove.
- Separate queue, Pipeline/CPS, agent/Remoting, controller/JVM/plugin, storage/network and external-service failure layers.
- Preserve sensitive evidence safely before disruptive intervention.
- Define recovery validation and post-incident evidence.
1. The practical problem: symptoms are not diagnoses
“The build is stuck,” “Jenkins is slow,” and “the agent disconnected” describe what a user sees. They do not identify why. A build can wait because no executor matches a label, a Pipeline can be waiting on a durable shell process, a controller thread can be blocked, an agent channel can be gone, disk can be saturated, or an external API can have accepted a request without returning a response.
The first engineering task is therefore to freeze identity and time: exact job full name, build number, queue item if present, agent/executor, source revision, first-observed timestamp, last-known-good timestamp and any external transaction ID.
2. Mental model: preserve → localize → verify → recover
flowchart TD
A["Symptom + time window"] --> B["Exact controller/job/build/queue/agent IDs"]
B --> C["Preserve logs/metrics/thread dump/support evidence"]
C --> D{"Causal layer?"}
D --> E["Queue/capacity"]
D --> F["Pipeline/CPS"]
D --> G["Agent/Remoting"]
D --> H["Controller/JVM/plugin"]
D --> I["Network/storage/external"]
E --> J["Least-invasive verification"]
F --> J
G --> J
H --> J
I --> J
J --> K["Contain/correct"]
K --> L["Recovery validation"]
L --> M["Incident handoff + runbook update"]
Each arrow narrows uncertainty. You do not jump from symptom to restart. You convert the symptom to exact execution objects, gather read-only evidence, form a layer hypothesis, test it with the smallest safe action, then recover and validate both Jenkins and any external side effects.
3. State to distinguish before changing anything
| Layer | Evidence to preserve | Question it answers |
|---|---|---|
| Incident identity | UTC window, symptom, controller URL, job/build/queue/agent IDs | Are all operators discussing the same incident? |
| Controller/JVM | LTS, Java, uptime, startup args, CPU/heap/GC, thread dumps | Is the controller process actually the failing layer? |
| Item/build | Job full name, build number/URL, source SHA, cause/result | Which exact execution is affected? |
| Queue | Queue item ID, why, labels, enqueue time |
Is work blocked before execution? |
| Agent/Remoting | Node, offline cause, Java, launcher, executor/workspace, logs | Did execution reach a healthy worker? |
| Pipeline/CPS | Stage/step, durable task, interruption/resumption context | Is orchestration waiting or agent work stuck? |
| Storage/network | Disk free/latency, DNS/TLS/HTTP evidence | Is infrastructure making Jenkins look unhealthy? |
| External state | Artifact/deployment/API transaction IDs and target state | Did a side effect already happen? |
| Recovery | Action, actor, timestamp, validation, residual risk | Was service restored without losing causal evidence? |
4. Read-only inspection first
Start with queue/build/node JSON, System Information, System Log, console timestamps, metrics, disk state and external service status. If a secured controller requires authentication, use an authorized narrowly scoped account and keep credentials out of shell history.
JENKINS_URL=http://127.0.0.1:8080
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,why,blocked,stuck,inQueueSince,task[name,url]]"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,busyExecutors,offlineCauseReason]"
curl -fsS "$JENKINS_URL/job/incident-lab/42/api/json?tree=number,url,result,building,timestamp,duration,estimatedDuration"
A queued item that says no matching label exists belongs to scheduling/capacity. A running build already assigned to an agent is a different diagnostic branch.
5. Thread dumps: instantaneous JVM execution evidence
Jenkins documents /threadDump for an authorized
controller capture when the UI is responsive. If the UI is
unavailable, use jstack as the same OS user running
Jenkins where practical; on Unix, kill -3 PID is a
lower-level fallback that writes to JVM standard output. Thread
dumps are useful for hangs, deadlocks, high CPU and builds that will
not respond to interruption.
PID=$(pgrep -f 'jenkins.war' | head -n 1)
ps -fp "$PID"
jstack "$PID" > "thread-dump-$(date +%Y%m%dT%H%M%S).txt"
sha256sum thread-dump-*.txt
One dump is one instant. Repeated dumps can show whether the same thread remains blocked on the same lock/network call. Interpret stacks with CPU, logs and timestamps rather than treating thread state names as verdicts.
6. Support bundles: broad evidence with a broad disclosure surface
Support Core can produce on-demand bundles and automatic bundles
under $JENKINS_HOME/support. The breadth is valuable
after a crash or multi-layer incident, but anonymization has
documented limitations. Plugin lists, internal URLs, exception text
and other metadata can remain recognizable. Keep the original
restricted; review/redact any copy shared outside the trusted
incident team.
7. Agent failures are not controller failures
Jenkins explicitly treats agents as workers that may be unreliable. Capture controller-side node/offline state, Remoting/channel messages, agent launcher/service/container logs, Java version, labels, executor and workspace. Do not “test Jenkins” by moving routine work onto the built-in node; keep controller executors at zero and reproduce on a disposable agent.
8. Causal layer map
| Evidence | Layer to test next | Bad shortcut |
|---|---|---|
| Queue says no matching label/executor | Queue/labels/capacity | Restart controller |
| Assigned build waits on outbound HTTP | Agent/network/external service | Rewrite Pipeline syntax |
| Many controller request threads repeat same blocked stack | Controller/JVM/plugin/dependency | Delete workspace |
| Only one agent channel fails | Agent/Remoting/network | Upgrade every plugin |
| JENKINS_HOME disk nearly full | Storage/retention/I/O | Add executors |
| Deployment API has transaction ID despite timeout | External side-effect reconciliation | Blind rerun |
9. Preserve first-failure evidence
A restart clears live thread state; reconnect changes node events; rerun creates a new build number; deleting a workspace destroys process/file evidence. Preserve and hash what matters before intervention.
mkdir -m 700 -p incident-evidence/INC-0040
curl -fsS "$JENKINS_URL/queue/api/json" > incident-evidence/INC-0040/queue-before.json
printf '%s
' 'job=incident-lab' 'build=42' 'agent=lab-agent-a' 'source=deadbeef' > incident-evidence/INC-0040/identity.txt
sha256sum incident-evidence/INC-0040/* > incident-evidence/INC-0040/SHA256SUMS
10. Recovery validation is not “the page loads”
After remediation, prove the affected layer is healthy. For an agent: channel, Java, labels, executor and bounded canary. For a controller restart: queue/run rehydration and representative Pipelines. For an uncertain deployment: query the external target before retry. Record what remains unverified.
11. Disposable orientation lab
- Record controller LTS/Java/plugin baseline and built-in executor count.
- Create one synthetic Pipeline on a non-controller agent.
- Record job/build/source/agent/workspace/timestamps.
- Inspect queue and node JSON before inducing any failure.
- Capture a controller thread dump and hash it.
- If Support Core exists, generate a lab bundle and inspect its manifest before sharing.
- Write a three-event incident timeline: symptom, evidence, next hypothesis.
Knowledge check
1. A build has no executor after 12 minutes. What do you inspect first?
Queue reason, labels, eligible online nodes and executor capacity—not controller restart or agent workspace.
2. Why can a support bundle remain sensitive after anonymization?
Some content types, URLs, plugin names and exception data may not be completely anonymized; review/redact before sharing.
3. What does one thread dump prove?
What JVM threads were doing at one instant; persistent diagnosis needs correlation with other evidence and sometimes multiple dumps.
4. Why record external transaction IDs before retrying?
A timeout may occur after the external action succeeded, so the ID supports reconciliation and prevents duplicate side effects.
5. Why keep built-in executors at zero?
Running builds on the controller changes security/performance boundaries and can hide an agent-specific issue.
12. Summary
Reliable troubleshooting starts by preserving identity and state, then moving through causal layers. Lesson 2 turns that model into a controlled queue/build/agent incident workflow.
Official references and version notes
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.