Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Diagnostics, Failure Modes, Security, and Performance
Incident response fails when operators change many layers before proving which one is broken. This lesson engineers those failure modes deliberately and repairs them with an evidence-first sequence.
Learning objectives
- Diagnose restart-first, unredacted-bundle, lost-agent-log, queue-capacity and destructive-retry mistakes.
- Apply the course-wide evidence-first diagnostic sequence.
- Interpret queue/thread/node/API evidence instead of merely collecting it.
- Separate authentication, authorization, transport, plugin and external-state failures.
- Repair an intentionally broken runbook without hiding the original cause.
1. Evidence-first diagnostic sequence
- Preserve job/build/queue/agent IDs and first-failure evidence.
- Confirm controller core/Java/plugin/security baseline.
- Confirm item/Jenkinsfile/source/cause.
- Inspect queue and node/label/executor eligibility.
- Confirm agent/Remoting/workspace/toolchain.
- Inspect Pipeline/CPS/step state.
- Inspect credential/plugin/network/external service.
- Inspect reports/artifacts/publication/deployment state.
- Apply the least destructive correction.
- Retry only the smallest safe scope after reconciling side effects.
2. Failure 1 — restart first
The controller freezes briefly and is restarted immediately. Availability returns, but live thread state and transient evidence are gone. Better: preserve timestamps/IDs, capture thread dumps and resource/log evidence when feasible, then restart only if availability requirements justify it. The restart itself belongs in the timeline.
3. Failure 2 — share an unreviewed support bundle
“Anonymized” is not equivalent to “safe for public sharing.” Support Core documents limitations. Keep the original restricted, inspect manifest/content and create a reviewed derivative for escalation.
4. Failure 3 — kill the agent before logs
An ephemeral agent is deleted after a Remoting disconnect. Replacement succeeds but nobody can tell whether Java, TLS/DNS, host eviction, resource pressure or launcher behavior caused the incident. Export controller channel messages, agent/container/service logs and provider lifecycle events before deletion when incident severity allows.
5. Failure 4 — blame Pipeline for capacity-bound queue
{
"id": 481,
"blocked": false,
"stuck": false,
"why": "There are no nodes with the label ‘browser-linux’",
"task": {"name": "ui-regression"}
}
No executor has started the Pipeline. Repair label/agent capacity and prove it with a canary. Controller restart, workspace deletion and Pipeline syntax changes do not address this evidence.
6. Failure 5 — retry externally destructive work
A deployment call times out after submission. Three blind reruns create three target changes. Repair the workflow by persisting idempotency/transaction IDs, querying the external target after uncertainty and resuming with the already-built immutable artifact.
7. Intentionally broken “fix everything” runbook
1. Restart Jenkins.
2. Delete the workspace.
3. Upgrade every plugin.
4. Recreate the agent.
5. Rerun the failed deployment until green.
6. Attach full support bundle to a public issue.
This changes controller, plugin, workspace, agent and external target before establishing cause. It also risks disclosure.
Repair
- Freeze incident ID/time window and exact identities.
- Capture queue/node/build/log state.
- Capture thread dumps only when the symptom warrants them.
- Reconcile external side effects.
- Form one layer hypothesis and verify minimally.
- Apply one targeted change.
- Run a bounded canary and compare.
- Schedule broad upgrades/cleanup separately unless they are the proven corrective action.
8. Interpreting a thread dump
"Handling GET /job/ui-regression/42/console" RUNNABLE
at java.net.SocketInputStream.socketRead0(Native Method)
"Remoting I/O" RUNNABLE
at sun.nio.ch.EPoll.wait(Native Method)
"jenkins.util.Timer" TIMED_WAITING
at java.lang.Object.wait(Native Method)
RUNNABLE does not necessarily mean CPU-bound; a thread
may be waiting in native I/O. Compare multiple dumps, CPU, logs and
external latency. Thread names/stacks are clues, not diagnoses by
themselves.
9. Security-sensitive troubleshooting actions
| Action | Risk | Rule |
|---|---|---|
| Script Console | Arbitrary controller code | Do not use ad hoc snippets as generic diagnostics; prefer supported APIs/UI/logs. |
| Plugin mutation | Controller executable dependency change | Perform as reviewed maintenance/corrective action, not blind incident experimentation. |
| Support/heap dump sharing | Sensitive data disclosure | Restricted storage + review/redaction + approved recipients. |
| Agent secrets | Credential exposure | Never paste inbound secrets in tickets/logs; rotate if exposed. |
| API cancel/delete | Evidence/state loss | Target exact IDs and record before/after evidence. |
10. Do not overload an already sick controller
Broad support bundles, heap dumps, massive debug logging and unbounded traces can increase I/O and storage pressure. Prefer lightweight queue/build/node API evidence and targeted thread/log capture first, then escalate collection when the hypothesis warrants it.
Knowledge check
1. Queued, no executor: highest-value evidence?
Queue reason, requested labels, eligible online nodes and executor availability.
2. Why is “restart fixed it” incomplete?
It changed state without identifying cause; recurrence risk remains and first-failure evidence may be lost.
3. Why not retry deployment until Jenkins is green?
The external system may already have accepted earlier attempts, creating duplicate side effects.
4. Why can RUNNABLE be misleading?
Native I/O waits may appear RUNNABLE; correlate stack, CPU, time and external dependency evidence.
5. When should plugin upgrade be immediate remediation?
When evidence identifies that plugin/version as causal and the upgrade is tested/reviewed.
11. Summary
Bad troubleshooting expands the blast radius and destroys causality. Good troubleshooting narrows the hypothesis and changes only the proven layer. The checkpoint now combines queue pressure, a running build and agent/network failure into one incident drill.
Official references and version notes
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.