Chapter 40Lesson 01~210 minutes

Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Concepts, Architecture, and Mental Model

Production troubleshooting is not “restart Jenkins and see.” It is a controlled reduction of uncertainty: preserve the incident, localize the causal layer, intervene minimally, and prove recovery.

incident modelthread dumpssupport bundlesqueueRemotingevidence-first

Learning objectives

  • Translate vague symptoms into exact controller, job, build, queue and agent identities.
  • Explain what thread dumps, support bundles, logs and metrics can and cannot prove.
  • Separate queue, Pipeline/CPS, agent/Remoting, controller/JVM/plugin, storage/network and external-service failure layers.
  • Preserve sensitive evidence safely before disruptive intervention.
  • Define recovery validation and post-incident evidence.

1. The practical problem: symptoms are not diagnoses

“The build is stuck,” “Jenkins is slow,” and “the agent disconnected” describe what a user sees. They do not identify why. A build can wait because no executor matches a label, a Pipeline can be waiting on a durable shell process, a controller thread can be blocked, an agent channel can be gone, disk can be saturated, or an external API can have accepted a request without returning a response.

The first engineering task is therefore to freeze identity and time: exact job full name, build number, queue item if present, agent/executor, source revision, first-observed timestamp, last-known-good timestamp and any external transaction ID.

2. Mental model: preserve → localize → verify → recover

Mental model: preserve → localize → verify → recover
flowchart TD
  A["Symptom + time window"] --> B["Exact controller/job/build/queue/agent IDs"]
  B --> C["Preserve logs/metrics/thread dump/support evidence"]
  C --> D{"Causal layer?"}
  D --> E["Queue/capacity"]
  D --> F["Pipeline/CPS"]
  D --> G["Agent/Remoting"]
  D --> H["Controller/JVM/plugin"]
  D --> I["Network/storage/external"]
  E --> J["Least-invasive verification"]
  F --> J
  G --> J
  H --> J
  I --> J
  J --> K["Contain/correct"]
  K --> L["Recovery validation"]
  L --> M["Incident handoff + runbook update"]
  

Each arrow narrows uncertainty. You do not jump from symptom to restart. You convert the symptom to exact execution objects, gather read-only evidence, form a layer hypothesis, test it with the smallest safe action, then recover and validate both Jenkins and any external side effects.

3. State to distinguish before changing anything

Layer Evidence to preserve Question it answers
Incident identity UTC window, symptom, controller URL, job/build/queue/agent IDs Are all operators discussing the same incident?
Controller/JVM LTS, Java, uptime, startup args, CPU/heap/GC, thread dumps Is the controller process actually the failing layer?
Item/build Job full name, build number/URL, source SHA, cause/result Which exact execution is affected?
Queue Queue item ID, why, labels, enqueue time Is work blocked before execution?
Agent/Remoting Node, offline cause, Java, launcher, executor/workspace, logs Did execution reach a healthy worker?
Pipeline/CPS Stage/step, durable task, interruption/resumption context Is orchestration waiting or agent work stuck?
Storage/network Disk free/latency, DNS/TLS/HTTP evidence Is infrastructure making Jenkins look unhealthy?
External state Artifact/deployment/API transaction IDs and target state Did a side effect already happen?
Recovery Action, actor, timestamp, validation, residual risk Was service restored without losing causal evidence?

4. Read-only inspection first

Start with queue/build/node JSON, System Information, System Log, console timestamps, metrics, disk state and external service status. If a secured controller requires authentication, use an authorized narrowly scoped account and keep credentials out of shell history.

JENKINS_URL=http://127.0.0.1:8080
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,why,blocked,stuck,inQueueSince,task[name,url]]"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,busyExecutors,offlineCauseReason]"
curl -fsS "$JENKINS_URL/job/incident-lab/42/api/json?tree=number,url,result,building,timestamp,duration,estimatedDuration"

A queued item that says no matching label exists belongs to scheduling/capacity. A running build already assigned to an agent is a different diagnostic branch.

5. Thread dumps: instantaneous JVM execution evidence

Jenkins documents /threadDump for an authorized controller capture when the UI is responsive. If the UI is unavailable, use jstack as the same OS user running Jenkins where practical; on Unix, kill -3 PID is a lower-level fallback that writes to JVM standard output. Thread dumps are useful for hangs, deadlocks, high CPU and builds that will not respond to interruption.

PID=$(pgrep -f 'jenkins.war' | head -n 1)
ps -fp "$PID"
jstack "$PID" > "thread-dump-$(date +%Y%m%dT%H%M%S).txt"
sha256sum thread-dump-*.txt

One dump is one instant. Repeated dumps can show whether the same thread remains blocked on the same lock/network call. Interpret stacks with CPU, logs and timestamps rather than treating thread state names as verdicts.

6. Support bundles: broad evidence with a broad disclosure surface

Support Core can produce on-demand bundles and automatic bundles under $JENKINS_HOME/support. The breadth is valuable after a crash or multi-layer incident, but anonymization has documented limitations. Plugin lists, internal URLs, exception text and other metadata can remain recognizable. Keep the original restricted; review/redact any copy shared outside the trusted incident team.

7. Agent failures are not controller failures

Jenkins explicitly treats agents as workers that may be unreliable. Capture controller-side node/offline state, Remoting/channel messages, agent launcher/service/container logs, Java version, labels, executor and workspace. Do not “test Jenkins” by moving routine work onto the built-in node; keep controller executors at zero and reproduce on a disposable agent.

8. Causal layer map

Evidence Layer to test next Bad shortcut
Queue says no matching label/executor Queue/labels/capacity Restart controller
Assigned build waits on outbound HTTP Agent/network/external service Rewrite Pipeline syntax
Many controller request threads repeat same blocked stack Controller/JVM/plugin/dependency Delete workspace
Only one agent channel fails Agent/Remoting/network Upgrade every plugin
JENKINS_HOME disk nearly full Storage/retention/I/O Add executors
Deployment API has transaction ID despite timeout External side-effect reconciliation Blind rerun

9. Preserve first-failure evidence

A restart clears live thread state; reconnect changes node events; rerun creates a new build number; deleting a workspace destroys process/file evidence. Preserve and hash what matters before intervention.

mkdir -m 700 -p incident-evidence/INC-0040
curl -fsS "$JENKINS_URL/queue/api/json" > incident-evidence/INC-0040/queue-before.json
printf '%s
' 'job=incident-lab' 'build=42' 'agent=lab-agent-a' 'source=deadbeef' > incident-evidence/INC-0040/identity.txt
sha256sum incident-evidence/INC-0040/* > incident-evidence/INC-0040/SHA256SUMS

10. Recovery validation is not “the page loads”

After remediation, prove the affected layer is healthy. For an agent: channel, Java, labels, executor and bounded canary. For a controller restart: queue/run rehydration and representative Pipelines. For an uncertain deployment: query the external target before retry. Record what remains unverified.

11. Disposable orientation lab

  1. Record controller LTS/Java/plugin baseline and built-in executor count.
  2. Create one synthetic Pipeline on a non-controller agent.
  3. Record job/build/source/agent/workspace/timestamps.
  4. Inspect queue and node JSON before inducing any failure.
  5. Capture a controller thread dump and hash it.
  6. If Support Core exists, generate a lab bundle and inspect its manifest before sharing.
  7. Write a three-event incident timeline: symptom, evidence, next hypothesis.
Next lesson

Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Guided Hands-On Workflow and Core Operations

Continue with the next lesson and preserve the evidence, safety boundaries, and verification habits established here.

Knowledge check

1. A build has no executor after 12 minutes. What do you inspect first?

2. Why can a support bundle remain sensitive after anonymization?

3. What does one thread dump prove?

4. Why record external transaction IDs before retrying?

5. Why keep built-in executors at zero?

12. Summary

Reliable troubleshooting starts by preserving identity and state, then moving through causal layers. Lesson 2 turns that model into a controlled queue/build/agent incident workflow.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.