Chapter 40Lesson 04~225 minutes

Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Diagnostics, Failure Modes, Security, and Performance

Incident response fails when operators change many layers before proving which one is broken. This lesson engineers those failure modes deliberately and repairs them with an evidence-first sequence.

failure modesdiagnosticssecuritystuck buildsagent failuresminimal repair

Learning objectives

  • Diagnose restart-first, unredacted-bundle, lost-agent-log, queue-capacity and destructive-retry mistakes.
  • Apply the course-wide evidence-first diagnostic sequence.
  • Interpret queue/thread/node/API evidence instead of merely collecting it.
  • Separate authentication, authorization, transport, plugin and external-state failures.
  • Repair an intentionally broken runbook without hiding the original cause.

1. Evidence-first diagnostic sequence

  1. Preserve job/build/queue/agent IDs and first-failure evidence.
  2. Confirm controller core/Java/plugin/security baseline.
  3. Confirm item/Jenkinsfile/source/cause.
  4. Inspect queue and node/label/executor eligibility.
  5. Confirm agent/Remoting/workspace/toolchain.
  6. Inspect Pipeline/CPS/step state.
  7. Inspect credential/plugin/network/external service.
  8. Inspect reports/artifacts/publication/deployment state.
  9. Apply the least destructive correction.
  10. Retry only the smallest safe scope after reconciling side effects.

2. Failure 1 — restart first

The controller freezes briefly and is restarted immediately. Availability returns, but live thread state and transient evidence are gone. Better: preserve timestamps/IDs, capture thread dumps and resource/log evidence when feasible, then restart only if availability requirements justify it. The restart itself belongs in the timeline.

3. Failure 2 — share an unreviewed support bundle

“Anonymized” is not equivalent to “safe for public sharing.” Support Core documents limitations. Keep the original restricted, inspect manifest/content and create a reviewed derivative for escalation.

4. Failure 3 — kill the agent before logs

An ephemeral agent is deleted after a Remoting disconnect. Replacement succeeds but nobody can tell whether Java, TLS/DNS, host eviction, resource pressure or launcher behavior caused the incident. Export controller channel messages, agent/container/service logs and provider lifecycle events before deletion when incident severity allows.

5. Failure 4 — blame Pipeline for capacity-bound queue

{
  "id": 481,
  "blocked": false,
  "stuck": false,
  "why": "There are no nodes with the label ‘browser-linux’",
  "task": {"name": "ui-regression"}
}

No executor has started the Pipeline. Repair label/agent capacity and prove it with a canary. Controller restart, workspace deletion and Pipeline syntax changes do not address this evidence.

6. Failure 5 — retry externally destructive work

A deployment call times out after submission. Three blind reruns create three target changes. Repair the workflow by persisting idempotency/transaction IDs, querying the external target after uncertainty and resuming with the already-built immutable artifact.

7. Intentionally broken “fix everything” runbook

1. Restart Jenkins.
2. Delete the workspace.
3. Upgrade every plugin.
4. Recreate the agent.
5. Rerun the failed deployment until green.
6. Attach full support bundle to a public issue.

This changes controller, plugin, workspace, agent and external target before establishing cause. It also risks disclosure.

Repair

  1. Freeze incident ID/time window and exact identities.
  2. Capture queue/node/build/log state.
  3. Capture thread dumps only when the symptom warrants them.
  4. Reconcile external side effects.
  5. Form one layer hypothesis and verify minimally.
  6. Apply one targeted change.
  7. Run a bounded canary and compare.
  8. Schedule broad upgrades/cleanup separately unless they are the proven corrective action.

8. Interpreting a thread dump

"Handling GET /job/ui-regression/42/console" RUNNABLE
  at java.net.SocketInputStream.socketRead0(Native Method)
"Remoting I/O" RUNNABLE
  at sun.nio.ch.EPoll.wait(Native Method)
"jenkins.util.Timer" TIMED_WAITING
  at java.lang.Object.wait(Native Method)

RUNNABLE does not necessarily mean CPU-bound; a thread may be waiting in native I/O. Compare multiple dumps, CPU, logs and external latency. Thread names/stacks are clues, not diagnoses by themselves.

9. Security-sensitive troubleshooting actions

Action Risk Rule
Script Console Arbitrary controller code Do not use ad hoc snippets as generic diagnostics; prefer supported APIs/UI/logs.
Plugin mutation Controller executable dependency change Perform as reviewed maintenance/corrective action, not blind incident experimentation.
Support/heap dump sharing Sensitive data disclosure Restricted storage + review/redaction + approved recipients.
Agent secrets Credential exposure Never paste inbound secrets in tickets/logs; rotate if exposed.
API cancel/delete Evidence/state loss Target exact IDs and record before/after evidence.

10. Do not overload an already sick controller

Broad support bundles, heap dumps, massive debug logging and unbounded traces can increase I/O and storage pressure. Prefer lightweight queue/build/node API evidence and targeted thread/log capture first, then escalate collection when the hypothesis warrants it.

Next lesson

Checkpoint Lab — Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks

Continue with the next lesson and preserve the evidence, safety boundaries, and verification habits established here.

Knowledge check

1. Queued, no executor: highest-value evidence?

2. Why is “restart fixed it” incomplete?

3. Why not retry deployment until Jenkins is green?

4. Why can RUNNABLE be misleading?

5. When should plugin upgrade be immediate remediation?

11. Summary

Bad troubleshooting expands the blast radius and destroys causality. Good troubleshooting narrows the hypothesis and changes only the proven layer. The checkpoint now combines queue pressure, a running build and agent/network failure into one incident drill.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.