Chapter 40Lesson 03~205 minutes

Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Configuration, Design Choices, and Tradeoffs

Troubleshooting design means choosing the least destructive evidence and remediation for the causal layer while preserving enough state to prove the decision later.

tradeoffsheap vs threadsupport scopeagent lifecyclerestart criteriaauditability

Learning objectives

  • Choose abort versus preserve based on evidence and external side-effect certainty.
  • Compare agent reconnect with immutable-agent recreation.
  • Choose thread dump versus heap dump.
  • Scope support bundles according to diagnostic need and disclosure risk.
  • Choose controller restart versus targeted remediation.

1. Design principle: preserve optionality

The best incident action leaves multiple recovery and learning paths open. A hard restart may restore service but destroys live thread state. Recreating an agent can be operationally correct but may erase the only workspace/process evidence. A heap dump can answer memory-retention questions but adds I/O, storage and disclosure risk. Choose based on the hypothesis, evidence already captured and recovery objective.

2. Decision table

Choice Prefer when State changed Evidence gate
Abort build vs preserve Preserve while evidence/side-effect certainty is incomplete; abort once containment requires it and evidence is sufficient. Build/executor + external state IDs, last log timestamp, process/thread state, external transaction query
Reconnect vs recreate agent Reconnect when host/workspace continuity matters; recreate ephemeral agents after logs/metadata are exported. Node/Remoting/workspace Offline cause, agent log, image/template identity, canary
Thread dump vs heap dump Thread dump for hangs/deadlocks/CPU; heap dump for retained-object memory diagnosis. Controller JVM Thread stacks vs heap/GC evidence; sensitivity/capture cost
Narrow vs broad support bundle Use the smallest scope that answers the hypothesis; broaden when causal layer is uncertain. Controller/plugin/log evidence Bundle manifest, timestamps, redaction review
Restart vs targeted remediation Target agent/plugin/external service when controller is healthy; restart with a controller-level hypothesis and recovery plan. Controller/Pipeline/queue Thread dump, uptime, queue/run state, canary

3. Abort versus preserve

Preserve long enough to classify the execution, then stop if it consumes scarce capacity, blocks a lock or continues unsafe work. Jenkins interruption is not instantaneous for every operation. Escalate from normal stop carefully. For remote side effects, an aborted Jenkins build does not prove the external action did not happen.

4. Reconnect versus recreate an agent

A static agent may hold the only local service log/process state, so preserve and diagnose before reconnecting. An ephemeral agent should normally be replaced from a reviewed image/template, but only after controller channel messages, agent/container logs, image identity, labels and timestamps are exported. “Ephemeral” is not permission to discard evidence.

Agent model Recovery default Evidence before change Primary risk
Static VM/bare metal Diagnose then reconnect Service logs, Java/Remoting, OS/network, workspace/process Drift/stale workspace
Ephemeral container/pod Export evidence then replace Container/pod logs, image digest, node metadata First failure disappears on deletion
Autoscaled cloud Correlate provider ID then recycle Provisioning/provider events + Jenkins node record Provider lifecycle outruns retention

5. Thread dump versus heap dump

Thread dumps are relatively lightweight and show thread states/stacks: use them for hangs, deadlocks, request stalls, high CPU and unresponsive aborts. Heap dumps record object memory and can be huge: use them for suspected memory leaks/OOM retention when storage and privacy controls are prepared. A heap dump may contain credential/user/request material.

6. Support bundle scope

Content Value Sensitivity
Plugin inventory Version/dependency/security baseline Can reveal custom/proprietary plugin names
System config Runtime/network context Internal hostnames/URLs
Node data Agent capacity/state Infrastructure topology
Logs Timeline and exceptions Parameters, paths, identities, transformed secrets
Thread dumps Blocked execution/deadlock Request strings and operational details

Keep the original restricted and hash it. If outside escalation is needed, create a reviewed/redacted derivative instead of altering the original evidence.

7. Controller restart versus targeted remediation

If only one agent channel fails, fix that agent. If one artifact service times out, reconcile/recover that dependency. If evidence points to a controller JVM/plugin condition that cannot be safely corrected in place, plan the restart and record pre/post evidence. Restart is an intervention with state consequences, not a neutral diagnostic.

8. Worked scenario: upload timed out after remote acceptance

Console output stops after an artifact PUT. Agent/controller are healthy. Repository shows the immutable object and expected digest, but Jenkins never received the response.

  1. Preserve build/source/artifact digest and repository transaction/request ID.
  2. Hypothesis: response-path/network failure after successful side effect.
  3. Verify repository metadata by immutable coordinate/digest.
  4. Stop the hanging client if necessary after evidence capture.
  5. Do not rebuild/republish. Continue using the already-verified bytes.
  6. Record the reconciled outcome in Jenkins incident evidence.

9. Operational cost matters

Action Diagnostic value Risk/cost
Repeated thread dumps High for hangs Low overhead, still sensitive
Heap dump High for retained-memory problems Large I/O/storage; highly sensitive
Broad support bundle High cross-layer context Collection cost + disclosure surface
Controller restart Can clear process-level fault Outage + destroys live evidence
Agent replacement Fast immutable recovery Can erase ephemeral logs/workspace
Next lesson

Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Diagnostics, Failure Modes, Security, and Performance

Continue with the next lesson and preserve the evidence, safety boundaries, and verification habits established here.

Knowledge check

1. When is agent recreation preferable to reconnect?

2. What is a thread dump better for than a heap dump?

3. Why keep original and sanitized support bundles separately?

4. A deployment timed out but has a remote transaction ID. Rerun first?

5. What supports a controller restart rather than agent remediation?

10. Summary

Incident design is a tradeoff between information, availability, security and recovery. Lesson 4 applies these choices to failure modes and repairs deliberately bad troubleshooting habits.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.