Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Configuration, Design Choices, and Tradeoffs
Troubleshooting design means choosing the least destructive evidence and remediation for the causal layer while preserving enough state to prove the decision later.
Learning objectives
- Choose abort versus preserve based on evidence and external side-effect certainty.
- Compare agent reconnect with immutable-agent recreation.
- Choose thread dump versus heap dump.
- Scope support bundles according to diagnostic need and disclosure risk.
- Choose controller restart versus targeted remediation.
1. Design principle: preserve optionality
The best incident action leaves multiple recovery and learning paths open. A hard restart may restore service but destroys live thread state. Recreating an agent can be operationally correct but may erase the only workspace/process evidence. A heap dump can answer memory-retention questions but adds I/O, storage and disclosure risk. Choose based on the hypothesis, evidence already captured and recovery objective.
2. Decision table
| Choice | Prefer when | State changed | Evidence gate |
|---|---|---|---|
| Abort build vs preserve | Preserve while evidence/side-effect certainty is incomplete; abort once containment requires it and evidence is sufficient. | Build/executor + external state | IDs, last log timestamp, process/thread state, external transaction query |
| Reconnect vs recreate agent | Reconnect when host/workspace continuity matters; recreate ephemeral agents after logs/metadata are exported. | Node/Remoting/workspace | Offline cause, agent log, image/template identity, canary |
| Thread dump vs heap dump | Thread dump for hangs/deadlocks/CPU; heap dump for retained-object memory diagnosis. | Controller JVM | Thread stacks vs heap/GC evidence; sensitivity/capture cost |
| Narrow vs broad support bundle | Use the smallest scope that answers the hypothesis; broaden when causal layer is uncertain. | Controller/plugin/log evidence | Bundle manifest, timestamps, redaction review |
| Restart vs targeted remediation | Target agent/plugin/external service when controller is healthy; restart with a controller-level hypothesis and recovery plan. | Controller/Pipeline/queue | Thread dump, uptime, queue/run state, canary |
3. Abort versus preserve
Preserve long enough to classify the execution, then stop if it consumes scarce capacity, blocks a lock or continues unsafe work. Jenkins interruption is not instantaneous for every operation. Escalate from normal stop carefully. For remote side effects, an aborted Jenkins build does not prove the external action did not happen.
4. Reconnect versus recreate an agent
A static agent may hold the only local service log/process state, so preserve and diagnose before reconnecting. An ephemeral agent should normally be replaced from a reviewed image/template, but only after controller channel messages, agent/container logs, image identity, labels and timestamps are exported. “Ephemeral” is not permission to discard evidence.
| Agent model | Recovery default | Evidence before change | Primary risk |
|---|---|---|---|
| Static VM/bare metal | Diagnose then reconnect | Service logs, Java/Remoting, OS/network, workspace/process | Drift/stale workspace |
| Ephemeral container/pod | Export evidence then replace | Container/pod logs, image digest, node metadata | First failure disappears on deletion |
| Autoscaled cloud | Correlate provider ID then recycle | Provisioning/provider events + Jenkins node record | Provider lifecycle outruns retention |
5. Thread dump versus heap dump
Thread dumps are relatively lightweight and show thread states/stacks: use them for hangs, deadlocks, request stalls, high CPU and unresponsive aborts. Heap dumps record object memory and can be huge: use them for suspected memory leaks/OOM retention when storage and privacy controls are prepared. A heap dump may contain credential/user/request material.
6. Support bundle scope
| Content | Value | Sensitivity |
|---|---|---|
| Plugin inventory | Version/dependency/security baseline | Can reveal custom/proprietary plugin names |
| System config | Runtime/network context | Internal hostnames/URLs |
| Node data | Agent capacity/state | Infrastructure topology |
| Logs | Timeline and exceptions | Parameters, paths, identities, transformed secrets |
| Thread dumps | Blocked execution/deadlock | Request strings and operational details |
Keep the original restricted and hash it. If outside escalation is needed, create a reviewed/redacted derivative instead of altering the original evidence.
7. Controller restart versus targeted remediation
If only one agent channel fails, fix that agent. If one artifact service times out, reconcile/recover that dependency. If evidence points to a controller JVM/plugin condition that cannot be safely corrected in place, plan the restart and record pre/post evidence. Restart is an intervention with state consequences, not a neutral diagnostic.
8. Worked scenario: upload timed out after remote acceptance
Console output stops after an artifact PUT. Agent/controller are healthy. Repository shows the immutable object and expected digest, but Jenkins never received the response.
- Preserve build/source/artifact digest and repository transaction/request ID.
- Hypothesis: response-path/network failure after successful side effect.
- Verify repository metadata by immutable coordinate/digest.
- Stop the hanging client if necessary after evidence capture.
- Do not rebuild/republish. Continue using the already-verified bytes.
- Record the reconciled outcome in Jenkins incident evidence.
9. Operational cost matters
| Action | Diagnostic value | Risk/cost |
|---|---|---|
| Repeated thread dumps | High for hangs | Low overhead, still sensitive |
| Heap dump | High for retained-memory problems | Large I/O/storage; highly sensitive |
| Broad support bundle | High cross-layer context | Collection cost + disclosure surface |
| Controller restart | Can clear process-level fault | Outage + destroys live evidence |
| Agent replacement | Fast immutable recovery | Can erase ephemeral logs/workspace |
Knowledge check
1. When is agent recreation preferable to reconnect?
After evidence export, when the immutable ephemeral image/template is authoritative and stateful workspace continuity is not required.
2. What is a thread dump better for than a heap dump?
Hangs, blocked/deadlocked threads, high CPU stacks and unresponsive executor/requests.
3. Why keep original and sanitized support bundles separately?
The restricted original preserves forensic integrity; the sanitized derivative minimizes disclosure.
4. A deployment timed out but has a remote transaction ID. Rerun first?
No. Reconcile remote state first because the side effect may already exist.
5. What supports a controller restart rather than agent remediation?
Controller/JVM/plugin evidence showing the process is causal plus a defined recovery/validation plan.
10. Summary
Incident design is a tradeoff between information, availability, security and recovery. Lesson 4 applies these choices to failure modes and repairs deliberately bad troubleshooting habits.
Official references and version notes
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.