Chapter 40Lesson 02~230 minutes

Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Guided Hands-On Workflow and Core Operations

A useful incident lab bounds the failure and preserves the evidence. This workflow creates queue pressure, a long-running step and an agent-channel loss without touching production infrastructure.

guided workflowqueue IDsagent lossthread dumpsupport bundletimeline

Learning objectives

  • Create a reproducible disposable controller/project/agent incident scenario.
  • Diagnose queue pressure, a running long step and an agent disconnect as different layers.
  • Capture thread/support evidence before intervention.
  • Use stop/reconnect/restart options only after evidence capture.
  • Produce a compact incident timeline and evidence packet.

1. Scenario and lab boundary

Use local controller jenkins-incident-lab and one disposable agent labeled incident-lab-agent with one executor. Built-in controller executors stay at zero. All “external” latency is simulated with sleep; no real cloud, SCM organization, artifact repository or production credential is needed.

2. Preflight

Check Expected evidence Stop condition
Controller 2.568.3 / Java 21 / uptime Undocumented baseline difference
Support Core 1863.vdddd6f8d12a_8 if bundle exercise is used Unknown plugin/security baseline
Controller node 0 executors Routine work can land on controller
Agent Online, Java 21, one executor, expected label Unknown execution identity
Storage Space for logs/dumps/bundle Existing disk pressure
External targets None; local simulation only Any real endpoint configured

3. Create the synthetic Pipeline

Every build records identity first, then performs one bounded mode. The control mode proves the harness. The other modes produce deterministic delay without external side effects.

pipeline {
  agent { label 'incident-lab-agent' }
  options { timestamps(); disableConcurrentBuilds() }
  parameters {
    choice(name: 'MODE', choices: ['normal', 'slow-external', 'local-stall'], description: 'Synthetic incident mode')
  }
  stages {
    stage('Identity') {
      steps {
        sh '''set -eu
          mkdir -p evidence
          printf 'job=%s\nbuild=%s\nnode=%s\nworkspace=%s\n' \
            "$JOB_NAME" "$BUILD_NUMBER" "$NODE_NAME" "$WORKSPACE" > evidence/identity.txt
          git rev-parse HEAD 2>/dev/null > evidence/source-sha.txt || printf 'synthetic-no-scm\n' > evidence/source-sha.txt
        '''
      }
    }
    stage('Bounded work') {
      steps {
        script {
          if (params.MODE == 'normal') {
            sh 'sleep 5; echo ok > evidence/work-result.txt'
          } else if (params.MODE == 'slow-external') {
            sh 'echo synthetic-external-start; sleep 45; echo synthetic-external-done'
          } else {
            sh 'echo synthetic-stall-marker; sleep 120'
          }
        }
      }
    }
  }
  post { always { archiveArtifacts artifacts: 'evidence/**', allowEmptyArchive: true, fingerprint: true } }
}

4. Run a normal control build

Run MODE=normal and record build number, queue item if visible, assigned agent, workspace and source identity. Confirm the archived evidence exists. Do not inject failures until the control case is stable.

5. Case A — queue pressure

Occupy the agent's only executor with MODE=local-stall, then enqueue another build. Inspect queue/node JSON.

curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,why,blocked,stuck,inQueueSince,task[name,url]]" | tee evidence/queue-pressure.json
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,numExecutors,busyExecutors,offlineCauseReason]" | tee evidence/nodes-pressure.json

The expected diagnosis is capacity/eligibility. No second Pipeline executor thread exists yet because the second build has not started.

6. Case B — a running build appears stuck

With the long step active, record the exact build URL/executor, console tail and node state. Capture a controller thread dump before changing the build.

curl -fsS "$JENKINS_URL/job/incident-lab/$BUILD_NUMBER/api/json?tree=number,url,building,result,timestamp,duration" > evidence/build-before-stop.json
curl -fsS "$JENKINS_URL/threadDump" > evidence/controller-threadDump.txt
sha256sum evidence/controller-threadDump.txt >> evidence/SHA256SUMS

If you need to stop the exact disposable build, use normal Pipeline /stop first. Jenkins documents /term and then /kill as progressively more forceful options; hard kill is a last resort, not a routine diagnostic.

7. Case C — lose the disposable agent channel

While MODE=slow-external runs, stop only the lab agent process/network. Preserve the controller-side offline cause and agent-side launcher/service log before reconnecting. The controller can remain healthy while the running step loses its channel.

curl -fsS "$JENKINS_URL/computer/incident-lab-agent/api/json?tree=displayName,offline,temporarilyOffline,offlineCauseReason,executors[currentExecutable[url]]" > evidence/agent-loss.json
java -version 2> evidence/agent-java.txt
# Preserve the disposable agent service/container log before restart.

8. Review a support bundle before sharing

If Support Core is installed, generate a lab-only bundle through the Jenkins Support action or documented CLI. List contents, extract only in a restricted directory, and search likely sensitive terms. This is a review aid, not proof of complete redaction.

mkdir -m 700 -p evidence/support-review
unzip -l support-bundle.zip > evidence/support-review/member-list.txt
unzip -q support-bundle.zip -d evidence/support-review/extracted
grep -RniE 'token|password|secret|authorization|cookie|private|internal\.' evidence/support-review/extracted | head -n 100 || true

9. Broken incident response: restart first

BAD: immediately hard-kill build, restart controller, delete workspace, rerun deployment.
LOST: live thread state, queue/agent evidence, original workspace/process state, external side-effect certainty.

The repair is to capture IDs and read-only state first, reconcile external effects, and record every intervention as a timestamped event.

10. Write the incident timeline

15:20:04Z build #17 queued; item 381 recorded
15:21:10Z #17 assigned incident-lab-agent executor 0
15:21:42Z agent channel lost; node JSON + agent log preserved
15:22:05Z controller thread dump captured; SHA-256 recorded
15:23:12Z agent restarted after evidence capture
15:24:08Z bounded canary healthy; no real external side effect existed

11. Layer-selection challenge

A PR build waits 20 minutes. Controller CPU/heap are normal. Three agents are online but none has label browser-linux. The next action belongs to queue/node-label capacity, not Pipeline CPS, controller heap or support-bundle generation. State the evidence you would preserve before adding/changing capacity.

12. Verification and cleanup

  1. Verify built-in executors remain zero.
  2. Reconnect the lab agent and run one normal canary.
  3. Preserve the original failed/queued build evidence until review.
  4. Hash the final timeline/manifest.
  5. Delete only the synthetic job/agent/evidence copies after the exercise is closed.
Next lesson

Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks: Configuration, Design Choices, and Tradeoffs

Continue with the next lesson and preserve the evidence, safety boundaries, and verification habits established here.

Knowledge check

1. Why run a control build before failure injection?

2. A queue item exists but no executor is assigned. Is a running shell process relevant?

3. When is /kill appropriate?

4. Why preserve agent logs before reconnecting?

5. What do you do if a support bundle contains internal URLs?

13. Summary

Queue pressure, long-running work and agent loss can all look like “stuck Jenkins,” but they are different causal layers. Lesson 3 now turns these operations into explicit design tradeoffs.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.