Chapter 40Lesson 05~260 minutes

Checkpoint Lab — Production Troubleshooting: Thread Dumps, Support Bundles, Stuck Builds, Agent Failures, and Incident Playbooks

The checkpoint is a controlled production-style incident drill: collect evidence before intervention, separate primary cause from downstream symptoms, recover minimally, and leave a handoff another operator can continue.

checkpointincident drillqueue pressureagent lossthread dumprunbook

Learning objectives

  • Run bounded work on a disposable non-controller agent.
  • Create queue pressure and one agent/network symptom.
  • Preserve exact build/queue/agent/source IDs before intervention.
  • Capture and interpret thread/support evidence.
  • Recover only the failed layer and prove it.
  • Produce an incident handoff and runbook improvement.

1. Setup and assumptions

Use disposable controller jenkins-incident-drill and agent incident-lab-agent with one executor and matching label. Built-in node executors remain zero. No production external endpoint or credential is permitted.

Component Assumption Evidence
Controller 2.568.3 / Java 21 System Information/About + uptime
Support Core 1863.vdddd6f8d12a_8 if installed Plugin version + reviewed bundle manifest
Agent Java 21 / one executor / disposable Node JSON + launcher log + workspace
Security 16 Sep 2026 advisory reviewed Plugin inventory/notes
External Local simulation only Explicit statement of no production side effect

2. Predict state changes first

  1. When the only eligible executor is occupied, the second build will remain queued with a recorded queue reason while the controller remains healthy.
  2. When the disposable agent channel is lost, controller node/offline evidence and build console state will change; the controller process need not fail.
  3. Thread-dump capture will preserve build/queue identities rather than mutate them.
  4. Agent reconnect creates new connection events, so pre-reconnect logs must be preserved.

3. Preflight

  • Record core/Java/plugin inventory and controller uptime.
  • Confirm built-in executors are zero.
  • Confirm lab agent identity/Java/label/executor.
  • Record disk space and clock.
  • Create restricted local evidence directory.
  • Verify job has no real external endpoint/credential.
mkdir -m 700 -p incident-evidence/INC-0040
cd incident-evidence/INC-0040
date -u +%FT%TZ > preflight-time.txt
printf '%s
' 'controller=jenkins-incident-drill' 'scope=disposable-local-only' > scope.txt

4. Exact checkpoint Pipeline

pipeline {
  agent { label 'incident-lab-agent' }
  options { timestamps(); disableConcurrentBuilds() }
  parameters { booleanParam(name: 'INJECT_STALL', defaultValue: false, description: 'Synthetic long step') }
  stages {
    stage('Evidence seed') {
      steps {
        sh '''set -eu
          mkdir -p evidence
          printf 'job=%s\nbuild=%s\nnode=%s\nworkspace=%s\n' \
            "$JOB_NAME" "$BUILD_NUMBER" "$NODE_NAME" "$WORKSPACE" > evidence/identity.txt
          date -u +%FT%TZ > evidence/start-utc.txt
        '''
      }
    }
    stage('Synthetic work') {
      steps {
        script {
          if (params.INJECT_STALL) { sh 'echo stall-begin; sleep 120; echo stall-end' }
          else { sh 'sleep 3; echo healthy > evidence/result.txt' }
        }
      }
    }
  }
  post { always { archiveArtifacts artifacts: 'evidence/**', allowEmptyArchive: true, fingerprint: true } }
}

Run one normal control build. Then start INJECT_STALL=true. Once that build occupies the only executor, start a second build so the queue/capacity state is observable.

5. Evidence packet

Evidence Required fields Purpose
Controller baseline Core, Java, plugin manifest, uptime Platform identity
Build Job, build URL/number, cause, source/Jenkinsfile revision Exact execution identity
Queue Queue item ID, why, time, label Scheduling evidence
Agent Node, executor, Java/Remoting, workspace, offline cause Execution-environment evidence
Thread dump Timestamp + SHA-256 Controller/JVM state before intervention
Support review Manifest + sensitivity note Cross-layer evidence with disclosure control
External note No production target; simulation only Side-effect boundary
Actions UTC time, actor, action, reason Audit/replay

6. Induce the incident

  1. Record queued item and reason while stall build occupies executor.
  2. Capture build/node/queue state and console tail.
  3. Capture controller thread dump and hash it.
  4. Stop/disconnect only the disposable agent process/network.
  5. Preserve agent-side log and controller-side offline/channel evidence before reconnect.
  6. If Support Core is installed, create/review a restricted bundle.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,why,blocked,stuck,inQueueSince,task[name,url]]" > queue-before.json
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,numExecutors,busyExecutors,offlineCauseReason]" > nodes-before.json
curl -fsS "$JENKINS_URL/threadDump" > controller-threadDump-before.txt
sha256sum controller-threadDump-before.txt > SHA256SUMS

7. Classify the evidence

Observation Layer Interpretation
Second build waited before agent loss Queue/capacity Expected single-executor pressure
Agent offline during running build Agent/Remoting/network Execution channel changed
Controller UI/API/thread dump responsive Controller/JVM No controller failure evidence
Running step stops after channel loss Pipeline/agent boundary Inspect channel/durable-step behavior
No external transaction exists External state Lab retry risk is bounded by design

Root cause should identify the injected agent/channel loss. Queue starvation after that is a downstream consequence, not the primary cause.

8. Recover minimally

Reconnect only the disposable agent after preserving logs. Verify Java/label/executor. Preserve the interrupted build result; if needed, run a new bounded control canary rather than rewriting history. Do not restart the controller because the evidence does not identify it as causal.

9. Verification checklist

  • Pre-failure queue item/reason retained.
  • Pre-reconnect agent log/offline cause retained.
  • Thread-dump SHA-256 verifies.
  • Controller availability history documented.
  • Agent reconnects with expected Java/labels/executor.
  • New normal canary succeeds on that agent.
  • No build ran on built-in node.
  • No production external side effect occurred.
  • Every intervention has timestamp/reason.

10. Incident handoff

# INC-0040 handoff

## Impact
Disposable Jenkins lab only. Build #17 interrupted; build #18 queued.

## Identity
- controller: jenkins-incident-drill / 2.568.3 / Java 21
- job: incident-lab
- affected build: #17
- queue item: #481 for build #18
- agent: incident-lab-agent / executor 0

## Timeline
- 16:00:00Z #17 began synthetic stall
- 16:00:08Z #18 queued; capacity reason captured
- 16:00:25Z controller thread dump captured and hashed
- 16:00:40Z agent channel intentionally stopped
- 16:00:43Z agent/controller channel logs preserved
- 16:02:00Z agent restarted after evidence capture
- 16:03:10Z normal canary succeeded

## Root cause
Injected loss of the only eligible agent channel; queue starvation followed.

## Recovery
Reconnect same disposable agent after evidence capture; verify Java/label; run bounded canary.

## Runbook update
Export ephemeral-agent logs before teardown and correlate job/build/node/timestamp automatically.

11. Mandatory runbook improvement

Choose one evidence-derived improvement: automatic ephemeral-agent log export, a standard queue/node JSON capture command, documented thread-dump capture, support-bundle review ownership, or external transaction-ID reconciliation before retries. Give it an owner and review date.

12. Cleanup

  1. Confirm canary and evidence packet.
  2. Remove only synthetic job and disposable agent.
  3. Delete extracted support working copies after approved evidence is secured.
  4. Keep the failed build until exercise review completes.
  5. Remove the controller only after confirming no other lab depends on it.
Next chapter

Chapter 41 — Production Capstone: Design, Automate, Secure, Scale, Upgrade, and Recover an Enterprise Jenkins Platform

Carry the verified evidence and operating discipline from this chapter into the next chapter.

Knowledge check

1. Primary root cause when the only eligible agent is disconnected?

2. Why preserve the failed build after canary recovery?

3. Controller remains responsive after agent loss. Restart anyway?

4. Which evidence is especially sensitive?

5. What makes recovery minimal?

6. What does the canary prove?

13. What Chapter 40 adds

You now have an incident method that preserves state, correlates controller/job/build/queue/agent/external evidence, protects diagnostic data, chooses intervention by causal layer, validates recovery and feeds learning back into runbooks. Chapter 41 brings these practices together in the final enterprise Jenkins platform capstone.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.