Chapter 35Lesson 04~210 minutes

Controller Resilience, Queue Recovery, Agent Loss, Pipeline Durability, Maintenance Windows, and Failure Modes: Diagnostics, Failure Modes, Security, and Performance

Diagnose resilience failures from preserved evidence, distinguish controller/queue/agent/Pipeline/external layers, and avoid “restart and retry” fixes that hide causes or duplicate side effects.

diagnosticsfailure modessecurityperformanceincident responsefirst failure

Learning objectives

  • Apply an evidence-first diagnostic sequence to interrupted Jenkins work.
  • Interpret nonresumable, agent-loss and queue-capacity failures separately.
  • Reconcile external side effects before retries.
  • Preserve sensitive operational evidence without leaking credentials.
  • Tune resilience only after measuring queue/controller/agent bottlenecks.

1. Evidence-first diagnostic sequence

  1. Freeze the story: job full name, build number, queue item, source SHA, timestamps, controller version, agent/node and external operation ID.
  2. Preserve console output and relevant controller/agent logs before restart/retry.
  3. Classify whether the failure is controller process, queue/capacity, agent/Remoting, workspace/tool, Pipeline/CPS/step, credential/network/plugin, artifact/report or external target.
  4. Determine whether the interrupted step is resumable, conditionally retryable or manual-reconcile.
  5. Apply the least destructive correction.
  6. Verify both Jenkins state and external target state independently.

2. Failure 1: blindly rerunning an uncertain deployment

# WRONG: the previous request timed out, so immediately send it again.
curl --fail --request POST https://deploy.example.invalid/releases \
  --data-binary @release.json

The timeout may have happened after the server committed the release. The repair is to preserve the original operation/release key and query it first.

# SAFE PATTERN: query exact operation identity before deciding.
operation_id="jenkins-labs-resilience#${BUILD_NUMBER}:deploy"
# pseudo-command for a disposable/mock API:
./deployctl get --operation-id "$operation_id"
# only create if the provider proves the operation is absent and create is idempotent.

3. Failure 2: assuming the agent workspace survived

After an ephemeral agent disappears, a replacement may have no checkout, stash, cache or generated files. If the Pipeline silently assumes a file still exists, the resulting error is a workspace/state problem—not controller data corruption.

set -euo pipefail
# Validate inputs, do not assume them.
test -f recovery-evidence/source-sha.txt || {
  echo 'required evidence missing; recover from Jenkins archive/SCM, not guesswork' >&2
  exit 66
}
expected="$(cat recovery-evidence/source-sha.txt)"
actual="$(git rev-parse HEAD)"
[ "$actual" = "$expected" ] || { echo 'source mismatch' >&2; exit 67; }

4. Failure 3: restart lands inside a nonresumable step

Jenkins can resume the Pipeline but fail the particular step because it cannot be resumed. Current Pipeline Basic Steps exposes nonresumable() as a retry condition for this scenario.

retry(count: 2, conditions: [nonresumable()]) {
  junit testResults: 'reports/junit/*.xml'
}

Before applying this pattern, confirm that re-running the step is safe and that the report path refers to still-valid evidence. Do not regenerate test results merely to make ingestion succeed unless that is explicitly the intended recovery.

5. Failure 4: killing the controller before evidence capture

A restart can clear transient symptoms or alter timing without proving root cause. For a stuck build, preserve build thread/step context, controller/agent logs, queue reason, plugin versions and external state before restarting. If the controller is in an active security incident, containment may outrank evidence collection—but that is an incident-response decision, not routine troubleshooting.

6. Failure 5: maintenance with no queue drain

Starting maintenance while new jobs continue entering execution increases the number of interrupted work items and side effects. Quiet-down reduces that uncertainty. It does not eliminate the need to identify long-running jobs, input steps, deployments and external operations that may extend the drain indefinitely.

7. Queue diagnosis before adding executors

A recovered controller may show a growing queue. Do not automatically add executors. Inspect why each item is blocked: label mismatch, offline nodes, throttling/locks, no available executor, cloud provisioning failure or quiet-down. More executors can overload the agent host or run jobs on less-isolated infrastructure without fixing label/policy causes.

8. Correlate evidence across layers

Evidence Correlate with Question
Build console Build number + timestamp Which step was executing when interruption occurred?
Controller log Controller restart/session timestamp Was shutdown graceful? Did Pipeline reload report errors?
Agent log/Remoting Node name + channel reconnect time Did the same agent reconnect or was it replaced?
Queue reason Queue item ID + label Was work waiting because of capacity, quiet-down or eligibility?
External API/ledger Operation ID + artifact digest Did the side effect complete independently of Jenkins?
Metrics Queue wait, heap/GC, disk I/O, executor utilization Was resource pressure causal or merely correlated?

9. Operational evidence can be sensitive

Support bundles, controller logs, environment data, thread dumps and agent logs may reveal internal paths, hostnames, usernames, credential IDs or secret-adjacent data. Do not paste them into public tickets/chat without review. Preserve originals in authorized storage and share redacted subsets.

10. Performance: measure before changing durability or executor count

Recovery problems can be worsened by controller disk saturation, excessive Pipeline state, huge logs/artifacts, slow external storage or overloaded agents. Measure queue wait, controller CPU/heap/GC, disk I/O, workspace/artifact/log growth, agent utilization and external latency. Change one hypothesis at a time and rerun the same controlled workload.

Symptom Weak reaction Better investigation
Long queue Add executors everywhere Inspect queue reasons, labels and agent utilization.
Slow controller Lower durability immediately Measure disk I/O, GC, logs/artifacts, plugin load first.
Agent disconnects Blind retry many times Check network/host lifecycle/Remoting logs and safe retry scope.
Restart takes long Delete old state Measure startup/plugin/job load; preserve recoverable data.
Repeated deployment timeouts Increase retry count Reconcile target state and measure external latency/error semantics.

11. Minimal recovery incident record

Incident / maintenance ID:
Controller + version + restart timestamp:
Job full name / build number / queue item:
Source SHA / Jenkinsfile / library refs:
Agent name / label / workspace / reconnect or replacement:
Interrupted stage / step:
Resumability classification:
External operation IDs and observed state:
First-failure evidence locations:
Correction applied:
Retry/rerun/manual-reconcile decision:
Post-recovery verification:
Follow-up preventive action:
Next

Run the checkpoint failure drill

Lesson 5 combines controller restart, agent loss and an uncertain side effect, then requires an explicit resumable/retryable/manual-reconcile classification for every interrupted action.

Knowledge check

Answer before revealing the explanation.

1. A deployment request timed out. What is the first recovery question?

2. Why can adding executors make resilience worse?

3. What does nonresumable() tell retry to do?

4. Why are thread dumps/support bundles treated carefully?

5. What should remain unchanged when merely retrying a bounded step?

Official references and version notes

Resilience behavior depends on Jenkins core, Pipeline plugins, agent launchers and individual steps. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.