Controller Resilience, Queue Recovery, Agent Loss, Pipeline Durability, Maintenance Windows, and Failure Modes: Diagnostics, Failure Modes, Security, and Performance
Diagnose resilience failures from preserved evidence, distinguish controller/queue/agent/Pipeline/external layers, and avoid “restart and retry” fixes that hide causes or duplicate side effects.
Learning objectives
- Apply an evidence-first diagnostic sequence to interrupted Jenkins work.
- Interpret nonresumable, agent-loss and queue-capacity failures separately.
- Reconcile external side effects before retries.
- Preserve sensitive operational evidence without leaking credentials.
- Tune resilience only after measuring queue/controller/agent bottlenecks.
1. Evidence-first diagnostic sequence
- Freeze the story: job full name, build number, queue item, source SHA, timestamps, controller version, agent/node and external operation ID.
- Preserve console output and relevant controller/agent logs before restart/retry.
- Classify whether the failure is controller process, queue/capacity, agent/Remoting, workspace/tool, Pipeline/CPS/step, credential/network/plugin, artifact/report or external target.
- Determine whether the interrupted step is resumable, conditionally retryable or manual-reconcile.
- Apply the least destructive correction.
- Verify both Jenkins state and external target state independently.
2. Failure 1: blindly rerunning an uncertain deployment
# WRONG: the previous request timed out, so immediately send it again.
curl --fail --request POST https://deploy.example.invalid/releases \
--data-binary @release.json
The timeout may have happened after the server committed the release. The repair is to preserve the original operation/release key and query it first.
# SAFE PATTERN: query exact operation identity before deciding.
operation_id="jenkins-labs-resilience#${BUILD_NUMBER}:deploy"
# pseudo-command for a disposable/mock API:
./deployctl get --operation-id "$operation_id"
# only create if the provider proves the operation is absent and create is idempotent.
3. Failure 2: assuming the agent workspace survived
After an ephemeral agent disappears, a replacement may have no checkout, stash, cache or generated files. If the Pipeline silently assumes a file still exists, the resulting error is a workspace/state problem—not controller data corruption.
set -euo pipefail
# Validate inputs, do not assume them.
test -f recovery-evidence/source-sha.txt || {
echo 'required evidence missing; recover from Jenkins archive/SCM, not guesswork' >&2
exit 66
}
expected="$(cat recovery-evidence/source-sha.txt)"
actual="$(git rev-parse HEAD)"
[ "$actual" = "$expected" ] || { echo 'source mismatch' >&2; exit 67; }
4. Failure 3: restart lands inside a nonresumable step
Jenkins can resume the Pipeline but fail the particular step because
it cannot be resumed. Current Pipeline Basic Steps exposes
nonresumable() as a retry condition for this scenario.
retry(count: 2, conditions: [nonresumable()]) {
junit testResults: 'reports/junit/*.xml'
}
Before applying this pattern, confirm that re-running the step is safe and that the report path refers to still-valid evidence. Do not regenerate test results merely to make ingestion succeed unless that is explicitly the intended recovery.
5. Failure 4: killing the controller before evidence capture
A restart can clear transient symptoms or alter timing without proving root cause. For a stuck build, preserve build thread/step context, controller/agent logs, queue reason, plugin versions and external state before restarting. If the controller is in an active security incident, containment may outrank evidence collection—but that is an incident-response decision, not routine troubleshooting.
6. Failure 5: maintenance with no queue drain
Starting maintenance while new jobs continue entering execution increases the number of interrupted work items and side effects. Quiet-down reduces that uncertainty. It does not eliminate the need to identify long-running jobs, input steps, deployments and external operations that may extend the drain indefinitely.
7. Queue diagnosis before adding executors
A recovered controller may show a growing queue. Do not automatically add executors. Inspect why each item is blocked: label mismatch, offline nodes, throttling/locks, no available executor, cloud provisioning failure or quiet-down. More executors can overload the agent host or run jobs on less-isolated infrastructure without fixing label/policy causes.
8. Correlate evidence across layers
| Evidence | Correlate with | Question |
|---|---|---|
| Build console | Build number + timestamp | Which step was executing when interruption occurred? |
| Controller log | Controller restart/session timestamp | Was shutdown graceful? Did Pipeline reload report errors? |
| Agent log/Remoting | Node name + channel reconnect time | Did the same agent reconnect or was it replaced? |
| Queue reason | Queue item ID + label | Was work waiting because of capacity, quiet-down or eligibility? |
| External API/ledger | Operation ID + artifact digest | Did the side effect complete independently of Jenkins? |
| Metrics | Queue wait, heap/GC, disk I/O, executor utilization | Was resource pressure causal or merely correlated? |
9. Operational evidence can be sensitive
Support bundles, controller logs, environment data, thread dumps and agent logs may reveal internal paths, hostnames, usernames, credential IDs or secret-adjacent data. Do not paste them into public tickets/chat without review. Preserve originals in authorized storage and share redacted subsets.
10. Performance: measure before changing durability or executor count
Recovery problems can be worsened by controller disk saturation, excessive Pipeline state, huge logs/artifacts, slow external storage or overloaded agents. Measure queue wait, controller CPU/heap/GC, disk I/O, workspace/artifact/log growth, agent utilization and external latency. Change one hypothesis at a time and rerun the same controlled workload.
| Symptom | Weak reaction | Better investigation |
|---|---|---|
| Long queue | Add executors everywhere | Inspect queue reasons, labels and agent utilization. |
| Slow controller | Lower durability immediately | Measure disk I/O, GC, logs/artifacts, plugin load first. |
| Agent disconnects | Blind retry many times | Check network/host lifecycle/Remoting logs and safe retry scope. |
| Restart takes long | Delete old state | Measure startup/plugin/job load; preserve recoverable data. |
| Repeated deployment timeouts | Increase retry count | Reconcile target state and measure external latency/error semantics. |
11. Minimal recovery incident record
Incident / maintenance ID:
Controller + version + restart timestamp:
Job full name / build number / queue item:
Source SHA / Jenkinsfile / library refs:
Agent name / label / workspace / reconnect or replacement:
Interrupted stage / step:
Resumability classification:
External operation IDs and observed state:
First-failure evidence locations:
Correction applied:
Retry/rerun/manual-reconcile decision:
Post-recovery verification:
Follow-up preventive action:
Knowledge check
Answer before revealing the explanation.
1. A deployment request timed out. What is the first recovery question?
Whether the external target committed the operation; query by exact operation/release identity before retrying.
2. Why can adding executors make resilience worse?
It can overload hosts or expand the execution/trust surface while leaving the real queue reason unchanged.
3. What does nonresumable() tell retry to do?
Retry only when the failure is classified as a step that could not resume after controller restart, not on arbitrary failures.
4. Why are thread dumps/support bundles treated carefully?
They can contain sensitive operational/internal information and should be reviewed/redacted before sharing.
5. What should remain unchanged when merely retrying a bounded step?
The Jenkins build identity and immutable source identity; a whole new build is a different recovery action.
Official references and version notes
Resilience behavior depends on Jenkins core, Pipeline plugins, agent launchers and individual steps. Re-check current primary documentation before applying these patterns to a real controller.
- Jenkins LTS changelog
- Jenkins Security Advisories
- Managing Jenkins — Prepare for Shutdown and restart guidance
- Scaling Pipelines — speed/durability settings
- Pipeline: Basic Steps — retry conditions
- Running Pipelines — restart from stage
- Durable Task plugin
- Pipeline: Groovy plugin
- Using Jenkins agents
- Controller Isolation
- Remote Access API
- Jenkins CLI
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.