Controller Resilience, Queue Recovery, Agent Loss, Pipeline Durability, Maintenance Windows, and Failure Modes: Guided Hands-On Workflow and Core Operations
Practice quiet-down, controller restart, agent interruption and side-effect reconciliation on a disposable Jenkins lab while keeping build, queue, agent and external state independently observable.
Learning objectives
- Establish an evidence baseline before planned maintenance.
- Enter and leave quiet-down deliberately.
- Observe a Pipeline across controller restart.
- Observe and classify agent loss without deleting first-failure evidence.
- Use an idempotency key to reconcile an uncertain external side effect.
1. Disposable scenario
Use the existing local Jenkins lab from earlier chapters. The
controller must be disposable and the build must run on a disposable
agent labelled resilience-lab, never on the built-in
node. Use synthetic Git content and a local side-effect ledger. If
your earlier lab uses containers, record controller and agent
container names; if it uses VMs/processes, record those exact
identities instead.
| Identity | Example | Purpose |
|---|---|---|
| Controller | jenkins-lab-controller |
Exact restart target; never use a production controller. |
| Job | jenkins-labs/resilience |
Stable build history. |
| Agent label | resilience-lab |
Keeps routine build work off controller. |
| Source | Synthetic repo + full SHA | Immutable input identity. |
| Side-effect ledger |
/tmp/jenkins-resilience-ledger on disposable
shared lab storage
|
Simulates an external system with durable operation IDs. |
| Evidence | recovery-evidence/ |
Archives state before/after disruptions. |
2. Preflight: prove what you are allowed to disrupt
set -euo pipefail
printf 'job=%s build=%s node=%s\n' \
"${JOB_NAME:-local}" "${BUILD_NUMBER:-0}" "${NODE_NAME:-local}"
if [ "${NODE_NAME:-local}" = "built-in" ]; then
echo 'Refusing lab on built-in controller node' >&2
exit 70
fi
mkdir -p recovery-evidence
printf '%s\n' "$(git rev-parse HEAD)" > recovery-evidence/source-sha.txt
java -version > recovery-evidence/java-version.txt 2>&1 || true
python3 --version > recovery-evidence/python-version.txt 2>&1 || true
At the controller, also record Jenkins core/LTS, Pipeline plugin versions, current quiet-down state, queue contents, agent online/offline cause and build URL. Do not begin if any target is production or shared with unrelated workloads.
3. Create a synthetic idempotent external side effect
The ledger stands in for an external deployment/package/ticket system. It persists by operation ID and returns the existing record on repeated calls. In a real integration the reconciliation query would go to the external provider API.
# ci/ledger.py
from pathlib import Path
import json, sys, time
root=Path('/tmp/jenkins-resilience-ledger')
root.mkdir(parents=True, exist_ok=True)
cmd=sys.argv[1]
operation=sys.argv[2]
path=root/f'{operation}.json'
if cmd == 'apply':
if path.exists():
print(path.read_text(), end='')
raise SystemExit(0)
data={'operation_id':operation,'state':'applied','created_epoch':int(time.time())}
path.write_text(json.dumps(data, sort_keys=True)+'\n')
print(json.dumps(data, sort_keys=True))
elif cmd == 'query':
if not path.exists():
print(json.dumps({'operation_id':operation,'state':'absent'}))
raise SystemExit(3)
print(path.read_text(), end='')
else:
raise SystemExit('use: apply|query OPERATION_ID')
The operation ID should derive from immutable Jenkins identity, such
as job-full-name#build-number:publish. Do not use
“latest”.
4. Use a Pipeline with explicit recovery boundaries
pipeline {
agent none
options {
timestamps()
disableConcurrentBuilds()
}
stages {
stage('Resumable wait') {
steps {
echo 'Controller may be restarted while this Pipeline is suspended.'
sleep time: 90, unit: 'SECONDS'
}
}
stage('Agent work') {
steps {
retry(count: 2, conditions: [agent()]) {
node('resilience-lab') {
checkout scm
sh '''#!/usr/bin/env bash
set -euo pipefail
mkdir -p recovery-evidence
printf 'node=%s workspace=%s\\n' "$NODE_NAME" "$WORKSPACE" \
| tee recovery-evidence/agent.txt
printf 'before=%s\\n' "$(date -u +%FT%TZ)" \
| tee recovery-evidence/durable-task.txt
sleep 75
printf 'after=%s\\n' "$(date -u +%FT%TZ)" \
>> recovery-evidence/durable-task.txt
'''
}
}
}
}
stage('Synthetic side effect') {
steps {
node('resilience-lab') {
sh '''#!/usr/bin/env bash
set -euo pipefail
op="${JOB_NAME}#${BUILD_NUMBER}:publish"
python3 ci/ledger.py apply "$op" | tee recovery-evidence/side-effect.json
'''
}
}
}
}
post {
always {
node('resilience-lab') {
archiveArtifacts artifacts: 'recovery-evidence/**', allowEmptyArchive: true, fingerprint: true
}
}
}
}
The stages are intentionally small. A production Pipeline should not wrap a non-idempotent deployment inside broad infrastructure retry logic just because the test stage is safe to repeat.
5. Enter quiet-down and inspect the queue
Use Manage Jenkins → Prepare for Shutdown on the disposable controller, or the documented administrative API/CLI with proper authentication and CSRF handling. Add a maintenance message. Then trigger a second lab build and observe that new work does not start normally while the existing run is allowed to drain.
Record the queued item ID and why it is waiting. Do not cancel it just to make the screenshot/log cleaner.
6. Restart the disposable controller during a resumable boundary
For a container lab, restart only the exact disposable controller
container while the Pipeline is in the
Resumable wait or long sh stage. If your
lab is not containerized, use its documented service restart.
Preserve the build URL and current stage before disruption.
set -euo pipefail
controller='jenkins-lab-controller'
actual="$(docker inspect --format '{{.Name}}' "$controller" 2>/dev/null | sed 's#^/##')"
[ "$actual" = "$controller" ] || { echo 'wrong controller target' >&2; exit 70; }
docker restart "$controller"
After Jenkins returns, verify controller version, build number, source SHA and stage. A resumed run should remain the same build number; a manually rerun build is a new build and is different evidence.
7. Interrupt the disposable agent during safe work
While the Agent work stage is running, stop or
disconnect only the exact disposable agent. Record the agent offline
cause and console line where the interruption appears. If the block
is retried under agent(), Jenkins may
allocate/reconnect an eligible agent and rerun the bounded block.
set -euo pipefail
agent='jenkins-lab-agent'
actual="$(docker inspect --format '{{.Name}}' "$agent" 2>/dev/null | sed 's#^/##')"
[ "$actual" = "$agent" ] || { echo 'wrong agent target' >&2; exit 70; }
docker stop "$agent"
# Preserve Jenkins console/agent evidence now; then restore only this lab agent.
docker start "$agent"
If the workspace is gone after replacement, that is not a Jenkins controller data-loss event. It is an execution-environment change. Re-checkout/recreate bounded build inputs rather than pretending workspace persistence is guaranteed.
8. Reconcile an uncertain side effect before retry
Suppose the agent/controller connection disappears immediately after
the ledger write. Do not call apply again first. Query
by the exact operation ID.
set -euo pipefail
op="${JOB_NAME}#${BUILD_NUMBER}:publish"
if python3 ci/ledger.py query "$op" | tee recovery-evidence/reconcile.json; then
echo 'Side effect already exists; do not duplicate it.'
else
rc=$?
if [ "$rc" -eq 3 ]; then
echo 'Side effect absent; a guarded apply may be considered.'
else
exit "$rc"
fi
fi
Production equivalents include querying a package coordinate/digest, deployment/release ID, change request, cloud operation, or notification event ID.
9. Expected observations
| Event | Expected Jenkins evidence | What it proves / does not prove |
|---|---|---|
| Quiet-down | Maintenance banner/state; new queue item waiting | Scheduler drain state; not process shutdown. |
| Controller restart | Same Pipeline build record reappears and continues or emits a resumability failure | Controller state recovery; not external target reconciliation. |
| Agent stop | Node offline/disconnect evidence; safe block may retry | Infrastructure failure path; not source-code failure. |
| Workspace replacement | New/empty workspace possible | Workspace is execution state, not durable artifact storage. |
| Ledger query | Exact operation ID is applied or absent | External state can be reconciled independently from Jenkins result. |
10. Cleanup
Exit quiet-down, restore only the disposable lab agent, preserve archived evidence until review is complete, then remove the synthetic ledger and lab resources by exact identity. Do not delete “all stopped containers” or broad temporary directories.
set -euo pipefail
case '/tmp/jenkins-resilience-ledger' in /tmp/jenkins-resilience-ledger) ;; *) exit 70;; esac
rm -rf -- /tmp/jenkins-resilience-ledger
Knowledge check
Answer before revealing the explanation.
1. Why trigger a second build during quiet-down?
To observe scheduler/queue behavior separately from the already-running build and controller process.
2. Why does agent() retry wrap only safe build work?
Because infrastructure retry repeats the enclosed block; non-idempotent side effects could be duplicated.
3. What proves a Pipeline resumed rather than reran?
The same Jenkins build number/run identity continues after restart.
4. What should you do after an uncertain publish/deploy acknowledgment?
Query the external system by immutable operation/artifact/release identity before deciding to retry.
5. Why archive evidence before cleanup?
Recovery actions can overwrite workspace/log context and make root-cause reconstruction harder.
Official references and version notes
Resilience behavior depends on Jenkins core, Pipeline plugins, agent launchers and individual steps. Re-check current primary documentation before applying these patterns to a real controller.
- Jenkins LTS changelog
- Jenkins Security Advisories
- Managing Jenkins — Prepare for Shutdown and restart guidance
- Scaling Pipelines — speed/durability settings
- Pipeline: Basic Steps — retry conditions
- Running Pipelines — restart from stage
- Durable Task plugin
- Pipeline: Groovy plugin
- Using Jenkins agents
- Controller Isolation
- Remote Access API
- Jenkins CLI
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.