Checkpoint Lab — Controller Resilience, Queue Recovery, Agent Loss, Pipeline Durability, Maintenance Windows, and Failure Modes
Perform a controlled failure drill: quiet the disposable controller, restart it, interrupt an agent, reconcile an uncertain side effect, and classify every interruption with retained evidence.
Learning objectives
- Predict controller, queue, Pipeline, agent and external-state changes before disruption.
- Prove whether the same build resumes after controller restart.
- Prove agent loss is diagnosed separately from controller restart.
- Reconcile an external side effect before any repeat operation.
- Produce a reviewable recovery decision packet and clean up only exact lab resources.
1. Safety rules
- Controller identity must be explicitly recorded.
- Built-in node must have no routine build workload.
-
Agent label must be
resilience-labor another explicitly disposable label. - All side effects use the local idempotent ledger from Lesson 2.
- Preserve evidence before each disruption and cleanup step.
2. Required predictions before starting
- When the controller restarts during the resumable wait, the same Jenkins build number should return and continue if persistence/plugins are healthy.
- When the disposable agent is stopped during the safe test block, the build may wait/fail/retry according to the agent failure and retry condition, but the external ledger must remain untouched until the side-effect stage.
- If the connection is lost after the ledger write, the ledger query—not Jenkins result color—determines whether the side effect already exists.
Write these predictions into
recovery-evidence/predictions.txt before the drill.
3. Preflight packet
set -euo pipefail
mkdir -p recovery-evidence
{
printf 'utc=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
printf 'job=%s\n' "${JOB_NAME:-unknown}"
printf 'build=%s\n' "${BUILD_NUMBER:-unknown}"
printf 'build_url=%s\n' "${BUILD_URL:-unknown}"
printf 'node=%s\n' "${NODE_NAME:-unknown}"
printf 'workspace=%s\n' "${WORKSPACE:-unknown}"
printf 'source_sha=%s\n' "$(git rev-parse HEAD)"
} | tee recovery-evidence/preflight.txt
[ "${NODE_NAME:-built-in}" != 'built-in' ] || {
echo 'Refusing checkpoint on built-in node' >&2; exit 70;
}
From Manage Jenkins / API, add: Jenkins LTS/core, Java, installed Pipeline/Durable Task versions, controller uptime, quiet-down state, queue snapshot and exact lab agent online state.
4. Checkpoint Pipeline
Use the Lesson 2 Pipeline, but make the interruption markers explicit. The external operation ID must be derived from this build, and its query must be available independently of workspace state.
pipeline {
agent none
options { timestamps(); disableConcurrentBuilds() }
stages {
stage('Controller restart window') {
steps {
echo "RESTART_WINDOW build=${env.BUILD_NUMBER}"
sleep time: 120, unit: 'SECONDS'
}
}
stage('Agent-loss window') {
steps {
retry(count: 2, conditions: [agent()]) {
node('resilience-lab') {
checkout scm
sh '''#!/usr/bin/env bash
set -euo pipefail
mkdir -p recovery-evidence
printf 'attempt_start=%s node=%s\\n' "$(date -u +%FT%TZ)" "$NODE_NAME" \
| tee -a recovery-evidence/agent-attempts.txt
sleep 120
printf 'attempt_end=%s node=%s\\n' "$(date -u +%FT%TZ)" "$NODE_NAME" \
| tee -a recovery-evidence/agent-attempts.txt
'''
}
}
}
}
stage('Idempotent external operation') {
steps {
node('resilience-lab') {
sh '''#!/usr/bin/env bash
set -euo pipefail
op="${JOB_NAME}#${BUILD_NUMBER}:publish"
python3 ci/ledger.py apply "$op" | tee recovery-evidence/apply.json
printf '%s\\n' "$op" > recovery-evidence/operation-id.txt
'''
}
}
}
stage('Verify external operation') {
steps {
node('resilience-lab') {
sh '''#!/usr/bin/env bash
set -euo pipefail
op="$(cat recovery-evidence/operation-id.txt)"
python3 ci/ledger.py query "$op" | tee recovery-evidence/query.json
'''
}
}
}
}
post {
always {
node('resilience-lab') {
archiveArtifacts artifacts: 'recovery-evidence/**', allowEmptyArchive: true, fingerprint: true
}
}
}
}
5. Required drill sequence
| Step | Action | Evidence to preserve before next step |
|---|---|---|
| A | Start Build N and confirm exact source SHA/agent label. | Build URL/number/cause, source SHA, controller/plugin baseline. |
| B | Enter quiet-down; trigger Build N+1 so it waits. | Quiet-down message/state, queue ID/reason for N+1. |
| C | During Build N restart window, restart only disposable controller. | Console position, restart timestamp, controller log excerpt. |
| D | After return, prove Build N is same run and continuing. | Same build number, resumed stage/step, controller uptime reset. |
| E | During agent-loss window, stop exact disposable agent. | Agent name/channel/offline cause, first failure console lines. |
| F | Restore or provide eligible lab agent; observe retry/reconnect behavior. | Agent attempt file, node/workspace identity, retry evidence. |
| G | Allow side-effect apply; then simulate uncertainty by temporarily disrupting observation, not deleting ledger. | Operation ID + apply record. |
| H | Query ledger before any repeat apply. | Query result proving applied/absent. |
| I | Exit quiet-down and allow queued Build N+1 only after recovery review. | Cancel-quiet-down state and queue transition. |
| J | Archive evidence, then clean exact lab resources. | Final verification checklist + limitations. |
6. Required recovery classification
For every interrupted activity, fill this table in your evidence packet.
| Activity | Expected class | Reason |
|---|---|---|
Pipeline sleep around controller restart |
Resumable | Pipeline step is designed to suspend/resume around controller restart. |
| Long durable shell with controller restart while agent survives | Resumable in supported conditions | Durable Task persists task coordination; verify same build/task evidence. |
| Checkout/report step interrupted exactly by restart | Retryable only if classified nonresumable and repeat is safe | Use targeted retry rather than blind rerun. |
| Test block lost with agent | Retryable infrastructure scope |
agent() can target infrastructure failure when
work is repeatable.
|
| Workspace on destroyed ephemeral agent | Recreate input, not “resume workspace” | Workspace is disposable execution state. |
| External publish with lost acknowledgment | Manual/automated reconcile first | External target may already have committed the operation. |
7. Required evidence packet
| Evidence group | Minimum contents |
|---|---|
| Controller | URL/name, LTS/core, Java, uptime before/after, restart timestamp, quiet-down state. |
| Plugins | Pipeline/Durable Task relevant versions and 17 Sep 2026 advisory review note. |
| Job/build | Full name, Build N number/URL/cause/result and queued Build N+1 queue item. |
| Source | Exact SHA, Jenkinsfile/library refs if used. |
| Agent | Node/label, online/offline cause, reconnect/replacement identity, workspace path. |
| Pipeline | Stage/step at each interruption; resumability classification. |
| External | Operation ID, apply/query response, proof no duplicate apply was needed. |
| Logs | First failure console lines + relevant controller/agent timestamps. |
| Decision | Why resume/retry/recreate/manual-reconcile was chosen. |
| Limitations | What the lab does not prove about HA storage, cloud agents or production deploy APIs. |
8. Verification checklist
- Build N after restart is the same build number, not a fresh rerun.
- Queued Build N+1 did not start while quiet-down was active.
- Agent loss evidence names the exact disposable node and first failure timestamp.
- No routine build ran on the built-in node.
- No release artifact was rebuilt as a recovery shortcut.
- The external operation was queried by exact operation ID before any retry decision.
- Archived evidence contains no real credentials or production URLs/data.
- Quiet-down was cancelled only after the recovery state was understood.
- Cleanup removed only exact lab resources.
9. Cleanup / rollback
Cancel quiet-down on the disposable controller, restore the lab agent only if you still need it, and retain Jenkins archived evidence until the checkpoint is reviewed. Then remove the synthetic ledger by exact path and optionally delete the lab job/agent by exact name.
set -euo pipefail
ledger='/tmp/jenkins-resilience-ledger'
[ "$ledger" = '/tmp/jenkins-resilience-ledger' ] || exit 70
rm -rf -- "$ledger"
If a container image/plugin/controller version was temporarily changed for the drill, restore only through the lab’s documented configuration source. Do not perform unsupported core/plugin downgrades as a casual rollback.
10. What the checkpoint proves—and does not prove
| Supported claim | Not supported by this lab alone |
|---|---|
| Pipeline/controller state can be observed across a controlled disposable restart. | That every plugin/step on a production controller is resumable. |
| Agent loss can be classified separately and safe work can be retried narrowly. | That every cloud/ephemeral agent failure should be retried. |
| An idempotent operation ID supports external reconciliation. | That a real deployment API is idempotent without provider-specific validation. |
| Quiet-down provides a controlled drain mechanism. | That all maintenance windows finish quickly or without manual intervention. |
| The recovery packet supports evidence-driven decisions. | That backups/DR are validated; Chapter 36 covers restore and disaster recovery. |
11. What Chapter 35 adds to the Jenkins operating model
You now have a recovery vocabulary grounded in state: controller process, persisted Pipeline/queue state, agent/workspace state and external side effects are inspected independently. A restart is no longer a generic troubleshooting button, and retry is no longer a synonym for resilience.
Chapter 36 turns this into disaster-recovery engineering: backups, restores, controller migration, secret/key recovery and recovery testing. Resilience keeps work recoverable through ordinary failures; disaster recovery proves you can rebuild the controller when the controller state itself is lost or moved.
Knowledge check
Answer before revealing the explanation.
1. What observation proves controller restart recovery rather than rerun?
The same Build N run/number continues after the controller restarts.
2. Why is Build N+1 intentionally queued during quiet-down?
It proves admission/scheduling is paused while the controller itself remains running.
3. What is the correct classification for an external publish with unknown acknowledgment?
Manual/automated reconciliation first; query the exact operation identity before retry.
4. Why is a destroyed workspace not restored from memory?
Workspace is agent execution state; recreate from immutable SCM/artifact inputs rather than assuming persistence.
5. What does Chapter 36 add beyond this checkpoint?
Validated backup/restore, controller migration, secret/key recovery and disaster-recovery testing when controller state itself must be reconstructed.
Official references and version notes
Resilience behavior depends on Jenkins core, Pipeline plugins, agent launchers and individual steps. Re-check current primary documentation before applying these patterns to a real controller.
- Jenkins LTS changelog
- Jenkins Security Advisories
- Managing Jenkins — Prepare for Shutdown and restart guidance
- Scaling Pipelines — speed/durability settings
- Pipeline: Basic Steps — retry conditions
- Running Pipelines — restart from stage
- Durable Task plugin
- Pipeline: Groovy plugin
- Using Jenkins agents
- Controller Isolation
- Remote Access API
- Jenkins CLI
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.