Chapter 35Lesson 02~230 minutes

Controller Resilience, Queue Recovery, Agent Loss, Pipeline Durability, Maintenance Windows, and Failure Modes: Guided Hands-On Workflow and Core Operations

Practice quiet-down, controller restart, agent interruption and side-effect reconciliation on a disposable Jenkins lab while keeping build, queue, agent and external state independently observable.

hands-onquiet-downrestartagent lossdurable taskreconcile

Learning objectives

  • Establish an evidence baseline before planned maintenance.
  • Enter and leave quiet-down deliberately.
  • Observe a Pipeline across controller restart.
  • Observe and classify agent loss without deleting first-failure evidence.
  • Use an idempotency key to reconcile an uncertain external side effect.

1. Disposable scenario

Use the existing local Jenkins lab from earlier chapters. The controller must be disposable and the build must run on a disposable agent labelled resilience-lab, never on the built-in node. Use synthetic Git content and a local side-effect ledger. If your earlier lab uses containers, record controller and agent container names; if it uses VMs/processes, record those exact identities instead.

Identity Example Purpose
Controller jenkins-lab-controller Exact restart target; never use a production controller.
Job jenkins-labs/resilience Stable build history.
Agent label resilience-lab Keeps routine build work off controller.
Source Synthetic repo + full SHA Immutable input identity.
Side-effect ledger /tmp/jenkins-resilience-ledger on disposable shared lab storage Simulates an external system with durable operation IDs.
Evidence recovery-evidence/ Archives state before/after disruptions.

2. Preflight: prove what you are allowed to disrupt

set -euo pipefail
printf 'job=%s build=%s node=%s\n' \
  "${JOB_NAME:-local}" "${BUILD_NUMBER:-0}" "${NODE_NAME:-local}"
if [ "${NODE_NAME:-local}" = "built-in" ]; then
  echo 'Refusing lab on built-in controller node' >&2
  exit 70
fi
mkdir -p recovery-evidence
printf '%s\n' "$(git rev-parse HEAD)" > recovery-evidence/source-sha.txt
java -version > recovery-evidence/java-version.txt 2>&1 || true
python3 --version > recovery-evidence/python-version.txt 2>&1 || true

At the controller, also record Jenkins core/LTS, Pipeline plugin versions, current quiet-down state, queue contents, agent online/offline cause and build URL. Do not begin if any target is production or shared with unrelated workloads.

3. Create a synthetic idempotent external side effect

The ledger stands in for an external deployment/package/ticket system. It persists by operation ID and returns the existing record on repeated calls. In a real integration the reconciliation query would go to the external provider API.

# ci/ledger.py
from pathlib import Path
import json, sys, time
root=Path('/tmp/jenkins-resilience-ledger')
root.mkdir(parents=True, exist_ok=True)
cmd=sys.argv[1]
operation=sys.argv[2]
path=root/f'{operation}.json'
if cmd == 'apply':
    if path.exists():
        print(path.read_text(), end='')
        raise SystemExit(0)
    data={'operation_id':operation,'state':'applied','created_epoch':int(time.time())}
    path.write_text(json.dumps(data, sort_keys=True)+'\n')
    print(json.dumps(data, sort_keys=True))
elif cmd == 'query':
    if not path.exists():
        print(json.dumps({'operation_id':operation,'state':'absent'}))
        raise SystemExit(3)
    print(path.read_text(), end='')
else:
    raise SystemExit('use: apply|query OPERATION_ID')

The operation ID should derive from immutable Jenkins identity, such as job-full-name#build-number:publish. Do not use “latest”.

4. Use a Pipeline with explicit recovery boundaries

pipeline {
  agent none
  options {
    timestamps()
    disableConcurrentBuilds()
  }
  stages {
    stage('Resumable wait') {
      steps {
        echo 'Controller may be restarted while this Pipeline is suspended.'
        sleep time: 90, unit: 'SECONDS'
      }
    }
    stage('Agent work') {
      steps {
        retry(count: 2, conditions: [agent()]) {
          node('resilience-lab') {
            checkout scm
            sh '''#!/usr/bin/env bash
              set -euo pipefail
              mkdir -p recovery-evidence
              printf 'node=%s workspace=%s\\n' "$NODE_NAME" "$WORKSPACE" \
                | tee recovery-evidence/agent.txt
              printf 'before=%s\\n' "$(date -u +%FT%TZ)" \
                | tee recovery-evidence/durable-task.txt
              sleep 75
              printf 'after=%s\\n' "$(date -u +%FT%TZ)" \
                >> recovery-evidence/durable-task.txt
            '''
          }
        }
      }
    }
    stage('Synthetic side effect') {
      steps {
        node('resilience-lab') {
          sh '''#!/usr/bin/env bash
            set -euo pipefail
            op="${JOB_NAME}#${BUILD_NUMBER}:publish"
            python3 ci/ledger.py apply "$op" | tee recovery-evidence/side-effect.json
          '''
        }
      }
    }
  }
  post {
    always {
      node('resilience-lab') {
        archiveArtifacts artifacts: 'recovery-evidence/**', allowEmptyArchive: true, fingerprint: true
      }
    }
  }
}

The stages are intentionally small. A production Pipeline should not wrap a non-idempotent deployment inside broad infrastructure retry logic just because the test stage is safe to repeat.

5. Enter quiet-down and inspect the queue

Use Manage Jenkins → Prepare for Shutdown on the disposable controller, or the documented administrative API/CLI with proper authentication and CSRF handling. Add a maintenance message. Then trigger a second lab build and observe that new work does not start normally while the existing run is allowed to drain.

Record the queued item ID and why it is waiting. Do not cancel it just to make the screenshot/log cleaner.

6. Restart the disposable controller during a resumable boundary

For a container lab, restart only the exact disposable controller container while the Pipeline is in the Resumable wait or long sh stage. If your lab is not containerized, use its documented service restart. Preserve the build URL and current stage before disruption.

set -euo pipefail
controller='jenkins-lab-controller'
actual="$(docker inspect --format '{{.Name}}' "$controller" 2>/dev/null | sed 's#^/##')"
[ "$actual" = "$controller" ] || { echo 'wrong controller target' >&2; exit 70; }
docker restart "$controller"

After Jenkins returns, verify controller version, build number, source SHA and stage. A resumed run should remain the same build number; a manually rerun build is a new build and is different evidence.

7. Interrupt the disposable agent during safe work

While the Agent work stage is running, stop or disconnect only the exact disposable agent. Record the agent offline cause and console line where the interruption appears. If the block is retried under agent(), Jenkins may allocate/reconnect an eligible agent and rerun the bounded block.

set -euo pipefail
agent='jenkins-lab-agent'
actual="$(docker inspect --format '{{.Name}}' "$agent" 2>/dev/null | sed 's#^/##')"
[ "$actual" = "$agent" ] || { echo 'wrong agent target' >&2; exit 70; }
docker stop "$agent"
# Preserve Jenkins console/agent evidence now; then restore only this lab agent.
docker start "$agent"

If the workspace is gone after replacement, that is not a Jenkins controller data-loss event. It is an execution-environment change. Re-checkout/recreate bounded build inputs rather than pretending workspace persistence is guaranteed.

8. Reconcile an uncertain side effect before retry

Suppose the agent/controller connection disappears immediately after the ledger write. Do not call apply again first. Query by the exact operation ID.

set -euo pipefail
op="${JOB_NAME}#${BUILD_NUMBER}:publish"
if python3 ci/ledger.py query "$op" | tee recovery-evidence/reconcile.json; then
  echo 'Side effect already exists; do not duplicate it.'
else
  rc=$?
  if [ "$rc" -eq 3 ]; then
    echo 'Side effect absent; a guarded apply may be considered.'
  else
    exit "$rc"
  fi
fi

Production equivalents include querying a package coordinate/digest, deployment/release ID, change request, cloud operation, or notification event ID.

9. Expected observations

Event Expected Jenkins evidence What it proves / does not prove
Quiet-down Maintenance banner/state; new queue item waiting Scheduler drain state; not process shutdown.
Controller restart Same Pipeline build record reappears and continues or emits a resumability failure Controller state recovery; not external target reconciliation.
Agent stop Node offline/disconnect evidence; safe block may retry Infrastructure failure path; not source-code failure.
Workspace replacement New/empty workspace possible Workspace is execution state, not durable artifact storage.
Ledger query Exact operation ID is applied or absent External state can be reconciled independently from Jenkins result.

10. Cleanup

Exit quiet-down, restore only the disposable lab agent, preserve archived evidence until review is complete, then remove the synthetic ledger and lab resources by exact identity. Do not delete “all stopped containers” or broad temporary directories.

set -euo pipefail
case '/tmp/jenkins-resilience-ledger' in /tmp/jenkins-resilience-ledger) ;; *) exit 70;; esac
rm -rf -- /tmp/jenkins-resilience-ledger
Next

Choose the recovery strategy deliberately

Lesson 3 compares restart versus reload, stage retry versus full rerun, workspace reuse versus replacement, drain versus abrupt stop, and durability versus I/O cost.

Knowledge check

Answer before revealing the explanation.

1. Why trigger a second build during quiet-down?

2. Why does agent() retry wrap only safe build work?

3. What proves a Pipeline resumed rather than reran?

4. What should you do after an uncertain publish/deploy acknowledgment?

5. Why archive evidence before cleanup?

Official references and version notes

Resilience behavior depends on Jenkins core, Pipeline plugins, agent launchers and individual steps. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.