Chapter 35Lesson 01~190 minutes

Controller Resilience, Queue Recovery, Agent Loss, Pipeline Durability, Maintenance Windows, and Failure Modes: Concepts, Architecture, and Mental Model

Understand what Jenkins persists across controller restarts, what belongs to an agent workspace, what Pipeline can resume, and why external side effects always require explicit reconciliation.

resiliencequeuerestartdurabilityagentsrecovery

Learning objectives

  • Trace queued/running work through controller restart, agent loss and recovery.
  • Separate controller process state, queue/run records, Pipeline CPS state, durable task state and workspace state.
  • Explain quiet-down, safe restart, abrupt restart and configuration reload as different operations.
  • Classify work as resumable, retryable, re-runnable or manual-reconcile.
  • Recognize why external side effects are not automatically rolled back by Jenkins.

1. The problem: “Jenkins came back” is not the same as “the delivery recovered”

Jenkins is a stateful automation controller. A process restart can preserve a queued item, a Pipeline flow graph and a durable shell task while still losing an agent connection or workspace. Conversely, the controller may return healthy while a deployment call completed externally just before the connection was lost, leaving Jenkins uncertain whether retrying is safe.

The resilience engineer therefore asks a narrower question than “did Jenkins restart?”: which state was persisted, which state still exists outside Jenkins, what evidence survived, and what is the smallest safe recovery action?

2. Mental model: persisted orchestration meets disposable execution

Mental model: persisted orchestration meets disposable execution
flowchart TD
  A["Queued or running Jenkins workload"] --> B["Persisted controller / Pipeline state"]
  A --> C["Agent process + workspace + durable task"]
  A --> D["External side effect"]
  B --> E{"Controller restart / crash"}
  C --> F{"Agent or network loss"}
  E --> G["Controller starts and reloads jobs / queue / runs"]
  F --> H["Agent reconnects or replacement agent is allocated"]
  G --> I{"Step resumable?"}
  H --> I
  I -->|yes| J["Resume / continue"]
  I -->|infra retry safe| K["Retry bounded scope"]
  I -->|unknown side effect| L["Manual or automated reconcile"]
  D --> L
  J --> M["Verify retained evidence"]
  K --> M
  L --> M
  

Jenkins owns orchestration state, but it does not own every filesystem, process or external system touched by a build. Resilience depends on correctly matching the recovery action to the layer that failed.

3. Keep these state layers separate

Layer Typical state What restart/loss means
Controller process JVM PID, uptime, HTTP listener Process disappears and returns; not itself a build result.
JENKINS_HOME Jobs, run records, queue persistence, plugin/config data Durability depends on consistent disk/state and compatible core/plugins.
Pipeline CPS Flow graph, suspended Groovy continuation, step metadata May resume after restart according to step and durability behavior.
Queue Queue item ID, why-blocked reason, label request May be reloaded; still needs eligible executor/agent.
Agent channel Remoting connection, node online/offline state Can disappear independently of controller.
Workspace Checkout, generated files, tool caches May be lost/replaced; do not treat as durable release storage.
Durable task Long-running shell/process metadata and result files Designed to survive controller restart while agent/workspace remain viable.
External target Package publish, deployment, ticket, notification Jenkins restart cannot undo or prove this automatically.
Recovery evidence Build/queue IDs, timestamps, logs, API IDs Must be preserved before retry/restart/cleanup decisions.

4. Quiet-down is a scheduling control, not a shutdown

Jenkins exposes Prepare for Shutdown (quiet-down). Current documentation describes it as preventing new builds from starting while allowing the controller to keep running. It is useful for draining work before planned maintenance. A safe restart is a separate operation. An abrupt container/process restart is different again because it may interrupt steps before they reach a resumable boundary.

Operation Primary intent Do not assume
Quiet-down / Prepare for Shutdown Stop new builds from starting while existing work drains That the controller will stop by itself.
Safe restart Restart after Jenkins has tried to reach a safe point That every external step is reversible or every plugin is resumable.
Force/abrupt restart Immediate process interruption That in-flight state has been flushed or external effects are known.
Reload configuration from disk Reload controller configuration into memory That it is equivalent to process restart or a backup/restore test.

5. Pipeline durability is a trade-off

Pipeline writes orchestration state so execution can survive controller restarts. Jenkins exposes speed/durability settings because writing more state to disk costs I/O. Current Jenkins documentation still describes MAX_SURVIVABILITY as the most durable and slowest mode, while performance-optimized behavior reduces disk I/O and can lose more state after an ungraceful shutdown.

The setting applies to a Pipeline run according to when that run starts; changing a durability setting is not a magic repair for an already-corrupted run. Critical release/deployment Pipelines usually justify stronger durability than disposable short-lived verification work.

6. Not every Pipeline step resumes the same way

Some steps are designed to survive controller restart. Current Pipeline Basic Steps documentation gives sh and input as examples of resumable steps, while short operations such as checkout or junit can be nonresumable if Jenkins restarts while they are executing. The retry step supports a nonresumable() condition specifically for this case.

retry(count: 2, conditions: [nonresumable()]) {
  checkout scm
}

This does not mean “retry every failure twice.” It means retry only when Jenkins identifies the failure as the documented nonresumable-after-restart class.

7. Agent loss is an infrastructure failure, not automatically a code failure

An online node can vanish while a build is running because of network failure, VM termination, container/pod deletion or host maintenance. Jenkins can distinguish some infrastructure failures and the Pipeline retry step supports an agent() condition that can allocate a fresh agent and rerun the enclosed block. The safety question is whether the block contains repeatable work.

retry(count: 2, conditions: [agent()]) {
  node('resilience-lab') {
    checkout scm
    sh './ci/run-tests.sh'
  }
}

Wrapping a deployment with this pattern without idempotency can duplicate external changes. Retry classification belongs to the side effects, not just the exception type.

8. Read-only inspection before touching recovery state

set -euo pipefail
printf 'job=%s\n' "${JOB_NAME:-unknown}"
printf 'build=%s\n' "${BUILD_NUMBER:-unknown}"
printf 'build_url=%s\n' "${BUILD_URL:-unknown}"
printf 'node=%s\n' "${NODE_NAME:-unknown}"
printf 'workspace=%s\n' "${WORKSPACE:-unknown}"
printf 'source_sha=%s\n' "$(git rev-parse HEAD 2>/dev/null || printf unavailable)"
printf 'timestamp_utc=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"

Also preserve the queue item ID/reason, controller version/uptime, agent offline cause, exact Pipeline stage/step and any external operation ID. Capture before restarting, reconnecting, replaying or deleting anything.

9. Four recovery classes

Class Question Example action
Resumable Can the same running step continue from persisted state? Allow durable task/Pipeline to resume and verify output.
Retryable Can the bounded operation safely run again? Retry checkout/test on a fresh agent.
Re-runnable Is the whole build free of unsafe duplicate effects? Trigger a new build with same immutable source and parameters.
Manual-reconcile Could an external effect have happened without Jenkins knowing? Query target by operation/artifact ID before any retry.

10. Common misconceptions

  • “Pipeline is durable, so every step is resumable.” False; resumability is step-specific.
  • “The agent came back, so its workspace is intact.” Not guaranteed, especially with ephemeral agents.
  • “Safe restart means no failed builds.” It reduces disruption but cannot make all plugins/external effects atomic.
  • “If Jenkins marked a deployment step failed, the target is unchanged.” False; the request may have committed before the connection failed.
  • “Retry is always resilience.” A blind retry can be the mechanism that causes duplicate production effects.
Next

Run a controlled recovery workflow

Lesson 2 turns the model into a disposable drill: quiet the controller, restart it, interrupt an agent, observe evidence and reconcile a synthetic side effect before retry.

Knowledge check

Answer before revealing the explanation.

1. What does quiet-down prove?

2. Why can a Pipeline resume while a workspace is lost?

3. When is retry with agent() appropriate?

4. Why is an external deployment different from a durable shell step?

5. What should happen before a restart during incident response?

Official references and version notes

Resilience behavior depends on Jenkins core, Pipeline plugins, agent launchers and individual steps. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.