Chapter 35Lesson 03~185 minutes

Controller Resilience, Queue Recovery, Agent Loss, Pipeline Durability, Maintenance Windows, and Failure Modes: Configuration, Design Choices, and Tradeoffs

Choose restart, retry, workspace and durability policies from failure semantics—not from convenience—and document the evidence and rollback constraints for each choice.

designtradeoffsdurabilitymaintenanceidempotencyRTO/RPO

Learning objectives

  • Compare restart, reload, retry and rerun without collapsing them into one recovery action.
  • Choose workspace and agent strategies that match trust and reproducibility needs.
  • Choose Pipeline durability with explicit performance/recovery tradeoffs.
  • Design maintenance windows around queue drain, external side effects and rollback.
  • Write a decision record using observable evidence and recovery objectives.

1. Restart versus reload

A controller process restart reinitializes the JVM and plugins around persisted controller state. Reload configuration from disk discards in-memory configuration and reloads files, but it is not a JVM restart, plugin upgrade, backup restore or queue-drain mechanism. Use the operation that matches the actual change.

Need Preferred direction Reason
Apply a core/JVM/container change Planned controller restart/cutover The process/runtime itself changed.
Apply plugin changes that require restart Planned restart after compatibility validation Plugin classloading/runtime state changed.
Recover intentionally edited on-disk config in disposable lab Reload may be relevant Reload targets configuration, not all runtime faults.
Fix an unknown stuck build Diagnose first Restart can destroy first-failure evidence without fixing cause.

2. Retry same stage/block versus rerun the whole build

Retry keeps the same build identity and repeats a bounded block. A full rerun creates a new build identity and may repeat every side effect. Declarative Restart from Stage is yet another mechanism: it creates a new execution derived from a completed Pipeline and can preserve selected stashes when configured. None of these actions is automatically safe for deployment.

Choice Good fit Risk
retry(... agent()) Agent/network infrastructure failure around repeatable checkout/test Repeats everything inside block.
retry(... nonresumable()) Controller restart interrupted a short nonresumable step Step must itself be safe to repeat.
Restart from Stage Known transient failure in a completed Declarative Pipeline Downstream side effects can be repeated; preserved inputs must be verified.
New build Need complete fresh execution/evidence New source/ref resolution or side effects may differ unless pinned.

3. Workspace reuse versus clean replacement

Workspace reuse can improve speed through caches, but it also preserves untracked state and increases coupling to a specific agent. Clean disposable workspaces improve isolation and reproducibility but cost checkout/tool/cache time. Neither approach makes the workspace a durable artifact store.

Design Advantages Controls needed
Reusable static workspace Warm caches, lower setup latency Cleanliness checks, cache integrity, trust separation, disk lifecycle.
Ephemeral workspace/agent Isolation, easier horizontal scaling Immutable source/artifact inputs, external cache strategy, evidence export before teardown.
Persistent shared workspace Rarely justified for untrusted concurrent jobs Strong locking/isolation; usually prefer better artifact/cache design.

4. Quiet-down/drain versus abrupt stop

Planned maintenance should normally stop admitting new work and allow known critical work to reach safe boundaries. Abrupt stop is reserved for emergencies where continuing is more dangerous than interruption. The tradeoff is between maintenance latency and recovery uncertainty.

A maintenance runbook should name: controller, start/end window, change owner, queue policy, jobs that must drain, jobs safe to abort, external operations in progress, backup/snapshot status, health checks, rollback trigger and communication path.

5. Durability versus performance

Jenkins Pipeline durability settings trade disk persistence for I/O performance. The correct answer is workload-specific. A critical multi-hour release orchestration usually values recovery more than a ten-minute disposable branch test. Benchmark rather than assuming a faster durability setting will solve controller performance problems.

Pipeline class Typical bias Why
Critical release/deployment Higher durability Long duration and expensive uncertain side effects.
Long integration suite Higher/normal durability Recomputing hours of work may be costly.
Short disposable validation Potentially performance-oriented after measurement Failure can often be rerun cheaply if no side effects.
Controller already I/O-bound Measure first Durability may contribute, but logs/artifacts/plugins/workspaces can dominate I/O.

6. Connect resilience settings to RPO/RTO

RPO asks how much controller/build state you can afford to lose. RTO asks how long the service can be unavailable or degraded. Pipeline durability helps with in-flight run recovery, but it does not replace controller backups, artifact repositories or external system reconciliation. A low RPO for release evidence needs all of those layers.

7. Trusted static versus ephemeral execution

Resilience and security interact. A static signing/deployment agent may retain keys/caches but becomes a high-value persistent trust boundary. Ephemeral agents reduce persistence but require stronger external identity, reproducible toolchains and evidence export. Do not sacrifice isolation just to make workspace recovery easier.

8. Worked decision: release Pipeline before monthly maintenance

Question Decision Evidence required
Build still queued? Enter quiet-down and let current release drain; do not start another. Queue IDs/reasons; release owner confirmation.
Deployment API call in flight? Do not restart until operation ID can be reconciled or reaches known terminal state. Deployment/release ID + target API state.
Long test shell running? Safe restart may allow durable step continuation if execution infrastructure remains. Same build/step ID and agent/task evidence after restart.
Agent being replaced? Assume workspace can disappear; restore from SCM/artifacts, not local assumptions. Source SHA + artifact digest + clean workspace.
Core/plugin update planned? Snapshot/backup and compatibility test first; define rollback limits. Version inventory, backup/restore test, upgrade notes.

9. Small decision record template

Decision: Jenkins maintenance/recovery policy for <job/folder>
Date / owner:
Controller + core/Java/plugin baseline:
Workload class:
Durability setting and rationale:
Agent/workspace strategy:
External side effects and idempotency key:
Allowed retry conditions:
Quiet-down/drain rule:
RPO / RTO expectation:
Evidence required before retry/restart:
Rollback / manual reconciliation trigger:
Review date:

10. Cost and operational tradeoffs

Higher durability increases controller disk I/O. Ephemeral agents increase provisioning and cache misses. Static agents increase patching/trust cost. Long maintenance drains can delay releases. More standby agents reduce queue time but increase infrastructure cost. Make each tradeoff measurable: queue wait, build critical path, controller I/O/heap, recovery time, failure recurrence and side-effect reconciliation time.

Next

Diagnose the ugly cases

Lesson 4 breaks the system deliberately: uncertain deployment acknowledgment, missing workspace, nonresumable restart failure, blind restart and maintenance without drain.

Knowledge check

Answer before revealing the explanation.

1. When is full rerun riskier than retrying a bounded block?

2. Why can high durability be inappropriate as a universal default?

3. Does an ephemeral agent improve security automatically?

4. What is the difference between RPO and Pipeline durability?

5. Why is reload configuration not a substitute for restart?

Official references and version notes

Resilience behavior depends on Jenkins core, Pipeline plugins, agent launchers and individual steps. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.