Chapter 36Lesson 01~190 minutes

Backup, Restore, Disaster Recovery, Controller Migration, Configuration Recovery, and Recovery Testing: Concepts, Architecture, and Mental Model

Build the recovery mental model before touching backup tools: define the recovery objective, identify authoritative controller and external state, protect the decryption key separately, and prove restore viability in isolation.

backuprestoreJENKINS_HOMERPO/RTOcontroller keysDR

Learning objectives

  • Define recovery scope, RPO and RTO for Jenkins instead of treating “backup” as a single checkbox.
  • Separate JENKINS_HOME state, controller decryption key material, version/plugin inventory, configuration-as-code repositories and external systems.
  • Explain why a copied backup is not proven until it has been restored and validated.
  • Distinguish configuration reconstruction from stateful recovery.
  • Design an evidence chain from backup timestamp through isolated restore and eventual cutover.

1. The practical problem: a backup file is only a claim

A Jenkins controller is not merely a web process. Its durable operational history lives primarily under JENKINS_HOME: job configuration, build records, plugin state, credentials metadata, fingerprints and other controller data. Pipelines can also depend on configuration repositories, artifact repositories, secret providers, SCM systems, agent images and deployment targets that live outside that directory.

Disaster recovery succeeds only when the organization can reconstruct a compatible controller, decrypt the state it must decrypt, reconnect the right external dependencies, verify expected jobs/build evidence, and do so within stated recovery objectives. Until that has been rehearsed, “we have backups” means very little.

2. Mental model: recovery is an evidence chain

Mental model: recovery is an evidence chain
flowchart TD
  A["Recovery scope + RPO/RTO"] --> B["Consistent controller backup"]
  B --> C["Protected off-controller backup storage"]
  A --> D["Separately protected controller key"]
  B --> E["Core / Java / plugin inventory"]
  F["JCasC / Job DSL / Shared Library refs"] --> E
  G["External artifact / secret / SCM dependencies"] --> H["Dependency inventory"]
  C --> I["Isolated clean restore target"]
  D --> I
  E --> I
  H --> I
  I --> J["Restore validation"]
  J --> K{"Meets recovery objective?"}
  K -->|no| L["Fix backup / runbook / dependency gaps"]
  L --> B
  K -->|yes| M["Controlled migration or cutover"]
  M --> N["Post-cutover verification + recovery evidence"]
  

Notice that the backup archive is only one node. Recovery also needs the correct interpretation environment and external dependency state.

3. Recovery state layers

Layer Examples Recovery question
Controller files JENKINS_HOME XML, jobs, runs, plugins, user/config data Was the snapshot internally consistent and complete enough for the stated objective?
Controller cryptographic state master.key and data protected by Jenkins secret machinery Can encrypted credentials be recovered without storing key and backup together?
Runtime baseline Jenkins core/LTS, Java, startup parameters, OS/container Can the restored data be interpreted by a compatible runtime?
Plugin baseline Exact plugin versions/dependencies Will restored configuration and job state deserialize correctly?
Configuration-as-code JCasC, Job DSL, Shared Library refs Which settings/items are reconstructed from SCM versus restored state?
Build evidence Run numbers, logs, archived artifacts, fingerprints Which history must meet audit/release requirements?
External dependencies SCM, artifact repo, secret manager, IdP, webhooks, registries Which state is outside Jenkins and must be reconnected or independently restored?
Execution state Agents/workspaces/caches What should be recreated rather than backed up?
Recovery proof Checksums, restore test, timestamps, checklist What demonstrates the backup is usable and within RPO/RTO?

4. What JENKINS_HOME means during recovery

Current Jenkins documentation describes JENKINS_HOME as the root directory that stores controller configuration, jobs and archives. Backing up the entire directory is conceptually simple, but not every subdirectory has the same business value. Workspaces, caches and exploded runtime content can often be recreated; build history, job definitions, plugin packages and controller configuration may be essential depending on your recovery scope.

The correct backup set therefore follows the recovery objective, not a generic file list. If audit policy requires five years of archived release evidence, excluding builds/archive can be unacceptable even though a developer could rebuild source.

5. The controller key must be recoverable—but separated from the routine backup

This creates two independent recovery dependencies: the state backup and the separately protected key. Losing either can make credential recovery impossible; compromising both can expose the controller’s secrets. The lab uses only fake credentials and records metadata—not secret values.

6. RPO and RTO turn “backup” into an engineering requirement

Objective Question Example decision
RPO — recovery point objective How much controller state can we afford to lose? Hourly config backup may be enough for jobs but not for high-value build audit records generated every minute.
RTO — recovery time objective How long can Jenkins be unavailable/degraded? A 30-minute RTO requires pre-provisioned runtime/install media and rehearsed restore steps.
Recovery scope Which state must return? Controller config + jobs + credential usability + selected build records; workspaces rebuilt.
Dependency scope Which external services must also work? SCM and artifact repository may have separate DR and credentials/network prerequisites.

7. Inspect before you copy anything

Read-only inspection comes first. Record the controller identity, current core/Java version, JENKINS_HOME path, plugin inventory, job/build sample, configuration sources and external dependencies. A backup taken without this inventory may be impossible to interpret later.

set -euo pipefail
printf 'UTC: '; date -u +%Y-%m-%dT%H:%M:%SZ
printf 'Java: '; java -version 2>&1 | head -n 1
printf 'JENKINS_HOME: %s\n' "${JENKINS_HOME:-not-set-in-this-shell}"
# Run against the disposable controller filesystem only.
# Never print secret files or credential values.

8. Build a recovery inventory, not a secret dump

Controller: lab-jenkins-dr
Jenkins core: 2.568.3
Controller Java: 21
JENKINS_HOME: /var/jenkins_home
Backup policy: config + jobs + required build records
Controller key: stored separately (location recorded, content never printed)
JCasC source: synthetic SCM ref <commit>
Job DSL source: none / <commit>
External dependencies: local mock SCM + local artifact directory
Target RPO: <= 15 minutes for checkpoint lab
Target RTO: <= 30 minutes for checkpoint lab
Last restore validation: <timestamp / result>

9. JCasC is valuable recovery input, not a full backup

JCasC expresses supported controller/plugin configuration and works well as a reviewed reconstruction source. It does not inherently contain every job build record, archived artifact, plugin binary, credential decryption dependency or external system state. Treat “rebuild from code” and “restore state” as different recovery strategies that can be combined deliberately.

10. Validation is a restore, not a checksum alone

A checksum proves that bytes have not changed since you computed the checksum. It does not prove the backup contains the right files or can start Jenkins. Current Jenkins guidance explicitly recommends restoring a full backup into a temporary location and starting it on a different port. A credible validation then checks jobs, configuration, build records, credential metadata and selected workflows without touching production dependencies.

11. What different evidence actually proves

Evidence Supports Does not prove
Backup archive exists Some bytes were copied Completeness, consistency or restorability.
SHA-256 matches Backup bytes are unchanged from recorded digest Jenkins can interpret them.
Restored login/UI works Controller can start with some restored state Every job/plugin/credential/external integration is correct.
Fake credential metadata is present and usable Key + encrypted state worked for tested credential All production secrets/providers are available.
Sample Pipeline succeeds Selected execution path is functional All agents/plugins/external targets satisfy DR.
Measured restore meets RTO This rehearsal completed within target Future disasters will match the same conditions.
Next

Perform a disposable backup and isolated restore

Lesson 2 turns the model into a small, evidence-driven workflow: inventory, quiesce or snapshot, copy, checksum, separately protect the key, restore to an alternate JENKINS_HOME/port, and verify state without production side effects.

Knowledge check

Answer before revealing the explanation.

1. Why is a backup archive alone insufficient evidence?

2. Why should the controller key be stored separately from routine backups?

3. What is the difference between RPO and RTO?

4. Does JCasC replace JENKINS_HOME backup?

5. What is the strongest basic proof that a backup is usable?

Official references and version notes

Recovery procedures are version-sensitive. Re-check current primary documentation and your own controller/plugin inventory before using these patterns on a real system.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.