Chapter 36Lesson 03~195 minutes

Backup, Restore, Disaster Recovery, Controller Migration, Configuration Recovery, and Recovery Testing: Configuration, Design Choices, and Tradeoffs

Select backup and restore mechanisms from recovery semantics: consistency, scope, key separation, compatibility, recovery objectives and testability matter more than tool convenience.

designsnapshotsthinBackupselective backupmigrationtradeoffs

Learning objectives

  • Compare filesystem snapshots, controlled file copies and maintained backup plugins.
  • Decide when full versus selective backup is justified.
  • Explain hot versus quiesced consistency risks.
  • Choose between same-host restore, migration and rebuild-from-code.
  • Document the decision with RPO/RTO, security and rollback constraints.

1. Filesystem snapshot versus file copy

Jenkins documentation favors filesystem snapshots for consistency where available. Snapshot support from LVM, ZFS, btrfs, cloud disks or storage arrays can capture a point-in-time view quickly and reduce mixed-time copies. A plain recursive copy is portable, but if Jenkins is mutating files while the copy runs, the resulting set can represent multiple moments.

Approach Strength Main risk/control
Filesystem/storage snapshot Fast, point-in-time consistency, often low pause Snapshot technology and retention are external dependencies; replicate off the same failure domain.
Quiesced file copy Simple, transparent, easy to inspect Requires planned pause/drain; copy duration can increase RTO/RPO pressure.
Live file copy Low operational interruption Mixed-time/inconsistent state risk; validate heavily or avoid for critical recovery.
Backup plugin Scheduling/UI convenience Scope may be selective; plugin health/security/version becomes part of recovery baseline.

2. Where thinBackup fits

Current Jenkins backup guidance says thinBackup is the actively maintained open-source backup plugin. The plugin intentionally focuses on vital global/job configuration and does not simply clone every archive/workspace. Its current 2.1.5 release requires Jenkins 2.541.3 and fixes the prior access-control advisory affecting 2.1.4 and earlier.

That makes it an option for some configuration-recovery designs, not proof that your entire disaster-recovery scope is covered. If archived build evidence or specific controller-secret dependencies are required, document them separately.

3. Full versus selective backup

Choice Use when Tradeoff
Full JENKINS_HOME You need broad controller reconstruction and can afford storage/time Simpler restore model, larger backup and potentially more sensitive data.
Selective config/jobs You deliberately reconstruct caches/tools/history elsewhere Smaller/faster, but omissions become recovery dependencies.
Tiered schedules Config frequently; large build archives less frequently or externally Matches different RPOs but complicates runbook and dependency matrix.
Rebuild controller from code + restore state subset Controller baseline is reproducibly provisioned Requires version-pinned IaC/JCasC/plugin catalog and proven state import.

4. Hot versus quiesced backup

A hot backup minimizes downtime but must account for files changing during the capture. A quiesced/snapshot-consistent backup creates a clearer point in time. For a production system, compare maintenance cost against the business cost of an inconsistent restore. “The copy command exited 0” is not a consistency proof.

5. Same-host restore versus migration

Restoring on the original host can be fast after accidental deletion, but it may share the failure domain that caused the incident. A clean alternate host/container validates that hidden local assumptions have been removed. Migration also introduces networking, DNS/reverse proxy, filesystem ownership, agent connectivity and external allow-list changes. Treat those as explicit cutover dependencies.

6. Rebuild-from-code versus stateful recovery

State Good reconstruction source Why backup may still matter
Controller global config JCasC where supported Not every state/history/plugin binary is JCasC-owned.
Jobs/items Job DSL / SCM Jenkinsfiles where designed Build history and manually-owned items may remain.
Plugins Pinned plugin catalog/image/package baseline Exact dependency graph must remain reproducible.
Credentials External secret provider + minimal Jenkins metadata where possible Controller-encrypted credentials may still require key/state recovery.
Build artifacts External artifact repository Jenkins archive may still contain audit evidence not promoted externally.
Run history JENKINS_HOME backup Usually not reconstructed from source code.

7. Key separation changes the recovery design

Separating the controller decryption key from routine backups improves security but creates an operational dependency. The runbook must name who can retrieve it, how access is audited, how it is transported to the isolated restore target, how it is removed after use, and how recovery is tested without exposing it in logs or tickets.

8. Different data can have different RPOs

Data class Example RPO Design implication
Controller configuration 15 minutes Frequent snapshot/config capture.
Build audit records 5 minutes for release jobs Higher-frequency persistence or external evidence export.
Archived binaries Near zero loss Prefer external immutable artifact repository.
Workspace/cache No backup objective Recreate from SCM/artifacts/tool cache.
Controller key Changes rarely Separate protected copy with periodic access/retrieval test.

9. Restore compatibility: do not turn recovery into an accidental upgrade

The safest first restore normally matches the source Jenkins core, Java and plugin baseline closely enough to interpret stored data. If the disaster also requires an upgrade, stage that as a separate validated change. Read every relevant LTS upgrade guide, plugin minimum core requirement and migration note before cutover.

10. Worked decision: medium-sized controller with external artifacts

Question Decision Reason/evidence
Artifacts already immutable in Nexus/Artifactory? Do not duplicate all release binaries in every controller backup Repository has separate DR and digest evidence.
Need 90 days build history? Include required job build records Audit requirement is controller-state scope.
Storage supports snapshots? Use scheduled snapshots + off-host replication Consistency and short capture window.
JCasC owns global config? Retain pinned JCasC ref + backup state Reconstruction plus state recovery; avoid competing ownership.
RTO 45 minutes? Pre-stage compatible Jenkins runtime/install media and automated validation Downloading/upgrading during incident risks exceeding RTO.
Credentials local? Protect controller key separately and test synthetic credential restore Prevents “controller starts but credentials unusable” surprise.

11. Recovery decision record template

Decision: Jenkins backup and recovery architecture
Date / owner:
Controller(s) in scope:
Business services depending on Jenkins:
Recovery scope:
RPO by data class:
RTO:
Backup mechanism and consistency model:
Routine backup storage / failure domain:
Separate controller-key custody:
Core / Java / plugin baseline source:
JCasC / Job DSL / Shared Library refs:
External dependency DR assumptions:
Restore validation frequency:
Cutover / rollback criteria:
Known exclusions and accepted risk:
Next

Diagnose restore failures from evidence

Lesson 4 intentionally breaks backups and restores: missing key, incompatible plugin/core baseline, inconsistent live copy, same-disk backup loss and false confidence in JCasC.

Knowledge check

Answer before revealing the explanation.

1. Why can a filesystem snapshot be preferable to a long live file copy?

2. Does thinBackup necessarily satisfy full disaster recovery?

3. Why might a clean alternate host be a stronger recovery test than same-host restore?

4. What is the risk of upgrading core/plugins during the first restore attempt?

5. Why can different Jenkins data classes have different RPOs?

Official references and version notes

Recovery procedures are version-sensitive. Re-check current primary documentation and your own controller/plugin inventory before using these patterns on a real system.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.