Chapter 36Lesson 05~260 minutes

Checkpoint Lab — Backup, Restore, Disaster Recovery, Controller Migration, Configuration Recovery, and Recovery Testing

Prove disaster recovery by destroying a disposable source controller after a verified backup, restoring it into a clean alternate location, validating required operational state, measuring RPO/RTO, and documenting what remains external.

checkpointclean-room restoredisaster recoverymigrationRPO/RTOevidence packet

Learning objectives

  • Predict which controller, key, runtime, plugin and external states must survive independently.
  • Verify the backup before destructive action.
  • Destroy only the disposable source after guard checks.
  • Restore to a clean alternate location/port using a compatible runtime baseline.
  • Prove jobs/configuration/build evidence and synthetic credential usability.
  • Measure achieved RPO/RTO and document residual dependency gaps.

1. Safety rules

  • The source path must equal the exact lab path before deletion.
  • The backup and checksum must be verified before deletion.
  • The separate key dependency must be present in the synthetic key vault before deletion.
  • The Jenkins core/Java/plugin baseline must be recorded.
  • The restore target must use a different path and HTTP port.
  • No cleanup occurs until the final evidence packet is copied outside the lab paths.

2. Required predictions before destruction

  1. The restored controller will recover the synthetic job full name and required build record only if those files were inside the validated backup scope.
  2. The fake Jenkins-managed credential will be unusable if the separately protected controller key dependency is not restored correctly.
  3. The workspace does not need to survive; it should be recreated from the synthetic source.
  4. The local “artifact repository” directory is an external dependency and will not magically return from a controller-only backup.
  5. Changing Jenkins core/plugins during restore may invalidate the checkpoint because it mixes disaster recovery with upgrade.

Record predictions before the destructive step.

3. Create the disposable recovery target state

Your source controller should contain: job dr-lab/pipeline, at least two builds, one archived text artifact, and one fake credential used only against a local mock service. Record exact source SHA/Jenkinsfile ref if the job comes from SCM.

Required identities
-------------------
controller=jenkins-dr-source
job=dr-lab/pipeline
expected_build=<number>
source_sha=<sha>
fake_credential_id=dr-lab-token
external_artifact=/tmp/jenkins-dr-external/artifacts/<digest>.txt
restore_home=/tmp/jenkins-dr-restore
restore_port=9999

4. Preflight and version inventory

set -euo pipefail
SOURCE_HOME='/tmp/jenkins-dr-source'
BACKUP_ROOT='/tmp/jenkins-dr-backups'
KEY_ROOT='/tmp/jenkins-dr-key-vault'
RESTORE_HOME='/tmp/jenkins-dr-restore'
EVIDENCE='/tmp/jenkins-dr-evidence'
mkdir -p "$EVIDENCE"
[ "$SOURCE_HOME" = '/tmp/jenkins-dr-source' ] || exit 70
[ -d "$SOURCE_HOME" ] || exit 66
{
  printf 'preflight_utc=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
  printf 'source_home=%s\n' "$SOURCE_HOME"
  printf 'restore_home=%s\n' "$RESTORE_HOME"
  printf 'restore_port=9999\n'
  printf 'expected_core=2.568.3\n'
  printf 'expected_java=21\n'
} | tee "$EVIDENCE/preflight.txt"

Add an exported plugin inventory from the disposable controller through an authorized read-only method. Do not use Script Console solely for this checkpoint.

5. Take and verify the backup

set -euo pipefail
SOURCE_HOME='/tmp/jenkins-dr-source'
BACKUP_ROOT='/tmp/jenkins-dr-backups'
BACKUP_ID="jenkins-dr-checkpoint-$(date -u +%Y%m%dT%H%M%SZ)"
BACKUP_DIR="$BACKUP_ROOT/$BACKUP_ID"
mkdir -p "$BACKUP_DIR/jenkins_home"

# Precondition: disposable source is stopped or snapshot-consistent.
cp -a "$SOURCE_HOME/." "$BACKUP_DIR/jenkins_home/"
(
  cd "$BACKUP_DIR"
  find jenkins_home -type f -print0 | sort -z | xargs -0 sha256sum > SHA256SUMS
  sha256sum -c SHA256SUMS
)
printf '%s\n' "$BACKUP_ID" > /tmp/jenkins-dr-evidence/backup-id.txt

6. Verify the separate lab key dependency

set -euo pipefail
KEY_ROOT='/tmp/jenkins-dr-key-vault'
[ -d "$KEY_ROOT" ] || { echo 'key vault missing' >&2; exit 66; }
[ -f "$KEY_ROOT/master.key" ] || { echo 'lab controller key dependency missing' >&2; exit 67; }
# Never print the key. Record only presence, mode/ownership metadata as appropriate.
printf 'key_dependency_present=yes\n' >> /tmp/jenkins-dr-evidence/preflight.txt

7. Stop and ask: are we allowed to destroy the source?

Before deletion, verify all of these are true: backup checksum passes, backup ID is explicit, key dependency exists separately, version/plugin inventory is captured, external dependency identities are recorded, and the path guard is exact. If any check fails, the checkpoint stops.

8. Destroy only the disposable source

set -euo pipefail
SOURCE_HOME='/tmp/jenkins-dr-source'
[ "$SOURCE_HOME" = '/tmp/jenkins-dr-source' ] || exit 70
[ -d /tmp/jenkins-dr-backups ] || exit 71
[ -f /tmp/jenkins-dr-key-vault/master.key ] || exit 72
rm -rf -- "$SOURCE_HOME"
[ ! -e "$SOURCE_HOME" ] || exit 73
printf 'source_destroyed_utc=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" | tee -a /tmp/jenkins-dr-evidence/timeline.txt

9. Restore into a clean alternate location

set -euo pipefail
BACKUP_DIR="$(find /tmp/jenkins-dr-backups -mindepth 1 -maxdepth 1 -type d | sort | tail -n 1)"
RESTORE_HOME='/tmp/jenkins-dr-restore'
KEY_ROOT='/tmp/jenkins-dr-key-vault'
[ -n "$BACKUP_DIR" ] || exit 66
[ "$RESTORE_HOME" = '/tmp/jenkins-dr-restore' ] || exit 70
rm -rf -- "$RESTORE_HOME"
mkdir -p "$RESTORE_HOME"
cp -a "$BACKUP_DIR/jenkins_home/." "$RESTORE_HOME/"
install -m 0600 "$KEY_ROOT/master.key" "$RESTORE_HOME/secrets/master.key"
printf 'restore_copy_complete_utc=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" | tee -a /tmp/jenkins-dr-evidence/timeline.txt

10. Start the compatible Jenkins runtime

export JENKINS_HOME='/tmp/jenkins-dr-restore'
java -jar ./jenkins-2.568.3.war --httpPort=9999

Do not bind the restored test to the production hostname or enable real outgoing webhooks. Keep it isolated until validation is complete.

11. Validate independently

Validation Pass criterion
Core/runtime 2.568.3 on Java 21 for this dated checkpoint; no unsupported runtime warning.
Plugins Expected required versions load; no unexplained dependency failures.
Job dr-lab/pipeline exists at exact full name.
Build history Expected build number/result exists if history was in recovery scope.
Archived evidence Expected synthetic artifact/fingerprint record is readable where included.
Fake credential Synthetic authentication works without printing the secret.
Workspace May be absent; new build recreates from source rather than relying on backup.
External artifact Checked separately under /tmp/jenkins-dr-external; absence is recorded as external dependency failure, not controller corruption.
No production side effects No real SCM webhook/chat/deploy/secret provider receives traffic.
RPO/RTO Newest recovered required record and end-to-end validation timestamps meet targets.

12. Measure RPO and RTO

Use recorded UTC timestamps, not memory. RPO is the time between the latest required state successfully recovered and the incident/destruction point. RTO is measured according to your stated service definition—for this lab, from destruction start until the restored controller passes all required validation checks.

Checkpoint metrics
------------------
source_destroyed_utc=
latest_recovered_required_state_utc=
achieved_rpo=
restore_start_utc=
validation_complete_utc=
achieved_rto=
rpo_target=15m
rto_target=30m
pass_or_fail=

13. Required evidence packet

Group Minimum contents
Backup Backup ID, capture timestamps, SHA-256 verification result, storage/failure-domain note.
Controller key Presence/retrieval event reference only; never key contents.
Runtime Core, Java, startup mode/path/port.
Plugins Source and restored required plugin inventory/differences.
Configuration JCasC/Job DSL/Shared Library refs where applicable.
Jobs/builds Exact full name, sample build number/result/URL or local equivalent.
Credential test Fake credential ID and pass/fail without value.
External dependencies SCM/artifact/secret/IdP identities and separate recovery status.
Timeline Backup, destruction, restore, validation timestamps.
Objectives RPO/RTO targets versus achieved.
Limitations What the local lab does not prove about storage snapshots, DNS, cloud IAM, HA, production scale.

14. Cleanup after review

First copy the evidence packet to a path outside the disposable resources. Then remove only exact lab directories. Keep no synthetic key material unnecessarily.

set -euo pipefail
for p in /tmp/jenkins-dr-restore /tmp/jenkins-dr-backups /tmp/jenkins-dr-key-vault /tmp/jenkins-dr-external; do
  case "$p" in
    /tmp/jenkins-dr-restore|/tmp/jenkins-dr-backups|/tmp/jenkins-dr-key-vault|/tmp/jenkins-dr-external) rm -rf -- "$p" ;;
    *) echo "refusing unexpected cleanup path" >&2; exit 70 ;;
  esac
done

15. What this checkpoint proves—and what it does not

Supported claim Not proven by this lab alone
A bounded JENKINS_HOME backup can be restored into an isolated target. That every production plugin/integration is recoverable.
The separate synthetic key dependency can be re-applied without logging its contents. That production key custody/rotation process is adequate.
Required sample job/build evidence can be validated after destruction. That every historical artifact meets retention/audit requirements.
Measured local RPO/RTO can be calculated from timestamps. That production-scale network/storage recovery meets the same objectives.
External dependencies can be classified separately from controller state. That SCM, IdP, artifact repository or secret manager DR has been tested.

16. What Chapter 36 adds to the production Jenkins operating model

Chapter 35 taught you to recover through ordinary controller restart and agent loss. Chapter 36 proves something stronger: when controller state is lost or moved, you can rebuild a compatible Jenkins service from protected state, independently protected key material, versioned configuration sources and explicit external dependencies—and you can demonstrate that claim with a clean-room restore.

Chapter 37 adds observability. Recovery is much faster when metrics, logs, health checks, audit trails and queue telemetry make the pre-failure and post-restore state measurable instead of anecdotal.

Next chapter

Chapter 37 — Monitoring Jenkins with Metrics, Logs, Health Checks, Audit Trails, Queue Telemetry, and Observability

Carry the verified evidence and operating discipline from this chapter into the next chapter.

Knowledge check

Answer before revealing the explanation.

1. Why must the backup be verified before deleting the disposable source?

2. What should happen if the controller key dependency is missing before destruction?

3. Why can a missing workspace still be an acceptable recovery result?

4. How is achieved RTO measured in this checkpoint?

5. What is the main bridge to Chapter 37?

Official references and version notes

Recovery procedures are version-sensitive. Re-check current primary documentation and your own controller/plugin inventory before using these patterns on a real system.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.