Chapter 25Lesson 05260–360 min

Checkpoint Lab — Backup, Restore, Configuration Export, Blob Recovery, Disaster Scenarios, and Recovery Validation

Run a complete disaster drill on disposable data: declare objectives, capture a consistent recovery set, verify it, inject a controlled failure, restore to isolation, prove critical bytes and configuration, measure elapsed recovery work, and record gaps that would block production readiness.

Checkpoint labDisaster drillRPO/RTORestore validationRunbook

Learning objectives

  • Define a concrete recovery scenario, RPO/RTO, critical artifacts, and acceptance criteria before the drill starts.
  • Create and verify a complete synthetic recovery set containing metadata, blobs, configuration, node identity, and an external manifest.
  • Inject controlled data loss and restore into isolation without overwriting source evidence.
  • Produce machine-verifiable checksum/configuration results and a human-readable recovery report.
  • Identify gaps between a successful fixture restore and a production Nexus recovery architecture, then bridge to Chapter 26 migration planning.
Dated baseline (27 August 2026). These lessons use Nexus Repository 3.95.2-01 with Java 21 as the reference line. Recovery procedures are version/database/storage sensitive: before a real restore, verify the exact release notes, database support, storage backend, and restore documentation for the version that created the backup.
Recovery-set invariant. A Nexus backup is not “the database” or “the blobs.” Sonatype requires database/configuration state and blob content to be protected together. Preserve the node identity under $data-dir/keystores/node/ as part of the recovery set as well.
Destructive-lab boundary. Every loss/restore exercise in this chapter targets synthetic data in an isolated directory or disposable Nexus instance. Never delete, overwrite, or restore an employer database, blob store, object-store bucket, production secret, or production data directory for a training exercise.

1. Scenario and objectives

Your team operates a small internal artifact service. At 10:20, an operator discovers that a destructive maintenance action removed the current release artifact. The approved recovery objective is:

  • RPO: 30 minutes — recovery point must be 09:50 or newer.
  • RTO: 120 minutes — validated service must be available by 12:20.
  • Critical content: learner-example/app/2.0.0/app-2.0.0.bin.
  • Required configuration: repository ch25-release-hosted maps to blob store ch25-file; node identity and synthetic role metadata must survive.

The mandatory lab is a local fixture. An optional disposable Nexus H2 implementation may repeat the drill using the current documented backup/restore procedure.

2. Write predictions before action

  1. If only the live blob is deleted, metadata validation must detect a missing binary even though the database record still exists.
  2. Restoring from a verified 10:00 recovery set should meet the 30-minute RPO for a 10:20 incident and should restore the same SHA-256.
  3. Copying only the database directory into the recovery target must fail acceptance because the critical blob will be absent.

Save these to ch25-drill/evidence/01-predictions.md.

3. Create the source instance fixture

from pathlib import Path
import hashlib, json, shutil
root=Path('ch25-drill/source')
if root.parent.exists(): shutil.rmtree(root.parent)
(root/'db').mkdir(parents=True); (root/'blobs/learner-example/app/2.0.0').mkdir(parents=True)
(root/'keystores/node').mkdir(parents=True); (root/'etc').mkdir()
artifact=root/'blobs/learner-example/app/2.0.0/app-2.0.0.bin'
artifact.write_bytes(b'chapter25-release-2.0.0
')
sha=hashlib.sha256(artifact.read_bytes()).hexdigest()
(root/'db/state.json').write_text(json.dumps({
  'repositories':[{'name':'ch25-release-hosted','blobStore':'ch25-file'}],
  'roles':[{'id':'ch25-reader','privileges':['read:ch25-release-hosted']}],
  'assets':[{'path':'learner-example/app/2.0.0/app-2.0.0.bin','sha256':sha}]
}, indent=2)+'
')
(root/'keystores/node/id').write_text('NODE-CH25-DRILL
')
(root/'etc/runtime.properties').write_text('version=3.95.2-01
java=21
')
print('critical sha256:', sha)

4. Create the 10:00 recovery set and immutable-style manifest

from pathlib import Path
import hashlib, json, shutil
src=Path('ch25-drill/source'); b=Path('ch25-drill/backup-1000')
shutil.copytree(src, b/'payload')
files=[]
for p in sorted((b/'payload').rglob('*')):
    if p.is_file(): files.append({'path':p.relative_to(b/'payload').as_posix(),'sha256':hashlib.sha256(p.read_bytes()).hexdigest()})
manifest={'recoveryPoint':'2026-08-27T10:00:00Z','rpoMinutes':30,'rtoMinutes':120,
          'nexusVersion':'3.95.2-01','database':'H2-model','blobStores':['ch25-file'],'files':files}
(b/'manifest.json').write_text(json.dumps(manifest,indent=2)+'
')
print('manifest entries:',len(files))

In production, backup media and manifests should be protected by access controls, encryption, and preferably deletion-resistant retention appropriate to your threat model.

5. Gate: verify the recovery set before declaring it usable

from pathlib import Path
import hashlib, json
b=Path('ch25-drill/backup-1000'); m=json.loads((b/'manifest.json').read_text())
issues=[]
for item in m['files']:
    p=b/'payload'/item['path']
    if not p.exists(): issues.append([item['path'],'missing'])
    elif hashlib.sha256(p.read_bytes()).hexdigest()!=item['sha256']: issues.append([item['path'],'hash'])
assert not issues, issues
print('recovery-set gate: PASS')

6. Inject the controlled incident

from pathlib import Path
victim=Path('ch25-drill/source/blobs/learner-example/app/2.0.0/app-2.0.0.bin')
assert victim.exists(); victim.unlink()
print('INCIDENT: deleted only disposable critical blob')

Now run the source validator from Lesson 2 logic. It should report a missing binary. Capture that output as 02-incident-validation.txt.

7. Demonstrate the wrong recovery: database only

from pathlib import Path
import shutil
bad=Path('ch25-drill/bad-target')
if bad.exists(): shutil.rmtree(bad)
(bad/'db').mkdir(parents=True)
for p in Path('ch25-drill/backup-1000/payload/db').iterdir(): shutil.copy2(p,bad/'db'/p.name)
print('bad target intentionally contains DB only')

Do not accept this target. The repository definition may be represented, but the critical artifact cannot be served. Record why the attempt fails the acceptance criteria, then leave it intact as evidence.

8. Perform the correct isolated restore

from pathlib import Path
import shutil
src=Path('ch25-drill/backup-1000/payload'); target=Path('ch25-drill/restored-target')
if target.exists(): shutil.rmtree(target)
shutil.copytree(src,target)
print('full recovery set restored')

9. Execute the acceptance validator

from pathlib import Path
import hashlib, json, time
start=time.monotonic(); target=Path('ch25-drill/restored-target')
state=json.loads((target/'db/state.json').read_text())
assert state['repositories']==[{'name':'ch25-release-hosted','blobStore':'ch25-file'}]
assert state['roles'][0]['id']=='ch25-reader'
assert (target/'keystores/node/id').read_text().strip()=='NODE-CH25-DRILL'
for a in state['assets']:
    p=target/'blobs'/a['path']; assert p.exists(), a['path']
    assert hashlib.sha256(p.read_bytes()).hexdigest()==a['sha256'], a['path']
print('ACCEPTANCE: PASS')
print('validationSeconds:', round(time.monotonic()-start,3))

For a live Nexus drill, add controlled client requests using a least-privilege reader, repository browse/search evidence, task/log inspection, and the pre-incident artifact checksum.

10. Evaluate RPO explicitly

Incident detection is 10:20; recovery point is 10:00. The potential loss window is 20 minutes, which is within the 30-minute RPO. This does not mean every write between 10:00 and 10:20 is acceptable to lose operationally; it means the declared objective allows it. Record any business-critical publication that occurred after the recovery point as a gap requiring replay/republication from trusted build evidence.

11. Evaluate RTO with measured timestamps

For the fixture, elapsed time is tiny and therefore not representative. In production, record incident declaration, backup selection, infrastructure provisioning, DB restore, blob restore, Nexus startup, reconciliation, validation, DNS/traffic cutover, and service acceptance timestamps. The sum—not the copy command duration—is your demonstrated recovery time.

12. Optional live Nexus drill

On a disposable H2 archive installation only:

  1. Publish a harmless raw asset and record its SHA-256.
  2. Capture repository/blob configuration, node-ID backup, and the exact 3.95.2-01/Java 21 baseline.
  3. Create/run Admin - Backup H2 Database; back up the associated file blob store and required configuration/node ID as one declared recovery set. Prefer a maintenance-window/offline embedded DB backup as the stronger periodic method per Sonatype guidance.
  4. Stop Nexus and preserve the original disposable data directory.
  5. Restore into a separate disposable target using the documented H2 restore procedure and corresponding blob backup.
  6. Start the restored target, confirm repository/security state, download the known asset, and compare SHA-256.
  7. If you intentionally created bounded DB/blob skew, generate/review the current repair plan before executing anything.

13. Required evidence bundle

  • 01-predictions.md
  • 02-source-manifest.json
  • 03-backup-verification.txt
  • 04-incident-validation.txt
  • 05-bad-restore-result.md
  • 06-good-restore-log.txt
  • 07-acceptance-validation.txt
  • 08-rpo-rto.md
  • 09-security-and-secrets-review.md
  • 10-runbook-gaps.md

14. Record what this lab does not prove

The fixture proves reasoning and validation mechanics, not production-scale throughput, PostgreSQL PITR, object-store recovery, KMS/secret restoration, HA behavior, reverse-proxy/DNS cutover, or multi-terabyte restore times. A mature runbook turns each of those unknowns into a tested dependency or an explicit risk.

15. Cleanup / rollback

Keep the recovery report, then delete only the disposable ch25-drill tree. For the optional live drill, decommission only the isolated target after saving sanitized evidence. Never destroy the only known-good backup to “clean up” the exercise. Production backup retention follows policy, not lab convenience.

16. Knowledge check

Did the database-only restore satisfy the checkpoint?

The incident was at 10:20 and the recovery point at 10:00. Does that meet a 30-minute RPO?

Why keep the intentionally bad restore?

What should happen if the recovered artifact downloads but has a different SHA-256?

What does Chapter 26 add after this recovery discipline?

17. Chapter summary and bridge to Chapter 26

You now have a production-oriented recovery model: define RPO/RTO, capture database/configuration and blob content as one recovery set, protect node identity and custom configuration, verify backups before loss, restore into isolation, prove critical bytes and authorization/client paths, and record gaps. Chapter 26 builds on that discipline to handle database and instance migration without confusing migration with backup or rollback.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.