Backup, Restore, Configuration Export, Blob Recovery, Disaster Scenarios, and Recovery Validation: Guided Hands-On Workflow and Core Operations
Run a recovery workflow where every state transition is observable. The mandatory path uses a deterministic local fixture so every learner can practice consistency points, loss injection, restore sequencing, and checksum validation for free; an optional disposable H2 Nexus extension maps the same reasoning to current Sonatype procedures.
Learning objectives
- Create a synthetic recovery set with database-like metadata, blob bytes, configuration, node identity, and a signed-by-hash manifest.
- Inject controlled loss only into disposable working state, restore to a separate target, and verify byte/metadata consistency.
- Map the fixture workflow to the current Admin - Backup H2 Database and H2 restore procedure without treating the task as a complete backup.
- Capture before/after evidence, RPO/RTO assumptions, exact backup timestamps, and restore validation.
- Recognize when reconciliation may be required and why a repair plan must follow current documented tasks rather than old task names.
$data-dir/keystores/node/ as part of the
recovery set as well.
1. Preflight and disposable scope
Mandatory path: Python 3, a temporary working directory, no Nexus server required. Optional live extension: a throwaway single-node Nexus Repository 3.95.2-01 archive install using H2 and a file blob store on loopback/private networking. Do not use a container-based H2 deployment; current Sonatype system requirements do not support that production/deployment combination.
Disposable names: ch25-recovery-source,
ch25-recovery-backup,
ch25-recovery-target, repository
ch25-raw-hosted, and blob store ch25-file.
2. Build the synthetic source state
The fixture deliberately separates “database metadata” from blob bytes and custom state:
from pathlib import Path
import hashlib, json, shutil, time
root = Path('ch25-recovery-source')
if root.exists(): shutil.rmtree(root)
(root/'db').mkdir(parents=True)
(root/'blobs').mkdir()
(root/'keystores/node').mkdir(parents=True)
(root/'etc').mkdir()
assets = {
'releases/learner-example-1.0.txt': b'release-1.0
',
'releases/learner-example-1.1.txt': b'release-1.1
'
}
rows=[]
for name, data in assets.items():
p=root/'blobs'/name; p.parent.mkdir(parents=True, exist_ok=True); p.write_bytes(data)
rows.append({'path':name,'sha256':hashlib.sha256(data).hexdigest(),'size':len(data)})
(root/'db/assets.json').write_text(json.dumps(rows, indent=2)+'
')
(root/'db/repos.json').write_text(json.dumps([{'name':'ch25-raw-hosted','blobStore':'ch25-file'}], indent=2)+'
')
(root/'keystores/node/node-id').write_text('node-CH25-SYNTHETIC
')
(root/'etc/nexus.properties').write_text('# synthetic fixture; no secrets
')
print('source ready:', len(rows), 'assets')
This is not a reverse-engineered Nexus data directory. It is an intentionally simple teaching fixture that models the relationship between metadata and bytes.
3. Create a recovery manifest and backup set
Freeze the fixture by making no further writes, then copy it to a separate backup directory while calculating a manifest:
from pathlib import Path
import hashlib, json, shutil, datetime
src=Path('ch25-recovery-source'); backup=Path('ch25-recovery-backup')
if backup.exists(): shutil.rmtree(backup)
shutil.copytree(src, backup/'payload')
def digest(p): return hashlib.sha256(p.read_bytes()).hexdigest()
files=[]
for p in sorted((backup/'payload').rglob('*')):
if p.is_file(): files.append({'path':p.relative_to(backup/'payload').as_posix(),'sha256':digest(p),'size':p.stat().st_size})
manifest={
'createdAt': datetime.datetime.now(datetime.timezone.utc).isoformat(),
'nexusVersion':'3.95.2-01', 'java':'21', 'database':'H2-model',
'blobStores':['ch25-file'], 'rpoMinutes':60, 'rtoMinutes':120,
'files':files
}
(backup/'recovery-manifest.json').write_text(json.dumps(manifest, indent=2)+'
')
print('backup files:', len(files))
Store the manifest beside—but logically distinct from—the backed-up payload. In production, protect manifests and backup metadata against unauthorized modification as well.
4. Verify the backup before destroying anything
from pathlib import Path
import hashlib, json
backup=Path('ch25-recovery-backup')
m=json.loads((backup/'recovery-manifest.json').read_text())
for item in m['files']:
p=backup/'payload'/item['path']
assert p.exists(), item['path']
assert hashlib.sha256(p.read_bytes()).hexdigest()==item['sha256'], item['path']
print('backup verification: PASS')
A copy that cannot be read back is not a useful backup. This step catches missing files and obvious corruption before the loss exercise.
5. Inject controlled loss in the source only
Predict first: if 1.1 is deleted from source storage
but metadata remains, source consistency should fail while the
backup remains valid.
from pathlib import Path
victim=Path('ch25-recovery-source/blobs/releases/learner-example-1.1.txt')
assert victim.exists()
victim.unlink()
print('controlled loss injected:', victim)
Do not “fix” the source by editing metadata to hide the missing file. The goal is to practice restore from the known recovery point.
6. Restore into an isolated target
from pathlib import Path
import shutil
backup=Path('ch25-recovery-backup/payload')
target=Path('ch25-recovery-target')
if target.exists(): shutil.rmtree(target)
shutil.copytree(backup, target)
print('restored into:', target)
Notice that restore goes to a new directory. Production recovery should likewise prefer an isolated validation target when feasible; overwriting the only remaining copy destroys forensic and rollback options.
7. Validate metadata-to-blob consistency and checksums
from pathlib import Path
import hashlib, json
target=Path('ch25-recovery-target')
rows=json.loads((target/'db/assets.json').read_text())
fail=[]
for row in rows:
p=target/'blobs'/row['path']
if not p.exists(): fail.append((row['path'],'missing'))
elif hashlib.sha256(p.read_bytes()).hexdigest()!=row['sha256']: fail.append((row['path'],'checksum'))
if fail:
raise SystemExit('restore validation failed: '+repr(fail))
print('restore validation: PASS', len(rows), 'assets')
This models a core production assertion: every sampled/critical metadata record resolves to bytes with the expected immutable hash.
8. Validate configuration and identity separately
Check that db/repos.json still maps
ch25-raw-hosted to ch25-file, that the
synthetic node ID exists, and that configuration files are present.
In a real Nexus restore, also inspect repositories,
roles/privileges/realms, blob stores, base URL/reverse-proxy
expectations, TLS trust, scheduled tasks, and external integrations.
Do not assume “the artifact downloaded” proves the
security/configuration layer is correct.
9. Optional live H2 extension: map the fixture to current Nexus tasks
On a disposable archive-based H2 Nexus instance:
-
Create
ch25-fileandch25-raw-hosted; upload two harmless text assets. - Record Nexus version, Java runtime, data directory, repository→blob mapping, asset checksums, and node-ID backup location.
- Create/run Admin - Backup H2 Database to a configured relative backup location.
- Back up the file blob store and required custom configuration/node identity as the same declared recovery set. For a stronger embedded-database backup posture, schedule a maintenance window and take the periodic offline database backup Sonatype recommends.
- Stop Nexus before restore. Restore the complete H2 database backup and the corresponding blob backup; preserve the original source until validation completes.
- Restart and verify repository/configuration/assets/checksums/client requests.
Exact filesystem commands are intentionally not hard-coded because
installation/data paths differ and a blind rm -rf or
copy command is unsafe. The operator must identify the disposable
paths from the instance’s own configuration.
10. What if database and blob backup times differ?
Record the skew and inspect the affected window. Current Sonatype data-repair tooling uses a plan-oriented workflow to reconcile discrepancies between database rows, blob metadata, and binaries for supported formats. Do not immediately execute repair. Generate/review the plan, bound the timespan to the recovery window, and understand each proposed action. Earlier task names and folklore may be obsolete.
11. Small challenge: choose the least destructive recovery control
You restored an H2 backup from 02:00 and a blob snapshot from 02:03. One artifact uploaded at 02:02 exists in blobs but not in the restored database. What should you do?
- Do not delete the “extra” blob manually.
- Preserve the source/backup evidence and identify the expected RPO point.
- If the 02:02 artifact should survive according to the chosen recovery objective, use the current documented reconciliation-plan workflow for the supported format and review the planned action.
- If the chosen recovery point is strictly 02:00 and the artifact should not exist, follow the documented restore/reconciliation policy rather than manipulating internal files.
12. Required evidence packet
-
source-state.txt— exact version/database/blob topology or fixture description. -
recovery-manifest.json— recovery point and file/hash inventory. -
backup-verification.txt— pre-loss verification. -
loss-event.md— what was intentionally removed and why. -
restore-log.txt— isolated target and restore sequence. -
validation.txt— metadata/blob/hash/config/client checks. gaps.md— anything the exercise did not prove.
13. Cleanup / rollback
After preserving learning evidence, delete only
ch25-recovery-source,
ch25-recovery-backup, and
ch25-recovery-target from the local fixture. For the
optional Nexus extension, remove only the explicitly named
disposable repository through supported Nexus administration after
validation. Do not delete shared blob directories or production
backups.
14. Knowledge check
Why verify the backup before injecting loss?
Because recovery testing should not destroy the source only to discover that the backup was already incomplete or corrupt.
Why restore into an isolated target?
It preserves source/forensic evidence and lets you validate recovered state before replacing the only functioning or remaining copy.
What does the Admin - Backup H2 Database task not back up for you?
The repository blob stores; they must be backed up separately as part of the same recovery set, along with required configuration/node identity.
Why is a reconciliation plan preferable to direct blob/database edits when backup times differ?
It uses supported Nexus semantics, exposes proposed actions for review, and avoids corrupting internal relationships with manual edits.
What is the strongest proof in this lab that the restored artifact is the intended one?
The restored metadata resolves to a present file whose cryptographic checksum matches the checksum captured before the loss.
15. Summary and bridge
You created, verified, damaged, restored, and independently validated a recovery set without touching production state. The live H2 mapping shows where Sonatype’s task fits and where it stops. Lesson 3 now turns this workflow into design choices: online versus offline protection, H2 versus PostgreSQL, local versus object storage, and full recovery versus content transfer.
Official references and version notes
- Sonatype: Backup and Restore — embedded H2 backup-task behavior and the role of database snapshots.
- Sonatype: Prepare a Backup — blob-store, node-ID, database, and custom-configuration backup requirements.
- Sonatype: Configure the Backup Task — current H2 task name and the recommendation for periodic offline embedded-database backups.
- Sonatype: Restore an H2 Database — stop/restore/restart sequence and same-point blob-store requirement.
- Sonatype: Nexus Repository Database — H2/PostgreSQL boundaries and PostgreSQL backup responsibilities.
- PostgreSQL: Backup and Restore — database-native backup methods for external PostgreSQL.
- Sonatype: Storage Guide — blob-store layout and the warning not to modify internal blob files manually.
- Sonatype: Data Repair Tasks — current plan-based reconciliation for database/blob inconsistencies and supported formats.
- Sonatype: Repository Export — Pro-only content export and its operational limits.
- Sonatype: Repository Import — Pro-only content import, generated metadata, and what is not preserved.
- Sonatype: Backup/Same-Site Restore — RPO/RTO and test-restore expectations.
- Sonatype: Resiliency — recovery expectations after database failure and the distinction from HA.
- Sonatype: Nexus Repository 3.95.x Release Notes — dated version line used as the chapter reference.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.