Chapter 25Lesson 04220–300 min

Backup, Restore, Configuration Export, Blob Recovery, Disaster Scenarios, and Recovery Validation: Diagnostics, Failure Modes, Security, and Performance

Recovery failures are often consistency failures disguised as missing files, startup errors, authorization surprises, or client 404s. Diagnose from preserved evidence, identify which state plane is wrong, and apply the least destructive supported correction rather than editing database rows or blob files.

Recovery diagnosticsData repairCorruptionVersion mismatchSecurity

Learning objectives

  • Diagnose database-only, blob-only, configuration-only, and version-incompatible restore failures using a repeatable evidence sequence.
  • Differentiate a missing binary, missing database row, soft-delete state, checksum mismatch, and client/routing/authentication failure.
  • Use current plan-based data-repair concepts only when the evidence matches their purpose and the Nexus version is supported.
  • Recognize backup-time concurrency, disk exhaustion, missing secrets, and untested backups as distinct operational failure modes.
  • Protect forensic evidence and avoid destructive “cleanup” during incident recovery.
Dated baseline (27 August 2026). These lessons use Nexus Repository 3.95.2-01 with Java 21 as the reference line. Recovery procedures are version/database/storage sensitive: before a real restore, verify the exact release notes, database support, storage backend, and restore documentation for the version that created the backup.
Recovery-set invariant. A Nexus backup is not “the database” or “the blobs.” Sonatype requires database/configuration state and blob content to be protected together. Preserve the node identity under $data-dir/keystores/node/ as part of the recovery set as well.
Destructive-lab boundary. Every loss/restore exercise in this chapter targets synthetic data in an isolated directory or disposable Nexus instance. Never delete, overwrite, or restore an employer database, blob store, object-store bucket, production secret, or production data directory for a training exercise.

1. Recovery diagnostic sequence

  1. Preserve logs, backup manifests, timestamps, checksums, task history, and the original failed restore.
  2. Confirm Nexus version/edition/Java and whether the restore target is compatible with the backup procedure.
  3. Confirm database type and exact restored database point.
  4. Confirm every configured blob store, backend, and restored storage point.
  5. Confirm node-ID/custom configuration/secret prerequisites.
  6. Start in isolation if the documented procedure permits; inspect startup/database/blob errors.
  7. Test repository definitions, authorization, known component metadata, and byte hashes.
  8. Only then choose reconciliation, configuration correction, or a fresh restore from a different point.

2. Failure: only the database was restored

Symptom: repositories/configuration appear, search may show components, but downloads fail because referenced blobs are missing. Cause: metadata was restored without matching content. Correction: stop escalating changes, identify the corresponding blob recovery point, restore it according to supported procedure, then validate/reconcile as documented. Do not delete database rows to make the errors disappear.

3. Failure: only blobs were copied

Symptom: storage contains files, but Nexus does not expose the expected component because the database lacks the row or configuration mapping. Cause: blob storage is not a self-describing replacement for complete Nexus state. Correction: restore the expected database/configuration point; if the selected recovery design intentionally allows a bounded mismatch, use current supported reconciliation planning for eligible formats.

4. Failure: corrupted blob is discovered only on client request

Backups can copy corruption perfectly. That is why restore validation must include cryptographic sampling/critical-artifact checks. If pre-backup evidence recorded a SHA-256 and the restored download differs, quarantine the restore target, preserve both versions, identify which recovery point first diverged, and restore from a known-good copy. A checksum mismatch is not fixed by rewriting the expected hash in metadata.

5. Failure: restore attempted on an unsupported/different Nexus line

A backup is not an upgrade mechanism. Database schemas, runtime requirements, migration state, plugins, and storage semantics change across versions. Restore to the version/procedure supported for the backup point, validate recovery, then execute the documented upgrade path separately. Chapter 27 handles upgrade planning; Chapter 26 handles database migration. Combining restore, migration, and upgrade into one untested emergency step multiplies failure modes.

6. Failure: database/blobs restore, but service integration fails

Missing TLS key material, trust stores, reverse-proxy configuration, external DB credentials, LDAP/OIDC/SAML dependencies, or object-store credentials can make a technically intact database/blob restore unusable. Treat these as separate prerequisites. Restore secrets from the organization’s approved secret system; never archive real credentials in plain-text lesson evidence.

7. Failure: backup overlapped cleanup, compaction, publish, or maintenance

Capture task history and timestamps. A database snapshot taken while aggressive content deletion or storage movement is occurring can increase inconsistency risk or make the recovery window harder to reason about. In production, define maintenance concurrency policy: which tasks may overlap backup, which must be paused, and how writers are controlled. The correct answer is not always “stop everything,” but it must be documented and tested.

8. Failure: backup or restore fills disk

Disk exhaustion can corrupt in-flight writes and create a second incident. Preflight both backup destination and restore target with headroom for database files, blob copies, decompression, temporary files, logs, and reconciliation. Use metrics/quotas/alerts before starting a restore. If a disk-full event occurs, preserve logs and do not start deleting internal Nexus files to create emergency space.

9. Data repair is not a generic recovery button

Current Sonatype data repair uses a plan-oriented process. It compares database rows, storage metadata, and binaries; plans can be reviewed before execution and constrained to a timespan. That is valuable when backup/storage points differ or metadata is missing. It is not permission to run repair after every restore.

Version history matters: Sonatype documents a Verify/Repair bug in 3.83.0–3.89.1 that could delete valid assets, fixed in 3.90.0. Always verify the exact task behavior and release line before execution.

10. Intentionally broken fixture: detect both missing and altered blobs

from pathlib import Path
import hashlib, json, shutil
root=Path('ch25-broken')
if root.exists(): shutil.rmtree(root)
(root/'blobs').mkdir(parents=True)
(root/'blobs/a.bin').write_bytes(b'A-original')
(root/'blobs/b.bin').write_bytes(b'B-original')
rows=[]
for name in ['a.bin','b.bin']:
    p=root/'blobs'/name
    rows.append({'path':name,'sha256':hashlib.sha256(p.read_bytes()).hexdigest()})
(root/'db.json').write_text(json.dumps(rows))
# Failure injection: one missing blob, one corrupted blob.
(root/'blobs/a.bin').unlink()
(root/'blobs/b.bin').write_bytes(b'B-corrupt')
issues=[]
for row in json.loads((root/'db.json').read_text()):
    p=root/'blobs'/row['path']
    if not p.exists(): issues.append((row['path'],'MISSING'))
    elif hashlib.sha256(p.read_bytes()).hexdigest()!=row['sha256']: issues.append((row['path'],'CHECKSUM_MISMATCH'))
print(issues)

Expected result: one MISSING and one CHECKSUM_MISMATCH. They are different problems and may require different recovery evidence.

11. Do not misdiagnose client failures as recovery corruption

A post-restore 401 can be authentication/realm configuration; 403 can be privilege/content-selector policy; 404 can be a wrong repository/path or missing asset; 429 can be rate limiting; TLS failures can be reverse-proxy/certificate trust. Before rebuilding data, compare the exact client URL, auth, repository type, and server-side asset state.

12. Support bundles and privacy

Logs and support ZIPs can be extremely useful during recovery, but inspect them for credentials, internal URLs, usernames, repository names, and proprietary metadata before external sharing. Preserve original evidence internally under appropriate access controls.

13. Failure decision table

Observation Likely plane First safe action
Metadata exists, binary missing DB newer than blob / blob loss Preserve evidence; identify matching blob recovery point
Binary exists, DB row missing Blob newer than DB / DB loss Review selected RPO and current reconciliation plan behavior
Hash differs Corruption/wrong backup Quarantine target; locate known-good recovery point
Repository missing entirely Configuration/database Verify restored DB/config point and version
401/403 only Identity/authorization Inspect realms/roles/credentials—not blobs
Nexus fails at startup Version/runtime/DB/config Inspect exact startup logs and compatibility before mutations

14. Knowledge check

Why is deleting database rows a bad response to missing blobs?

What does a checksum mismatch tell you that a 404 does not?

Why should restore and upgrade be separate operations?

When might current reconciliation tooling be appropriate?

Why preserve the failed restore instead of repeatedly modifying it?

15. Summary and bridge

Recovery diagnosis is state-plane diagnosis. Database, blobs, configuration, identity, client routing, and runtime compatibility can fail independently. Evidence first; supported repair second; destructive edits never. Lesson 5 now combines the chapter into a timed disaster drill with RPO/RTO, loss injection, isolated restore, validation, and a written operational handoff.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.