Backup, Restore, Configuration Export, Blob Recovery, Disaster Scenarios, and Recovery Validation: Diagnostics, Failure Modes, Security, and Performance
Recovery failures are often consistency failures disguised as missing files, startup errors, authorization surprises, or client 404s. Diagnose from preserved evidence, identify which state plane is wrong, and apply the least destructive supported correction rather than editing database rows or blob files.
Learning objectives
- Diagnose database-only, blob-only, configuration-only, and version-incompatible restore failures using a repeatable evidence sequence.
- Differentiate a missing binary, missing database row, soft-delete state, checksum mismatch, and client/routing/authentication failure.
- Use current plan-based data-repair concepts only when the evidence matches their purpose and the Nexus version is supported.
- Recognize backup-time concurrency, disk exhaustion, missing secrets, and untested backups as distinct operational failure modes.
- Protect forensic evidence and avoid destructive “cleanup” during incident recovery.
$data-dir/keystores/node/ as part of the
recovery set as well.
1. Recovery diagnostic sequence
- Preserve logs, backup manifests, timestamps, checksums, task history, and the original failed restore.
- Confirm Nexus version/edition/Java and whether the restore target is compatible with the backup procedure.
- Confirm database type and exact restored database point.
- Confirm every configured blob store, backend, and restored storage point.
- Confirm node-ID/custom configuration/secret prerequisites.
- Start in isolation if the documented procedure permits; inspect startup/database/blob errors.
- Test repository definitions, authorization, known component metadata, and byte hashes.
- Only then choose reconciliation, configuration correction, or a fresh restore from a different point.
2. Failure: only the database was restored
Symptom: repositories/configuration appear, search may show components, but downloads fail because referenced blobs are missing. Cause: metadata was restored without matching content. Correction: stop escalating changes, identify the corresponding blob recovery point, restore it according to supported procedure, then validate/reconcile as documented. Do not delete database rows to make the errors disappear.
3. Failure: only blobs were copied
Symptom: storage contains files, but Nexus does not expose the expected component because the database lacks the row or configuration mapping. Cause: blob storage is not a self-describing replacement for complete Nexus state. Correction: restore the expected database/configuration point; if the selected recovery design intentionally allows a bounded mismatch, use current supported reconciliation planning for eligible formats.
4. Failure: corrupted blob is discovered only on client request
Backups can copy corruption perfectly. That is why restore validation must include cryptographic sampling/critical-artifact checks. If pre-backup evidence recorded a SHA-256 and the restored download differs, quarantine the restore target, preserve both versions, identify which recovery point first diverged, and restore from a known-good copy. A checksum mismatch is not fixed by rewriting the expected hash in metadata.
5. Failure: restore attempted on an unsupported/different Nexus line
A backup is not an upgrade mechanism. Database schemas, runtime requirements, migration state, plugins, and storage semantics change across versions. Restore to the version/procedure supported for the backup point, validate recovery, then execute the documented upgrade path separately. Chapter 27 handles upgrade planning; Chapter 26 handles database migration. Combining restore, migration, and upgrade into one untested emergency step multiplies failure modes.
6. Failure: database/blobs restore, but service integration fails
Missing TLS key material, trust stores, reverse-proxy configuration, external DB credentials, LDAP/OIDC/SAML dependencies, or object-store credentials can make a technically intact database/blob restore unusable. Treat these as separate prerequisites. Restore secrets from the organization’s approved secret system; never archive real credentials in plain-text lesson evidence.
7. Failure: backup overlapped cleanup, compaction, publish, or maintenance
Capture task history and timestamps. A database snapshot taken while aggressive content deletion or storage movement is occurring can increase inconsistency risk or make the recovery window harder to reason about. In production, define maintenance concurrency policy: which tasks may overlap backup, which must be paused, and how writers are controlled. The correct answer is not always “stop everything,” but it must be documented and tested.
8. Failure: backup or restore fills disk
Disk exhaustion can corrupt in-flight writes and create a second incident. Preflight both backup destination and restore target with headroom for database files, blob copies, decompression, temporary files, logs, and reconciliation. Use metrics/quotas/alerts before starting a restore. If a disk-full event occurs, preserve logs and do not start deleting internal Nexus files to create emergency space.
9. Data repair is not a generic recovery button
Current Sonatype data repair uses a plan-oriented process. It compares database rows, storage metadata, and binaries; plans can be reviewed before execution and constrained to a timespan. That is valuable when backup/storage points differ or metadata is missing. It is not permission to run repair after every restore.
Version history matters: Sonatype documents a Verify/Repair bug in 3.83.0–3.89.1 that could delete valid assets, fixed in 3.90.0. Always verify the exact task behavior and release line before execution.
10. Intentionally broken fixture: detect both missing and altered blobs
from pathlib import Path
import hashlib, json, shutil
root=Path('ch25-broken')
if root.exists(): shutil.rmtree(root)
(root/'blobs').mkdir(parents=True)
(root/'blobs/a.bin').write_bytes(b'A-original')
(root/'blobs/b.bin').write_bytes(b'B-original')
rows=[]
for name in ['a.bin','b.bin']:
p=root/'blobs'/name
rows.append({'path':name,'sha256':hashlib.sha256(p.read_bytes()).hexdigest()})
(root/'db.json').write_text(json.dumps(rows))
# Failure injection: one missing blob, one corrupted blob.
(root/'blobs/a.bin').unlink()
(root/'blobs/b.bin').write_bytes(b'B-corrupt')
issues=[]
for row in json.loads((root/'db.json').read_text()):
p=root/'blobs'/row['path']
if not p.exists(): issues.append((row['path'],'MISSING'))
elif hashlib.sha256(p.read_bytes()).hexdigest()!=row['sha256']: issues.append((row['path'],'CHECKSUM_MISMATCH'))
print(issues)
Expected result: one MISSING and one
CHECKSUM_MISMATCH. They are different problems and may
require different recovery evidence.
11. Do not misdiagnose client failures as recovery corruption
A post-restore 401 can be authentication/realm
configuration; 403 can be privilege/content-selector
policy; 404 can be a wrong repository/path or missing
asset; 429 can be rate limiting; TLS failures can be
reverse-proxy/certificate trust. Before rebuilding data, compare the
exact client URL, auth, repository type, and server-side asset
state.
12. Support bundles and privacy
Logs and support ZIPs can be extremely useful during recovery, but inspect them for credentials, internal URLs, usernames, repository names, and proprietary metadata before external sharing. Preserve original evidence internally under appropriate access controls.
13. Failure decision table
| Observation | Likely plane | First safe action |
|---|---|---|
| Metadata exists, binary missing | DB newer than blob / blob loss | Preserve evidence; identify matching blob recovery point |
| Binary exists, DB row missing | Blob newer than DB / DB loss | Review selected RPO and current reconciliation plan behavior |
| Hash differs | Corruption/wrong backup | Quarantine target; locate known-good recovery point |
| Repository missing entirely | Configuration/database | Verify restored DB/config point and version |
| 401/403 only | Identity/authorization | Inspect realms/roles/credentials—not blobs |
| Nexus fails at startup | Version/runtime/DB/config | Inspect exact startup logs and compatibility before mutations |
14. Knowledge check
Why is deleting database rows a bad response to missing blobs?
It hides symptoms by destroying metadata evidence and can make recovery harder; restore/reconcile through supported procedures instead.
What does a checksum mismatch tell you that a 404 does not?
A checksum mismatch proves bytes were found but differ from expected identity; a 404 can arise from routing/path/auth/metadata absence and does not itself prove corruption.
Why should restore and upgrade be separate operations?
They have different compatibility gates and failure modes. First recover the known supported state, validate it, then follow the documented upgrade path.
When might current reconciliation tooling be appropriate?
When evidence shows database/blob/metadata inconsistency for supported formats—especially bounded backup skew—after reviewing a generated plan and version-specific behavior.
Why preserve the failed restore instead of repeatedly modifying it?
It retains forensic evidence, allows comparison against later attempts, and prevents troubleshooting actions from erasing the original cause.
15. Summary and bridge
Recovery diagnosis is state-plane diagnosis. Database, blobs, configuration, identity, client routing, and runtime compatibility can fail independently. Evidence first; supported repair second; destructive edits never. Lesson 5 now combines the chapter into a timed disaster drill with RPO/RTO, loss injection, isolated restore, validation, and a written operational handoff.
Official references and version notes
- Sonatype: Backup and Restore — embedded H2 backup-task behavior and the role of database snapshots.
- Sonatype: Prepare a Backup — blob-store, node-ID, database, and custom-configuration backup requirements.
- Sonatype: Configure the Backup Task — current H2 task name and the recommendation for periodic offline embedded-database backups.
- Sonatype: Restore an H2 Database — stop/restore/restart sequence and same-point blob-store requirement.
- Sonatype: Nexus Repository Database — H2/PostgreSQL boundaries and PostgreSQL backup responsibilities.
- PostgreSQL: Backup and Restore — database-native backup methods for external PostgreSQL.
- Sonatype: Storage Guide — blob-store layout and the warning not to modify internal blob files manually.
- Sonatype: Data Repair Tasks — current plan-based reconciliation for database/blob inconsistencies and supported formats.
- Sonatype: Repository Export — Pro-only content export and its operational limits.
- Sonatype: Repository Import — Pro-only content import, generated metadata, and what is not preserved.
- Sonatype: Backup/Same-Site Restore — RPO/RTO and test-restore expectations.
- Sonatype: Resiliency — recovery expectations after database failure and the distinction from HA.
- Sonatype: Nexus Repository 3.95.x Release Notes — dated version line used as the chapter reference.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.