Chapter 25Lesson 03200–270 min

Backup, Restore, Configuration Export, Blob Recovery, Disaster Scenarios, and Recovery Validation: Configuration, Design Choices, and Tradeoffs

Choose recovery mechanisms from failure modes and service objectives, not from convenience. This lesson compares embedded and external database strategies, filesystem/object-store protection, full-instance restore versus repository content transfer, online versus offline consistency, and the operational cost of tighter RPO/RTO targets.

Recovery designH2 vs PostgreSQLObject storageExport/importTradeoffs

Learning objectives

  • Select backup mechanisms based on RPO/RTO, database type, blob backend, repository criticality, and restore-test evidence.
  • Compare online snapshots, maintenance-window/offline backups, and database-native methods without pretending they provide identical consistency guarantees.
  • Separate full instance recovery from Pro-only repository export/import and from replication/HA patterns.
  • Design backup retention, encryption, off-site copies, and validation schedules around realistic capacity and security constraints.
  • Build a decision table that makes recovery assumptions explicit before production deployment.
Dated baseline (27 August 2026). These lessons use Nexus Repository 3.95.2-01 with Java 21 as the reference line. Recovery procedures are version/database/storage sensitive: before a real restore, verify the exact release notes, database support, storage backend, and restore documentation for the version that created the backup.
Recovery-set invariant. A Nexus backup is not “the database” or “the blobs.” Sonatype requires database/configuration state and blob content to be protected together. Preserve the node identity under $data-dir/keystores/node/ as part of the recovery set as well.
Destructive-lab boundary. Every loss/restore exercise in this chapter targets synthetic data in an isolated directory or disposable Nexus instance. Never delete, overwrite, or restore an employer database, blob store, object-store bucket, production secret, or production data directory for a training exercise.

1. Start from failure domains

Ask what you need to survive: accidental component deletion, cleanup-policy mistakes, database corruption, blob-store loss, host loss, region loss, secret loss, operator error, or a bad upgrade. Different failures require different evidence and restore points. A backup strategy that handles host loss may not handle malicious deletion if every backup is immediately mutable from the same compromised account.

2. Design matrix

Choice Advantages Risks/cost Use when
H2 task + blob backup Simple for small/disposable H2 instances Separate blob operation; online embedded backup has corruption risk; H2 workload limits Small supported H2 footprint with tested restore
Periodic offline H2 DB + blob backup Stronger embedded DB consistency posture Maintenance downtime; still need blob/config/node-ID protection When short planned downtime is acceptable
PostgreSQL logical dump Portable logical recovery; familiar DBA tooling Restore time can be high; not a blob backup Moderate DBs and suitable RPO/RTO
PostgreSQL physical/base + WAL strategy Better fit for larger DBs/PITR strategies More operational complexity and storage Production external DB with tighter objectives
File/blob storage snapshot Fast point capture where storage supports it Must align with DB point; snapshot is not validation Storage platform has reliable snapshots
Object-store versioning/replication Useful durability and rollback primitives Deletion/replication semantics, cost, account compromise risk Cloud blob stores with governed policies
Repository export/import (Pro) Content transfer/consolidation Not full configuration/identity restore; additional disk/time Migration/content movement, not DR substitute

3. Online versus offline: availability is part of the tradeoff

Online backup minimizes service interruption, but consistency across an embedded database and active blob writes can be harder to reason about. Offline maintenance windows reduce write concurrency and make the recovery point easier to describe, at the cost of planned downtime. For H2, Sonatype explicitly recommends regular offline embedded-database backups even though the H2 backup task exists.

For PostgreSQL, database-native online backup/PITR mechanisms can provide strong database consistency while Nexus remains online, but the operator still needs a coherent blob-store point and a procedure for reconciling bounded skew.

4. Backup frequency follows RPO, not habit

Suppose releases are published continuously and the business accepts at most 15 minutes of lost publication. A nightly database dump cannot satisfy the RPO. Conversely, a development proxy cache may tolerate a longer RPO because some content can be re-fetched. Classify repositories by business criticality instead of applying one schedule blindly.

recovery_classes:
  internal-releases:
    rpo_minutes: 15
    rto_minutes: 120
    criticality: high
  internal-snapshots:
    rpo_minutes: 240
    rto_minutes: 480
    criticality: medium
  public-proxy-cache:
    rpo_minutes: 1440
    rto_minutes: 480
    criticality: lower

5. Backup retention and repository retention are different controls

Cleanup policy may intentionally remove old components from the live repository. Backup retention determines how long older recovery points remain available. If cleanup immediately propagates to every backup copy, you have not created a useful rollback window. Define retention independently and protect recovery sets from accidental/malicious deletion through separate identities, immutable/WORM options where appropriate, and off-site/account separation.

6. Backups are a concentrated secret and IP boundary

Repository backups can contain proprietary binaries, internal coordinates, security configuration, credentials/tokens, certificate material, and metadata revealing software architecture. Apply encryption at rest/in transit, least-privilege backup identities, access logging, retention/destruction policy, and periodic restore access tests. Do not collect real secrets in a course evidence bundle.

7. Local file blobs versus object storage

File systems often make snapshots and offline copies easy to understand, but large datasets can make full copies slow. Object stores provide versioning, replication, lifecycle, and cross-region primitives, but those controls have provider-specific semantics. Record bucket/container versioning state, retention policy, encryption keys, replication lag, and restore process. “Replicated” is not the same as “backed up”: a mistaken deletion can replicate too.

8. When repository export/import is the right tool

Use Pro repository export/import for content migration, consolidation, or disconnected transfer when its supported format/behavior fits the requirement. Do not select it to recover users, roles, realms, repository definitions, server settings, or tags. Current import generates metadata/checksums and sets imported blob timestamps at import time, which is another reason it is not a byte-for-byte full-instance time-machine.

9. Backup, replication, and HA solve different problems

HA improves availability during some node failures. Replication maintains another copy with some lag. Backup preserves recoverable historical points. None automatically substitutes for the others. Chapter 28 handles HA/shared-state architecture in depth; this chapter’s rule is simpler: preserve independent recovery points even if the service is highly available.

10. RTO is often a throughput problem

Estimate restore duration before promising it:

size_tb = 8
usable_mib_s = 220
seconds = size_tb * 1024 * 1024 / usable_mib_s
print('hours:', round(seconds / 3600, 1))

This simplistic estimate ignores small-file overhead, object API request rate, database restore, verification, decompression, and contention. Measure actual restore drills and use the evidence to revise RTO.

11. Recovery ownership must cross team boundaries

Define who owns Nexus, PostgreSQL, storage snapshots, object-store policies, KMS keys, DNS/TLS, reverse proxy/load balancer, identity provider, and secrets. A runbook that says “DBA restores PostgreSQL” without an escalation contact, recovery point identifier, and validation handoff is incomplete.

12. Worked design scenario

A company has 4 TB of hosted internal releases, 10 TB of proxy cache, external PostgreSQL, and an RPO/RTO of 30 minutes/4 hours for releases. A sound design might:

  • Use PostgreSQL PITR-capable backups with a declared recovery timestamp.
  • Use storage snapshots/versioned object storage for hosted blobs aligned to the database point.
  • Give proxy caches a looser backup objective or rebuild strategy, but preserve configuration/routing controls.
  • Protect node ID/custom configuration and secret-recovery procedures.
  • Perform quarterly isolated restores plus more frequent automated checksum sampling.
  • Keep recovery credentials and delete permissions separated from normal Nexus administration.

The exact implementation depends on storage/database platforms and Sonatype support constraints, but the decision is driven by failure objectives and measurable restore time.

13. Knowledge check

Why can an hourly database backup still fail a 15-minute RPO?

Why is replication not automatically a backup?

When is repository export/import appropriate?

Why does object-store versioning not eliminate restore testing?

What design evidence should revise RTO?

14. Summary and bridge

Recovery design is a set of explicit tradeoffs among consistency, downtime, storage cost, security, and time-to-recover. H2 and PostgreSQL require different database procedures; file and object blobs require different storage controls; export/import solves content movement, not full DR. Lesson 4 now focuses on failures that expose bad assumptions and how to diagnose them without destroying remaining evidence.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.