Chapter 28Lesson 03220–300 min

High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Configuration, Design Choices, and Tradeoffs

HA architecture is a set of deliberate tradeoffs, not a universal “three-node” recipe. This lesson compares active-active application nodes with database single-writer semantics, shared-file versus object storage, multi-AZ cost, session persistence, blob-store groups, availability versus consistency, and the boundary between redundant infrastructure and recoverable state.

Design tradeoffsMulti-AZBlob groupsConsistencyRTO/RPO

Learning objectives

  • Choose application, database, storage, and load-balancer patterns based on explicit failure objectives.
  • Explain why local blob storage is incompatible with active HA nodes unless it is actually shared through a supported backend.
  • Use sticky sessions only for documented behavior such as current rolling-upgrade UI consistency.
  • Distinguish blob-store groups and storage redundancy from point-in-time backup.
  • Justify cost/RTO choices with observable system state.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live version-status, feature-matrix, HA, storage, and release-note pages before implementing a real cluster.
License boundary. Nexus Repository high availability is a Pro-only capability. The mandatory chapter path is therefore architecture- and fixture-driven and can be completed without a Pro license, cloud account, managed database, Kubernetes cluster, or multi-node Nexus deployment.
HA is not backup or disaster recovery. Multiple active nodes improve service continuity for selected failures. They do not create a point-in-time recovery copy, undo deletion/corruption, or make a cross-region cluster supported. Chapter 25 remains the recovery foundation.

1. Active-active application tier, single-writer database contract

Nexus application nodes can all accept client traffic, but the database is not an active-active set of independent writers. A supported PostgreSQL HA service may have replicas and automatic failover, yet Nexus still connects through a coherent writable endpoint. This preserves transactional ordering for repository/security/configuration metadata.

Wrong design: configure each Nexus node to round-robin SQL queries across a writer and read replicas. Current Sonatype requirements explicitly reject pgpool/generic database load balancing.

2. Local versus shared blob storage

Choice HA consequence Operational tradeoff
Private local disk per node Not a coherent active-node blob design Fast locally, but nodes can see different bytes
Supported NFSv4/network file store Shared namespace across nodes Latency/IO and storage-server availability must be engineered
Supported cloud object store Shared object namespace across nodes Provider latency, IAM, endpoint, SDK/version, cost and regional availability matter

The storage layer should be close enough to the application/database to meet latency and throughput requirements. “Shared” is not equivalent to “fast enough.”

3. Multi-AZ versus multi-region

A production cloud HA design can distribute application nodes across multiple availability zones/failure domains within one region, with a managed multi-AZ PostgreSQL service and regional shared object storage. Current Sonatype HA requirements do not support stretching one HA cluster across regions. Cross-region continuity belongs to DR/federation patterns.

4. Availability versus consistency

Artifact repositories are systems of record. Serving stale or divergent metadata merely to stay “green” can be worse than a short controlled outage. When the shared database or blob layer loses correctness, the safe response may be to stop writes/read-only/fail closed rather than let nodes invent divergent state.

5. Sticky sessions: do not cargo-cult them

Normal application HA should not assume session persistence unless the current deployment behavior requires it. Sonatype’s current rolling-upgrade guidance specifically requires sticky sessions on the load balancer during zero-downtime upgrades because UI behavior can differ while nodes run temporarily in mixed versions. That is a documented special case, not proof that every package-client request needs affinity forever.

6. Group blob stores: capacity tool, not availability magic

Group blob stores are Pro-only. Two current fill policies are important:

  • Round Robin: writes alternate across members; it does not choose the “least loaded” member and does not mirror the same blob to every member.
  • Write to First: writes target the first writable member, useful for controlled storage migration/capacity redirection.

Current UI health marks a group unhealthy when any member is unreachable. A group can be useful operationally, but you must separately reason about member failure, content placement, backup, and recovery.

7. Blob redundancy versus backup

Availability copy is not recovery history
flowchart LR
N[Nexus cluster] --> S[Shared blob service]
S --> R[Provider redundancy / replicas]
N --> DB[(PostgreSQL HA)]
DB --> DR[DB backup / PITR]
S --> BK[Blob backup / versioned recovery copy]
DEL[Accidental deletion or corruption] --> S
DEL --> DB
R -. may replicate bad change .-> X[Same logical damage]
BK --> REC[Point-in-time recovery]
DR --> REC

Infrastructure redundancy can faithfully replicate a bad delete. Recovery history must be independent enough to return to a known-good consistency point.

8. Cost versus RTO

Adding failure domains reduces some outage risks but increases infrastructure, testing, operational, and database/storage costs. Define service objectives first:

Objective Example design response What to measure
Survive one Nexus host loss 2+ active nodes on separate hosts + LB LB drain/failover seconds; client errors
Survive one AZ loss Nodes and managed DB/storage across AZs in one region Remaining capacity and DB failover time
Recover from deletion Independent tested backup/DR Restore RTO and backup RPO
Cross-region continuity Separate DR/federation pattern Replication lag and DR cutover, not HA node count

9. Worked decision scenario

Team: 4,000 developers, release builds must keep resolving artifacts through one application-node/AZ failure; a 5-minute DB failover is acceptable; accidental deletion must be recoverable to the prior 30-minute checkpoint.

Reasoned design: Pro HA application nodes across at least two failure domains in one region; managed supported PostgreSQL HA endpoint; shared regional object storage; redundant load-balancer service; independent DB/blob backup aligned to 30-minute RPO. A group blob store is optional and only added for a concrete capacity/migration requirement.

10. Configuration ownership table

State Owner Do not confuse with
Nexus cluster properties Nexus deployment Load-balancer policy
PostgreSQL primary/failover DB platform Nexus node membership
Blob replication/durability Storage platform Group blob-store fill policy
TLS/session persistence Reverse proxy/LB Nexus repository routing
Backups/PITR Recovery platform/runbook HA replicas

Knowledge check

Why is a read replica not another Nexus database writer?

When are sticky sessions explicitly required by current Sonatype guidance?

Does round-robin group blob storage replicate each blob to every member?

Why can storage replication fail to protect against accidental deletion?

Why keep a Nexus HA cluster within one region?

Summary and next step

HA design is a controlled balance: active Nexus nodes improve application availability, but PostgreSQL keeps one coherent writer contract, blobs remain shared state, and redundancy cannot replace backup.

Lesson 4 diagnoses the most dangerous ways an HA design can look redundant while still being incorrect or fragile.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.