Chapter 28Lesson 01220–300 min

High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Concepts, Architecture, and Mental Model

High availability is not “run Nexus twice.” A supported HA deployment is one logical Nexus Repository service whose active application nodes coordinate around the same authoritative PostgreSQL database and the same accessible blob content, with a load balancer steering clients away from failed nodes. This lesson builds that state model before any cluster operation is attempted.

HA mental modelPro boundaryPostgreSQLShared blobsFailure domains

Learning objectives

  • Explain why multiple independent Nexus instances are not an HA cluster.
  • Identify the application, database, blob, load-balancer, node-local, and recovery state in a supported HA topology.
  • Distinguish application-node failover from PostgreSQL failover and storage durability.
  • Explain group blob stores without confusing distribution with replication or backup.
  • Use read-only node/status/storage evidence before changing HA configuration.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live version-status, feature-matrix, HA, storage, and release-note pages before implementing a real cluster.
License boundary. Nexus Repository high availability is a Pro-only capability. The mandatory chapter path is therefore architecture- and fixture-driven and can be completed without a Pro license, cloud account, managed database, Kubernetes cluster, or multi-node Nexus deployment.
HA is not backup or disaster recovery. Multiple active nodes improve service continuity for selected failures. They do not create a point-in-time recovery copy, undo deletion/corruption, or make a cross-region cluster supported. Chapter 25 remains the recovery foundation.

1. The practical problem: one service, several failure domains

A single Nexus node can be well backed up and still be unavailable while its host, process, network path, or maintenance window is down. HA reduces that service interruption by keeping other application nodes ready to serve requests. The hard part is not adding more JVMs; it is ensuring every active node sees the same authoritative metadata and bytes.

That is why the cluster boundary includes more than Nexus itself. The load balancer, PostgreSQL service, blob-storage service, network, DNS/TLS path, license/secrets, and per-node capacity all participate in availability.

2. Supported current architecture

One Nexus service, shared persistent state
flowchart TB
C[Package clients / CI] --> LB[Application load balancer]
LB --> A[Nexus node A]
LB --> B[Nexus node B]
LB --> D[Nexus node C]
A --> PG[(Shared external PostgreSQL)]
B --> PG
D --> PG
A --> BS[(Shared supported blob storage)]
B --> BS
D --> BS
MON[Monitoring] --> LB
MON --> A
MON --> B
MON --> D
MON --> PG
MON --> BS
BK[Backup / DR system] -. separate recovery path .-> PG
BK -. separate recovery path .-> BS

The load balancer chooses a healthy application node. The nodes do not each own a private copy of repository metadata or artifacts: active nodes must reach the same external PostgreSQL state and compatible shared blob storage. Backup/DR is drawn separately because it solves a different failure class.

3. Current edition and topology gates

Gate Current rule Why it exists
Edition Nexus Repository Pro HA clustering is a Pro feature.
Node version/config Same Nexus version and same nexus.properties in normal operation Avoid unsupported mixed behavior; temporary mixed versions are reserved for supported rolling upgrades.
Database External PostgreSQL reachable by every active node Repository/configuration state must be coherent across nodes.
Blob storage Supported shared/network/object storage reachable by every active node Every node must resolve the same binary content.
Traffic Application load balancer Route new requests away from failed/draining application nodes.
Geography One region/data-center scope Cross-region database latency and failure modes are handled by separate federation/DR patterns.

4. Application nodes are not fully stateless

The authoritative repository state is externalized, but a Nexus node still has node-local operational state: JVM/process state, local logs, temporary files, local work directories, and node identity/diagnostic context. Treating nodes as disposable compute is a deployment goal, not permission to ignore node-specific logs or configuration.

In a manual HA deployment the cluster property is enabled on every node:

nexus.datastore.clustered.enabled=true

The same setting does not magically create HA if the nodes point at different databases or private blob stores.

5. The PostgreSQL boundary: active-active Nexus, single writer database

The application tier can be active-active, but Nexus still requires coherent transactional database state. Database high availability therefore belongs to PostgreSQL or the managed PostgreSQL provider: one writable primary endpoint with provider-managed replication/failover. Nexus does not gain correctness by spraying SQL requests across read replicas.

Do not put pgpool or generic database load balancing in front of Nexus PostgreSQL. Current Sonatype system requirements explicitly do not support it. Use a supported PostgreSQL HA service/endpoint whose failover preserves the writable primary contract.

6. Shared blob storage is state, not a cache

Repository artifacts, generated metadata, hashes, and related blob content must remain reachable after an application-node failure. Supported HA storage includes current documented network/object choices such as S3, EFS/NFSv4, Azure Blob Storage, and Google Cloud Storage, subject to version/environment support and latency requirements.

A private local file blob store on node A plus a different local blob store on node B is not shared state. A package uploaded through node A could be missing when the load balancer later selects node B.

7. Group blob stores solve placement, not replication

A group blob store is a Pro storage abstraction that lets one repository write across multiple member blob stores. Its fill policy can be roundRobin or writeToFirst. That is useful for capacity expansion and migration, but it is not an automatic mirror and not a backup.

Mechanism Primary purpose Does it recover deleted/corrupted content?
Shared object/network blob store Make the same bytes reachable by all active Nexus nodes No, not by itself
Group blob store Distribute/redirect blob placement among members No
Storage-provider replication Storage availability/durability Maybe for infrastructure loss; not automatically for logical deletion
Backup/DR copy Point-in-time recovery Yes, when tested and within RPO

8. Node coordination and observation

Read-only inspection should establish the real cluster before you change it. In a licensed HA instance, the Nodes UI/API can show cluster node identities and friendly names. The Support status/system-information screens provide node-specific health and configuration evidence.

# Status endpoints are useful to a load balancer/monitoring probe.
curl -i https://repo.example.invalid/service/rest/v1/status
curl -i https://repo.example.invalid/service/rest/v1/status/writable

# Pro HA node inventory (authenticate with a scoped admin identity in a real lab).
curl -u 'admin-lab:password-FAKE_DO_NOT_USE' \
  https://repo.example.invalid/service/rest/v1/nodes

The status endpoints report Nexus internal read/write readiness; current documentation warns that they are not hardware checks and do not prove the external database or disk/storage service is healthy. Pair them with infrastructure-native monitoring.

9. Failure-domain map

Failure Expected HA effect What must still be healthy
One Nexus JVM/host fails Load balancer routes around it Another node + DB + blobs + LB
PostgreSQL primary fails Service pauses/degrades until DB failover Supported DB HA endpoint and replacement writer
Shared blob endpoint fails Artifact operations fail/degrade across all nodes Storage service/failover
Load balancer fails Clients lose the service front door Redundant LB/DNS path
Logical deletion/corruption All healthy nodes may serve the same bad state Backup/DR recovery set

10. DevOps connection

HA becomes useful only when delivery teams can state which failure is being tolerated, what evidence proves traffic shifted safely, and what recovery boundary remains. “Three nodes” is not an SLO. A production operating model needs explicit dependencies, health signals, capacity, failover ownership, and tested recovery.

Knowledge check

Why are two independent Nexus servers not an HA cluster?

What is the database topology boundary?

What does a group blob store guarantee?

Why can /status return useful information without proving the whole platform healthy?

What failure class does HA not solve?

Summary and next step

You now have the correct HA state model: active Nexus nodes share one coherent PostgreSQL state and shared blob content behind a load balancer; group blob stores are a separate Pro storage abstraction; and HA remains distinct from backup/DR.

Lesson 2 turns that model into a free, observable failure-simulation workflow.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.