High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Concepts, Architecture, and Mental Model
High availability is not “run Nexus twice.” A supported HA deployment is one logical Nexus Repository service whose active application nodes coordinate around the same authoritative PostgreSQL database and the same accessible blob content, with a load balancer steering clients away from failed nodes. This lesson builds that state model before any cluster operation is attempted.
Learning objectives
- Explain why multiple independent Nexus instances are not an HA cluster.
- Identify the application, database, blob, load-balancer, node-local, and recovery state in a supported HA topology.
- Distinguish application-node failover from PostgreSQL failover and storage durability.
- Explain group blob stores without confusing distribution with replication or backup.
- Use read-only node/status/storage evidence before changing HA configuration.
1. The practical problem: one service, several failure domains
A single Nexus node can be well backed up and still be unavailable while its host, process, network path, or maintenance window is down. HA reduces that service interruption by keeping other application nodes ready to serve requests. The hard part is not adding more JVMs; it is ensuring every active node sees the same authoritative metadata and bytes.
That is why the cluster boundary includes more than Nexus itself. The load balancer, PostgreSQL service, blob-storage service, network, DNS/TLS path, license/secrets, and per-node capacity all participate in availability.
2. Supported current architecture
flowchart TB C[Package clients / CI] --> LB[Application load balancer] LB --> A[Nexus node A] LB --> B[Nexus node B] LB --> D[Nexus node C] A --> PG[(Shared external PostgreSQL)] B --> PG D --> PG A --> BS[(Shared supported blob storage)] B --> BS D --> BS MON[Monitoring] --> LB MON --> A MON --> B MON --> D MON --> PG MON --> BS BK[Backup / DR system] -. separate recovery path .-> PG BK -. separate recovery path .-> BS
The load balancer chooses a healthy application node. The nodes do not each own a private copy of repository metadata or artifacts: active nodes must reach the same external PostgreSQL state and compatible shared blob storage. Backup/DR is drawn separately because it solves a different failure class.
3. Current edition and topology gates
| Gate | Current rule | Why it exists |
|---|---|---|
| Edition | Nexus Repository Pro | HA clustering is a Pro feature. |
| Node version/config |
Same Nexus version and same nexus.properties in
normal operation
|
Avoid unsupported mixed behavior; temporary mixed versions are reserved for supported rolling upgrades. |
| Database | External PostgreSQL reachable by every active node | Repository/configuration state must be coherent across nodes. |
| Blob storage | Supported shared/network/object storage reachable by every active node | Every node must resolve the same binary content. |
| Traffic | Application load balancer | Route new requests away from failed/draining application nodes. |
| Geography | One region/data-center scope | Cross-region database latency and failure modes are handled by separate federation/DR patterns. |
4. Application nodes are not fully stateless
The authoritative repository state is externalized, but a Nexus node still has node-local operational state: JVM/process state, local logs, temporary files, local work directories, and node identity/diagnostic context. Treating nodes as disposable compute is a deployment goal, not permission to ignore node-specific logs or configuration.
In a manual HA deployment the cluster property is enabled on every node:
nexus.datastore.clustered.enabled=true
The same setting does not magically create HA if the nodes point at different databases or private blob stores.
5. The PostgreSQL boundary: active-active Nexus, single writer database
The application tier can be active-active, but Nexus still requires coherent transactional database state. Database high availability therefore belongs to PostgreSQL or the managed PostgreSQL provider: one writable primary endpoint with provider-managed replication/failover. Nexus does not gain correctness by spraying SQL requests across read replicas.
6. Shared blob storage is state, not a cache
Repository artifacts, generated metadata, hashes, and related blob content must remain reachable after an application-node failure. Supported HA storage includes current documented network/object choices such as S3, EFS/NFSv4, Azure Blob Storage, and Google Cloud Storage, subject to version/environment support and latency requirements.
A private local file blob store on node A plus a different local blob store on node B is not shared state. A package uploaded through node A could be missing when the load balancer later selects node B.
7. Group blob stores solve placement, not replication
A group blob store is a Pro storage abstraction
that lets one repository write across multiple member blob stores.
Its fill policy can be roundRobin or
writeToFirst. That is useful for capacity expansion and
migration, but it is not an automatic mirror and not a backup.
| Mechanism | Primary purpose | Does it recover deleted/corrupted content? |
|---|---|---|
| Shared object/network blob store | Make the same bytes reachable by all active Nexus nodes | No, not by itself |
| Group blob store | Distribute/redirect blob placement among members | No |
| Storage-provider replication | Storage availability/durability | Maybe for infrastructure loss; not automatically for logical deletion |
| Backup/DR copy | Point-in-time recovery | Yes, when tested and within RPO |
8. Node coordination and observation
Read-only inspection should establish the real cluster before you change it. In a licensed HA instance, the Nodes UI/API can show cluster node identities and friendly names. The Support status/system-information screens provide node-specific health and configuration evidence.
# Status endpoints are useful to a load balancer/monitoring probe.
curl -i https://repo.example.invalid/service/rest/v1/status
curl -i https://repo.example.invalid/service/rest/v1/status/writable
# Pro HA node inventory (authenticate with a scoped admin identity in a real lab).
curl -u 'admin-lab:password-FAKE_DO_NOT_USE' \
https://repo.example.invalid/service/rest/v1/nodes
The status endpoints report Nexus internal read/write readiness; current documentation warns that they are not hardware checks and do not prove the external database or disk/storage service is healthy. Pair them with infrastructure-native monitoring.
9. Failure-domain map
| Failure | Expected HA effect | What must still be healthy |
|---|---|---|
| One Nexus JVM/host fails | Load balancer routes around it | Another node + DB + blobs + LB |
| PostgreSQL primary fails | Service pauses/degrades until DB failover | Supported DB HA endpoint and replacement writer |
| Shared blob endpoint fails | Artifact operations fail/degrade across all nodes | Storage service/failover |
| Load balancer fails | Clients lose the service front door | Redundant LB/DNS path |
| Logical deletion/corruption | All healthy nodes may serve the same bad state | Backup/DR recovery set |
10. DevOps connection
HA becomes useful only when delivery teams can state which failure is being tolerated, what evidence proves traffic shifted safely, and what recovery boundary remains. “Three nodes” is not an SLO. A production operating model needs explicit dependencies, health signals, capacity, failover ownership, and tested recovery.
Knowledge check
Why are two independent Nexus servers not an HA cluster?
Because they do not share one coherent database/blob state and can diverge; a supported HA deployment is one clustered service with shared persistent state.
What is the database topology boundary?
Nexus nodes use a shared external PostgreSQL service with a writable primary/failover endpoint; generic SQL load balancing across replicas is not supported.
What does a group blob store guarantee?
It controls placement across member blob stores; it is not automatic mirroring and not a point-in-time backup.
Why can /status return useful information without proving the whole platform healthy?
It reports Nexus internal readiness, not hardware/external-database/storage health; infrastructure monitoring must cover those dependencies.
What failure class does HA not solve?
Logical deletion/corruption and site-level disaster still require backup/disaster-recovery mechanisms.
Summary and next step
You now have the correct HA state model: active Nexus nodes share one coherent PostgreSQL state and shared blob content behind a load balancer; group blob stores are a separate Pro storage abstraction; and HA remains distinct from backup/DR.
Lesson 2 turns that model into a free, observable failure-simulation workflow.
Official references and version notes
- Sonatype: High Availability Deployment — current Pro-only clustered application model.
- Sonatype: Requirements for High Availability — same-version nodes, single-region scope, load balancer, shared blob storage, and external PostgreSQL requirements.
- Sonatype: Manual High Availability Deployment — clustered mode, shared storage, and per-node operational state.
- Sonatype: Validating Your HA Deployment — node status, system information, and per-node support evidence.
- Sonatype: Status API — read/writable status endpoints and their monitoring limitations.
- Sonatype: Nodes API — Pro cluster-node inventory and node identities.
- Sonatype: System Requirements — Java 21, PostgreSQL guidance, supported providers, and the explicit pgpool/load-balancing restriction.
- Sonatype: Blob Stores — shared/object storage behavior, health indicators, and group blob-store semantics.
- Sonatype: Storage Planning — group blob stores, storage layout tradeoffs, and performance implications.
- Sonatype: Nexus Repository Professional Features — HA, cloud blob stores, and group blob-store licensing boundaries.
- Sonatype: Rolling Upgrades in High Availability — mixed-version limits, sticky-session requirement during rolling upgrades, schema-finalization behavior, and restore-based rollback after finalization.
- Sonatype: Prepare a Backup — recovery-set consistency across database/configuration and blob content.
- Sonatype: Cross-Region Disaster Recovery — a separate DR pattern rather than an extension of a single-region HA cluster.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.