High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Configuration, Design Choices, and Tradeoffs
HA architecture is a set of deliberate tradeoffs, not a universal “three-node” recipe. This lesson compares active-active application nodes with database single-writer semantics, shared-file versus object storage, multi-AZ cost, session persistence, blob-store groups, availability versus consistency, and the boundary between redundant infrastructure and recoverable state.
Learning objectives
- Choose application, database, storage, and load-balancer patterns based on explicit failure objectives.
- Explain why local blob storage is incompatible with active HA nodes unless it is actually shared through a supported backend.
- Use sticky sessions only for documented behavior such as current rolling-upgrade UI consistency.
- Distinguish blob-store groups and storage redundancy from point-in-time backup.
- Justify cost/RTO choices with observable system state.
1. Active-active application tier, single-writer database contract
Nexus application nodes can all accept client traffic, but the database is not an active-active set of independent writers. A supported PostgreSQL HA service may have replicas and automatic failover, yet Nexus still connects through a coherent writable endpoint. This preserves transactional ordering for repository/security/configuration metadata.
2. Local versus shared blob storage
| Choice | HA consequence | Operational tradeoff |
|---|---|---|
| Private local disk per node | Not a coherent active-node blob design | Fast locally, but nodes can see different bytes |
| Supported NFSv4/network file store | Shared namespace across nodes | Latency/IO and storage-server availability must be engineered |
| Supported cloud object store | Shared object namespace across nodes | Provider latency, IAM, endpoint, SDK/version, cost and regional availability matter |
The storage layer should be close enough to the application/database to meet latency and throughput requirements. “Shared” is not equivalent to “fast enough.”
3. Multi-AZ versus multi-region
A production cloud HA design can distribute application nodes across multiple availability zones/failure domains within one region, with a managed multi-AZ PostgreSQL service and regional shared object storage. Current Sonatype HA requirements do not support stretching one HA cluster across regions. Cross-region continuity belongs to DR/federation patterns.
4. Availability versus consistency
Artifact repositories are systems of record. Serving stale or divergent metadata merely to stay “green” can be worse than a short controlled outage. When the shared database or blob layer loses correctness, the safe response may be to stop writes/read-only/fail closed rather than let nodes invent divergent state.
5. Sticky sessions: do not cargo-cult them
Normal application HA should not assume session persistence unless the current deployment behavior requires it. Sonatype’s current rolling-upgrade guidance specifically requires sticky sessions on the load balancer during zero-downtime upgrades because UI behavior can differ while nodes run temporarily in mixed versions. That is a documented special case, not proof that every package-client request needs affinity forever.
6. Group blob stores: capacity tool, not availability magic
Group blob stores are Pro-only. Two current fill policies are important:
- Round Robin: writes alternate across members; it does not choose the “least loaded” member and does not mirror the same blob to every member.
- Write to First: writes target the first writable member, useful for controlled storage migration/capacity redirection.
Current UI health marks a group unhealthy when any member is unreachable. A group can be useful operationally, but you must separately reason about member failure, content placement, backup, and recovery.
7. Blob redundancy versus backup
flowchart LR N[Nexus cluster] --> S[Shared blob service] S --> R[Provider redundancy / replicas] N --> DB[(PostgreSQL HA)] DB --> DR[DB backup / PITR] S --> BK[Blob backup / versioned recovery copy] DEL[Accidental deletion or corruption] --> S DEL --> DB R -. may replicate bad change .-> X[Same logical damage] BK --> REC[Point-in-time recovery] DR --> REC
Infrastructure redundancy can faithfully replicate a bad delete. Recovery history must be independent enough to return to a known-good consistency point.
8. Cost versus RTO
Adding failure domains reduces some outage risks but increases infrastructure, testing, operational, and database/storage costs. Define service objectives first:
| Objective | Example design response | What to measure |
|---|---|---|
| Survive one Nexus host loss | 2+ active nodes on separate hosts + LB | LB drain/failover seconds; client errors |
| Survive one AZ loss | Nodes and managed DB/storage across AZs in one region | Remaining capacity and DB failover time |
| Recover from deletion | Independent tested backup/DR | Restore RTO and backup RPO |
| Cross-region continuity | Separate DR/federation pattern | Replication lag and DR cutover, not HA node count |
9. Worked decision scenario
Team: 4,000 developers, release builds must keep resolving artifacts through one application-node/AZ failure; a 5-minute DB failover is acceptable; accidental deletion must be recoverable to the prior 30-minute checkpoint.
Reasoned design: Pro HA application nodes across at least two failure domains in one region; managed supported PostgreSQL HA endpoint; shared regional object storage; redundant load-balancer service; independent DB/blob backup aligned to 30-minute RPO. A group blob store is optional and only added for a concrete capacity/migration requirement.
10. Configuration ownership table
| State | Owner | Do not confuse with |
|---|---|---|
| Nexus cluster properties | Nexus deployment | Load-balancer policy |
| PostgreSQL primary/failover | DB platform | Nexus node membership |
| Blob replication/durability | Storage platform | Group blob-store fill policy |
| TLS/session persistence | Reverse proxy/LB | Nexus repository routing |
| Backups/PITR | Recovery platform/runbook | HA replicas |
Knowledge check
Why is a read replica not another Nexus database writer?
Nexus requires coherent transactional state through the supported PostgreSQL writer/failover endpoint; read-replica load balancing is not supported.
When are sticky sessions explicitly required by current Sonatype guidance?
During supported zero-downtime/rolling upgrades to avoid UI inconsistency while nodes temporarily run mixed versions.
Does round-robin group blob storage replicate each blob to every member?
No. It distributes write placement across members; it is not mirroring.
Why can storage replication fail to protect against accidental deletion?
The deletion/corruption can replicate too; point-in-time backup/DR is a different control.
Why keep a Nexus HA cluster within one region?
Current HA guidance is single-region/data-center scoped; cross-region continuity uses separate DR/federation patterns due to latency and consistency risks.
Summary and next step
HA design is a controlled balance: active Nexus nodes improve application availability, but PostgreSQL keeps one coherent writer contract, blobs remain shared state, and redundancy cannot replace backup.
Lesson 4 diagnoses the most dangerous ways an HA design can look redundant while still being incorrect or fragile.
Official references and version notes
- Sonatype: High Availability Deployment — current Pro-only clustered application model.
- Sonatype: Requirements for High Availability — same-version nodes, single-region scope, load balancer, shared blob storage, and external PostgreSQL requirements.
- Sonatype: Manual High Availability Deployment — clustered mode, shared storage, and per-node operational state.
- Sonatype: Validating Your HA Deployment — node status, system information, and per-node support evidence.
- Sonatype: Status API — read/writable status endpoints and their monitoring limitations.
- Sonatype: Nodes API — Pro cluster-node inventory and node identities.
- Sonatype: System Requirements — Java 21, PostgreSQL guidance, supported providers, and the explicit pgpool/load-balancing restriction.
- Sonatype: Blob Stores — shared/object storage behavior, health indicators, and group blob-store semantics.
- Sonatype: Storage Planning — group blob stores, storage layout tradeoffs, and performance implications.
- Sonatype: Nexus Repository Professional Features — HA, cloud blob stores, and group blob-store licensing boundaries.
- Sonatype: Rolling Upgrades in High Availability — mixed-version limits, sticky-session requirement during rolling upgrades, schema-finalization behavior, and restore-based rollback after finalization.
- Sonatype: Prepare a Backup — recovery-set consistency across database/configuration and blob content.
- Sonatype: Cross-Region Disaster Recovery — a separate DR pattern rather than an extension of a single-region HA cluster.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.