High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Guided Hands-On Workflow and Core Operations
A real HA cluster requires Pro and external infrastructure, so the mandatory lab models the same state transitions with deterministic fixtures. You will prove the prerequisites, simulate failures one dependency at a time, observe the expected service state, and learn which evidence belongs to Nexus versus the load balancer, PostgreSQL, and blob-storage layers.
Learning objectives
- Validate the current HA prerequisites before designing a topology.
- Build a reproducible topology fixture with independent failure domains.
- Simulate application-node loss, PostgreSQL failover, blob degradation, and load-balancer failure.
- Collect Nexus and infrastructure evidence without assuming one health signal covers every layer.
- Explain which failures HA masks automatically and which require operator recovery.
1. Mandatory free lab versus optional licensed lab
The required path runs only a local Python simulator and static configuration fixtures. It does not launch a fake “HA Nexus” and does not imply Community Edition can form a cluster. If you already have a licensed disposable Pro environment, you may map the same evidence steps to that instance, but the learning objective is the state transition—not cloud/Kubernetes administration.
2. Preflight: reject an invalid architecture before simulation
reference:
nexus: 3.95.2-01
java: 21
editionRequiredForLiveHA: Pro
cluster:
region: lab-region-1
nodes:
- {name: nx-a, failureDomain: az-a}
- {name: nx-b, failureDomain: az-b}
- {name: nx-c, failureDomain: az-c}
sameNexusVersion: true
sameNexusProperties: true
database:
type: PostgreSQL
access: stable-writer-endpoint
loadBalancedAcrossReadReplicas: false
blob:
type: shared-object-storage
reachableFromAllNodes: true
loadBalancer:
healthProbe: /service/rest/v1/status
Each field corresponds to a current HA gate. A topology checker should fail closed if one is false.
3. Run the topology checker
topology = {
"edition": "Pro",
"same_version": True,
"same_properties": True,
"single_region": True,
"postgres_external": True,
"db_generic_load_balancer": False,
"shared_blob": True,
"application_lb": True,
"failure_domains": {"az-a", "az-b", "az-c"},
}
def validate(t):
errors=[]
if t["edition"] != "Pro": errors.append("HA requires Nexus Repository Pro")
if not t["same_version"]: errors.append("nodes must match outside supported rolling upgrade")
if not t["same_properties"]: errors.append("nexus.properties mismatch")
if not t["single_region"]: errors.append("HA cluster must remain single-region")
if not t["postgres_external"]: errors.append("external PostgreSQL required")
if t["db_generic_load_balancer"]: errors.append("generic PostgreSQL load balancing is unsupported")
if not t["shared_blob"]: errors.append("shared blob storage required")
if not t["application_lb"]: errors.append("application load balancer required")
if len(t["failure_domains"]) < 2: errors.append("nodes share one failure domain")
return errors
errors=validate(topology)
print("GO" if not errors else "NO-GO")
for e in errors: print(" -", e)
Expected: GO. Change
db_generic_load_balancer to True or
shared_blob to False and confirm the
preflight blocks the design.
4. Service simulator: separate application availability from dependency health
state = {
"lb_up": True,
"nodes": {"nx-a": True, "nx-b": True, "nx-c": True},
"db_writer_endpoint_up": True,
"blob_up": True,
}
def artifact_service(s):
live_nodes=[n for n,up in s["nodes"].items() if up]
if not s["lb_up"]: return "UNAVAILABLE: front door"
if not live_nodes: return "UNAVAILABLE: no Nexus node"
if not s["db_writer_endpoint_up"]: return "UNAVAILABLE: PostgreSQL writer"
if not s["blob_up"]: return "DEGRADED: shared blob unavailable"
return "AVAILABLE via " + live_nodes[0]
print(artifact_service(state))
The simulator intentionally requires both database and blob dependencies for artifact service. A load-balancer HTTP health probe is only one input; it cannot replace dependency-native monitoring.
5. Failure injection A — lose one Nexus node
state["nodes"]["nx-a"] = False
print(artifact_service(state))
# Expected: AVAILABLE via nx-b (or another live node)
Expected state: no repository metadata or blob copy occurs. The failed node is simply removed from the serving set. PostgreSQL and blob content remain authoritative.
Evidence: load-balancer target health, remaining node request logs, cluster node status, client success, and unchanged artifact digest.
6. Failure injection B — database failover
state["db_writer_endpoint_up"] = False
print(artifact_service(state))
# Simulate provider promotion / stable endpoint recovery.
state["db_writer_endpoint_up"] = True
print(artifact_service(state))
The Nexus nodes do not elect a PostgreSQL writer. The database service performs replication/promotion and restores a supported writable endpoint. During that interval, Nexus can be unavailable or degraded even though all JVMs are alive.
7. Failure injection C — shared blob degradation
state["blob_up"] = False
print(artifact_service(state))
# Expected: DEGRADED: shared blob unavailable
Do not “fix” this by making each Nexus node write to private local storage. Diagnose the storage service, credentials, network, latency, quotas, and backend health. In a real environment, compare blob-specific metrics and client errors with Nexus logs.
8. Failure injection D — load balancer failure
state["blob_up"] = True
state["lb_up"] = False
print(artifact_service(state))
# Expected: UNAVAILABLE: front door
Three healthy Nexus nodes behind one failed front door still produce an outage. A production design must include load-balancer service redundancy and the DNS/network path in its failure-domain analysis.
9. What the Nexus status endpoint can and cannot prove
| Signal | Useful for | Not sufficient for |
|---|---|---|
/service/rest/v1/status |
Can this Nexus process respond to read requests? | External DB/storage hardware health |
/status/writable |
Can Nexus report read/write readiness and not read-only? | Object-store durability or DB replication lag |
| Load-balancer target health | Which Nexus nodes should receive traffic? | Backup validity |
| PostgreSQL provider metrics | Writer/failover/replication behavior | Nexus authorization/client behavior |
| Blob-provider metrics | Latency, errors, capacity, endpoint health | Repository metadata consistency by itself |
10. Before/after evidence packet
For each injected failure, record: timestamp, failed dependency, Nexus node set, load-balancer target set, DB writer state, blob state, client HTTP result, artifact checksum, and recovery action. This makes causality explicit and prevents a “green dashboard” from becoming the only proof.
11. Small challenge
Modify the fixture so all three Nexus nodes remain healthy but
db_writer_endpoint_up=False. Explain why adding a
fourth Nexus node does not improve availability. Then restore the
database endpoint and fail one Nexus node; explain why that failure
is masked by the application tier.
Knowledge check
What should happen when one Nexus node fails but DB, blobs, LB, and another node remain healthy?
The load balancer should route requests to a remaining active node; no repository-state copy is required.
Who performs PostgreSQL primary promotion?
The PostgreSQL/managed database HA layer, not the Nexus application nodes.
Why is a Nexus /status probe not enough for full HA monitoring?
The status API reports Nexus state and does not replace external database, storage, network, or hardware monitoring.
Why does adding more Nexus nodes not help when the shared DB writer is unavailable?
All nodes depend on the same authoritative database service, so the database remains the blocking dependency.
What is the mandatory lab environment?
A deterministic free simulation/fixture; live multi-node HA is optional and requires a licensed disposable Pro environment.
Summary and next step
The simulator proves that HA availability is the intersection of a healthy front door, at least one active Nexus node, a coherent writable PostgreSQL service, and accessible shared blob content.
Lesson 3 converts those observations into architecture choices and tradeoffs.
Official references and version notes
- Sonatype: High Availability Deployment — current Pro-only clustered application model.
- Sonatype: Requirements for High Availability — same-version nodes, single-region scope, load balancer, shared blob storage, and external PostgreSQL requirements.
- Sonatype: Manual High Availability Deployment — clustered mode, shared storage, and per-node operational state.
- Sonatype: Validating Your HA Deployment — node status, system information, and per-node support evidence.
- Sonatype: Status API — read/writable status endpoints and their monitoring limitations.
- Sonatype: Nodes API — Pro cluster-node inventory and node identities.
- Sonatype: System Requirements — Java 21, PostgreSQL guidance, supported providers, and the explicit pgpool/load-balancing restriction.
- Sonatype: Blob Stores — shared/object storage behavior, health indicators, and group blob-store semantics.
- Sonatype: Storage Planning — group blob stores, storage layout tradeoffs, and performance implications.
- Sonatype: Nexus Repository Professional Features — HA, cloud blob stores, and group blob-store licensing boundaries.
- Sonatype: Rolling Upgrades in High Availability — mixed-version limits, sticky-session requirement during rolling upgrades, schema-finalization behavior, and restore-based rollback after finalization.
- Sonatype: Prepare a Backup — recovery-set consistency across database/configuration and blob content.
- Sonatype: Cross-Region Disaster Recovery — a separate DR pattern rather than an extension of a single-region HA cluster.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.