Chapter 28Lesson 02240–320 min

High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Guided Hands-On Workflow and Core Operations

A real HA cluster requires Pro and external infrastructure, so the mandatory lab models the same state transitions with deterministic fixtures. You will prove the prerequisites, simulate failures one dependency at a time, observe the expected service state, and learn which evidence belongs to Nexus versus the load balancer, PostgreSQL, and blob-storage layers.

SimulationFailure injectionStatus evidenceLoad balancerFailover

Learning objectives

  • Validate the current HA prerequisites before designing a topology.
  • Build a reproducible topology fixture with independent failure domains.
  • Simulate application-node loss, PostgreSQL failover, blob degradation, and load-balancer failure.
  • Collect Nexus and infrastructure evidence without assuming one health signal covers every layer.
  • Explain which failures HA masks automatically and which require operator recovery.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live version-status, feature-matrix, HA, storage, and release-note pages before implementing a real cluster.
License boundary. Nexus Repository high availability is a Pro-only capability. The mandatory chapter path is therefore architecture- and fixture-driven and can be completed without a Pro license, cloud account, managed database, Kubernetes cluster, or multi-node Nexus deployment.
HA is not backup or disaster recovery. Multiple active nodes improve service continuity for selected failures. They do not create a point-in-time recovery copy, undo deletion/corruption, or make a cross-region cluster supported. Chapter 25 remains the recovery foundation.

1. Mandatory free lab versus optional licensed lab

The required path runs only a local Python simulator and static configuration fixtures. It does not launch a fake “HA Nexus” and does not imply Community Edition can form a cluster. If you already have a licensed disposable Pro environment, you may map the same evidence steps to that instance, but the learning objective is the state transition—not cloud/Kubernetes administration.

2. Preflight: reject an invalid architecture before simulation

reference:
  nexus: 3.95.2-01
  java: 21
  editionRequiredForLiveHA: Pro
cluster:
  region: lab-region-1
  nodes:
    - {name: nx-a, failureDomain: az-a}
    - {name: nx-b, failureDomain: az-b}
    - {name: nx-c, failureDomain: az-c}
  sameNexusVersion: true
  sameNexusProperties: true
database:
  type: PostgreSQL
  access: stable-writer-endpoint
  loadBalancedAcrossReadReplicas: false
blob:
  type: shared-object-storage
  reachableFromAllNodes: true
loadBalancer:
  healthProbe: /service/rest/v1/status

Each field corresponds to a current HA gate. A topology checker should fail closed if one is false.

3. Run the topology checker

topology = {
    "edition": "Pro",
    "same_version": True,
    "same_properties": True,
    "single_region": True,
    "postgres_external": True,
    "db_generic_load_balancer": False,
    "shared_blob": True,
    "application_lb": True,
    "failure_domains": {"az-a", "az-b", "az-c"},
}

def validate(t):
    errors=[]
    if t["edition"] != "Pro": errors.append("HA requires Nexus Repository Pro")
    if not t["same_version"]: errors.append("nodes must match outside supported rolling upgrade")
    if not t["same_properties"]: errors.append("nexus.properties mismatch")
    if not t["single_region"]: errors.append("HA cluster must remain single-region")
    if not t["postgres_external"]: errors.append("external PostgreSQL required")
    if t["db_generic_load_balancer"]: errors.append("generic PostgreSQL load balancing is unsupported")
    if not t["shared_blob"]: errors.append("shared blob storage required")
    if not t["application_lb"]: errors.append("application load balancer required")
    if len(t["failure_domains"]) < 2: errors.append("nodes share one failure domain")
    return errors

errors=validate(topology)
print("GO" if not errors else "NO-GO")
for e in errors: print(" -", e)

Expected: GO. Change db_generic_load_balancer to True or shared_blob to False and confirm the preflight blocks the design.

4. Service simulator: separate application availability from dependency health

state = {
  "lb_up": True,
  "nodes": {"nx-a": True, "nx-b": True, "nx-c": True},
  "db_writer_endpoint_up": True,
  "blob_up": True,
}

def artifact_service(s):
    live_nodes=[n for n,up in s["nodes"].items() if up]
    if not s["lb_up"]: return "UNAVAILABLE: front door"
    if not live_nodes: return "UNAVAILABLE: no Nexus node"
    if not s["db_writer_endpoint_up"]: return "UNAVAILABLE: PostgreSQL writer"
    if not s["blob_up"]: return "DEGRADED: shared blob unavailable"
    return "AVAILABLE via " + live_nodes[0]

print(artifact_service(state))

The simulator intentionally requires both database and blob dependencies for artifact service. A load-balancer HTTP health probe is only one input; it cannot replace dependency-native monitoring.

5. Failure injection A — lose one Nexus node

state["nodes"]["nx-a"] = False
print(artifact_service(state))
# Expected: AVAILABLE via nx-b (or another live node)

Expected state: no repository metadata or blob copy occurs. The failed node is simply removed from the serving set. PostgreSQL and blob content remain authoritative.

Evidence: load-balancer target health, remaining node request logs, cluster node status, client success, and unchanged artifact digest.

6. Failure injection B — database failover

state["db_writer_endpoint_up"] = False
print(artifact_service(state))
# Simulate provider promotion / stable endpoint recovery.
state["db_writer_endpoint_up"] = True
print(artifact_service(state))

The Nexus nodes do not elect a PostgreSQL writer. The database service performs replication/promotion and restores a supported writable endpoint. During that interval, Nexus can be unavailable or degraded even though all JVMs are alive.

7. Failure injection C — shared blob degradation

state["blob_up"] = False
print(artifact_service(state))
# Expected: DEGRADED: shared blob unavailable

Do not “fix” this by making each Nexus node write to private local storage. Diagnose the storage service, credentials, network, latency, quotas, and backend health. In a real environment, compare blob-specific metrics and client errors with Nexus logs.

8. Failure injection D — load balancer failure

state["blob_up"] = True
state["lb_up"] = False
print(artifact_service(state))
# Expected: UNAVAILABLE: front door

Three healthy Nexus nodes behind one failed front door still produce an outage. A production design must include load-balancer service redundancy and the DNS/network path in its failure-domain analysis.

9. What the Nexus status endpoint can and cannot prove

Signal Useful for Not sufficient for
/service/rest/v1/status Can this Nexus process respond to read requests? External DB/storage hardware health
/status/writable Can Nexus report read/write readiness and not read-only? Object-store durability or DB replication lag
Load-balancer target health Which Nexus nodes should receive traffic? Backup validity
PostgreSQL provider metrics Writer/failover/replication behavior Nexus authorization/client behavior
Blob-provider metrics Latency, errors, capacity, endpoint health Repository metadata consistency by itself

10. Before/after evidence packet

For each injected failure, record: timestamp, failed dependency, Nexus node set, load-balancer target set, DB writer state, blob state, client HTTP result, artifact checksum, and recovery action. This makes causality explicit and prevents a “green dashboard” from becoming the only proof.

11. Small challenge

Modify the fixture so all three Nexus nodes remain healthy but db_writer_endpoint_up=False. Explain why adding a fourth Nexus node does not improve availability. Then restore the database endpoint and fail one Nexus node; explain why that failure is masked by the application tier.

Knowledge check

What should happen when one Nexus node fails but DB, blobs, LB, and another node remain healthy?

Who performs PostgreSQL primary promotion?

Why is a Nexus /status probe not enough for full HA monitoring?

Why does adding more Nexus nodes not help when the shared DB writer is unavailable?

What is the mandatory lab environment?

Summary and next step

The simulator proves that HA availability is the intersection of a healthy front door, at least one active Nexus node, a coherent writable PostgreSQL service, and accessible shared blob content.

Lesson 3 converts those observations into architecture choices and tradeoffs.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.