Chapter 28Lesson 04240–320 min

High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Diagnostics, Failure Modes, Security, and Performance

HA failures become dangerous when redundancy hides the real dependency. Two JVMs may still share one failed switch; a “healthy” node may be unable to reach storage; a database proxy may silently route writes incorrectly; mixed Nexus versions may linger outside a supported rolling upgrade; and replicated bad data may remain perfectly available. Diagnose the shared state first, not the node count.

DiagnosticsSplit brainStorage latencyMixed modeCorruption

Learning objectives

  • Detect independent-instance/split-state designs that are incorrectly labeled HA.
  • Diagnose storage, database, load-balancer, version, capacity, and dependency failures in order.
  • Explain the limits of Nexus status checks and why per-layer telemetry is required.
  • Distinguish failover from recovery when corrupted state is shared or replicated.
  • Apply the least destructive supported correction without editing database/blob internals.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live version-status, feature-matrix, HA, storage, and release-note pages before implementing a real cluster.
License boundary. Nexus Repository high availability is a Pro-only capability. The mandatory chapter path is therefore architecture- and fixture-driven and can be completed without a Pro license, cloud account, managed database, Kubernetes cluster, or multi-node Nexus deployment.
HA is not backup or disaster recovery. Multiple active nodes improve service continuity for selected failures. They do not create a point-in-time recovery copy, undo deletion/corruption, or make a cross-region cluster supported. Chapter 25 remains the recovery foundation.

1. Diagnostic sequence for an HA incident

  1. Preserve client errors, timestamps, load-balancer target health, node identities/versions, and recent changes.
  2. Confirm every node’s Nexus version, Java runtime, edition/license, and nexus.properties.
  3. Verify the client URL/TLS/authentication path and which node actually served the request.
  4. Check cluster/node status and repository authorization.
  5. Check the PostgreSQL writer endpoint, connectivity, pool saturation, failover state, and latency.
  6. Check shared blob connectivity, latency, quota/capacity, credentials, and provider errors.
  7. Check node-local logs, tasks, JVM heap/direct memory, file handles, and temporary disk.
  8. Compare an artifact digest/metadata request across repeated load-balanced calls.
  9. Apply the smallest supported correction, then verify through the same client path.
  10. If shared state is corrupted, stop treating the incident as “HA failover” and enter the tested recovery runbook.

2. Failure: “two independent Nexus servers are HA”

Symptom: clients intermittently see different repositories/users/components depending on the node.

Cause: each Nexus server has its own database and/or blob store. A load balancer in front does not merge them.

Safe correction: stop write traffic to the invalid topology, preserve both states for forensics, and migrate to one supported clustered state using official migration/recovery procedures. Do not copy database directories or blob files between live servers.

3. Failure: shared filesystem is reachable but too slow

Symptom: package downloads time out, upload latency spikes, cleanup/repair tasks stretch for hours, and JVMs appear underutilized.

Cause: “shared” met the namespace requirement but not the IO/latency requirement.

Evidence: storage latency/IOPS, Nexus request latency, task duration, DB latency, client-cache state. Measure them separately before changing heap or adding nodes.

4. Failure: every node shares one failure domain

Symptom: one rack/host cluster/AZ/network outage removes every Nexus node simultaneously.

Cause: node count increased but failure-domain diversity did not.

Correction: distribute nodes across supported independent hosts/AZs while keeping the cluster in one region and retaining low-latency access to shared DB/blob services.

5. Failure: database requests load-balanced to read replicas

Symptom: write failures, stale reads, transaction anomalies, or unsupported topology errors.

Cause: an operator inserted pgpool/generic SQL load balancing because the application tier is active-active.

Correction: restore the supported PostgreSQL writer/failover endpoint model. Do not “solve” the issue by disabling constraints or forcing writes to replicas.

6. Failure: mixed Nexus versions linger

Normal HA requires the same Nexus version/configuration across nodes. Mixed mode is supported only during an explicitly enabled rolling upgrade. Current guidance says to keep that interval as short as possible; new features remain unavailable until the database schema is finalized, and mixed mode outside an active upgrade is unsupported.

After schema finalization, rollback is restore-based. The database schema upgrade is one-way. Do not simply start old binaries against the finalized database.

7. Failure: failover hides corrupted content

Symptom: every surviving node can serve the same wrong/missing artifact, so uptime dashboards stay green while clients fail checksum or business validation.

Cause: the corruption lives in shared authoritative state. HA replicated/shared the problem correctly.

Correction: preserve evidence, stop harmful writes if needed, identify the consistency point, and restore through Chapter 25’s tested database/blob recovery process.

8. Failure: health endpoint says 200 but storage is failing

Current Status API documentation explicitly says these endpoints report Nexus state and are not hardware checks; they do not replace dedicated checks for external database reachability or disk/storage access. A load balancer should use an appropriate application probe, while monitoring also observes the DB and blob services directly.

9. Performance: add nodes only after locating the bottleneck

Observed bottleneck Why another Nexus node may not help
PostgreSQL latency/pool saturation Every node increases DB concurrency and still depends on the same DB service.
Blob IO latency All nodes read/write the same storage backend.
Upstream proxy latency More nodes do not make the remote registry faster.
Client local cache miss Problem may be client-side, not cluster capacity.
CPU/JVM saturation on one node Additional healthy nodes may help if LB distribution and DB/storage headroom exist.

10. Security-sensitive HA operations

  • Do not expose node management endpoints publicly.
  • Keep database/blob credentials in supported secret mechanisms; never duplicate plaintext secrets in lesson files.
  • Review support ZIPs per node before sharing because logs/configuration can contain sensitive metadata.
  • Do not bypass TLS verification to “fix” cross-node or storage connectivity.
  • Do not edit shared blob files or PostgreSQL rows directly during failover.

11. Intentionally broken topology classifier

bad = {
  "nexus_db": {"nx-a":"db-a", "nx-b":"db-b"},
  "blob": {"nx-a":"/local/a", "nx-b":"/local/b"},
  "same_version": False,
  "rolling_upgrade_active": False,
  "db_mode": "round-robin-writer-and-replica",
}

problems=[]
if len(set(bad["nexus_db"].values())) != 1:
    problems.append("independent databases: not one HA service")
if len(set(bad["blob"].values())) != 1:
    problems.append("private blob stores: content can diverge")
if not bad["same_version"] and not bad["rolling_upgrade_active"]:
    problems.append("unsupported mixed Nexus versions")
if bad["db_mode"] != "stable-writer-endpoint":
    problems.append("unsupported database routing")

for p in problems: print("FAIL:", p)

Expected: four failures. The lesson is not to patch each symptom independently; the topology violates the shared-state contract and must be redesigned before production traffic continues.

Knowledge check

What is the clearest sign that two “HA” nodes are actually independent instances?

Why is pgpool/read-replica load balancing a bad Nexus HA design?

When may Nexus nodes temporarily run different versions?

Why can HA make corrupted content highly available?

What should you check before adding more Nexus nodes for performance?

Summary and next step

HA troubleshooting follows the shared dependency chain: front door → live Nexus nodes → coherent PostgreSQL writer → shared blob state → node-local resources. Failover is useful only while the shared state itself remains correct.

Lesson 5 turns the chapter into a production-style architecture and failure-validation checkpoint.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.