High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery: Diagnostics, Failure Modes, Security, and Performance
HA failures become dangerous when redundancy hides the real dependency. Two JVMs may still share one failed switch; a “healthy” node may be unable to reach storage; a database proxy may silently route writes incorrectly; mixed Nexus versions may linger outside a supported rolling upgrade; and replicated bad data may remain perfectly available. Diagnose the shared state first, not the node count.
Learning objectives
- Detect independent-instance/split-state designs that are incorrectly labeled HA.
- Diagnose storage, database, load-balancer, version, capacity, and dependency failures in order.
- Explain the limits of Nexus status checks and why per-layer telemetry is required.
- Distinguish failover from recovery when corrupted state is shared or replicated.
- Apply the least destructive supported correction without editing database/blob internals.
1. Diagnostic sequence for an HA incident
- Preserve client errors, timestamps, load-balancer target health, node identities/versions, and recent changes.
-
Confirm every node’s Nexus version, Java runtime, edition/license,
and
nexus.properties. - Verify the client URL/TLS/authentication path and which node actually served the request.
- Check cluster/node status and repository authorization.
- Check the PostgreSQL writer endpoint, connectivity, pool saturation, failover state, and latency.
- Check shared blob connectivity, latency, quota/capacity, credentials, and provider errors.
- Check node-local logs, tasks, JVM heap/direct memory, file handles, and temporary disk.
- Compare an artifact digest/metadata request across repeated load-balanced calls.
- Apply the smallest supported correction, then verify through the same client path.
- If shared state is corrupted, stop treating the incident as “HA failover” and enter the tested recovery runbook.
2. Failure: “two independent Nexus servers are HA”
Symptom: clients intermittently see different repositories/users/components depending on the node.
Cause: each Nexus server has its own database and/or blob store. A load balancer in front does not merge them.
Safe correction: stop write traffic to the invalid topology, preserve both states for forensics, and migrate to one supported clustered state using official migration/recovery procedures. Do not copy database directories or blob files between live servers.
3. Failure: shared filesystem is reachable but too slow
Symptom: package downloads time out, upload latency spikes, cleanup/repair tasks stretch for hours, and JVMs appear underutilized.
Cause: “shared” met the namespace requirement but not the IO/latency requirement.
Evidence: storage latency/IOPS, Nexus request latency, task duration, DB latency, client-cache state. Measure them separately before changing heap or adding nodes.
4. Failure: every node shares one failure domain
Symptom: one rack/host cluster/AZ/network outage removes every Nexus node simultaneously.
Cause: node count increased but failure-domain diversity did not.
Correction: distribute nodes across supported independent hosts/AZs while keeping the cluster in one region and retaining low-latency access to shared DB/blob services.
5. Failure: database requests load-balanced to read replicas
Symptom: write failures, stale reads, transaction anomalies, or unsupported topology errors.
Cause: an operator inserted pgpool/generic SQL load balancing because the application tier is active-active.
Correction: restore the supported PostgreSQL writer/failover endpoint model. Do not “solve” the issue by disabling constraints or forcing writes to replicas.
6. Failure: mixed Nexus versions linger
Normal HA requires the same Nexus version/configuration across nodes. Mixed mode is supported only during an explicitly enabled rolling upgrade. Current guidance says to keep that interval as short as possible; new features remain unavailable until the database schema is finalized, and mixed mode outside an active upgrade is unsupported.
7. Failure: failover hides corrupted content
Symptom: every surviving node can serve the same wrong/missing artifact, so uptime dashboards stay green while clients fail checksum or business validation.
Cause: the corruption lives in shared authoritative state. HA replicated/shared the problem correctly.
Correction: preserve evidence, stop harmful writes if needed, identify the consistency point, and restore through Chapter 25’s tested database/blob recovery process.
8. Failure: health endpoint says 200 but storage is failing
Current Status API documentation explicitly says these endpoints report Nexus state and are not hardware checks; they do not replace dedicated checks for external database reachability or disk/storage access. A load balancer should use an appropriate application probe, while monitoring also observes the DB and blob services directly.
9. Performance: add nodes only after locating the bottleneck
| Observed bottleneck | Why another Nexus node may not help |
|---|---|
| PostgreSQL latency/pool saturation | Every node increases DB concurrency and still depends on the same DB service. |
| Blob IO latency | All nodes read/write the same storage backend. |
| Upstream proxy latency | More nodes do not make the remote registry faster. |
| Client local cache miss | Problem may be client-side, not cluster capacity. |
| CPU/JVM saturation on one node | Additional healthy nodes may help if LB distribution and DB/storage headroom exist. |
10. Security-sensitive HA operations
- Do not expose node management endpoints publicly.
- Keep database/blob credentials in supported secret mechanisms; never duplicate plaintext secrets in lesson files.
- Review support ZIPs per node before sharing because logs/configuration can contain sensitive metadata.
- Do not bypass TLS verification to “fix” cross-node or storage connectivity.
- Do not edit shared blob files or PostgreSQL rows directly during failover.
11. Intentionally broken topology classifier
bad = {
"nexus_db": {"nx-a":"db-a", "nx-b":"db-b"},
"blob": {"nx-a":"/local/a", "nx-b":"/local/b"},
"same_version": False,
"rolling_upgrade_active": False,
"db_mode": "round-robin-writer-and-replica",
}
problems=[]
if len(set(bad["nexus_db"].values())) != 1:
problems.append("independent databases: not one HA service")
if len(set(bad["blob"].values())) != 1:
problems.append("private blob stores: content can diverge")
if not bad["same_version"] and not bad["rolling_upgrade_active"]:
problems.append("unsupported mixed Nexus versions")
if bad["db_mode"] != "stable-writer-endpoint":
problems.append("unsupported database routing")
for p in problems: print("FAIL:", p)
Expected: four failures. The lesson is not to patch each symptom independently; the topology violates the shared-state contract and must be redesigned before production traffic continues.
Knowledge check
What is the clearest sign that two “HA” nodes are actually independent instances?
Requests can observe different repository/security/content state because the nodes use different databases or blob stores.
Why is pgpool/read-replica load balancing a bad Nexus HA design?
Current Sonatype guidance does not support it; Nexus requires a coherent writable PostgreSQL endpoint with database-layer failover.
When may Nexus nodes temporarily run different versions?
Only during a supported, explicitly enabled rolling upgrade; mixed mode outside that window is unsupported.
Why can HA make corrupted content highly available?
All nodes share the same authoritative database/blob state, so they can all serve the same corruption. Recovery requires backup/DR.
What should you check before adding more Nexus nodes for performance?
Measure whether the bottleneck is actually application CPU/JVM versus shared DB, blob IO, network/upstream, tasks, or client-cache behavior.
Summary and next step
HA troubleshooting follows the shared dependency chain: front door → live Nexus nodes → coherent PostgreSQL writer → shared blob state → node-local resources. Failover is useful only while the shared state itself remains correct.
Lesson 5 turns the chapter into a production-style architecture and failure-validation checkpoint.
Official references and version notes
- Sonatype: High Availability Deployment — current Pro-only clustered application model.
- Sonatype: Requirements for High Availability — same-version nodes, single-region scope, load balancer, shared blob storage, and external PostgreSQL requirements.
- Sonatype: Manual High Availability Deployment — clustered mode, shared storage, and per-node operational state.
- Sonatype: Validating Your HA Deployment — node status, system information, and per-node support evidence.
- Sonatype: Status API — read/writable status endpoints and their monitoring limitations.
- Sonatype: Nodes API — Pro cluster-node inventory and node identities.
- Sonatype: System Requirements — Java 21, PostgreSQL guidance, supported providers, and the explicit pgpool/load-balancing restriction.
- Sonatype: Blob Stores — shared/object storage behavior, health indicators, and group blob-store semantics.
- Sonatype: Storage Planning — group blob stores, storage layout tradeoffs, and performance implications.
- Sonatype: Nexus Repository Professional Features — HA, cloud blob stores, and group blob-store licensing boundaries.
- Sonatype: Rolling Upgrades in High Availability — mixed-version limits, sticky-session requirement during rolling upgrades, schema-finalization behavior, and restore-based rollback after finalization.
- Sonatype: Prepare a Backup — recovery-set consistency across database/configuration and blob content.
- Sonatype: Cross-Region Disaster Recovery — a separate DR pattern rather than an extension of a single-region HA cluster.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.