Chapter 28Lesson 05280–360 min

Checkpoint Lab — High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery

The checkpoint lab is an architecture review and failure drill, not a cloud deployment exercise. You will design one regional Nexus service across three failure domains, feed synthetic telemetry into a deterministic evaluator, calculate scenario-specific RTO/RPO from declared assumptions, prove which failures HA masks, and show which incidents require backup or disaster recovery instead.

CheckpointThree domainsSynthetic telemetryRTO/RPOHA vs DR

Learning objectives

  • Design a supported three-failure-domain HA architecture without requiring paid infrastructure to run the lab.
  • Predict client-visible behavior for node, AZ, database, blob, load-balancer, and logical-corruption failures.
  • Calculate RTO/RPO from explicit synthetic telemetry rather than claiming universal guarantees.
  • Validate that artifact identity remains unchanged after infrastructure failover.
  • Produce an evidence matrix that distinguishes HA, redundancy, backup, and cross-region DR.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live version-status, feature-matrix, HA, storage, and release-note pages before implementing a real cluster.
License boundary. Nexus Repository high availability is a Pro-only capability. The mandatory chapter path is therefore architecture- and fixture-driven and can be completed without a Pro license, cloud account, managed database, Kubernetes cluster, or multi-node Nexus deployment.
HA is not backup or disaster recovery. Multiple active nodes improve service continuity for selected failures. They do not create a point-in-time recovery copy, undo deletion/corruption, or make a cross-region cluster supported. Chapter 25 remains the recovery foundation.

1. Scenario contract

You are reviewing a proposed production Nexus Repository service for a large engineering organization. The mandatory exercise is a simulation; no live Pro instance is required.

serviceObjectives:
  normalArtifactAvailability: 99.95-percent target
  maxNodeFailureRTOSeconds: 30
  databaseFailoverRTOMinutes: 5
  databaseFailoverRPOSeconds: 30
  blobEndpointFailoverRTOMinutes: 8
  logicalDeletionRecoveryRPOMinutes: 30
architecture:
  scope: one-region
  failureDomains: [az-a, az-b, az-c]
  nexusNodes: [nx-a, nx-b, nx-c]
  database: managed-postgresql-stable-writer-endpoint
  blob: shared-regional-object-store
  loadBalancer: redundant-managed-application-lb
  backup: independent-db-plus-blob-recovery-set

These numbers are synthetic lab assumptions, not Sonatype guarantees. Your evidence must show whether the design meets them.

2. Target architecture

Three failure domains in one region
flowchart TB
CL[CI + developer clients] --> LB[Regional redundant load balancer]
subgraph R[One region]
subgraph A[AZ / failure domain A]
NA[Nexus nx-a]
end
subgraph B[AZ / failure domain B]
NB[Nexus nx-b]
end
subgraph C[AZ / failure domain C]
NC[Nexus nx-c]
end
LB --> NA
LB --> NB
LB --> NC
NA --> PG[(Managed PostgreSQL writer endpoint)]
NB --> PG
NC --> PG
NA --> BS[(Shared regional object blob storage)]
NB --> BS
NC --> BS
end
PG -. backups / PITR .-> BK[Independent recovery system]
BS -. versioned / point-in-time copy .-> BK

The application nodes span three failure domains, but the cluster remains one regional service. The PostgreSQL and blob services must themselves have tested HA characteristics appropriate to the service objective.

3. Freeze artifact identity before failure injection

import hashlib
artifact=b"learner-example release 28.5\n"
expected=hashlib.sha256(artifact).hexdigest()
print("expected_sha256=", expected)

The same SHA-256 must be observed after application-node/AZ/database infrastructure failover. A changed digest is a content-integrity incident, not a successful availability test.

4. Synthetic telemetry model

assumptions = {
  "node_rto_s": 20,
  "db_failover_rto_s": 180,
  "db_replication_lag_s": 12,
  "blob_failover_rto_s": 240,
  "backup_interval_s": 1800,
}

base = {
  "lb": True,
  "nodes": {"nx-a": True, "nx-b": True, "nx-c": True},
  "db": True,
  "blob": True,
  "content_correct": True,
}

def classify(s):
    if not s["lb"]: return "OUTAGE-front-door"
    if not any(s["nodes"].values()): return "OUTAGE-no-app-node"
    if not s["db"]: return "OUTAGE-database"
    if not s["blob"]: return "OUTAGE-blob"
    if not s["content_correct"]: return "AVAILABLE-BUT-CORRUPT"
    return "AVAILABLE"

5. Predict before running

Scenario Your prediction Primary mechanism
nx-a stops Service remains available LB removes one node
AZ-a fails Service remains available if nx-b/c + shared dependencies retain capacity Failure-domain diversity
DB writer failover Temporary outage/degradation, then recovery PostgreSQL HA layer
Blob endpoint degradation Artifact operations fail/degrade across nodes Storage failover/recovery
Logical deletion Service may be “available” but content is wrong Backup/DR, not HA

6. Execute failure scenarios

from copy import deepcopy

scenarios={}

s=deepcopy(base); s["nodes"]["nx-a"]=False
scenarios["node-loss"]=classify(s)

s=deepcopy(base); s["nodes"]["nx-a"]=False  # represent AZ-a app loss
scenarios["az-a-loss"]=classify(s)

s=deepcopy(base); s["db"]=False
scenarios["db-before-promotion"]=classify(s)
s["db"]=True
scenarios["db-after-promotion"]=classify(s)

s=deepcopy(base); s["blob"]=False
scenarios["blob-outage"]=classify(s)

s=deepcopy(base); s["content_correct"]=False
scenarios["logical-corruption"]=classify(s)

for name,result in scenarios.items(): print(name, "=>", result)

Expected shape: node/AZ loss remain available; DB is unavailable until provider failover completes; blob outage affects the whole artifact path; corruption can remain highly available but wrong.

7. Calculate scenario RTO/RPO from declared assumptions

rows = [
  ("node-loss", assumptions["node_rto_s"], 0, "HA application tier"),
  ("db-failover", assumptions["db_failover_rto_s"], assumptions["db_replication_lag_s"], "DB HA"),
  ("blob-endpoint", assumptions["blob_failover_rto_s"], 0, "storage HA"),
  ("logical-corruption", assumptions["backup_interval_s"], assumptions["backup_interval_s"], "backup/restore"),
]
for name,rto,rpo,control in rows:
    print(f"{name:18} RTO<={rto:4}s RPO<={rpo:4}s via {control}")

The logical-corruption row intentionally uses the backup interval for recovery exposure. More Nexus nodes cannot reduce that RPO.

8. Validation matrix: HA versus backup/DR

Failure HA? Backup/DR? Required evidence
Single Nexus node Primary control Not normally invoked LB target drain + client success
One failure domain/AZ Primary control when design/capacity supports it Not normally invoked Remaining nodes/DB/blob capacity
PostgreSQL primary loss DB HA control Needed if failover cannot recover consistent state Writer promotion + transaction/client validation
Blob service outage Storage HA control Needed for unrecoverable loss/corruption Provider failover + artifact digest
Accidental delete/corruption No Yes Known-good DB/blob restore checkpoint
Whole region disaster Not this HA cluster Cross-region DR/federation pattern DR RPO/RTO and cutover evidence

9. Evidence bundle

Submit these files/records from the exercise:

  • Topology YAML and prerequisite validation output.
  • Failure prediction table written before simulation.
  • Scenario output with timestamps and dependency state.
  • Expected and post-failover artifact SHA-256.
  • RTO/RPO calculation with assumptions clearly labeled synthetic.
  • HA-versus-DR matrix.
  • Security note showing credentials/real endpoints were not used.
  • Runbook paragraph describing when the incident crosses from failover into restore.

10. Cleanup/rollback

The mandatory lab creates only local text/JSON/YAML/Python fixture state. Delete that disposable folder after preserving the evidence packet. If you used an optional licensed Pro cluster, remove only resources created for the lab through supported platform/Nexus mechanisms, and restore the lab environment from its documented checkpoint if you altered shared storage, database, load-balancer, or cluster settings.

11. Final operating statement

A production Nexus artifact platform is highly available only when the application tier, load balancer, database, blob storage, network, secrets, and capacity are designed as one tested service. It is resilient only when those HA mechanisms are paired with tested point-in-time recovery and explicit disaster scenarios.

Knowledge check

What should happen to artifact SHA-256 after a node or DB infrastructure failover?

Which control reduces RPO for accidental deletion?

Why label the checkpoint RTO/RPO values synthetic?

Can one HA cluster span regions to solve regional disaster?

What is the bridge to Chapter 29?

Summary and next step

The checkpoint distinguishes infrastructure failover from data recovery and ties every availability claim to explicit telemetry, RTO/RPO assumptions, and artifact-integrity evidence.

Chapter 29 now builds the observability layer required to operate this architecture: logs, metrics, support ZIPs, request auditing, performance tuning, and capacity diagnostics.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.