Checkpoint Lab — High Availability, Resiliency, Shared Storage, Blob Store Groups, Node Coordination, and Failure Recovery
The checkpoint lab is an architecture review and failure drill, not a cloud deployment exercise. You will design one regional Nexus service across three failure domains, feed synthetic telemetry into a deterministic evaluator, calculate scenario-specific RTO/RPO from declared assumptions, prove which failures HA masks, and show which incidents require backup or disaster recovery instead.
Learning objectives
- Design a supported three-failure-domain HA architecture without requiring paid infrastructure to run the lab.
- Predict client-visible behavior for node, AZ, database, blob, load-balancer, and logical-corruption failures.
- Calculate RTO/RPO from explicit synthetic telemetry rather than claiming universal guarantees.
- Validate that artifact identity remains unchanged after infrastructure failover.
- Produce an evidence matrix that distinguishes HA, redundancy, backup, and cross-region DR.
1. Scenario contract
You are reviewing a proposed production Nexus Repository service for a large engineering organization. The mandatory exercise is a simulation; no live Pro instance is required.
serviceObjectives:
normalArtifactAvailability: 99.95-percent target
maxNodeFailureRTOSeconds: 30
databaseFailoverRTOMinutes: 5
databaseFailoverRPOSeconds: 30
blobEndpointFailoverRTOMinutes: 8
logicalDeletionRecoveryRPOMinutes: 30
architecture:
scope: one-region
failureDomains: [az-a, az-b, az-c]
nexusNodes: [nx-a, nx-b, nx-c]
database: managed-postgresql-stable-writer-endpoint
blob: shared-regional-object-store
loadBalancer: redundant-managed-application-lb
backup: independent-db-plus-blob-recovery-set
These numbers are synthetic lab assumptions, not Sonatype guarantees. Your evidence must show whether the design meets them.
2. Target architecture
flowchart TB CL[CI + developer clients] --> LB[Regional redundant load balancer] subgraph R[One region] subgraph A[AZ / failure domain A] NA[Nexus nx-a] end subgraph B[AZ / failure domain B] NB[Nexus nx-b] end subgraph C[AZ / failure domain C] NC[Nexus nx-c] end LB --> NA LB --> NB LB --> NC NA --> PG[(Managed PostgreSQL writer endpoint)] NB --> PG NC --> PG NA --> BS[(Shared regional object blob storage)] NB --> BS NC --> BS end PG -. backups / PITR .-> BK[Independent recovery system] BS -. versioned / point-in-time copy .-> BK
The application nodes span three failure domains, but the cluster remains one regional service. The PostgreSQL and blob services must themselves have tested HA characteristics appropriate to the service objective.
3. Freeze artifact identity before failure injection
import hashlib
artifact=b"learner-example release 28.5\n"
expected=hashlib.sha256(artifact).hexdigest()
print("expected_sha256=", expected)
The same SHA-256 must be observed after application-node/AZ/database infrastructure failover. A changed digest is a content-integrity incident, not a successful availability test.
4. Synthetic telemetry model
assumptions = {
"node_rto_s": 20,
"db_failover_rto_s": 180,
"db_replication_lag_s": 12,
"blob_failover_rto_s": 240,
"backup_interval_s": 1800,
}
base = {
"lb": True,
"nodes": {"nx-a": True, "nx-b": True, "nx-c": True},
"db": True,
"blob": True,
"content_correct": True,
}
def classify(s):
if not s["lb"]: return "OUTAGE-front-door"
if not any(s["nodes"].values()): return "OUTAGE-no-app-node"
if not s["db"]: return "OUTAGE-database"
if not s["blob"]: return "OUTAGE-blob"
if not s["content_correct"]: return "AVAILABLE-BUT-CORRUPT"
return "AVAILABLE"
5. Predict before running
| Scenario | Your prediction | Primary mechanism |
|---|---|---|
| nx-a stops | Service remains available | LB removes one node |
| AZ-a fails | Service remains available if nx-b/c + shared dependencies retain capacity | Failure-domain diversity |
| DB writer failover | Temporary outage/degradation, then recovery | PostgreSQL HA layer |
| Blob endpoint degradation | Artifact operations fail/degrade across nodes | Storage failover/recovery |
| Logical deletion | Service may be “available” but content is wrong | Backup/DR, not HA |
6. Execute failure scenarios
from copy import deepcopy
scenarios={}
s=deepcopy(base); s["nodes"]["nx-a"]=False
scenarios["node-loss"]=classify(s)
s=deepcopy(base); s["nodes"]["nx-a"]=False # represent AZ-a app loss
scenarios["az-a-loss"]=classify(s)
s=deepcopy(base); s["db"]=False
scenarios["db-before-promotion"]=classify(s)
s["db"]=True
scenarios["db-after-promotion"]=classify(s)
s=deepcopy(base); s["blob"]=False
scenarios["blob-outage"]=classify(s)
s=deepcopy(base); s["content_correct"]=False
scenarios["logical-corruption"]=classify(s)
for name,result in scenarios.items(): print(name, "=>", result)
Expected shape: node/AZ loss remain available; DB is unavailable until provider failover completes; blob outage affects the whole artifact path; corruption can remain highly available but wrong.
7. Calculate scenario RTO/RPO from declared assumptions
rows = [
("node-loss", assumptions["node_rto_s"], 0, "HA application tier"),
("db-failover", assumptions["db_failover_rto_s"], assumptions["db_replication_lag_s"], "DB HA"),
("blob-endpoint", assumptions["blob_failover_rto_s"], 0, "storage HA"),
("logical-corruption", assumptions["backup_interval_s"], assumptions["backup_interval_s"], "backup/restore"),
]
for name,rto,rpo,control in rows:
print(f"{name:18} RTO<={rto:4}s RPO<={rpo:4}s via {control}")
The logical-corruption row intentionally uses the backup interval for recovery exposure. More Nexus nodes cannot reduce that RPO.
8. Validation matrix: HA versus backup/DR
| Failure | HA? | Backup/DR? | Required evidence |
|---|---|---|---|
| Single Nexus node | Primary control | Not normally invoked | LB target drain + client success |
| One failure domain/AZ | Primary control when design/capacity supports it | Not normally invoked | Remaining nodes/DB/blob capacity |
| PostgreSQL primary loss | DB HA control | Needed if failover cannot recover consistent state | Writer promotion + transaction/client validation |
| Blob service outage | Storage HA control | Needed for unrecoverable loss/corruption | Provider failover + artifact digest |
| Accidental delete/corruption | No | Yes | Known-good DB/blob restore checkpoint |
| Whole region disaster | Not this HA cluster | Cross-region DR/federation pattern | DR RPO/RTO and cutover evidence |
9. Evidence bundle
Submit these files/records from the exercise:
- Topology YAML and prerequisite validation output.
- Failure prediction table written before simulation.
- Scenario output with timestamps and dependency state.
- Expected and post-failover artifact SHA-256.
- RTO/RPO calculation with assumptions clearly labeled synthetic.
- HA-versus-DR matrix.
- Security note showing credentials/real endpoints were not used.
- Runbook paragraph describing when the incident crosses from failover into restore.
10. Cleanup/rollback
The mandatory lab creates only local text/JSON/YAML/Python fixture state. Delete that disposable folder after preserving the evidence packet. If you used an optional licensed Pro cluster, remove only resources created for the lab through supported platform/Nexus mechanisms, and restore the lab environment from its documented checkpoint if you altered shared storage, database, load-balancer, or cluster settings.
11. Final operating statement
A production Nexus artifact platform is highly available only when the application tier, load balancer, database, blob storage, network, secrets, and capacity are designed as one tested service. It is resilient only when those HA mechanisms are paired with tested point-in-time recovery and explicit disaster scenarios.
Knowledge check
What should happen to artifact SHA-256 after a node or DB infrastructure failover?
It should remain identical. A changed digest is an integrity incident, not acceptable failover.
Which control reduces RPO for accidental deletion?
Independent tested backup/point-in-time recovery, not additional Nexus application nodes.
Why label the checkpoint RTO/RPO values synthetic?
They come from the exercise assumptions; real objectives depend on measured infrastructure/provider behavior and tested runbooks.
Can one HA cluster span regions to solve regional disaster?
Current HA guidance is single-region/data-center scoped; cross-region continuity uses a separate DR/federation pattern.
What is the bridge to Chapter 29?
HA claims require evidence from logs, metrics, node/support data, request traces, and capacity telemetry—the observability topics of Chapter 29.
Summary and next step
The checkpoint distinguishes infrastructure failover from data recovery and ties every availability claim to explicit telemetry, RTO/RPO assumptions, and artifact-integrity evidence.
Chapter 29 now builds the observability layer required to operate this architecture: logs, metrics, support ZIPs, request auditing, performance tuning, and capacity diagnostics.
Official references and version notes
- Sonatype: High Availability Deployment — current Pro-only clustered application model.
- Sonatype: Requirements for High Availability — same-version nodes, single-region scope, load balancer, shared blob storage, and external PostgreSQL requirements.
- Sonatype: Manual High Availability Deployment — clustered mode, shared storage, and per-node operational state.
- Sonatype: Validating Your HA Deployment — node status, system information, and per-node support evidence.
- Sonatype: Status API — read/writable status endpoints and their monitoring limitations.
- Sonatype: Nodes API — Pro cluster-node inventory and node identities.
- Sonatype: System Requirements — Java 21, PostgreSQL guidance, supported providers, and the explicit pgpool/load-balancing restriction.
- Sonatype: Blob Stores — shared/object storage behavior, health indicators, and group blob-store semantics.
- Sonatype: Storage Planning — group blob stores, storage layout tradeoffs, and performance implications.
- Sonatype: Nexus Repository Professional Features — HA, cloud blob stores, and group blob-store licensing boundaries.
- Sonatype: Rolling Upgrades in High Availability — mixed-version limits, sticky-session requirement during rolling upgrades, schema-finalization behavior, and restore-based rollback after finalization.
- Sonatype: Prepare a Backup — recovery-set consistency across database/configuration and blob content.
- Sonatype: Cross-Region Disaster Recovery — a separate DR pattern rather than an extension of a single-region HA cluster.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.