Translate recovery objectives into quorum, archive cadence, dependency ordering, routing, key, schema, stream, and derived-store decisions, then measure the resulting recovery timeline.
RPO/RTO, Regional Failure, Quorum Loss, Restore Ordering, and Dependency Recovery
AtlasMart now turns backups and topology into measurable RPO/RTO, quorum-loss behavior, and a dependency-aware recovery sequence that includes identity, keys, metadata, streams, and routing.
Translate RPO and RTO into archive cadence, topology, spare capacity, runbook, and dependency decisions.
Model region loss and quorum loss without assuming surviving nodes can automatically accept writes.
Order recovery across identity, keys, metadata, schema, primary data, streams, derived views, and routing.
Measure whether AtlasMart actually meets its recovery objectives.
1. RPO and RTO are measured system properties, not slide-deck promises
Recovery point objective (RPO) bounds acceptable data loss measured in time or business state; recovery time objective (RTO) bounds acceptable restoration time. An RPO of five minutes influences WAL/change-log archival, snapshot/incremental frequency, cross-region replication, and validation. An RTO of twenty minutes influences restore throughput, automation, warm capacity, dependency readiness, DNS/routing, and operator practice.
Objectives must name the service scope. “Database recovered” can still leave AtlasMart unavailable if identity, encryption keys, schema metadata, routing, or downstream event processing is missing.
2. Regional failure and quorum loss
Topology determines which failures can preserve quorum. A five-voter metadata service spread 3+2 across two regions loses its majority when the three-voter region disappears; the surviving two nodes should not simply self-declare authority if that violates the consensus protocol. Recovery may require restoring voters, invoking a documented reconfiguration procedure, or failing over to an independently prepared control plane.
A manual “force the survivors writable” procedure can create conflicting authority if the supposedly failed side is merely partitioned. Recovery runbooks must distinguish confirmed destruction from communication loss and use product-supported fencing/reconfiguration semantics.
3. Dependency ordering is part of RTO
Common dependencies include identity/break-glass credentials, key management, cluster metadata, schema/configuration, primary data, change streams, derived search/cache/warehouse state, and traffic routing. Some can recover in parallel, but only after prerequisites. A derived view restored before its source schema/key state may fail or, worse, appear healthy while serving semantically invalid data.
4. AtlasMart lab: compute RPO and the critical recovery path
Minute values below are deterministic scenario inputs, not vendor benchmarks. Change them to model your own objectives and dependency graph.
from collections import defaultdict, deque
RPO_TARGET = 5
RTO_TARGET = 20
failure_minute = 100
last_archived_minute = 96
print("RPO evidence: failure at", failure_minute, "last durable archive at", last_archived_minute,
"=> possible loss", failure_minute-last_archived_minute, "minutes")
print("RPO objective met:", failure_minute-last_archived_minute <= RPO_TARGET)
print("\nREGIONAL / QUORUM FAILURE")
voters = {"a1":"east","a2":"east","a3":"east","b1":"west","b2":"west"}
lost_region="east"
survivors=[n for n,r in voters.items() if r != lost_region]
majority=len(voters)//2+1
print("survivors:", survivors, "need majority:", majority, "quorum available:", len(survivors)>=majority)
steps = {
"identity": {"mins":3, "deps":[]},
"keys": {"mins":4, "deps":["identity"]},
"metadata": {"mins":3, "deps":["identity","keys"]},
"schema": {"mins":2, "deps":["metadata"]},
"primary_data": {"mins":8, "deps":["schema","keys"]},
"streams": {"mins":3, "deps":["primary_data"]},
"derived_views": {"mins":4, "deps":["primary_data","streams"]},
"routing": {"mins":2, "deps":["primary_data","identity"]},
}
def topo_order():
indeg={k:0 for k in steps}; out=defaultdict(list)
for k,v in steps.items():
for d in v["deps"]: indeg[k]+=1; out[d].append(k)
q=deque(sorted(k for k,v in indeg.items() if v==0)); order=[]
while q:
x=q.popleft(); order.append(x)
for y in sorted(out[x]):
indeg[y]-=1
if indeg[y]==0: q.append(y)
return order
order=topo_order()
print("\nSAFE DEPENDENCY ORDER:", order)
finish={}
for step in order:
start=max([finish[d] for d in steps[step]["deps"]], default=0)
finish[step]=start+steps[step]["mins"]
print(f"{step:14} start={start:02d} finish={finish[step]:02d}")
actual_rto=max(finish.values())
print("actual RTO:", actual_rto, "minutes; objective:", RTO_TARGET, "met:", actual_rto<=RTO_TARGET)
print("\nWRONG ORDER EXAMPLE")
print("starting derived_views before keys/schema/primary_data => FAIL: dependencies unavailable")
print("recovery plan must include DNS/routing, identity, keys, schema, streams, and derived stores, not only database bytes")
The archived-log gap is four minutes, meeting the five-minute RPO. Losing the three-voter region leaves only two of five voters, so quorum is unavailable. The dependency schedule exposes the actual critical path and whether the twenty-minute RTO is achievable.
5. Recovery ordering beyond the database
DNS, load balancers, certificates, secret distribution, service discovery, KMS, CDC connectors, message brokers, search indexes, and caches can all extend real RTO. Define which components are authoritative and which are rebuildable. Rebuild derived views only after authoritative source state and replay positions are trustworthy; otherwise a fast recovery can publish stale or duplicated business state.
6. Production judgment
Track last recoverable point, archive lag, backup-copy region/account, quorum topology, spare restore capacity, key access, dependency readiness, restore bandwidth, and human decision time. Test region/quorum scenarios safely in staging or simulation. The next lesson focuses on the evidence after bytes are restored: integrity, reconciliation, and repair.
Check your understanding
- What does RPO measure?
- What does RTO measure?
- Why might two surviving nodes in a five-voter system remain unavailable?
- Why include keys and identity in the recovery dependency graph?
- Why can a database restore complete before the service RTO is complete?
Review the answers
1. The maximum acceptable data loss, often expressed as time between the failure and the latest recoverable business state.
2. The maximum acceptable time to restore the required service capability after disruption.
3. Two is not a majority of five; accepting authority without the protocol’s safe recovery/reconfiguration rules risks conflicting decisions.
4. Encrypted data and administrative recovery steps are unusable if operators/services cannot authenticate or decrypt required artifacts.
5. Routing, streams, derived stores, application dependencies, validation, and traffic cutover can still be pending.
References
Foundational claims use standards, specifications, primary research, or current official documentation where practical. Product references are optional implementation anchors; the mandatory labs are vendor-neutral.
- NIST SP 800-34 Rev. 1 — Contingency-planning guidance covering recovery strategies, testing, training, and plan maintenance.
- PostgreSQL 18 — Continuous Archiving and PITR — Current WAL/base-backup recovery mechanisms useful for reasoning about recoverable points.
- PostgreSQL 18 — High Availability, Load Balancing, and Replication — Official HA/replication context for distinguishing failover from backup recovery.
- Raft paper — Primary reference for majority/quorum-based replicated control-plane agreement assumptions.