Turn recovery design into an executable runbook with decision thresholds, owners, validation, cutover, rollback, communications, measured RPO/RTO, and concrete follow-up actions.
Write a Disaster-Recovery Runbook and Prove It with Game-Day Exercises
Finally AtlasMart runs the plan as a safe game day, measures RPO/RTO, fails its first RTO target, records bottlenecks, and proves the improved runbook on a second exercise.
Write a disaster-recovery runbook with decision thresholds, owners, exact steps, validation, cutover, rollback, and communications.
Inject a safe deterministic AtlasMart failure and measure actual RPO/RTO instead of declaring the exercise successful by intention.
Record failed objectives and convert them into owned improvements.
Use game-day evidence as input to observability and capacity planning.
1. A runbook is executable operational logic
A disaster-recovery (DR) runbook identifies the incident trigger, who has authority to declare disaster, which automation must stop, which data source is authoritative, the last recoverable point, exact restore/failover steps, key/credential retrieval, validation commands/queries, traffic-cutover criteria, rollback conditions, communications, and escalation contacts. Vague steps such as “restore the database” hide the dependencies that dominate real RTO.
2. Decision thresholds and evidence gates
Define the observations required before destructive or irreversible actions. For example, do not promote an isolated stale copy merely because the primary endpoint is unreachable; first classify network partition vs confirmed regional destruction, inspect quorum/control-plane state, fence old writers where necessary, and identify the recovery point. Each gate should emit an auditable decision record.
Traffic cutover is not the end. Define what metrics/invariants trigger rollback, whether the old site may safely accept writes, how to prevent dual writers, and how new recovery-site writes would be reconciled if you reverse course.
3. Safe game days measure the plan
A game day injects a controlled failure into a disposable/staging/simulated environment, follows the runbook, records timestamps and evidence, and compares measured RPO/RTO with objectives. Avoid performing destructive tests against production merely to make an exercise “realistic.” Mature programs can use progressively higher-fidelity experiments with explicit blast-radius controls.
4. AtlasMart lab: the first exercise misses RTO
This program changes no services. It executes a deterministic timeline so you can practice measuring recovery and recording follow-up changes.
from dataclasses import dataclass
@dataclass
class Step:
name: str
mins: int
result: str
RPO_OBJECTIVE=5
RTO_OBJECTIVE=20
failure_time=100
last_recoverable=97
runbook=[
Step("declare incident and freeze risky automation",2,"owner assigned; changes frozen"),
Step("confirm identity/key access",4,"break-glass identity and backup key usable"),
Step("restore control metadata/schema",3,"topology/schema validated"),
Step("restore base backup + replay logs",8,"primary data recovered to minute 97"),
Step("run invariants/checksums",4,"validation passed"),
Step("cut traffic to recovered service",3,"synthetic reads/writes pass"),
]
def run(steps):
elapsed=0
evidence=[]
for s in steps:
elapsed += s.mins
evidence.append((elapsed,s.name,s.result))
return elapsed,evidence
actual_rto,evidence=run(runbook)
actual_rpo=failure_time-last_recoverable
print("GAME DAY 1")
for row in evidence: print(row)
print("actual RPO:",actual_rpo,"objective:",RPO_OBJECTIVE,"met:",actual_rpo<=RPO_OBJECTIVE)
print("actual RTO:",actual_rto,"objective:",RTO_OBJECTIVE,"met:",actual_rto<=RTO_OBJECTIVE)
print("result: exercise FAILED its RTO objective even though recovery eventually worked")
print("\nFOLLOW-UP CHANGES")
print("1 pre-stage tested break-glass/key access; 2 automate metadata/schema restore; 3 parallelize validation where safe")
improved=[
Step("declare incident and freeze risky automation",2,"owner assigned; changes frozen"),
Step("confirm pre-tested identity/key access",1,"key path immediately verified"),
Step("restore control metadata/schema",2,"automated checksum validated"),
Step("restore base backup + replay logs",8,"primary data recovered to minute 97"),
Step("run parallel invariants/checksums",2,"validation passed"),
Step("cut traffic to recovered service",2,"synthetic reads/writes pass"),
]
rto2,evidence2=run(improved)
print("\nGAME DAY 2")
for row in evidence2: print(row)
print("actual RPO:",actual_rpo,"minutes")
print("actual RTO:",rto2,"minutes; objective met:",rto2<=RTO_OBJECTIVE)
print("record remaining risks, owners, due dates, and rollback/cutover evidence after every exercise")
Game Day 1 recovers data within the modeled three-minute RPO but requires 24 minutes, missing the 20-minute RTO. The follow-up actions pre-test key access, automate metadata/schema restore, and parallelize safe validation; Game Day 2 completes in 17 minutes. Improvement comes from evidence, not from redefining success.
5. Runbook content checklist
Include incident classification, authoritative status page/channel, system/data owners, backup inventory and locations, credentials/keys, dependency graph, commands or automation references, expected outputs, integrity checks, synthetic transactions, stream offsets, DNS/routing changes, fencing, cutover approval, rollback, stakeholder/customer communication, evidence storage, and post-incident follow-ups. Keep secrets out of the document; reference controlled secret-retrieval procedures instead.
6. Production judgment and bridge to observability
Schedule exercises after architecture, version, topology, backup, KMS, routing, or staffing changes. Measure human decision latency, restore throughput, archive lag, validation duration, traffic warm-up, and post-cutover error/freshness. These metrics bridge directly into Chapter 23: observability, capacity, benchmarking, tail latency, and headroom determine whether the next real recovery can meet its objectives.
Check your understanding
- What makes a DR runbook different from a high-level policy?
- Why can a game day be considered unsuccessful even when service eventually recovers?
- Why include rollback criteria before cutover?
- What should happen after a failed RTO exercise?
- How does DR connect to observability and capacity?
Review the answers
1. It contains executable decisions, owners, dependencies, evidence gates, validation, cutover, rollback, and communications rather than only objectives.
2. If measured RPO/RTO, integrity, or safety objectives are missed, the exercise has exposed a real design/runbook gap.
3. Because the recovered site may fail validation under real traffic and reversing direction can create dual-writer or reconciliation hazards.
4. Record the bottlenecks, assign owners/due dates, change architecture/runbook/automation, and re-test rather than lowering the objective without business justification.
5. Recovery depends on measurable archive lag, restore throughput, saturation, validation latency, routing state, and spare capacity/headroom.
References
Foundational claims use standards, specifications, primary research, or current official documentation where practical. Product references are optional implementation anchors; the mandatory labs are vendor-neutral.
- NIST SP 800-34 Rev. 1 — Contingency-planning lifecycle including testing, training, exercises, and maintenance.
- Google SRE Workbook — Non-Abstract Large System Design — Operational design reasoning around dependencies, capacity, failures, and recovery planning.
- PostgreSQL 18 — Backup and Restore — Current recovery mechanisms used as an optional implementation anchor.
- Apache Cassandra — Repair — Current repair behavior and operational cost considerations relevant after partial failure/recovery.