Chapter 25 · Snapshots, Backups, Commitlog Archiving, Restore, and Disaster Recovery

Perform a Multi-Node Restore Drill and Measure RPO / RTO, Repair Needs, and Application Cutover

Run a reversible multi-node restore drill, verify replica/data state, cut the client over, and measure evidence-based RPO/RTO.

Advanced140–200 minutesMulti-node DR acceptance drillApache Cassandra 5.0.9 · cqlsh/nodetool · Docker Desktop/WSL containers · RF=3 · LOCAL_QUORUM · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's runbook is not accepted until a team can start with a cataloged backup, detect a bad copy, rebuild a multi-node target, verify Cassandra convergence, point an application at it, and calculate the actual RPO/RTO. This lesson is the chapter acceptance test.

01

Run a reversible multi-node disaster-recovery drill without deleting the original source volumes.

02

Detect missing/corrupted backup artifacts before restore using catalog/checksum verification.

03

Restore schema and SSTables into a new RF=3 cluster, then verify rows and anti-entropy state.

04

Measure RPO and RTO from explicit timestamps and explain what is included in each.

05

Define application cutover, rollback, security/config dependencies, and post-restore ownership.

Chapter 25 lab baseline

The mandatory labs use Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 container image. The normal source cluster keeps the course conventions: cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy with replication factor (RF) 3, and LOCAL_QUORUM for the chapter's verification reads/writes. Tables explicitly use UnifiedCompactionStrategy (UCS); authentication, client/internode TLS, and remote JMX remain disabled only inside this isolated local learning topology. The chapter keyspace is atlasmart_backup. No managed service, paid backup product, or cloud account is required.

For restore drills, a separate target cluster is created only after the source containers are stopped, so a laptop does not need to run six Cassandra nodes simultaneously. A functional one-node fallback is acceptable on a resource-constrained machine, but it cannot reproduce RF=3 replica/repair behavior. Plan roughly 6–8 GiB of free RAM and several GiB of free disk for the three-node exercises, and capture actual container/host limits in your evidence. Windows learners should use Docker Desktop/WSL-style Linux containers; commands that manipulate Linux inodes run inside the Cassandra container.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and recovery mental model

An SSTable (Sorted String Table) is Cassandra's immutable on-disk representation of flushed table data. A snapshot is a node-local point-in-time set of hard links to SSTable component files plus metadata such as schema.cql; a hard link is another directory entry referencing the same filesystem inode/blocks. An incremental backup is a hard link placed in a table's backups/ directory whenever a new SSTable is flushed or streamed after incremental backup is enabled. Neither mechanism automatically places data in a different failure domain.

A backup catalog records exactly which node/table/SSTable files, schema/configuration versions, timestamps, checksums, and recovery dependencies belong to a restore point. An off-host copy is a separately stored copy outside the node/volume failure domain; production designs often add account/region separation, encryption, access controls, immutability, and retention policy. A commit log records mutations before they are applied to memtables; commit-log archiving copies completed segments so they can be replayed during a point-in-time-oriented restore. PITR means Point-in-Time Recovery.

sstableloader reads backed-up SSTables from an offline utility process and streams their token ranges to the replicas in the current cluster topology. nodetool refresh tells one running node to discover newly placed compatible SSTables in that table's local data directory. repair is Cassandra's anti-entropy process for comparing replica ranges and streaming differences. RPO (Recovery Point Objective) is the tolerated amount of data loss measured in time; RTO (Recovery Time Objective) is the tolerated service-recovery duration.

1. Define the drill before touching data

Timestamp/evidence Capture Why it matters
T_backup latest cataloged recoverable mutation or snapshot/incremental/PITR cutoff defines recovery point
T_incident simulated disaster declaration RPO comparison and RTO start
T_target_ready three restore nodes UN + schema agreement infrastructure milestone, not cutover
T_data_loaded sstableloader/import complete data-transfer milestone, not proof of convergence
T_verified business rows/hash + repair/convergence accepted technical recovery milestone
T_cutover application/client successfully uses target RTO end for service recovery

For this deterministic lab, use the ch25-base snapshot as the explicit recovery point. If you also include Lesson 2 incremental SSTables or Lesson 3 archived commit logs, update T_backup to the latest mutation you can actually prove recoverable. Do not claim an RPO smaller than your verified chain.

PowerShell · start a local drill record
$Drill = [ordered]@{  backup_id = "ch25-20260908"  T_backup = "CAPTURE_FROM_CATALOG"  T_incident = (Get-Date).ToUniversalTime().ToString("o")  T_target_ready = $null  T_data_loaded = $null  T_verified = $null  T_cutover = $null}$Drill | ConvertTo-Json | Set-Content .\atlasmart-backup-vault\ch25\restore-drill.jsonGet-Content .\atlasmart-backup-vault\ch25\restore-drill.json

2. Inject a safe backup-copy failure and reject it before restore

Never discover corruption after hours of loading. Copy one backup file to a separate corruption-test directory, alter the copy, and prove the checksum changes. Do not alter the canonical backup.

PowerShell · corruption detection against a throwaway copy
$source = Get-ChildItem .\atlasmart-backup-vault\ch25 -File -Recurse |  Where-Object { $_.Name -like '*-Data.db' } | Select-Object -First 1if (-not $source) { throw "No Data.db backup file found" }New-Item -ItemType Directory -Force .\atlasmart-backup-vault\corruption-test | Out-Null$copy = Join-Path .\atlasmart-backup-vault\corruption-test $source.NameCopy-Item $source.FullName $copy -Force$before = (Get-FileHash -Algorithm SHA256 $copy).HashAdd-Content -Encoding Byte -Path $copy -Value 0$after = (Get-FileHash -Algorithm SHA256 $copy).Hash"before=$before""after =$after"if ($before -eq $after) { throw "Expected checksum change was not detected" }"CORRUPTION DETECTION PASS"Remove-Item $copy
Wrong acceptance criterion: “The copy command returned exit code 0.”

Transport success is not integrity/completeness proof. The drill must verify the manifest/inventory and cryptographic hashes after retrieval, then continue only with a catalog whose schema/version/security dependencies are available.

3. Rebuild the multi-node target and restore the cataloged point

If you completed Lesson 4, reuse that target after clearing/recreating the table. Otherwise use the exact target-cluster and sstableloader sequence from Lesson 4. The source containers stay stopped but their volumes remain intact as rollback.

bash · target readiness and schema agreement evidence
docker exec atlasmart-restore-1 nodetool versiondocker exec atlasmart-restore-1 nodetool statusdocker exec atlasmart-restore-1 nodetool describeclusterdocker exec atlasmart-restore-1 cqlsh -e \"SELECT cluster_name,release_version,data_center,rack FROM system.local;"
bash · after sstableloader, verify data and repair/convergence
# Query at the application consistency level.docker exec atlasmart-restore-1 cqlsh -e \"CONSISTENCY LOCAL_QUORUM; SELECT customer_id,order_month,order_time,order_id,status,total,note FROM atlasmart_backup.orders_by_customer WHERE customer_id='cust-42' AND order_month='2026-09-01';"# Observe table and replica state.docker exec atlasmart-restore-1 nodetool tablestats atlasmart_backup.orders_by_customerdocker exec atlasmart-restore-1 nodetool repair --preview --full atlasmart_backup orders_by_customer# If preview/restore policy shows convergence work is required, run the controlled full repair.docker exec atlasmart-restore-1 nodetool repair --full atlasmart_backup orders_by_customer# Verify from all three coordinators after repair.for n in 1 2 3; do  docker exec atlasmart-restore-$n cqlsh -e \  "CONSISTENCY LOCAL_ONE; SELECT status,total,note FROM atlasmart_backup.orders_by_customer WHERE customer_id='cust-42' AND order_month='2026-09-01';"done

A successful repair command is still not enough by itself. Compare known business rows/IDs/totals, schema, RF, node/rack topology, table metrics, and application queries. For a large production dataset, avoid expensive global COUNT(*) as the only proof; use partition-aware inventories, sampled/business invariants, source-side export/catalog hashes where practical, and backup-system object/file counts.

4. Measure RPO/RTO and perform application cutover

RPO answers “how far back did we have to recover?” If the simulated incident is at 09:30 and the latest proven recoverable mutation is 09:17, the measured data-loss window is 13 minutes. RTO answers “how long until usable service?” and includes target provisioning, schema/security restoration, transfer, repair/index readiness, application credential/network changes, smoke tests, and approval—not just SSTable streaming.

PowerShell · close the drill timestamps and calculate durations
$path = ".\atlasmart-backup-vault\ch25\restore-drill.json"$Drill = Get-Content $path | ConvertFrom-Json# Fill T_target_ready/T_data_loaded at the moments those milestones occurred.$Drill.T_verified = (Get-Date).ToUniversalTime().ToString("o")# Lab cutover = successful cqlsh/application-equivalent query to the restore cluster.docker exec atlasmart-restore-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT status,total FROM atlasmart_backup.orders_by_customer WHERE customer_id='cust-42' AND order_month='2026-09-01';"if ($LASTEXITCODE -ne 0) { throw "Cutover smoke test failed" }$Drill.T_cutover = (Get-Date).ToUniversalTime().ToString("o")$Drill | ConvertTo-Json | Set-Content $path$incident = [datetimeoffset]$Drill.T_incident$cutover  = [datetimeoffset]$Drill.T_cutover$rto = $cutover - $incident"Measured RTO = $rto"# RPO requires T_backup to be an actual parseable timestamp from your backup catalog.if ($Drill.T_backup -ne "CAPTURE_FROM_CATALOG") {  $backup = [datetimeoffset]$Drill.T_backup  "Measured recovery gap at incident (RPO evidence) = $($incident - $backup)"}

In a real application, cutover also updates driver contact points/service discovery, local-DC policy, credentials, TLS trust, authorization roles, secrets, connection pools, request timeouts/retries/idempotency, and monitoring. Rollback must be explicit: either return to the original healthy source if still authoritative or freeze writes and choose one direction of truth. Never permit two independently writable recovered clusters without a reconciliation design.

5. Acceptance checklist and reversible cleanup

  • Backup inventory/checksums verified before load.
  • Cassandra patch/schema/table IDs and RF/DC/rack assumptions reviewed.
  • Three restore nodes are UN and schema agreement is healthy.
  • Known AtlasMart rows/business values match the selected recovery point.
  • Repair preview/full strategy and post-repair replica checks are recorded.
  • SAI/vector/index dependencies are either restored/queryable or explicitly rebuilt/validated if the backed-up components do not satisfy the target platform.
  • Authentication/TLS/secrets/JMX/network policy are restored before production exposure; the lab keeps them disabled only because it is isolated.
  • Measured RPO/RTO and bottlenecks are recorded, including transfer, compaction/repair, and application cutover.
  • Rollback and ownership after cutover are explicit.
bash · remove the target cluster and restart the untouched source rollback path
docker rm -f atlasmart-restore-1 atlasmart-restore-2 atlasmart-restore-3docker volume rm atlasmart-restore-1-data atlasmart-restore-2-data atlasmart-restore-3-datadocker network rm atlasmart-cassandra-restore 2>/dev/null || true# Original source volumes were never deleted.docker start atlasmart-cass-1docker exec atlasmart-cass-1 nodetool statusdocker start atlasmart-cass-2 atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status# Keep the backup vault if you want to repeat the restore drill.# Delete it only as an intentional retention/reset action.

Check your understanding

  1. When does the RTO clock stop?
  2. How should RPO be reported?
  3. Why keep source volumes during the lab drill?
  4. Why run repair/replica checks after restore?
  5. What makes a restore drill production-relevant?
Review the answers

1. When the recovered service is verified and the application can safely cut over, not when file copy or sstableloader finishes.

2. From the latest mutation that the tested backup/PITR chain can actually recover to the incident point; do not use an unverified theoretical interval.

3. They provide a reversible rollback path while the target restore process is tested.

4. To verify anti-entropy convergence when the restore method/topology can leave replicas inconsistent and to provide evidence beyond a successful loader exit code.

5. It includes catalog/integrity, schema/config/security, topology, data validation, convergence/index state, application cutover, measured RPO/RTO, rollback, and documented owners.

Production judgment

A backup design is an application-recovery design, not a file-copy checkbox. Record workload/partition shape, retention and TTL/delete rates, RF/consistency level (CL), node/DC/rack topology, SSTable format and compaction, repair cadence, schema/table IDs, authentication/authorization/TLS/JMX dependencies, driver/native-protocol compatibility, version/upgrade path, encryption/key-management dependencies, SAI/vector-index requirements, disk/network throughput, snapshot/incremental/commit-log cadence, off-host failure domains, immutable retention, checksum/catalog ownership, and restore permissions. A backup that cannot recreate schema/security/configuration or that exists only on the same host/volume is not sufficient disaster-recovery evidence.

Measure Recovery Point Objective (RPO) from the latest recoverable mutation to the incident/cutover point and Recovery Time Objective (RTO) from recovery declaration to verified application service—not from “copy finished.” Include schema creation, file transfer, SSTable load/streaming, index readiness, repair/convergence, application credentials/TLS, driver contact-point cutover, smoke tests, and rollback in RTO. Avoid universal backup intervals or retention numbers: they must derive from business loss tolerance, data volume, change rate, restore throughput, compliance, cost, and tested operational skill. Chapter 26 now treats the recovered data as sensitive production state and builds defense in depth around authentication, authorization, TLS, secrets, JMX/network boundaries, backup access, and recovery credentials.

Summary and next bridge

A backup is proven by recovery. Chapter 25 ends only after the team can retrieve verified artifacts, rebuild a compatible cluster, stream/import data, converge replicas, validate business state, cut the application over, and report measured RPO/RTO. Chapter 26 secures that production/recovery path.

Authoritative references

These are the version-sensitive source of truth for the mechanisms used in this lesson. Re-check them when regenerating or operating on a different Cassandra patch/distribution.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.