Chapter 25 · Snapshots, Backups, Commitlog Archiving, Restore, and Disaster Recovery
Perform a Multi-Node Restore Drill and Measure RPO / RTO, Repair Needs, and Application Cutover
Run a reversible multi-node restore drill, verify replica/data state, cut the client over, and measure evidence-based RPO/RTO.
Learning outcomes
AtlasMart's runbook is not accepted until a team can start with a cataloged backup, detect a bad copy, rebuild a multi-node target, verify Cassandra convergence, point an application at it, and calculate the actual RPO/RTO. This lesson is the chapter acceptance test.
Run a reversible multi-node disaster-recovery drill without deleting the original source volumes.
Detect missing/corrupted backup artifacts before restore using catalog/checksum verification.
Restore schema and SSTables into a new RF=3 cluster, then verify rows and anti-entropy state.
Measure RPO and RTO from explicit timestamps and explain what is included in each.
Define application cutover, rollback, security/config dependencies, and post-restore ownership.
The mandatory labs use Apache Cassandra 5.0.9 in
the pinned cassandra:5.0.9 container image. The
normal source cluster keeps the course conventions: cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy with replication
factor (RF) 3, and LOCAL_QUORUM for the chapter's
verification reads/writes. Tables explicitly use
UnifiedCompactionStrategy (UCS); authentication,
client/internode TLS, and remote JMX remain disabled only
inside this isolated local learning topology. The chapter
keyspace is atlasmart_backup. No managed service,
paid backup product, or cloud account is required.
For restore drills, a separate target cluster is created only after the source containers are stopped, so a laptop does not need to run six Cassandra nodes simultaneously. A functional one-node fallback is acceptable on a resource-constrained machine, but it cannot reproduce RF=3 replica/repair behavior. Plan roughly 6–8 GiB of free RAM and several GiB of free disk for the three-node exercises, and capture actual container/host limits in your evidence. Windows learners should use Docker Desktop/WSL-style Linux containers; commands that manipulate Linux inodes run inside the Cassandra container.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and recovery mental model
An SSTable (Sorted String Table) is Cassandra's
immutable on-disk representation of flushed table data. A
snapshot is a node-local point-in-time set of
hard links to SSTable component files plus metadata such as
schema.cql; a hard link is another directory entry
referencing the same filesystem inode/blocks. An
incremental backup is a hard link placed in a
table's backups/ directory whenever a new SSTable
is flushed or streamed after incremental backup is enabled.
Neither mechanism automatically places data in a different
failure domain.
A backup catalog records exactly which node/table/SSTable files, schema/configuration versions, timestamps, checksums, and recovery dependencies belong to a restore point. An off-host copy is a separately stored copy outside the node/volume failure domain; production designs often add account/region separation, encryption, access controls, immutability, and retention policy. A commit log records mutations before they are applied to memtables; commit-log archiving copies completed segments so they can be replayed during a point-in-time-oriented restore. PITR means Point-in-Time Recovery.
sstableloader reads backed-up SSTables from an offline utility process and streams their token ranges to the replicas in the current cluster topology. nodetool refresh tells one running node to discover newly placed compatible SSTables in that table's local data directory. repair is Cassandra's anti-entropy process for comparing replica ranges and streaming differences. RPO (Recovery Point Objective) is the tolerated amount of data loss measured in time; RTO (Recovery Time Objective) is the tolerated service-recovery duration.
1. Define the drill before touching data
| Timestamp/evidence | Capture | Why it matters |
|---|---|---|
| T_backup | latest cataloged recoverable mutation or snapshot/incremental/PITR cutoff | defines recovery point |
| T_incident | simulated disaster declaration | RPO comparison and RTO start |
| T_target_ready | three restore nodes UN + schema agreement | infrastructure milestone, not cutover |
| T_data_loaded | sstableloader/import complete | data-transfer milestone, not proof of convergence |
| T_verified | business rows/hash + repair/convergence accepted | technical recovery milestone |
| T_cutover | application/client successfully uses target | RTO end for service recovery |
For this deterministic lab, use the
ch25-base snapshot as the explicit recovery point.
If you also include Lesson 2 incremental SSTables or Lesson 3
archived commit logs, update T_backup to the latest
mutation you can actually prove recoverable. Do not claim an RPO
smaller than your verified chain.
$Drill = [ordered]@{ backup_id = "ch25-20260908" T_backup = "CAPTURE_FROM_CATALOG" T_incident = (Get-Date).ToUniversalTime().ToString("o") T_target_ready = $null T_data_loaded = $null T_verified = $null T_cutover = $null}$Drill | ConvertTo-Json | Set-Content .\atlasmart-backup-vault\ch25\restore-drill.jsonGet-Content .\atlasmart-backup-vault\ch25\restore-drill.json
2. Inject a safe backup-copy failure and reject it before restore
Never discover corruption after hours of loading. Copy one
backup file to a separate
corruption-test directory, alter the copy, and
prove the checksum changes. Do not alter the canonical backup.
$source = Get-ChildItem .\atlasmart-backup-vault\ch25 -File -Recurse | Where-Object { $_.Name -like '*-Data.db' } | Select-Object -First 1if (-not $source) { throw "No Data.db backup file found" }New-Item -ItemType Directory -Force .\atlasmart-backup-vault\corruption-test | Out-Null$copy = Join-Path .\atlasmart-backup-vault\corruption-test $source.NameCopy-Item $source.FullName $copy -Force$before = (Get-FileHash -Algorithm SHA256 $copy).HashAdd-Content -Encoding Byte -Path $copy -Value 0$after = (Get-FileHash -Algorithm SHA256 $copy).Hash"before=$before""after =$after"if ($before -eq $after) { throw "Expected checksum change was not detected" }"CORRUPTION DETECTION PASS"Remove-Item $copy
Transport success is not integrity/completeness proof. The drill must verify the manifest/inventory and cryptographic hashes after retrieval, then continue only with a catalog whose schema/version/security dependencies are available.
3. Rebuild the multi-node target and restore the cataloged point
If you completed Lesson 4, reuse that target after
clearing/recreating the table. Otherwise use the exact
target-cluster and sstableloader sequence from
Lesson 4. The source containers stay stopped but their volumes
remain intact as rollback.
docker exec atlasmart-restore-1 nodetool versiondocker exec atlasmart-restore-1 nodetool statusdocker exec atlasmart-restore-1 nodetool describeclusterdocker exec atlasmart-restore-1 cqlsh -e \"SELECT cluster_name,release_version,data_center,rack FROM system.local;"
# Query at the application consistency level.docker exec atlasmart-restore-1 cqlsh -e \"CONSISTENCY LOCAL_QUORUM; SELECT customer_id,order_month,order_time,order_id,status,total,note FROM atlasmart_backup.orders_by_customer WHERE customer_id='cust-42' AND order_month='2026-09-01';"# Observe table and replica state.docker exec atlasmart-restore-1 nodetool tablestats atlasmart_backup.orders_by_customerdocker exec atlasmart-restore-1 nodetool repair --preview --full atlasmart_backup orders_by_customer# If preview/restore policy shows convergence work is required, run the controlled full repair.docker exec atlasmart-restore-1 nodetool repair --full atlasmart_backup orders_by_customer# Verify from all three coordinators after repair.for n in 1 2 3; do docker exec atlasmart-restore-$n cqlsh -e \ "CONSISTENCY LOCAL_ONE; SELECT status,total,note FROM atlasmart_backup.orders_by_customer WHERE customer_id='cust-42' AND order_month='2026-09-01';"done
A successful repair command is still not enough by itself.
Compare known business rows/IDs/totals, schema, RF, node/rack
topology, table metrics, and application queries. For a large
production dataset, avoid expensive global
COUNT(*) as the only proof; use partition-aware
inventories, sampled/business invariants, source-side
export/catalog hashes where practical, and backup-system
object/file counts.
4. Measure RPO/RTO and perform application cutover
RPO answers “how far back did we have to recover?” If the simulated incident is at 09:30 and the latest proven recoverable mutation is 09:17, the measured data-loss window is 13 minutes. RTO answers “how long until usable service?” and includes target provisioning, schema/security restoration, transfer, repair/index readiness, application credential/network changes, smoke tests, and approval—not just SSTable streaming.
$path = ".\atlasmart-backup-vault\ch25\restore-drill.json"$Drill = Get-Content $path | ConvertFrom-Json# Fill T_target_ready/T_data_loaded at the moments those milestones occurred.$Drill.T_verified = (Get-Date).ToUniversalTime().ToString("o")# Lab cutover = successful cqlsh/application-equivalent query to the restore cluster.docker exec atlasmart-restore-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT status,total FROM atlasmart_backup.orders_by_customer WHERE customer_id='cust-42' AND order_month='2026-09-01';"if ($LASTEXITCODE -ne 0) { throw "Cutover smoke test failed" }$Drill.T_cutover = (Get-Date).ToUniversalTime().ToString("o")$Drill | ConvertTo-Json | Set-Content $path$incident = [datetimeoffset]$Drill.T_incident$cutover = [datetimeoffset]$Drill.T_cutover$rto = $cutover - $incident"Measured RTO = $rto"# RPO requires T_backup to be an actual parseable timestamp from your backup catalog.if ($Drill.T_backup -ne "CAPTURE_FROM_CATALOG") { $backup = [datetimeoffset]$Drill.T_backup "Measured recovery gap at incident (RPO evidence) = $($incident - $backup)"}
In a real application, cutover also updates driver contact points/service discovery, local-DC policy, credentials, TLS trust, authorization roles, secrets, connection pools, request timeouts/retries/idempotency, and monitoring. Rollback must be explicit: either return to the original healthy source if still authoritative or freeze writes and choose one direction of truth. Never permit two independently writable recovered clusters without a reconciliation design.
5. Acceptance checklist and reversible cleanup
- Backup inventory/checksums verified before load.
- Cassandra patch/schema/table IDs and RF/DC/rack assumptions reviewed.
- Three restore nodes are UN and schema agreement is healthy.
- Known AtlasMart rows/business values match the selected recovery point.
- Repair preview/full strategy and post-repair replica checks are recorded.
- SAI/vector/index dependencies are either restored/queryable or explicitly rebuilt/validated if the backed-up components do not satisfy the target platform.
- Authentication/TLS/secrets/JMX/network policy are restored before production exposure; the lab keeps them disabled only because it is isolated.
- Measured RPO/RTO and bottlenecks are recorded, including transfer, compaction/repair, and application cutover.
- Rollback and ownership after cutover are explicit.
docker rm -f atlasmart-restore-1 atlasmart-restore-2 atlasmart-restore-3docker volume rm atlasmart-restore-1-data atlasmart-restore-2-data atlasmart-restore-3-datadocker network rm atlasmart-cassandra-restore 2>/dev/null || true# Original source volumes were never deleted.docker start atlasmart-cass-1docker exec atlasmart-cass-1 nodetool statusdocker start atlasmart-cass-2 atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status# Keep the backup vault if you want to repeat the restore drill.# Delete it only as an intentional retention/reset action.
Check your understanding
- When does the RTO clock stop?
- How should RPO be reported?
- Why keep source volumes during the lab drill?
- Why run repair/replica checks after restore?
- What makes a restore drill production-relevant?
Review the answers
1. When the recovered service is verified and the application can safely cut over, not when file copy or sstableloader finishes.
2. From the latest mutation that the tested backup/PITR chain can actually recover to the incident point; do not use an unverified theoretical interval.
3. They provide a reversible rollback path while the target restore process is tested.
4. To verify anti-entropy convergence when the restore method/topology can leave replicas inconsistent and to provide evidence beyond a successful loader exit code.
5. It includes catalog/integrity, schema/config/security, topology, data validation, convergence/index state, application cutover, measured RPO/RTO, rollback, and documented owners.
Production judgment
A backup design is an application-recovery design, not a file-copy checkbox. Record workload/partition shape, retention and TTL/delete rates, RF/consistency level (CL), node/DC/rack topology, SSTable format and compaction, repair cadence, schema/table IDs, authentication/authorization/TLS/JMX dependencies, driver/native-protocol compatibility, version/upgrade path, encryption/key-management dependencies, SAI/vector-index requirements, disk/network throughput, snapshot/incremental/commit-log cadence, off-host failure domains, immutable retention, checksum/catalog ownership, and restore permissions. A backup that cannot recreate schema/security/configuration or that exists only on the same host/volume is not sufficient disaster-recovery evidence.
Measure Recovery Point Objective (RPO) from the latest recoverable mutation to the incident/cutover point and Recovery Time Objective (RTO) from recovery declaration to verified application service—not from “copy finished.” Include schema creation, file transfer, SSTable load/streaming, index readiness, repair/convergence, application credentials/TLS, driver contact-point cutover, smoke tests, and rollback in RTO. Avoid universal backup intervals or retention numbers: they must derive from business loss tolerance, data volume, change rate, restore throughput, compliance, cost, and tested operational skill. Chapter 26 now treats the recovered data as sensitive production state and builds defense in depth around authentication, authorization, TLS, secrets, JMX/network boundaries, backup access, and recovery credentials.
Summary and next bridge
A backup is proven by recovery. Chapter 25 ends only after the team can retrieve verified artifacts, rebuild a compatible cluster, stream/import data, converge replicas, validate business state, cut the application over, and report measured RPO/RTO. Chapter 26 secures that production/recovery path.
Authoritative references
These are the version-sensitive source of truth for the mechanisms used in this lesson. Re-check them when regenerating or operating on a different Cassandra patch/distribution.