Chapter 25 · Snapshots, Backups, Commitlog Archiving, Restore, and Disaster Recovery
Incremental Backups, Backup Catalogs, Remote Copies, Retention, and Integrity Checks
Build a cataloged incremental-backup chain with integrity checks, retention dependencies, and an explicit off-host-copy boundary.
Learning outcomes
AtlasMart now has a base snapshot, but orders continue to change every minute. Copying a full snapshot after every write would be wasteful; keeping only incremental files without a base/catalog would be unrecoverable. This lesson builds a coherent backup chain.
Explain what Cassandra incremental backups capture and why they are also hard-link/local mechanisms.
Enable/verify incremental backup and observe backups/ files created after flush or streaming.
Design a backup catalog containing schema/topology/version/file checksums and recovery boundaries.
Simulate an off-host copy without pretending the same laptop is a separate production failure domain.
Define retention as a recoverable chain policy rather than deleting old files by age alone.
The mandatory labs use Apache Cassandra 5.0.9 in
the pinned cassandra:5.0.9 container image. The
normal source cluster keeps the course conventions: cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy with replication
factor (RF) 3, and LOCAL_QUORUM for the chapter's
verification reads/writes. Tables explicitly use
UnifiedCompactionStrategy (UCS); authentication,
client/internode TLS, and remote JMX remain disabled only
inside this isolated local learning topology. The chapter
keyspace is atlasmart_backup. No managed service,
paid backup product, or cloud account is required.
For restore drills, a separate target cluster is created only after the source containers are stopped, so a laptop does not need to run six Cassandra nodes simultaneously. A functional one-node fallback is acceptable on a resource-constrained machine, but it cannot reproduce RF=3 replica/repair behavior. Plan roughly 6–8 GiB of free RAM and several GiB of free disk for the three-node exercises, and capture actual container/host limits in your evidence. Windows learners should use Docker Desktop/WSL-style Linux containers; commands that manipulate Linux inodes run inside the Cassandra container.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and recovery mental model
An SSTable (Sorted String Table) is Cassandra's
immutable on-disk representation of flushed table data. A
snapshot is a node-local point-in-time set of
hard links to SSTable component files plus metadata such as
schema.cql; a hard link is another directory entry
referencing the same filesystem inode/blocks. An
incremental backup is a hard link placed in a
table's backups/ directory whenever a new SSTable
is flushed or streamed after incremental backup is enabled.
Neither mechanism automatically places data in a different
failure domain.
A backup catalog records exactly which node/table/SSTable files, schema/configuration versions, timestamps, checksums, and recovery dependencies belong to a restore point. An off-host copy is a separately stored copy outside the node/volume failure domain; production designs often add account/region separation, encryption, access controls, immutability, and retention policy. A commit log records mutations before they are applied to memtables; commit-log archiving copies completed segments so they can be replayed during a point-in-time-oriented restore. PITR means Point-in-Time Recovery.
sstableloader reads backed-up SSTables from an offline utility process and streams their token ranges to the replicas in the current cluster topology. nodetool refresh tells one running node to discover newly placed compatible SSTables in that table's local data directory. repair is Cassandra's anti-entropy process for comparing replica ranges and streaming differences. RPO (Recovery Point Objective) is the tolerated amount of data loss measured in time; RTO (Recovery Time Objective) is the tolerated service-recovery duration.
1. Incremental backup is “new SSTables since enabled,” not a standalone database image
When incremental backup is enabled, Cassandra hard-links each
newly flushed or streamed SSTable into that table's
backups/ directory. Those files represent new
immutable generations after the base point. They do not
automatically include earlier SSTables and they do not carry the
table DDL. A restore chain therefore needs a compatible base
snapshot/schema plus the relevant later incremental SSTables, or
another complete source of historical SSTables.
Because these are hard links on the Cassandra node, incremental backups have the same local-failure-domain weakness as snapshots until copied elsewhere. They also require operator retention: deleting the only base while retaining later incremental files can leave a gap.
| Chain element | Purpose | Catalog fields that matter |
|---|---|---|
| Base snapshot | known complete point per selected scope | snapshot tag/time, node host ID, table ID/schema hash, file inventory |
| Incremental SSTables | new immutable files after base | first/last observed time, source node, table, checksum, bytes |
| Schema/config/security bundle | recreate compatible environment | Cassandra patch, DDL, RF/DC/racks, auth/TLS/key dependencies |
| Off-host object/copy | survive source failure domain | destination, encryption/immutability class, upload checksum/status |
| Restore drill result | prove operational usability | tool/path, duration, errors, repair outcome, RPO/RTO, application checks |
2. Enable backup, mutate data, and observe new SSTable links
# Verify an existing course cluster first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course cluster does not exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before adding peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three nodes are UN in dc1.docker exec atlasmart-cass-1 nodetool status
CREATE KEYSPACE IF NOT EXISTS atlasmart_backupWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_backup.orders_by_customer ( customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_backup.orders_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-08T08:00:00Z',00000000-0000-0000-0000-000000000101,'PAID',129.90,'base-A');INSERT INTO atlasmart_backup.orders_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-08T08:05:00Z',00000000-0000-0000-0000-000000000102,'SHIPPED',49.50,'base-B');SELECT customer_id,order_month,order_time,order_id,status,total,noteFROM atlasmart_backup.orders_by_customerWHERE customer_id='cust-42' AND order_month='2026-09-01';
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool enablebackup docker exec atlasmart-cass-$n nodetool statusbackupdone
CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_backup.orders_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-08T09:00:00Z',00000000-0000-0000-0000-000000000103,'PAID',79.00,'incremental-C');UPDATE atlasmart_backup.orders_by_customerSET status='DELIVERED', note='incremental-update'WHERE customer_id='cust-42' AND order_month='2026-09-01' AND order_time='2026-09-08T08:05:00Z' AND order_id=00000000-0000-0000-0000-000000000102;
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_backup orders_by_customer echo "=== node $n incremental backup files ===" docker exec atlasmart-cass-$n sh -lc \ 'find /var/lib/cassandra/data/atlasmart_backup -type f -path "*/backups/*" -printf "%s %p\n" | sort'done
The exact number of files/components is format-dependent. The
important observation is that new SSTable generations appear
under backups/ only after the feature is enabled
and a flush/stream creates new SSTables. These files are not a
schema backup.
3. Create an off-host-copy simulation and cryptographic inventory
The local lab uses the Docker host directory
atlasmart-backup-vault/ to demonstrate
copy/catalog/checksum mechanics. This is not a
real separate failure domain if the host and Cassandra data
share the same physical machine. In production, place the backup
in separately administered storage—often another
account/site/region—with encryption, least privilege,
immutability/versioning, monitoring, and tested retrieval.
New-Item -ItemType Directory -Force .\atlasmart-backup-vault\ch25 | Out-Null1..3 | ForEach-Object { $n = $_ New-Item -ItemType Directory -Force ".\atlasmart-backup-vault\ch25\node$n" | Out-Null docker cp "atlasmart-cass-$n`:/var/lib/cassandra/data/atlasmart_backup" ".\atlasmart-backup-vault\ch25\node$n\atlasmart_backup"}# Create a checksum inventory of every copied file.Get-ChildItem .\atlasmart-backup-vault\ch25 -File -Recurse | Sort-Object FullName | ForEach-Object { $h = Get-FileHash -Algorithm SHA256 $_.FullName "$($h.Hash.ToLower()) $($_.FullName)" } | Set-Content .\atlasmart-backup-vault\ch25\SHA256SUMS.txtGet-Content .\atlasmart-backup-vault\ch25\SHA256SUMS.txt | Select-Object -First 20
{ "backup_id": "ch25-20260908T090000Z", "cassandra_version": "5.0.9", "container_image": "cassandra:5.0.9", "source_cluster": "atlasmart-course", "keyspace": "atlasmart_backup", "tables": ["orders_by_customer"], "base_snapshot_tag": "ch25-base", "incremental_backups_enabled": true, "rf": {"dc1": 3}, "consistency": "LOCAL_QUORUM", "nodes": [ {"name":"atlasmart-cass-1","rack":"rack1","host_id":"CAPTURE_ME"}, {"name":"atlasmart-cass-2","rack":"rack2","host_id":"CAPTURE_ME"}, {"name":"atlasmart-cass-3","rack":"rack3","host_id":"CAPTURE_ME"} ], "schema_sha256": "CAPTURE_ME", "file_manifest": "SHA256SUMS.txt", "off_host_destination": "LOCAL_SIMULATION_ONLY", "recoverable_through": "CAPTURE_FROM_YOUR_RUN", "restore_test": {"status":"NOT_YET_TESTED"}}
Retention must preserve a usable recovery chain. A later SSTable does not magically reconstruct files that existed before it. Expiration/retention should be driven by restore points and catalog dependencies, then validated by periodic restore drills.
4. Integrity, retention, and reset
A checksum proves that the retrieved bytes equal the cataloged bytes; it does not prove semantic completeness, compatible schema, decryptable secrets, or recoverability. Keep an inventory of expected files and verify both count/bytes and hashes after transfer. For large backups, use the destination system's checksum/manifest features but understand multipart/encryption semantics instead of assuming an object-store ETag equals SHA-256.
# Recompute and compare rather than trusting a successful copy command.$bad = @()Get-Content .\atlasmart-backup-vault\ch25\SHA256SUMS.txt | ForEach-Object { if ($_ -match '^([0-9a-f]{64}) (.+)$') { $expected = $matches[1] $path = $matches[2] if (-not (Test-Path $path)) { $bad += "MISSING $path"; return } $actual = (Get-FileHash -Algorithm SHA256 $path).Hash.ToLower() if ($actual -ne $expected) { $bad += "HASH $path" } }}if ($bad.Count -eq 0) { "CHECKSUM VERIFICATION PASS" } else { $bad }
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool disablebackup docker exec atlasmart-cass-$n nodetool statusbackupdone# Removing existing backups/ hard links is a separate retention action; do not delete# them blindly while they belong to a recovery chain.
Check your understanding
- Why do incremental backups need a base/schema source?
- What does a SHA-256 verification prove?
- Why is the host-side vault only a simulation?
- What should retention expire: individual files or recoverable chains?
- What operational data belongs beside SSTable files?
Review the answers
1. They only capture newly created SSTables after the feature is enabled and do not include table DDL.
2. That retrieved file bytes match the cataloged bytes; it does not prove the backup set is complete or restorable.
3. It may share the same physical host/storage/site failure domain as the source Cassandra data.
4. Recoverable restore points/chains, because deleting a required base or intermediate artifact can break recovery.
5. At minimum schema/table IDs, Cassandra/version/topology/RF, config/security dependencies, node identity, inventory/checksums, timestamps, destination and restore-test evidence.
Production judgment
A backup design is an application-recovery design, not a file-copy checkbox. Record workload/partition shape, retention and TTL/delete rates, RF/consistency level (CL), node/DC/rack topology, SSTable format and compaction, repair cadence, schema/table IDs, authentication/authorization/TLS/JMX dependencies, driver/native-protocol compatibility, version/upgrade path, encryption/key-management dependencies, SAI/vector-index requirements, disk/network throughput, snapshot/incremental/commit-log cadence, off-host failure domains, immutable retention, checksum/catalog ownership, and restore permissions. A backup that cannot recreate schema/security/configuration or that exists only on the same host/volume is not sufficient disaster-recovery evidence.
Measure Recovery Point Objective (RPO) from the latest recoverable mutation to the incident/cutover point and Recovery Time Objective (RTO) from recovery declaration to verified application service—not from “copy finished.” Include schema creation, file transfer, SSTable load/streaming, index readiness, repair/convergence, application credentials/TLS, driver contact-point cutover, smoke tests, and rollback in RTO. Avoid universal backup intervals or retention numbers: they must derive from business loss tolerance, data volume, change rate, restore throughput, compliance, cost, and tested operational skill. Lesson 3 adds archived commit-log segments to reduce the recovery-point gap and shows why point-in-time recovery must be aligned with a compatible base snapshot, schema identity, and timestamp policy.
Summary and next bridge
Incremental backup is efficient because it captures new immutable SSTables, but the operator owns cataloging, off-host transfer, integrity, chain retention, and restore testing. Next, examine commit-log archiving for a finer-grained point-in-time-oriented recovery window.
Authoritative references
These are the version-sensitive source of truth for the mechanisms used in this lesson. Re-check them when regenerating or operating on a different Cassandra patch/distribution.