Chapter 25 · Snapshots, Backups, Commitlog Archiving, Restore, and Disaster Recovery
Restore with sstableloader / Refresh or Node-Level Replacement: Topology and Schema Considerations
Choose deliberately among sstableloader, refresh/import, and node replacement according to schema, token ownership, topology, and failure scope.
Learning outcomes
AtlasMart has a cataloged snapshot in a separate vault. The team now faces three possible restore paths: stream backed-up SSTables into a new topology, place files on a specific node and refresh/import them, or replace a failed node in an otherwise healthy cluster. These paths solve different failure modes.
Choose between sstableloader, nodetool refresh/import, and node replacement based on topology and failure scope.
Explain schema/table-ID and token-ownership requirements before loading any SSTable.
Run an offline utility-container sstableloader workflow into a fresh target cluster.
Explain why local refresh is node-specific and may require repair/cleanup/ownership validation.
Verify row/content, schema, topology, index readiness, and repair state after restore.
The mandatory labs use Apache Cassandra 5.0.9 in
the pinned cassandra:5.0.9 container image. The
normal source cluster keeps the course conventions: cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy with replication
factor (RF) 3, and LOCAL_QUORUM for the chapter's
verification reads/writes. Tables explicitly use
UnifiedCompactionStrategy (UCS); authentication,
client/internode TLS, and remote JMX remain disabled only
inside this isolated local learning topology. The chapter
keyspace is atlasmart_backup. No managed service,
paid backup product, or cloud account is required.
For restore drills, a separate target cluster is created only after the source containers are stopped, so a laptop does not need to run six Cassandra nodes simultaneously. A functional one-node fallback is acceptable on a resource-constrained machine, but it cannot reproduce RF=3 replica/repair behavior. Plan roughly 6–8 GiB of free RAM and several GiB of free disk for the three-node exercises, and capture actual container/host limits in your evidence. Windows learners should use Docker Desktop/WSL-style Linux containers; commands that manipulate Linux inodes run inside the Cassandra container.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and recovery mental model
An SSTable (Sorted String Table) is Cassandra's
immutable on-disk representation of flushed table data. A
snapshot is a node-local point-in-time set of
hard links to SSTable component files plus metadata such as
schema.cql; a hard link is another directory entry
referencing the same filesystem inode/blocks. An
incremental backup is a hard link placed in a
table's backups/ directory whenever a new SSTable
is flushed or streamed after incremental backup is enabled.
Neither mechanism automatically places data in a different
failure domain.
A backup catalog records exactly which node/table/SSTable files, schema/configuration versions, timestamps, checksums, and recovery dependencies belong to a restore point. An off-host copy is a separately stored copy outside the node/volume failure domain; production designs often add account/region separation, encryption, access controls, immutability, and retention policy. A commit log records mutations before they are applied to memtables; commit-log archiving copies completed segments so they can be replayed during a point-in-time-oriented restore. PITR means Point-in-Time Recovery.
sstableloader reads backed-up SSTables from an offline utility process and streams their token ranges to the replicas in the current cluster topology. nodetool refresh tells one running node to discover newly placed compatible SSTables in that table's local data directory. repair is Cassandra's anti-entropy process for comparing replica ranges and streaming differences. RPO (Recovery Point Objective) is the tolerated amount of data loss measured in time; RTO (Recovery Time Objective) is the tolerated service-recovery duration.
1. Restore-path decision model
| Restore path | Best fit | Topology behavior | Main risks |
|---|---|---|---|
| sstableloader | new/different cluster topology or logical table restore | streams token sections to current replicas | schema mismatch, high network/compaction load, missing dependencies |
| nodetool refresh | compatible SSTables manually placed on one node | loads files locally; no automatic redistribution | wrong token ownership, duplicate/obsolete files, missed replica convergence |
| nodetool import | modern local SSTable import with verification/copy options | node-local import with more controls than raw placement | still node-local; schema/topology compatibility required |
| node replacement | one dead node while surviving replicas form healthy cluster | replacement streams ownership from live replicas | wrong replacement identity, reusing stale data directories |
| full-cluster DR | original cluster unavailable | usually rebuild target schema/topology then bulk stream/import from backup | RPO/RTO, catalog completeness, credentials/TLS, repair/cutover |
sstableloader is the most topology-independent of these file restore tools: it reads offline SSTables and streams the relevant token sections to the current cluster's replicas. The target table must already exist. The directory hierarchy supplied to the tool must identify the keyspace/table. Run the tool from a utility environment that is not a live Cassandra process; Cassandra's SSTable-tool documentation warns against executing SSTable tools against a running Cassandra data directory.
refresh is different: after you place
compatible SSTables in a specific node's local table data
directory, nodetool refresh keyspace table makes
that node load them without restart. It does not magically
redistribute data to the proper RF replicas. Current
nodetool import offers validation/copy controls and
is often preferable when importing external SSTables locally.
Either path requires ownership and post-load convergence
evidence.
2. Build a fresh three-node restore target without running six nodes
Stop, but do not delete, the source containers so their volumes remain the rollback path and laptop memory is released. Then create a separate target network/cluster with new volumes and host IDs.
docker stop atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker network inspect atlasmart-cassandra-restore >/dev/null 2>&1 || docker network create atlasmart-cassandra-restorefor n in 1 2 3; do docker volume create atlasmart-restore-$n-data; donedocker run -d --name atlasmart-restore-1 --hostname atlasmart-restore-1 --network atlasmart-cassandra-restore -e CASSANDRA_CLUSTER_NAME=atlasmart-dr-restore -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-restore-1-data:/var/lib/cassandra cassandra:5.0.9# Wait for UN before peers.docker exec atlasmart-restore-1 nodetool statusdocker run -d --name atlasmart-restore-2 --hostname atlasmart-restore-2 --network atlasmart-cassandra-restore -e CASSANDRA_CLUSTER_NAME=atlasmart-dr-restore -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-restore-1 -v atlasmart-restore-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-restore-3 --hostname atlasmart-restore-3 --network atlasmart-cassandra-restore -e CASSANDRA_CLUSTER_NAME=atlasmart-dr-restore -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-restore-1 -v atlasmart-restore-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-restore-1 nodetool status
Before loading data, recreate the schema from the cataloged
schema.cql or reviewed DDL. Do not blindly restore
system keyspaces, node identity, secrets, or old topology
metadata. A target cluster can have new host IDs and token
ownership when using sstableloader.
CREATE KEYSPACE IF NOT EXISTS atlasmart_backupWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_backup.orders_by_customer ( customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;
3. Stream a backed-up snapshot with sstableloader
Prepare a host directory whose final two path components are the target keyspace and table, and copy only SSTable components from the cataloged snapshot into it. In this lab RF=3 with exactly three source nodes means every partition has a replica on node 1, so node 1's complete table snapshot can serve as the small logical source. That shortcut is not valid in a larger cluster where one node owns only part of the token ranges.
$root = ".\atlasmart-backup-vault\restore-input\atlasmart_backup\orders_by_customer"New-Item -ItemType Directory -Force $root | Out-Null$snap = Get-ChildItem ".\atlasmart-backup-vault\ch25\node1\atlasmart_backup" -Directory -Recurse | Where-Object { $_.FullName -match 'snapshots[\\/]ch25-base$' } | Select-Object -First 1if (-not $snap) { throw "ch25-base snapshot not found in the copied node-1 backup" }Get-ChildItem $snap.FullName -File | Where-Object { $_.Name -notin @('schema.cql','manifest.json') } | Copy-Item -Destination $root -ForceGet-ChildItem $root | Select-Object Name,Length
# The utility container does not run a Cassandra server.docker run --rm --network atlasmart-cassandra-restore \ -v "$PWD/atlasmart-backup-vault/restore-input:/restore:ro" \ --entrypoint sstableloader cassandra:5.0.9 \ --nodes atlasmart-restore-1 \ --throttle-mib 10 \ /restore/atlasmart_backup/orders_by_customerdocker exec atlasmart-restore-1 cqlsh -e \"CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_backup.orders_by_customer WHERE customer_id='cust-42' AND order_month='2026-09-01';"
The transfer summary—files/bytes/duration—is learner-captured evidence. The 10 MiB/s throttle is only a small-lab safety cap, not a production recommendation. In production choose throughput/connections based on network, disks, compaction headroom, application latency, and tested restore RTO.
4. Understand refresh/import and node-replacement boundaries
1. Recreate the exact compatible keyspace/table schema first.2. Identify the target node's table data directory with nodetool datapaths.3. Stop copying from live mutable directories; use cataloged backup SSTables.4. Verify component checksums before placement.5. Place/copy files with correct ownership/permissions.6. Run nodetool refresh <keyspace> <table> OR use nodetool import with version-appropriate verification/copy options.7. Query locally and at the intended CL.8. Validate token ownership and replica coverage.9. Run repair when the restore path/policy requires anti-entropy convergence.10. Clean obsolete non-owned ranges only after ownership/convergence is proven.
That mixes stale node identity/token/system metadata with a new membership operation. If one node dies while the cluster survives, use the supported node-replacement workflow from Chapter 23 so the replacement assumes the failed member's identity/ownership and streams current data from surviving replicas. Use backup-based whole-cluster restore when live replicas cannot provide recovery.
Check your understanding
- Why is sstableloader useful for a new topology?
- Why must the target table exist first?
- What is different about nodetool refresh?
- Why is one node1 snapshot complete in this tiny lab?
- When is node replacement preferable to backup restore?
Review the answers
1. It reads offline SSTables and streams the relevant token sections to replicas in the current target topology.
2. The loader restores SSTable data into an existing schema; incremental/snapshot files are not a schema service.
3. It loads compatible SSTables placed on one node locally; it does not redistribute them across the cluster.
4. RF=3 and exactly three nodes means every partition has a replica on every node; this is not true in a larger topology.
5. When a single node is dead but surviving replicas form a healthy cluster that can stream current data to the correctly identified replacement.
Production judgment
A backup design is an application-recovery design, not a file-copy checkbox. Record workload/partition shape, retention and TTL/delete rates, RF/consistency level (CL), node/DC/rack topology, SSTable format and compaction, repair cadence, schema/table IDs, authentication/authorization/TLS/JMX dependencies, driver/native-protocol compatibility, version/upgrade path, encryption/key-management dependencies, SAI/vector-index requirements, disk/network throughput, snapshot/incremental/commit-log cadence, off-host failure domains, immutable retention, checksum/catalog ownership, and restore permissions. A backup that cannot recreate schema/security/configuration or that exists only on the same host/volume is not sufficient disaster-recovery evidence.
Measure Recovery Point Objective (RPO) from the latest recoverable mutation to the incident/cutover point and Recovery Time Objective (RTO) from recovery declaration to verified application service—not from “copy finished.” Include schema creation, file transfer, SSTable load/streaming, index readiness, repair/convergence, application credentials/TLS, driver contact-point cutover, smoke tests, and rollback in RTO. Avoid universal backup intervals or retention numbers: they must derive from business loss tolerance, data volume, change rate, restore throughput, compliance, cost, and tested operational skill. Lesson 5 turns these mechanisms into an end-to-end multi-node restore drill with corruption detection, repair verification, measured RPO/RTO, application cutover, and rollback.
Summary and next bridge
Restore tooling is topology-specific. sstableloader streams data into the current layout, refresh/import are node-local, and node replacement is a membership/streaming operation from surviving replicas. The final lesson proves the entire DR workflow under a timed drill.
Authoritative references
These are the version-sensitive source of truth for the mechanisms used in this lesson. Re-check them when regenerating or operating on a different Cassandra patch/distribution.