Chapter 25 · Snapshots, Backups, Commitlog Archiving, Restore, and Disaster Recovery
nodetool snapshot: Hard Links, SSTable Immutability, Schema Capture, and Space Accounting
Prove what Cassandra snapshots actually store, how hard links affect disk accounting, and why local snapshots are only one ingredient of disaster recovery.
Learning outcomes
AtlasMart's operations dashboard says “backup succeeded” because every Cassandra node has a snapshot directory. Then a storage-volume failure destroys both live SSTables and the snapshot hard links. The lesson begins by proving what a snapshot really is and what it is not.
Explain how nodetool snapshot uses SSTable immutability and filesystem hard links.
Distinguish apparent snapshot size, true additional disk consumption, and blocks pinned by snapshots after compaction.
Inspect snapshot directories, inode/link-count evidence, manifest.json, and schema.cql safely.
Explain why a per-node local snapshot is a recovery ingredient, not a disaster-recovery backup.
Create a catalogable base snapshot that later lessons copy off-host and restore.
The mandatory labs use Apache Cassandra 5.0.9 in
the pinned cassandra:5.0.9 container image. The
normal source cluster keeps the course conventions: cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy with replication
factor (RF) 3, and LOCAL_QUORUM for the chapter's
verification reads/writes. Tables explicitly use
UnifiedCompactionStrategy (UCS); authentication,
client/internode TLS, and remote JMX remain disabled only
inside this isolated local learning topology. The chapter
keyspace is atlasmart_backup. No managed service,
paid backup product, or cloud account is required.
For restore drills, a separate target cluster is created only after the source containers are stopped, so a laptop does not need to run six Cassandra nodes simultaneously. A functional one-node fallback is acceptable on a resource-constrained machine, but it cannot reproduce RF=3 replica/repair behavior. Plan roughly 6–8 GiB of free RAM and several GiB of free disk for the three-node exercises, and capture actual container/host limits in your evidence. Windows learners should use Docker Desktop/WSL-style Linux containers; commands that manipulate Linux inodes run inside the Cassandra container.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and recovery mental model
An SSTable (Sorted String Table) is Cassandra's
immutable on-disk representation of flushed table data. A
snapshot is a node-local point-in-time set of
hard links to SSTable component files plus metadata such as
schema.cql; a hard link is another directory entry
referencing the same filesystem inode/blocks. An
incremental backup is a hard link placed in a
table's backups/ directory whenever a new SSTable
is flushed or streamed after incremental backup is enabled.
Neither mechanism automatically places data in a different
failure domain.
A backup catalog records exactly which node/table/SSTable files, schema/configuration versions, timestamps, checksums, and recovery dependencies belong to a restore point. An off-host copy is a separately stored copy outside the node/volume failure domain; production designs often add account/region separation, encryption, access controls, immutability, and retention policy. A commit log records mutations before they are applied to memtables; commit-log archiving copies completed segments so they can be replayed during a point-in-time-oriented restore. PITR means Point-in-Time Recovery.
sstableloader reads backed-up SSTables from an offline utility process and streams their token ranges to the replicas in the current cluster topology. nodetool refresh tells one running node to discover newly placed compatible SSTables in that table's local data directory. repair is Cassandra's anti-entropy process for comparing replica ranges and streaming differences. RPO (Recovery Point Objective) is the tolerated amount of data loss measured in time; RTO (Recovery Time Objective) is the tolerated service-recovery duration.
1. Snapshot mechanics: immutable SSTables make hard links safe
Cassandra does not modify an existing SSTable in place. A normal
nodetool snapshot first flushes relevant memtables
unless --skip-flush is chosen, then creates
snapshot-directory hard links to the current SSTable components.
The live directory entry and snapshot directory entry can
reference the same inode and physical blocks. That is why
snapshot creation is normally fast and initially cheap in
additional bytes.
The subtlety is retention: once compaction replaces a live
SSTable, Cassandra can remove the live directory entry, but the
snapshot's hard link can keep the old inode/blocks alive.
Snapshot space can therefore grow as live storage evolves.
nodetool listsnapshots reports both
apparent/on-disk snapshot size and a “true size” style estimate;
interpret it with hard-link behavior rather than simply adding
file sizes.
| Artifact | Location/owner | What it gives you | What it does not give you |
|---|---|---|---|
| SSTable live files | one node/table data directory | current immutable disk state | remote durability |
| snapshot/ |
same node/filesystem by default | point-in-time hard links + metadata | protection from node/volume/site loss |
| schema.cql | inside snapshot | table/keyspace DDL needed for restore | roles, TLS keys, secrets, all cluster configuration |
| manifest.json | inside snapshot | snapshot component inventory metadata | cryptographic off-host integrity catalog |
| off-host vault | operator-designed separate failure domain | survival beyond local disk/node | automatic proof that restore works |
2. Create the AtlasMart base snapshot on every replica
# Verify an existing course cluster first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course cluster does not exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before adding peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three nodes are UN in dc1.docker exec atlasmart-cass-1 nodetool status
CREATE KEYSPACE IF NOT EXISTS atlasmart_backupWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_backup.orders_by_customer ( customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_backup.orders_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-08T08:00:00Z',00000000-0000-0000-0000-000000000101,'PAID',129.90,'base-A');INSERT INTO atlasmart_backup.orders_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-08T08:05:00Z',00000000-0000-0000-0000-000000000102,'SHIPPED',49.50,'base-B');SELECT customer_id,order_month,order_time,order_id,status,total,noteFROM atlasmart_backup.orders_by_customerWHERE customer_id='cust-42' AND order_month='2026-09-01';
# Explicit flush makes the learning boundary obvious; snapshot also flushes by default.for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_backup orders_by_customerdone# nodetool targets one node. Snapshot each node for a cluster recovery point.for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool snapshot -t ch25-base -cf orders_by_customer atlasmart_backupdonefor n in 1 2 3; do echo "=== node $n snapshots ===" docker exec atlasmart-cass-$n nodetool listsnapshotsdone
Expected evidence is a ch25-base entry on each
node. Exact sizes differ by file format, compression, background
compaction, and dataset. The command's success proves that each
contacted node created a local snapshot; it does
not prove that another failure domain contains the
files or that a future cluster can restore them.
3. Prove the hard-link relationship and schema capture
docker exec atlasmart-cass-1 sh -lc 'TABLE_DIR=$(find /var/lib/cassandra/data/atlasmart_backup -maxdepth 1 -type d -name "orders_by_customer-*" | head -1)SNAP_DIR="$TABLE_DIR/snapshots/ch25-base"echo "TABLE_DIR=$TABLE_DIR"echo "SNAP_DIR=$SNAP_DIR"echo "--- snapshot metadata ---"ls -l "$SNAP_DIR/schema.cql" "$SNAP_DIR/manifest.json"echo "--- hard-link evidence for Data components ---"for f in "$SNAP_DIR"/*-Data.db; do base=$(basename "$f") if [ -f "$TABLE_DIR/$base" ]; then stat -c "%i links=%h bytes=%s %n" "$TABLE_DIR/$base" "$f" fidoneecho "--- schema excerpt ---"sed -n "1,120p" "$SNAP_DIR/schema.cql"'
Matching inode numbers for a live component and the snapshot component are direct local-filesystem evidence that they share physical blocks. A link count above one supports the same conclusion. If background compaction has already removed the live filename, the snapshot can remain as the last link; that is precisely why clearing old snapshots can reclaim blocks later. Filesystem/platform behavior is observed inside the Linux container, not inferred from the Windows host filesystem.
Delete the container and its volume in an isolated test and the snapshot disappears with the live data. The repaired design adds a cataloged, checksummed copy outside the Cassandra node/volume failure domain, preserves schema/config/security dependencies, and proves restore in a separate target cluster.
4. Verification and safe reset
-
All three nodes are
UNand report Cassandra 5.0.9. - Each node lists
ch25-base. -
Each snapshot contains SSTable components,
schema.cql, andmanifest.json. - At least one component shows hard-link inode/link-count evidence where the live file still exists.
- No snapshot is called an off-host/DR backup yet.
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool clearsnapshot -t ch25-base atlasmart_backupdone# Do not clear it if you are continuing directly to Lesson 2.
Check your understanding
- Why is a snapshot initially space-efficient?
- Why can snapshots later prevent disk space from being reclaimed?
- Does schema.cql contain every recovery dependency?
- Does nodetool snapshot on node 1 snapshot node 2 and node 3?
- What converts this snapshot into stronger DR evidence?
Review the answers
1. It creates hard links to immutable SSTable components instead of duplicating the file blocks immediately.
2. Compaction may remove a live directory entry, but the snapshot hard link still references the old inode/blocks.
3. No. It captures schema DDL, not every role, secret, TLS artifact, JMX setting, deployment parameter, driver policy, or external dependency.
4. No. nodetool operates on the targeted node; a cluster restore point needs an orchestrated per-node policy.
5. A cataloged/checksummed copy in a separate failure domain plus a tested restore and application verification.
Production judgment
A backup design is an application-recovery design, not a file-copy checkbox. Record workload/partition shape, retention and TTL/delete rates, RF/consistency level (CL), node/DC/rack topology, SSTable format and compaction, repair cadence, schema/table IDs, authentication/authorization/TLS/JMX dependencies, driver/native-protocol compatibility, version/upgrade path, encryption/key-management dependencies, SAI/vector-index requirements, disk/network throughput, snapshot/incremental/commit-log cadence, off-host failure domains, immutable retention, checksum/catalog ownership, and restore permissions. A backup that cannot recreate schema/security/configuration or that exists only on the same host/volume is not sufficient disaster-recovery evidence.
Measure Recovery Point Objective (RPO) from the latest recoverable mutation to the incident/cutover point and Recovery Time Objective (RTO) from recovery declaration to verified application service—not from “copy finished.” Include schema creation, file transfer, SSTable load/streaming, index readiness, repair/convergence, application credentials/TLS, driver contact-point cutover, smoke tests, and rollback in RTO. Avoid universal backup intervals or retention numbers: they must derive from business loss tolerance, data volume, change rate, restore throughput, compliance, cost, and tested operational skill. Lesson 2 turns the local base snapshot into a cataloged backup chain and adds incremental SSTables, retention, remote-copy simulation, and integrity verification.
Summary and next bridge
A Cassandra snapshot is an efficient node-local hard-link view of immutable SSTables plus schema metadata. It is valuable, but locality is its defining limitation. Next, build the catalog/off-host/integrity layer that makes local backup artifacts usable in a disaster-recovery process.
Authoritative references
These are the version-sensitive source of truth for the mechanisms used in this lesson. Re-check them when regenerating or operating on a different Cassandra patch/distribution.