Chapter 25 · Snapshots, Backups, Commitlog Archiving, Restore, and Disaster Recovery
Commitlog Archiving / Restore for Point-in-Time-Oriented Recovery Workflows
Use commit-log archiving to understand point-in-time-oriented recovery, timestamp cutoffs, and the requirement for a compatible base state.
Learning outcomes
AtlasMart's last full/incremental SSTable copy completed at 09:00, but an accidental delete happened at 09:17. The business asks to recover to 09:16:30. SSTable backups alone may leave a larger RPO; archived commit-log segments can replay mutations after the base point when the environment is designed for it.
Explain archive_command, restore_command, restore_directories, restore_point_in_time, and timestamp precision.
Relate archived commit-log segments to a compatible base SSTable/schema recovery point.
Build an isolated single-node PITR sandbox without changing the shared course cluster.
Explain why one tiny write may not immediately produce an archived segment and why segment rollover matters.
Identify timestamp, schema/table-ID, multi-node, and operational boundaries that make PITR harder than file restore.
The mandatory labs use Apache Cassandra 5.0.9 in
the pinned cassandra:5.0.9 container image. The
normal source cluster keeps the course conventions: cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy with replication
factor (RF) 3, and LOCAL_QUORUM for the chapter's
verification reads/writes. Tables explicitly use
UnifiedCompactionStrategy (UCS); authentication,
client/internode TLS, and remote JMX remain disabled only
inside this isolated local learning topology. The chapter
keyspace is atlasmart_backup. No managed service,
paid backup product, or cloud account is required.
For restore drills, a separate target cluster is created only after the source containers are stopped, so a laptop does not need to run six Cassandra nodes simultaneously. A functional one-node fallback is acceptable on a resource-constrained machine, but it cannot reproduce RF=3 replica/repair behavior. Plan roughly 6–8 GiB of free RAM and several GiB of free disk for the three-node exercises, and capture actual container/host limits in your evidence. Windows learners should use Docker Desktop/WSL-style Linux containers; commands that manipulate Linux inodes run inside the Cassandra container.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and recovery mental model
An SSTable (Sorted String Table) is Cassandra's
immutable on-disk representation of flushed table data. A
snapshot is a node-local point-in-time set of
hard links to SSTable component files plus metadata such as
schema.cql; a hard link is another directory entry
referencing the same filesystem inode/blocks. An
incremental backup is a hard link placed in a
table's backups/ directory whenever a new SSTable
is flushed or streamed after incremental backup is enabled.
Neither mechanism automatically places data in a different
failure domain.
A backup catalog records exactly which node/table/SSTable files, schema/configuration versions, timestamps, checksums, and recovery dependencies belong to a restore point. An off-host copy is a separately stored copy outside the node/volume failure domain; production designs often add account/region separation, encryption, access controls, immutability, and retention policy. A commit log records mutations before they are applied to memtables; commit-log archiving copies completed segments so they can be replayed during a point-in-time-oriented restore. PITR means Point-in-Time Recovery.
sstableloader reads backed-up SSTables from an offline utility process and streams their token ranges to the replicas in the current cluster topology. nodetool refresh tells one running node to discover newly placed compatible SSTables in that table's local data directory. repair is Cassandra's anti-entropy process for comparing replica ranges and streaming differences. RPO (Recovery Point Objective) is the tolerated amount of data loss measured in time; RTO (Recovery Time Objective) is the tolerated service-recovery duration.
1. Commit-log archiving is a replay stream, not a replacement for a base backup
Commit-log segments contain mutation records used for crash
recovery. Cassandra can call an operator-defined
archive_command when a segment becomes eligible for
archival. During recovery,
restore_directories identifies archived segments,
restore_command copies a selected segment to
Cassandra's live replay location, and
restore_point_in_time limits replay to mutations
whose timestamps are at or before the configured GMT time.
precision must match the timestamp precision being
interpreted.
Archived commit logs are meaningful only with a compatible starting state. Replaying them into the wrong schema/table identifiers, wrong Cassandra compatibility boundary, or unrelated dataset is not a general event-sourcing mechanism. They also do not generate embeddings, recreate roles/TLS keys, or repair replica divergence. In a multi-node Cassandra cluster, each node has its own commit-log history and replica set, so a production PITR design requires a tested per-node/cluster strategy rather than copying one node's segments and assuming global completeness.
| Property | Role | Failure if misunderstood |
|---|---|---|
| archive_command | copy completed/recycled segment to archive storage | no archive if command fails or storage is unavailable |
| restore_directories | locations Cassandra scans for archived segments | wrong directory means segments are not found |
| restore_command | copy selected archived segment into replay destination | permissions/path errors block recovery |
| restore_point_in_time | upper timestamp bound for applied mutations | bad clock/timestamp semantics recover wrong business state |
| precision | MILLISECONDS or MICROSECONDS interpretation | mismatch changes cutoff behavior |
2. Build a separate one-node PITR mechanics sandbox
This sandbox is intentionally separate because commit-log archiving is startup configuration. It demonstrates mechanism, not production-grade multi-node PITR. Create a tiny derived image with an archive directory owned by Cassandra and a pinned archiving file.
archive_command=/bin/cp -f %path /archive/%namerestore_command=restore_directories=restore_point_in_time=precision=MICROSECONDS
FROM cassandra:5.0.9USER rootRUN mkdir -p /archive && chown cassandra:cassandra /archiveCOPY commitlog-archiving.properties /etc/cassandra/commitlog-archiving.propertiesRUN chown cassandra:cassandra /etc/cassandra/commitlog-archiving.properties
# Put the two files above in ./ch25-pitr-source/docker build -t atlasmart-cassandra-pitr:5.0.9 ./ch25-pitr-sourcedocker network inspect atlasmart-pitr >/dev/null 2>&1 || docker network create atlasmart-pitrdocker volume create atlasmart-pitr-datadocker volume create atlasmart-pitr-archivedocker run -d --name atlasmart-pitr-1 --hostname atlasmart-pitr-1 \ --network atlasmart-pitr \ -e CASSANDRA_CLUSTER_NAME=atlasmart-pitr-source \ -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch \ -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-pitr-data:/var/lib/cassandra \ -v atlasmart-pitr-archive:/archive \ atlasmart-cassandra-pitr:5.0.9docker exec atlasmart-pitr-1 nodetool statusdocker exec atlasmart-pitr-1 nodetool version
3. Create a base, then mutations around a candidate recovery point
CREATE KEYSPACE IF NOT EXISTS atlasmart_pitrWITH replication = {'class':'NetworkTopologyStrategy','dc1':1};CREATE TABLE IF NOT EXISTS atlasmart_pitr.account_state ( account_id text PRIMARY KEY, balance decimal, note text) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY ONE;INSERT INTO atlasmart_pitr.account_state (account_id,balance,note)USING TIMESTAMP 1788854400000000VALUES ('acct-42',100.00,'base');SELECT writetime(balance),balance,note FROM atlasmart_pitr.account_stateWHERE account_id='acct-42';
docker exec atlasmart-pitr-1 nodetool flush atlasmart_pitr account_statedocker exec atlasmart-pitr-1 nodetool snapshot -t pitr-base -cf account_state atlasmart_pitrdocker exec atlasmart-pitr-1 nodetool listsnapshots
UPDATE atlasmart_pitr.account_stateUSING TIMESTAMP 1788854460000000SET balance=90.00,note='wanted-before-cutoff'WHERE account_id='acct-42';UPDATE atlasmart_pitr.account_stateUSING TIMESTAMP 1788854520000000SET balance=0.00,note='accidental-after-cutoff'WHERE account_id='acct-42';
These microsecond timestamps are deterministic lab metadata, not wall-clock proof. The restore cutoff must be chosen consistently with the timestamp policy. A small workload may not immediately archive a segment because commit-log archiving happens when segments roll/become recyclable. Generate enough controlled writes or perform a controlled drain/restart and then verify actual archive files; do not declare success because configuration text exists.
docker exec atlasmart-pitr-1 sh -lc 'ls -lh /archive || true'docker exec atlasmart-pitr-1 sh -lc 'grep -v "^[[:space:]]*#" /etc/cassandra/commitlog_archiving.properties | grep -v "^[[:space:]]*$"'# Optional lab-only drain after your writes to close/flush state; restart is required afterward.docker exec atlasmart-pitr-1 nodetool draindocker restart atlasmart-pitr-1docker exec atlasmart-pitr-1 sh -lc 'ls -lh /archive || true'
4. Restore-point semantics and the deliberate failure case
For a real restore, build a new sandbox from the base
snapshot with the original schema/table identity, place archived
segments in the restore directory, and start Cassandra with
restore_command, restore_directories,
and a GMT restore_point_in_time. Cassandra replays
mutations up to the cutoff. The exact operational sequence
depends on packaging/data-directory layout and should be
rehearsed against the target patch before an incident.
archive_command=restore_command=/bin/cp -f %from %torestore_directories=/archiverestore_point_in_time=2026:09:08 09:16:30precision=MICROSECONDS
Commit-log records are tied to Cassandra's mutation/schema identity and timestamp semantics. Restore must start from a compatible base, preserve the necessary table IDs/schema, use a cutoff consistent with timestamp precision, and then verify application state. In multi-node production, test the coordinated process for every replica/source and failure mode; this one-node sandbox only demonstrates the mechanism.
Check your understanding
- Why do archived commit logs not replace snapshots/incremental backups?
- Why might no archive file appear after one small write?
- What does restore_point_in_time compare against?
- Why is one-node PITR not proof of multi-node DR?
- What is the safe evidence of success?
Review the answers
1. They are replay records after/around a base state; recovery still needs a compatible starting SSTable/schema state.
2. A commit-log segment may still be active; archival occurs as segments roll/become eligible, so verify actual segment lifecycle.
3. Mutation timestamps interpreted according to the configured precision and GMT cutoff semantics.
4. Production replicas have independent commit-log histories, topology/repair requirements, and coordinated application cutover.
5. Recovered business rows/versions match the intended cutoff, logs show replay as expected, and post-restore convergence/application checks pass.
Production judgment
A backup design is an application-recovery design, not a file-copy checkbox. Record workload/partition shape, retention and TTL/delete rates, RF/consistency level (CL), node/DC/rack topology, SSTable format and compaction, repair cadence, schema/table IDs, authentication/authorization/TLS/JMX dependencies, driver/native-protocol compatibility, version/upgrade path, encryption/key-management dependencies, SAI/vector-index requirements, disk/network throughput, snapshot/incremental/commit-log cadence, off-host failure domains, immutable retention, checksum/catalog ownership, and restore permissions. A backup that cannot recreate schema/security/configuration or that exists only on the same host/volume is not sufficient disaster-recovery evidence.
Measure Recovery Point Objective (RPO) from the latest recoverable mutation to the incident/cutover point and Recovery Time Objective (RTO) from recovery declaration to verified application service—not from “copy finished.” Include schema creation, file transfer, SSTable load/streaming, index readiness, repair/convergence, application credentials/TLS, driver contact-point cutover, smoke tests, and rollback in RTO. Avoid universal backup intervals or retention numbers: they must derive from business loss tolerance, data volume, change rate, restore throughput, compliance, cost, and tested operational skill. Lesson 4 returns to SSTable restores and compares topology-aware sstableloader, node-local refresh/import, and node-replacement workflows so you can choose a restore path deliberately.
Summary and next bridge
Commit-log archiving can shrink the recovery-point gap, but only when paired with a compatible base, correct schema identity, disciplined timestamps, archived-segment integrity, and tested replay. Next, compare the concrete SSTable restore mechanisms and their topology assumptions.
Authoritative references
These are the version-sensitive source of truth for the mechanisms used in this lesson. Re-check them when regenerating or operating on a different Cassandra patch/distribution.