Chapter 09 · The Write Path: Coordinator, Commit Log, Memtables, Replicas, and Acknowledgment
Timestamps, Conflict Resolution, Tombstones, TTL, and Last-Write-Wins at the Cell Level
Prove Cassandra cell-level timestamp reconciliation, TTL expiry metadata, and tombstone semantics with bounded AtlasMart mutations rather than arrival-order intuition.
Learning outcomes
Two AtlasMart services race to update the same order status. One retry arrives late with an older timestamp; another writes a TTL; a delete follows. Cassandra must reconcile replicas and SSTables without in-place overwrites, so each live cell or deletion marker carries metadata that determines which version wins. The application therefore needs timestamp discipline, not “last HTTP request wins” intuition.
Explain last-write-wins (LWW) as timestamp-based cell reconciliation rather than arrival-order semantics.
Use WRITETIME and explicit USING TIMESTAMP values to prove that an older mutation can be acknowledged yet remain invisible.
Explain TTL as expiring cell metadata that eventually produces deletion/tombstone state.
Distinguish a tombstone from immediate physical erasure and connect gc_grace/repair to safe deletion convergence.
Design client timestamp ownership so clock skew and retries do not resurrect stale business state.
The mandatory labs use the pinned
cassandra:5.0.9 image and the Java 17,
cqlsh, and nodetool versions bundled
by that image. The shared course cluster is
atlasmart-course with three disposable nodes
(atlasmart-cass-1..3) on Docker network
atlasmart-cassandra, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. Chapter 09 uses keyspace
atlasmart_writepath with
NetworkTopologyStrategy, replication factor (RF)
3, and LOCAL_QUORUM unless an exercise
deliberately changes consistency level. New tables explicitly
use UnifiedCompactionStrategy (UCS); table TTL defaults to
zero and gc_grace_seconds is not changed.
Authentication, client TLS, internode TLS, and remote JMX are
disabled only inside the isolated learning network. Apache
Cassandra Java Driver 4.19.3 is optional for routing examples;
every mandatory exercise remains free/local with
cqlsh, nodetool, Docker, and shell
commands.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
1. Cassandra resolves versions at cell granularity
For ordinary CQL writes, if no timestamp is supplied, the coordinator assigns a timestamp in microseconds near the start of statement execution. Cassandra stores mutation timestamps and reconciles competing versions primarily by timestamp. That means last write wins means “highest relevant mutation timestamp wins,” not “the request that physically arrived last at this node wins.”
Different non-primary-key cells in one row can have different timestamps and TTLs. A retry with an older explicit timestamp can still receive a successful CL acknowledgment because replicas accepted the mutation operation, yet the older value will lose during reconciliation to an already stored newer cell version.
If multiple services generate their own timestamps, skew or retry behavior can make stale data appear newer. Prefer a clearly owned timestamp policy and business sequence/version logic when domain ordering must be stronger than physical clock order. Avoid intentionally equal timestamps; tie behavior should not become application correctness logic.
2. Lab: prove an older acknowledged write can lose
# Verify the shared lab if it already exists.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool status# Standalone recreation path. Skip resources that already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 answers before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait until all three nodes show UN in dc1 before continuing.docker exec atlasmart-cass-1 nodetool status# Open cqlsh on node 1 for the SQL/CQL blocks that follow.docker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_writepathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writepath.order_write_events ( order_id text, event_time timeuuid, status text, source text, note text, expires_at timestamp, PRIMARY KEY ((order_id), event_time)) WITH CLUSTERING ORDER BY (event_time DESC) AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name = 'atlasmart_writepath';
CONSISTENCY LOCAL_QUORUM;-- Run these statements inside cqlsh on atlasmart-cass-1.INSERT INTO atlasmart_writepath.order_write_events(order_id,event_time,status,source,note)VALUES ('order-clock', 8d2c0f30-6d91-11f0-8000-000000000001, 'PAID', 'checkout-api', 'newer timestamp')USING TIMESTAMP 2000000;SELECT status, WRITETIME(status), source, WRITETIME(source)FROM atlasmart_writepath.order_write_eventsWHERE order_id='order-clock';UPDATE atlasmart_writepath.order_write_eventsUSING TIMESTAMP 1000000SET status='CREATED', note='older retry should lose for overlapping cells'WHERE order_id='order-clock' AND event_time=8d2c0f30-6d91-11f0-8000-000000000001;SELECT status, WRITETIME(status), note, WRITETIME(note)FROM atlasmart_writepath.order_write_eventsWHERE order_id='order-clock';
The overlapping status cell should retain the
version with timestamp 2,000,000. The note outcome
depends on its own prior timestamp because reconciliation is
cell-level. This is why “row timestamp” shorthand can hide
important behavior.
3. TTL is future deletion metadata, not a scheduled physical erase
Time to live (TTL) attaches an expiration time
to written values. TTL(column) reports remaining
seconds while a value is live. Updating a cell can reset or
replace its TTL according to the new mutation. After expiration,
Cassandra must continue to suppress older replicas/SSTables that
still contain the pre-expiry value, so expiration participates
in tombstone/deletion semantics rather than simply removing
bytes at the deadline.
INSERT INTO atlasmart_writepath.order_write_events(order_id,event_time,status,source,note)VALUES ('order-ttl', now(), 'TEMPORARY', 'lab', 'expires for the exercise')USING TTL 120;SELECT status, TTL(status), WRITETIME(status), note, TTL(note)FROM atlasmart_writepath.order_write_eventsWHERE order_id='order-ttl';# Re-run the SELECT after a short interval and record the lower TTL.# Do not wait for compaction or modify gc_grace_seconds merely to make tombstones disappear.
The TTL countdown is direct evidence of expiry metadata. It is not evidence that bytes will be removed from every SSTable at the exact expiration second. Physical reclamation depends on compaction and deletion-safety conditions.
4. Deletes write tombstones so replicas cannot resurrect old data
An immutable SSTable cannot be edited in place to remove a cell. Cassandra therefore writes a tombstone, a deletion marker with a timestamp, which wins over older live values during reconciliation. Tombstones remain relevant long enough for replicas to learn about the deletion. Repair is critical because a replica that misses a delete and later presents an older live value can otherwise participate in resurrection risks after deletion metadata is garbage-collected.
UPDATE atlasmart_writepath.order_write_eventsUSING TIMESTAMP 3000000SET status='PACKED'WHERE order_id='order-clock' AND event_time=8d2c0f30-6d91-11f0-8000-000000000001;DELETE status FROM atlasmart_writepath.order_write_eventsUSING TIMESTAMP 4000000WHERE order_id='order-clock' AND event_time=8d2c0f30-6d91-11f0-8000-000000000001;SELECT status, note, WRITETIME(note)FROM atlasmart_writepath.order_write_eventsWHERE order_id='order-clock';
# Force flush only to make a disposable lab SSTable visible.docker exec atlasmart-cass-1 nodetool flush atlasmart_writepath order_write_eventsdocker exec atlasmart-cass-1 nodetool tablestats atlasmart_writepath.order_write_events# Version-dependent diagnostic tools can inspect SSTable contents, but exact file paths# and tool output vary. First locate data files rather than hard-coding a table UUID.docker exec atlasmart-cass-1 nodetool datapaths atlasmart_writepath.order_write_eventsdocker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_writepath -type f -name '*Data.db' -printf '%p\\n' 2>/dev/null | sort"
Do not lower gc_grace_seconds just to make a demo
shorter. Tombstone safety depends on repair and failure windows.
“I deleted it and SELECT no longer returns it” proves
client-visible reconciliation, not immediate physical purge
across all replicas and SSTables.
5. Production judgment
Timestamp ownership should be documented alongside retry/idempotency rules. Client-generated timestamps can be useful but turn clock quality and request replay ordering into data-correctness dependencies. Monitor tombstone scan warnings, read latency, partition size, TTL/delete rates, repair completion, clock synchronization, write timeouts, and application retry storms together.
Verification checklist
-
WRITETIME(status)proves the higher-timestamp value won. - The older timestamp mutation does not replace the newer overlapping cell.
-
TTL()visibly decreases on the expiring fixture. - A cell delete becomes invisible to SELECT without pretending the old SSTable bytes vanished immediately.
gc_grace_secondsremains unchanged.
Check your understanding
- What does last-write-wins mean in Cassandra?
- Can an older explicit timestamp mutation receive a successful CL acknowledgment yet not become the visible value?
- Does TTL physically erase SSTable bytes at the expiration second?
- Why are tombstones necessary in an immutable replicated store?
- Why is clock skew dangerous with client timestamps?
Review the answers
1. For ordinary cell reconciliation, versions are ordered primarily by mutation timestamp, not request arrival order.
2. Yes. Replicas can accept/process it, but a newer stored timestamp wins reconciliation.
3. No. Expiry creates deletion semantics; physical reclamation happens later through compaction under deletion-safety rules.
4. They communicate that older live versions are deleted so reads/replicas do not resurrect them merely because old SSTables still exist.
5. A stale business event from a fast clock can carry a higher timestamp and overwrite logically newer state from another service.
Summary and next bridge
Cassandra’s write path carries more than values: timestamps decide winners, TTL schedules expiry, and tombstones preserve deletion across immutable files and replicas. The final lesson combines these storage semantics with tracing and controlled replica failure so every acknowledgment is related to a concrete consistency requirement and recovery path.
Authoritative references
- Apache Cassandra 5.0 — Storage Engine / write path, commit log, memtables, flushes
- Apache Cassandra 5.0 — Hinted handoff
- Apache Cassandra CQL — DML, TIMESTAMP, TTL, WRITETIME
- nodetool tablestats
- nodetool listpendinghints
- Apache Cassandra native protocol — write timeout/failure metadata
- Apache Cassandra downloads and current releases
- Apache Cassandra Java Driver