Chapter 09 · The Write Path: Coordinator, Commit Log, Memtables, Replicas, and Acknowledgment

Commit Log for Durability and Memtables for In-Memory Sorted State

Separate Cassandra replica commit-log durability, memtable admission, fsync policy, and later SSTable creation with direct node-local evidence.

Intermediate100–140 minutesCommit log + memtable labApache Cassandra 5.0.9 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart has just proved that a coordinator waits for replica acknowledgments. Now the platform team asks the harder question: what must happen inside an acknowledging replica before it responds? Cassandra’s write path deliberately avoids in-place random updates. A replica records the mutation in a commit log for crash recovery and applies it to the table’s in-memory memtable; the immutable SSTable arrives later during flush.

01

Separate the append-only commit log from the per-table memtable and the later immutable SSTable.

02

Explain why the commit log is shared across tables while active memtables are table-specific.

03

Inspect the current Cassandra 5.0 commitlog_sync mode and connect periodic versus batch behavior to acknowledgment durability.

04

Use nodetool tablestats plus filesystem evidence to observe memtable data before and after a controlled flush.

05

Avoid the dangerous belief that every acknowledged write was already fsynced or already stored in an SSTable.

Chapter 09 lab baseline

The mandatory labs use the pinned cassandra:5.0.9 image and the Java 17, cqlsh, and nodetool versions bundled by that image. The shared course cluster is atlasmart-course with three disposable nodes (atlasmart-cass-1..3) on Docker network atlasmart-cassandra, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 09 uses keyspace atlasmart_writepath with NetworkTopologyStrategy, replication factor (RF) 3, and LOCAL_QUORUM unless an exercise deliberately changes consistency level. New tables explicitly use UnifiedCompactionStrategy (UCS); table TTL defaults to zero and gc_grace_seconds is not changed. Authentication, client TLS, internode TLS, and remote JMX are disabled only inside the isolated learning network. Apache Cassandra Java Driver 4.19.3 is optional for routing examples; every mandatory exercise remains free/local with cqlsh, nodetool, Docker, and shell commands.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

1. Two writes happen locally for one logical mutation

On each replica that accepts a normal table mutation, Cassandra appends the mutation to the node’s commit log and updates the table’s memtable. The commit log is sequential, append-oriented recovery data shared by many tables; the memtable is an in-memory, table-specific sorted structure that serves as the mutable representation of recent partitions.

This pairing is why Cassandra can acknowledge a write without performing an in-place update to an SSTable. SSTables are immutable. If the process crashes before a dirty memtable is flushed, Cassandra can replay eligible commit-log mutations during startup and reconstruct the in-memory state.

Component Scope Purpose Lifecycle
Commit log Node-wide segments shared by tables Crash recovery for mutations not yet safely represented in SSTables Segments can be recycled after their covered dirty data is flushed.
Memtable Per table Mutable sorted in-memory state; also participates in reads Switches/flushes when pressure or configured triggers require it.
SSTable Per table, immutable files Persistent sorted table storage used by the read path and compaction Created by flush/streaming and later compacted into new SSTables.

2. “Committed” is configuration-sensitive: periodic versus batch sync

Cassandra 5.0 documents commitlog_sync modes with different acknowledgment semantics. In periodic mode—the documented default in current 5.0 storage-engine configuration—the server may acknowledge after the mutation reaches the commit-log buffer, with fsync occurring periodically. In batch mode Cassandra does not acknowledge until the commit log has been fsynced. Therefore a production durability statement must name the mode, filesystem/device assumptions, RF, CL, and correlated-failure model.

Dangerous shortcut

Do not write “CL QUORUM means two replicas fsynced the write” unless you have verified the replica commit-log sync mode and storage semantics. CL counts acknowledgments; the local durability condition behind each acknowledgment is configuration-dependent.

Mode Acknowledgment relation to fsync Operational consequence
periodic Acknowledgment can precede the periodic fsync. Lower write latency, but a simultaneous crash/power-loss failure model must include the unsynced window.
batch Acknowledgment waits for commit-log fsync. Stronger per-ack stable-storage condition with fsync latency in the write path.

3. Lab: observe commit-log configuration and memtable state

bash · verify or recreate the disposable three-node cluster
# Verify the shared lab if it already exists.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool status# Standalone recreation path. Skip resources that already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 answers before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait until all three nodes show UN in dc1 before continuing.docker exec atlasmart-cass-1 nodetool status# Open cqlsh on node 1 for the SQL/CQL blocks that follow.docker exec -it atlasmart-cass-1 cqlsh
sql · create the isolated write-path keyspace and table
CREATE KEYSPACE IF NOT EXISTS atlasmart_writepathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writepath.order_write_events (    order_id text,    event_time timeuuid,    status text,    source text,    note text,    expires_at timestamp,    PRIMARY KEY ((order_id), event_time)) WITH CLUSTERING ORDER BY (event_time DESC)  AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name = 'atlasmart_writepath';
bash · inspect commit-log configuration and table statistics
# Configuration file location is from the official Docker image layout.docker exec atlasmart-cass-1 sh -lc "grep -E '^[[:space:]]*commitlog_sync:|^[[:space:]]*commitlog_sync_period:' /etc/cassandra/cassandra.yaml || true"# Inspect current commit-log files and disk footprint. Exact names/sizes vary.docker exec atlasmart-cass-1 sh -lc 'du -sh /var/lib/cassandra/commitlog; ls -lh /var/lib/cassandra/commitlog | tail -n +1'# Baseline table statistics before adding this lesson's rows.docker exec atlasmart-cass-1 nodetool tablestats atlasmart_writepath.order_write_events
sql · write several rows without manually flushing
CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_writepath.order_write_events(order_id,event_time,status,source,note)VALUES ('order-2001',now(),'CREATED','checkout-api','memtable evidence A');INSERT INTO atlasmart_writepath.order_write_events(order_id,event_time,status,source,note)VALUES ('order-2002',now(),'CREATED','checkout-api','memtable evidence B');INSERT INTO atlasmart_writepath.order_write_events(order_id,event_time,status,source,note)VALUES ('order-2003',now(),'CREATED','checkout-api','memtable evidence C');
bash · capture memtable and commit-log evidence before flush
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_writepath.order_write_eventsdocker exec atlasmart-cass-1 sh -lc 'du -sh /var/lib/cassandra/commitlog'# In tablestats, record memtable_cell_count, memtable_data_size,# memtable_switch_count, pending_flushes, local_write_count/latency, and sstable_count.# Do not copy sample numbers: capture your own values.

Because RF=3, a specific row may have been coordinated by one node and replicated to all three. Local tablestats is node-local evidence; compare nodes if you need replica-level visibility. A small fixture can be affected by background flushes, so treat “memtable count is exactly N” as a hypothesis, not a guaranteed constant.

4. Flush converts mutable memory into immutable SSTables

A manual nodetool flush is appropriate here only because the cluster and keyspace are disposable and the lesson needs deterministic evidence. It is not routine performance tuning. Flush writes the current memtable to an SSTable, creates associated index/metadata components, and allows commit-log segments to become recyclable only when all relevant dirty data in those shared segments is covered.

bash · force one lab flush and inspect before/after state
# Record the data directories first.docker exec atlasmart-cass-1 nodetool datapaths atlasmart_writepath.order_write_events# Disposable-lab-only deterministic transition.docker exec atlasmart-cass-1 nodetool flush atlasmart_writepath order_write_events# Observe the new state.docker exec atlasmart-cass-1 nodetool tablestats atlasmart_writepath.order_write_eventsdocker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_writepath -type f -name '*Data.db' -printf '%p %s bytes\\n' 2>/dev/null | sort"docker exec atlasmart-cass-1 sh -lc 'du -sh /var/lib/cassandra/commitlog' 

Expect the active memtable for that table to become smaller or reset and the SSTable count/files to change. Do not expect every commit-log file to disappear: commit-log segments are shared across tables and their reclamation depends on all mutations covered by each segment, not only this table.

5. Why disabling the commit log is not a casual latency trick

Cassandra exposes durable_writes at the keyspace level. Turning durability off changes the recovery contract and is not equivalent to “use a faster disk.” A memtable exists only in process memory until flush; without the recovery log, process or host failure before flush can discard acknowledged data. This chapter does not disable durability because the failure mechanism is obvious without risking confusing recovery state.

In production, pair commit-log mode with dedicated storage design, filesystem/device behavior, write SLOs, RF/CL, node-failure correlation, and recovery tests. Watch local write latency, commit-log device latency, pending flushes, memtable memory, dropped mutations, JVM/GC pressure, and commit-log space together.

Verification checklist

  • You recorded the actual commitlog_sync setting from the running image.
  • You observed nonzero write activity in tablestats.
  • You captured memtable statistics before the manual flush.
  • You observed SSTable/data-directory evidence after the flush.
  • You can explain why commit-log disk usage need not drop immediately after flushing one table.

Check your understanding

  1. Why does Cassandra write both the commit log and memtable?
  2. Is an SSTable required before a normal write can be acknowledged?
  3. Under periodic commitlog_sync, must an acknowledgment wait for fsync?
  4. Why might a commit-log segment remain after flushing one table?
  5. Why is nodetool flush a poor routine tuning tactic?
Review the answers

1. The commit log provides crash-recovery history; the memtable is the mutable in-memory table state used before flush and by reads.

2. No. The write path acknowledges based on replica/commit-log/memtable semantics and CL; SSTables are produced later by flush.

3. No. Current 5.0 documentation says periodic mode can acknowledge before periodic fsync.

4. Segments can contain mutations for multiple tables; reclamation waits until the relevant data in the segment has been flushed.

5. Frequent forced flushes can create many small SSTables and shift cost into reads/compaction instead of solving the underlying pressure.

Summary and next bridge

Replica acknowledgment is backed by a local commit-log/memtable transition, but fsync guarantees depend on commit-log sync mode and SSTables arrive later. Next we examine the triggers that switch memtables, the disk transition itself, and the backpressure signals that appear when flushing cannot keep up with incoming writes.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.