Chapter 09 · The Write Path: Coordinator, Commit Log, Memtables, Replicas, and Acknowledgment
Flush Triggers, Memtable-to-SSTable Transition, and Backpressure Under Write Load
Observe Cassandra memtable flush triggers, SSTable creation, pending flushes, and write backpressure without turning manual flush into a tuning recipe.
Learning outcomes
AtlasMart’s traffic surge is not failing at token routing; local write latency rises while memtable and disk work accumulate. The platform team needs to distinguish normal asynchronous flush from flush pressure, and then distinguish both from compaction. This lesson follows the memtable-to-SSTable boundary and treats backpressure as an observable capacity signal rather than a reason to force more flushes manually.
Name the major memtable flush triggers without assuming one universal byte threshold.
Use pending_flushes, memtable statistics, SSTable count, and local write latency as correlated evidence.
Explain how commit-log space pressure can force dirty memtables to flush so old segments can be recycled.
Distinguish flush from compaction and explain why frequent manual flushes can make downstream work worse.
Design a safe write-load investigation without trying to saturate a learner laptop or invent throughput numbers.
The mandatory labs use the pinned
cassandra:5.0.9 image and the Java 17,
cqlsh, and nodetool versions bundled
by that image. The shared course cluster is
atlasmart-course with three disposable nodes
(atlasmart-cass-1..3) on Docker network
atlasmart-cassandra, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. Chapter 09 uses keyspace
atlasmart_writepath with
NetworkTopologyStrategy, replication factor (RF)
3, and LOCAL_QUORUM unless an exercise
deliberately changes consistency level. New tables explicitly
use UnifiedCompactionStrategy (UCS); table TTL defaults to
zero and gc_grace_seconds is not changed.
Authentication, client TLS, internode TLS, and remote JMX are
disabled only inside the isolated learning network. Apache
Cassandra Java Driver 4.19.3 is optional for routing examples;
every mandatory exercise remains free/local with
cqlsh, nodetool, Docker, and shell
commands.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
1. Flush is lifecycle management, not a durability synonym
An active memtable eventually becomes a flushing memtable and is
written as immutable SSTable components. Triggers include memory
pressure, commit-log pressure, explicit table flush periods if
configured, shutdown/drain, and administrator commands such as
nodetool flush. Exact memory thresholds depend on
node configuration and current memory accounting, so a course
should not hard-code a magic “flush at X MB” rule.
The table option
memtable_flush_period_in_ms defaults to zero,
meaning there is no periodic time-based table flush solely from
that option. A zero value does not mean “never flush”; memory
and commit-log triggers still apply.
| Signal | Interpretation | What to correlate |
|---|---|---|
| memtable_data_size / memtable_cell_count | How much mutable table state this node currently holds. | Write rate, partition distribution, heap/off-heap configuration. |
| pending_flushes | Flush work waiting for this table. | Disk latency/throughput, flush writer saturation, commit-log pressure. |
| sstable_count | Immutable table files currently participating in reads/compaction. | Compaction strategy, flush cadence, read amplification. |
| local_write_latency_ms / write latency histograms | Node-local storage-path latency. | Commit-log fsync/device latency, queueing, mutation size, CPU/GC. |
2. Backpressure is the system refusing infinite ingestion
If dirty memory or commit-log pressure grows faster than storage can drain it, Cassandra cannot admit writes indefinitely. The storage engine can block or slow write admission while flush work catches up. That behavior protects bounded memory and recovery-log space; removing the pressure mechanism without fixing disk, mutation size, or workload shape would only move the failure elsewhere.
Do not publish a single throughput number from this chapter. A meaningful write-load result records payload size, partition-key distribution, RF/CL, commit-log mode/device, compaction strategy, concurrency, warmup, JVM/GC, disk/network, and p50/p95/p99/max latency. The mandatory lab therefore demonstrates the state transition with a small fixture and tells you where backpressure would appear rather than intentionally saturating a shared workstation.
3. Lab: capture the transition before and after flush
# Verify the shared lab if it already exists.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool status# Standalone recreation path. Skip resources that already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 answers before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait until all three nodes show UN in dc1 before continuing.docker exec atlasmart-cass-1 nodetool status# Open cqlsh on node 1 for the SQL/CQL blocks that follow.docker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_writepathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writepath.order_write_events ( order_id text, event_time timeuuid, status text, source text, note text, expires_at timestamp, PRIMARY KEY ((order_id), event_time)) WITH CLUSTERING ORDER BY (event_time DESC) AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name = 'atlasmart_writepath';
CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_writepath.order_write_events (order_id,event_time,status,source,note) VALUES ('burst-01',now(),'QUEUED','load-lab','bounded fixture');INSERT INTO atlasmart_writepath.order_write_events (order_id,event_time,status,source,note) VALUES ('burst-02',now(),'QUEUED','load-lab','bounded fixture');INSERT INTO atlasmart_writepath.order_write_events (order_id,event_time,status,source,note) VALUES ('burst-03',now(),'QUEUED','load-lab','bounded fixture');INSERT INTO atlasmart_writepath.order_write_events (order_id,event_time,status,source,note) VALUES ('burst-04',now(),'QUEUED','load-lab','bounded fixture');INSERT INTO atlasmart_writepath.order_write_events (order_id,event_time,status,source,note) VALUES ('burst-05',now(),'QUEUED','load-lab','bounded fixture');
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_writepath.order_write_eventsdocker exec atlasmart-cass-1 nodetool tpstats# Save memtable_data_size, memtable_cell_count, pending_flushes, sstable_count,# local_write_latency_ms and any flush-related pool names shown by your exact version.
docker exec atlasmart-cass-1 nodetool flush atlasmart_writepath order_write_eventsdocker exec atlasmart-cass-1 nodetool tablestats atlasmart_writepath.order_write_eventsdocker exec atlasmart-cass-1 nodetool tpstatsdocker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_writepath -type f -name '*Data.db' -printf '%p %s bytes\\n' 2>/dev/null | sort"
The correct conclusion is not “flush made writes faster.” You forced a lifecycle boundary. Compare memtable and SSTable evidence, then relate any real latency change to your environment. On a tiny fixture the counters may be too small to show pressure, which is itself preferable to manufacturing an unsafe saturation test.
4. Deliberately wrong approach: force flush after every micro-batch
A practitioner may see a dirty memtable and script
nodetool flush every few seconds. That reduces each
memtable’s residence time but creates more small SSTables,
increases metadata/file churn, and can increase read and
compaction work. Manual flushing treats the symptom, not the
capacity mismatch.
| Observed problem | Bad response | Mechanism-aware response |
|---|---|---|
| pending_flushes rising | Run more manual flush commands. | Check disk latency/throughput, mutation size, memtable/commit-log configuration, concurrency, and resource headroom. |
| many tiny SSTables | Force compaction repeatedly. | Find why flushes are too frequent; validate UCS/compaction behavior and write pattern before intervention. |
| write p99 spikes | Increase client timeout only. | Correlate commit-log fsync/device latency, flush queues, GC, network, dropped mutations, and replica health. |
| commit-log space pressure | Delete commit-log files. | Never delete active recovery files; find why dirty data cannot flush/recycle fast enough. |
5. Production judgment and capacity evidence
Backpressure is not a defect to disable; it is the boundary between accepted load and finite storage resources. Capacity plans need sustained write rate plus bursts, failure headroom, compaction/repair bandwidth, commit-log device characteristics, and free disk—not only steady-state averages. A node that is healthy with all replicas up may fail its SLO when one node is rebuilding or when compaction and repair compete for disk.
Verification checklist
- You captured table statistics before and after exactly one deliberate flush.
-
You can identify
pending_flushesandmemtable_data_sizein currenttablestatsoutput. - You can explain why a new SSTable is evidence of a flush, not a compaction result by itself.
- You did not saturate the host or use unsafe disk/clock manipulation.
- You can name at least three external factors required before a write-throughput number is meaningful.
Check your understanding
- Does memtable_flush_period_in_ms=0 disable all flushing?
- Why can commit-log pressure trigger memtable flushes?
- What does pending_flushes measure?
- Why can frequent manual flushes hurt later?
- Is backpressure itself the root cause?
Review the answers
1. No. It disables that periodic table timer; memory, commit-log, shutdown/drain, and manual triggers still exist.
2. Flushing dirty table data lets Cassandra make old commit-log segments recyclable once their covered mutations are safely represented in SSTables.
3. Queued/pending flush work for the table on that node; it should be interpreted with disk and write-latency signals.
4. They can create many small SSTables and increase read/compaction/file-management work.
5. Usually no. It is a protective response to a mismatch between incoming work and the rate at which bounded memory/log/storage resources can drain it.
Summary and next bridge
Flush converts mutable memtable state into immutable SSTables, and pressure around that boundary can push latency back toward writers. Next we look inside each cell mutation, where timestamps decide winners, TTL creates future deletion state, and tombstones make “delete” a replicated write rather than an immediate erase.
Authoritative references
- Apache Cassandra 5.0 — Storage Engine / write path, commit log, memtables, flushes
- Apache Cassandra 5.0 — Hinted handoff
- Apache Cassandra CQL — DML, TIMESTAMP, TTL, WRITETIME
- nodetool tablestats
- nodetool listpendinghints
- Apache Cassandra native protocol — write timeout/failure metadata
- Apache Cassandra downloads and current releases
- Apache Cassandra Java Driver