Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation
Read Amplification Across SSTables and Merging Multiple Versions / Tombstones
Create multiple immutable versions and tombstones, measure SSTables-per-read, and compare evidence before/after controlled compaction.
Learning outcomes
AtlasMart updates the same order repeatedly and deletes obsolete fields. Immutable SSTables mean newer cells and tombstones can coexist with older versions in several files. This lesson makes the merge cost measurable.
Explain why updates/deletes create newer immutable versions rather than in-place edits.
Create several SSTables safely with deterministic keys and forced flushes only in the lab.
Measure SSTables/read, Bloom, partition and tombstone evidence before/after a controlled compaction.
Relate read amplification to compaction state and query-first partition design.
Separate logical correctness from physical I/O cost.
The mandatory labs continue the established AtlasMart
disposable cluster: cassandra:5.0.9, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. Chapter 10 uses atlasmart_readpath with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. New tables
explicitly use UnifiedCompactionStrategy (UCS); table TTL
defaults to zero and gc_grace_seconds is not
changed. Authentication, client TLS, internode TLS, and remote
JMX remain disabled only inside the isolated learning network.
The Apache Cassandra Java Driver 4.19.3 is optional for
client-routing discussion; all mandatory evidence uses free
local cqlsh/nodetool. Capture
nodetool version, cqlsh --version,
and java -version on your machine rather than
treating a prose baseline as runtime proof.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms before tracing a read
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.
1. Immutable files create read-time merge work
An UPDATE writes a newer cell version; DELETE writes a tombstone. Neither edits an old SSTable. A read merges memtable state and candidate SSTables, compares timestamps, applies deletion/TTL markers, and returns the newest visible value. Compaction later rewrites files and can remove shadowed versions when safe.
Read amplification is not merely “SSTables on disk.” It is storage work for the logical read: Bloom candidates, overlap, compaction strategy/state, clustering slice, tombstones, cache/page-cache state, and file format all matter.
2. Build three generations for one row
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.read_merge_probe ( probe_key text, event_time timestamp, event_id uuid, status text, note text, PRIMARY KEY (probe_key,event_time,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_readpath.read_merge_probe VALUES('merge-42','2026-09-07T17:00:00Z',00000000-0000-0000-0000-000000000042,'CREATED','v1');
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath read_merge_probe; done
UPDATE atlasmart_readpath.read_merge_probe SET status='PAID',note='v2'WHERE probe_key='merge-42' AND event_time='2026-09-07T17:00:00Z'AND event_id=00000000-0000-0000-0000-000000000042;
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath read_merge_probe; done
UPDATE atlasmart_readpath.read_merge_probe SET status='SHIPPED'WHERE probe_key='merge-42' AND event_time='2026-09-07T17:00:00Z'AND event_id=00000000-0000-0000-0000-000000000042;DELETE note FROM atlasmart_readpath.read_merge_probeWHERE probe_key='merge-42' AND event_time='2026-09-07T17:00:00Z'AND event_id=00000000-0000-0000-0000-000000000042;
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath read_merge_probe; donedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.read_merge_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath read_merge_probe
3. Read the merged result, then compact deliberately
TRACING ON;SELECT status,note,writetime(status),writetime(note)FROM atlasmart_readpath.read_merge_probe WHERE probe_key='merge-42';TRACING OFF;
Expected logical state is status='SHIPPED' and no
visible note. Do not assume a specific number of
SSTables is touched: background compaction may already have
changed files. Record tablehistograms immediately
around the experiment.
docker exec atlasmart-cass-1 nodetool compact atlasmart_readpath read_merge_probedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.read_merge_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath read_merge_probe
Compaction consumes I/O, CPU, and disk headroom. Production root causes may be unbounded partitions, delete/TTL churn, insufficient compaction capacity, or the wrong read shape. Treat manual compaction as state-changing maintenance, not a universal tuning button.
4. Evidence and checks
| Signal | Before | After | Meaning |
|---|---|---|---|
| SSTables/read | capture p50/p95/max | capture again | logical read amplification distribution |
| Bloom false positives | cumulative counters | compare carefully | extra candidate work, not correctness |
| Partition rows/bytes | same fixture | same fixture | schema shape unchanged |
| Read latency | distribution | distribution | must control warmup/background work |
| Tombstones | capture | capture | deletion merge cost; purge safety depends on repair/gc_grace |
Check your understanding
- Why does an UPDATE create later merge work?
- What wins against an older live cell?
- Do three forced flushes guarantee three SSTables/read?
- Why use tablehistograms?
- Is manual compaction a universal fix?
Review the answers
1. SSTables are immutable; new cell versions can coexist with older versions until compaction rewrites them.
2. A newer tombstone shadows it according to timestamp/deletion semantics.
3. No. Background compaction, Bloom decisions, overlap, and caches can change the path.
4. It reports distributions including SSTables touched per logical read and partition/read latency statistics.
5. No; it is resource-intensive and may only hide a deeper schema/workload problem.
Production judgment
Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.
Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Lesson 4 moves from versions within a replica to divergent versions across replicas and distinguishes reconciliation, read repair, and scheduled repair.
Summary and next bridge
Immutable storage makes updates/deletes accumulate as physical versions. Reads merge those versions, so overlap and tombstones become measurable amplification. Next, create a stale replica and observe request-scoped reconciliation/read repair.
Authoritative references
Re-check these version-sensitive sources when regenerating the lesson.