Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation

Read Amplification Across SSTables and Merging Multiple Versions / Tombstones

Create multiple immutable versions and tombstones, measure SSTables-per-read, and compare evidence before/after controlled compaction.

Intermediate110–150 minutesRead-amplification + compaction labApache Cassandra 5.0.9 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart updates the same order repeatedly and deletes obsolete fields. Immutable SSTables mean newer cells and tombstones can coexist with older versions in several files. This lesson makes the merge cost measurable.

01

Explain why updates/deletes create newer immutable versions rather than in-place edits.

02

Create several SSTables safely with deterministic keys and forced flushes only in the lab.

03

Measure SSTables/read, Bloom, partition and tombstone evidence before/after a controlled compaction.

04

Relate read amplification to compaction state and query-first partition design.

05

Separate logical correctness from physical I/O cost.

Chapter 10 lab baseline

The mandatory labs continue the established AtlasMart disposable cluster: cassandra:5.0.9, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 10 uses atlasmart_readpath with NetworkTopologyStrategy, replication factor (RF) 3, and normally LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS); table TTL defaults to zero and gc_grace_seconds is not changed. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside the isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional for client-routing discussion; all mandatory evidence uses free local cqlsh/nodetool. Capture nodetool version, cqlsh --version, and java -version on your machine rather than treating a prose baseline as runtime proof.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms before tracing a read

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.

1. Immutable files create read-time merge work

An UPDATE writes a newer cell version; DELETE writes a tombstone. Neither edits an old SSTable. A read merges memtable state and candidate SSTables, compares timestamps, applies deletion/TTL markers, and returns the newest visible value. Compaction later rewrites files and can remove shadowed versions when safe.

Read amplification is not merely “SSTables on disk.” It is storage work for the logical read: Bloom candidates, overlap, compaction strategy/state, clustering slice, tombstones, cache/page-cache state, and file format all matter.

2. Build three generations for one row

bash · verify or recreate the disposable three-node cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CQL · deterministic merge fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.read_merge_probe (  probe_key text, event_time timestamp, event_id uuid,  status text, note text,  PRIMARY KEY (probe_key,event_time,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_readpath.read_merge_probe VALUES('merge-42','2026-09-07T17:00:00Z',00000000-0000-0000-0000-000000000042,'CREATED','v1');
bash · flush generation 1 on all replicas
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath read_merge_probe; done
CQL · generation 2
UPDATE atlasmart_readpath.read_merge_probe SET status='PAID',note='v2'WHERE probe_key='merge-42' AND event_time='2026-09-07T17:00:00Z'AND event_id=00000000-0000-0000-0000-000000000042;
bash · flush generation 2
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath read_merge_probe; done
CQL · generation 3 plus tombstone
UPDATE atlasmart_readpath.read_merge_probe SET status='SHIPPED'WHERE probe_key='merge-42' AND event_time='2026-09-07T17:00:00Z'AND event_id=00000000-0000-0000-0000-000000000042;DELETE note FROM atlasmart_readpath.read_merge_probeWHERE probe_key='merge-42' AND event_time='2026-09-07T17:00:00Z'AND event_id=00000000-0000-0000-0000-000000000042;
bash · flush and capture baseline
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath read_merge_probe; donedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.read_merge_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath read_merge_probe

3. Read the merged result, then compact deliberately

CQL · trace reconciliation across physical versions
TRACING ON;SELECT status,note,writetime(status),writetime(note)FROM atlasmart_readpath.read_merge_probe WHERE probe_key='merge-42';TRACING OFF;

Expected logical state is status='SHIPPED' and no visible note. Do not assume a specific number of SSTables is touched: background compaction may already have changed files. Record tablehistograms immediately around the experiment.

bash · controlled compaction and re-measurement
docker exec atlasmart-cass-1 nodetool compact atlasmart_readpath read_merge_probedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.read_merge_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath read_merge_probe
Wrong conclusion: “Manual compaction improved my laptop read, so run it whenever p99 rises.”

Compaction consumes I/O, CPU, and disk headroom. Production root causes may be unbounded partitions, delete/TTL churn, insufficient compaction capacity, or the wrong read shape. Treat manual compaction as state-changing maintenance, not a universal tuning button.

4. Evidence and checks

Signal Before After Meaning
SSTables/read capture p50/p95/max capture again logical read amplification distribution
Bloom false positives cumulative counters compare carefully extra candidate work, not correctness
Partition rows/bytes same fixture same fixture schema shape unchanged
Read latency distribution distribution must control warmup/background work
Tombstones capture capture deletion merge cost; purge safety depends on repair/gc_grace

Check your understanding

  1. Why does an UPDATE create later merge work?
  2. What wins against an older live cell?
  3. Do three forced flushes guarantee three SSTables/read?
  4. Why use tablehistograms?
  5. Is manual compaction a universal fix?
Review the answers

1. SSTables are immutable; new cell versions can coexist with older versions until compaction rewrites them.

2. A newer tombstone shadows it according to timestamp/deletion semantics.

3. No. Background compaction, Bloom decisions, overlap, and caches can change the path.

4. It reports distributions including SSTables touched per logical read and partition/read latency statistics.

5. No; it is resource-intensive and may only hide a deeper schema/workload problem.

Production judgment

Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.

Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Lesson 4 moves from versions within a replica to divergent versions across replicas and distinguishes reconciliation, read repair, and scheduled repair.

Summary and next bridge

Immutable storage makes updates/deletes accumulate as physical versions. Reads merge those versions, so overlap and tombstones become measurable amplification. Next, create a stale replica and observe request-scoped reconciliation/read repair.

Authoritative references

Re-check these version-sensitive sources when regenerating the lesson.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.