Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation
Diagnose a Slow Read by Partition Size, SSTable Count, Tombstones, Cache, and Disk Evidence
Diagnose a slow read using partition size, SSTable count, tombstones, Bloom/cache, replica, and disk evidence.
Learning outcomes
AtlasMart receives a p99 alert for an order-read endpoint. Instead of immediately enlarging caches or blaming disk, the team will collect evidence from the distributed request, partition shape, SSTable amplification, tombstones, cache/Bloom state, and host resources.
Apply an evidence-first slow-read workflow.
Interpret tracing, tablestats, tablehistograms, Bloom/cache and partition evidence together.
Separate warm-cache laptop artifacts from production steady-state behavior.
Compare controlled physical compaction with a query-first schema improvement.
Produce an incident record with version/topology/workload/rollback assumptions.
The mandatory labs continue the established AtlasMart
disposable cluster: cassandra:5.0.9, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. Chapter 10 uses atlasmart_readpath with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. New tables
explicitly use UnifiedCompactionStrategy (UCS); table TTL
defaults to zero and gc_grace_seconds is not
changed. Authentication, client TLS, internode TLS, and remote
JMX remain disabled only inside the isolated learning network.
The Apache Cassandra Java Driver 4.19.3 is optional for
client-routing discussion; all mandatory evidence uses free
local cqlsh/nodetool. Capture
nodetool version, cqlsh --version,
and java -version on your machine rather than
treating a prose baseline as runtime proof.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms before tracing a read
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.
1. Diagnose in layers
| Layer | Question | Evidence |
|---|---|---|
| Distributed request | coordinator/replicas/CL/mismatch? | tracing, status, timeout counters |
| Partition | rows/bytes/skew/bounded slice? | tablehistograms + schema/query |
| SSTables | files touched/overlap? | tablehistograms, tablestats |
| Tombstones | delete/TTL merge burden? | tablestats, warnings/traces |
| Bloom/cache | false positives or reusable hot data? | tablestats, nodetool info/config |
| Host/disk | queueing/CPU/GC/network? | container/OS metrics + Cassandra latency |
Each layer can explain the next. A large partition with tombstones can cause disk work despite excellent Bloom behavior. Many overlapping SSTables can make a small logical read expensive. A tiny warm working set can make local cache measurements look unrealistically good.
2. Build a bounded but amplification-prone fixture
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.slow_read_probe ( bucket text, event_time timestamp, event_id uuid, payload text, obsolete text, PRIMARY KEY (bucket,event_time,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_readpath.slow_read_probe VALUES ('b-2026-09-07','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'p1','x1');INSERT INTO atlasmart_readpath.slow_read_probe VALUES ('b-2026-09-07','2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'p2','x2');INSERT INTO atlasmart_readpath.slow_read_probe VALUES ('b-2026-09-07','2026-09-07T18:02:00Z',00000000-0000-0000-0000-000000000003,'p3','x3');
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath slow_read_probe; done
UPDATE atlasmart_readpath.slow_read_probe SET payload='p2-new'WHERE bucket='b-2026-09-07' AND event_time='2026-09-07T18:01:00Z'AND event_id=00000000-0000-0000-0000-000000000002;DELETE obsolete FROM atlasmart_readpath.slow_read_probeWHERE bucket='b-2026-09-07' AND event_time='2026-09-07T18:00:00Z'AND event_id=00000000-0000-0000-0000-000000000001;
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath slow_read_probe; donedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.slow_read_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath slow_read_probe
3. Trace, then separate warmup from evidence
CONSISTENCY LOCAL_QUORUM;TRACING ON;SELECT event_time,event_id,payload FROM atlasmart_readpath.slow_read_probeWHERE bucket='b-2026-09-07'AND event_time >= '2026-09-07T18:00:00Z'AND event_time < '2026-09-07T18:03:00Z';TRACING OFF;
docker exec atlasmart-cass-1 nodetool info | grep -i -A4 -B1 cache || truedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.slow_read_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath slow_read_probedocker stats --no-stream atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3
Repeat untraced reads only to establish a distribution; report first/cold-ish and later/warm observations separately. OS page cache, Cassandra caches, background compaction, and container scheduling can all change between attempts.
4. Compare physical rewrite with model improvement
docker exec atlasmart-cass-1 nodetool compact atlasmart_readpath slow_read_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath slow_read_probedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.slow_read_probe
Compaction changes the physical file set. A query-first redesign changes the unit of data a request addresses. For an ever-growing time partition, an hour/day bucket can bound partition growth if the application can route to the right bucket and accept bounded fan-out.
CREATE TABLE IF NOT EXISTS atlasmart_readpath.slow_read_probe_hourly ( bucket_day date, bucket_hour tinyint, event_time timestamp, event_id uuid, payload text, PRIMARY KEY ((bucket_day,bucket_hour),event_time,event_id)) WITH CLUSTERING ORDER BY (event_time DESC,event_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_readpath.slow_read_probe_hourly VALUES('2026-09-07',18,'2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'p2-new');SELECT * FROM atlasmart_readpath.slow_read_probe_hourlyWHERE bucket_day='2026-09-07' AND bucket_hour=18;
Prove disk rather than replica wait, partition size, SSTable overlap, tombstone scanning, compaction pressure, or query shape is dominant. Cache changes redistribute memory pressure and can improve a warm micro-fixture without fixing production p99.
5. Incident record and acceptance checks
| Record | Minimum context |
|---|---|
| Version/topology | Cassandra patch, Java/runtime, SSTable format, node/DC/rack, RF/CL, driver |
| Query | partition/clustering predicates, page size, result rows/bytes |
| Partition distribution | rows/bytes p50/p95/max, skew/hot-key evidence |
| SSTable/compaction | strategy, SSTables/read distribution, pending compactions/repaired state |
| Deletes/TTL | tombstones/read, gc_grace, TTL/delete rate, repair cadence |
| Bloom/cache | false positives/ratio, key/chunk cache state, warmup |
| Host | disk latency/queue, CPU, memory/GC, network |
| Outcome | before/after distributions, rollback, schema/operational change |
Check your understanding
- Why not tune cache from one slow query?
- Why is the minimum repeated-read latency misleading?
- Which histogram directly helps quantify SSTable read amplification?
- Why compare compaction and bucketing separately?
- When is the incident actually fixed?
Review the answers
1. The dominant cost may be replica wait, partition size, SSTable amplification, tombstones, compaction, disk, GC, or query shape.
2. Later runs can benefit from a tiny warm working set and multiple cache layers unlike production steady state.
3. The SSTables-per-read column in nodetool tablehistograms, interpreted with tablestats and tracing.
4. Compaction rewrites physical files for the same model; bucketing changes partition boundaries and the read unit.
5. Representative latency distributions meet the SLO while partition/SSTable/tombstone/resource/failure metrics remain acceptable and rollback is understood.
Production judgment
Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.
Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Chapter 11 opens the SSTable itself and examines its Data, index/trie, filter, statistics, compression, TOC, and streaming details.
Summary and next bridge
Slow-read diagnosis should move from distributed request evidence to partition shape, SSTable/tombstone amplification, caches/Bloom filters, and host I/O. Chapter 11 now examines SSTable components and on-disk storage internals directly.
Authoritative references
Re-check these version-sensitive sources when regenerating the lesson.