Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation

Diagnose a Slow Read by Partition Size, SSTable Count, Tombstones, Cache, and Disk Evidence

Diagnose a slow read using partition size, SSTable count, tombstones, Bloom/cache, replica, and disk evidence.

Intermediate120–160 minutesSlow-read diagnosis labApache Cassandra 5.0.9 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart receives a p99 alert for an order-read endpoint. Instead of immediately enlarging caches or blaming disk, the team will collect evidence from the distributed request, partition shape, SSTable amplification, tombstones, cache/Bloom state, and host resources.

01

Apply an evidence-first slow-read workflow.

02

Interpret tracing, tablestats, tablehistograms, Bloom/cache and partition evidence together.

03

Separate warm-cache laptop artifacts from production steady-state behavior.

04

Compare controlled physical compaction with a query-first schema improvement.

05

Produce an incident record with version/topology/workload/rollback assumptions.

Chapter 10 lab baseline

The mandatory labs continue the established AtlasMart disposable cluster: cassandra:5.0.9, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 10 uses atlasmart_readpath with NetworkTopologyStrategy, replication factor (RF) 3, and normally LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS); table TTL defaults to zero and gc_grace_seconds is not changed. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside the isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional for client-routing discussion; all mandatory evidence uses free local cqlsh/nodetool. Capture nodetool version, cqlsh --version, and java -version on your machine rather than treating a prose baseline as runtime proof.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms before tracing a read

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.

1. Diagnose in layers

Layer Question Evidence
Distributed request coordinator/replicas/CL/mismatch? tracing, status, timeout counters
Partition rows/bytes/skew/bounded slice? tablehistograms + schema/query
SSTables files touched/overlap? tablehistograms, tablestats
Tombstones delete/TTL merge burden? tablestats, warnings/traces
Bloom/cache false positives or reusable hot data? tablestats, nodetool info/config
Host/disk queueing/CPU/GC/network? container/OS metrics + Cassandra latency

Each layer can explain the next. A large partition with tombstones can cause disk work despite excellent Bloom behavior. Many overlapping SSTables can make a small logical read expensive. A tiny warm working set can make local cache measurements look unrealistically good.

2. Build a bounded but amplification-prone fixture

bash · verify or recreate the disposable three-node cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CQL · diagnostic fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.slow_read_probe (  bucket text, event_time timestamp, event_id uuid,  payload text, obsolete text,  PRIMARY KEY (bucket,event_time,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_readpath.slow_read_probe VALUES ('b-2026-09-07','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'p1','x1');INSERT INTO atlasmart_readpath.slow_read_probe VALUES ('b-2026-09-07','2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'p2','x2');INSERT INTO atlasmart_readpath.slow_read_probe VALUES ('b-2026-09-07','2026-09-07T18:02:00Z',00000000-0000-0000-0000-000000000003,'p3','x3');
bash · first flush
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath slow_read_probe; done
CQL · create newer versions and a tombstone
UPDATE atlasmart_readpath.slow_read_probe SET payload='p2-new'WHERE bucket='b-2026-09-07' AND event_time='2026-09-07T18:01:00Z'AND event_id=00000000-0000-0000-0000-000000000002;DELETE obsolete FROM atlasmart_readpath.slow_read_probeWHERE bucket='b-2026-09-07' AND event_time='2026-09-07T18:00:00Z'AND event_id=00000000-0000-0000-0000-000000000001;
bash · second flush and baseline statistics
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_readpath slow_read_probe; donedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.slow_read_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath slow_read_probe

3. Trace, then separate warmup from evidence

CQL · bounded traced slice
CONSISTENCY LOCAL_QUORUM;TRACING ON;SELECT event_time,event_id,payload FROM atlasmart_readpath.slow_read_probeWHERE bucket='b-2026-09-07'AND event_time >= '2026-09-07T18:00:00Z'AND event_time < '2026-09-07T18:03:00Z';TRACING OFF;
bash · capture cache/table/container state
docker exec atlasmart-cass-1 nodetool info | grep -i -A4 -B1 cache || truedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.slow_read_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath slow_read_probedocker stats --no-stream atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3

Repeat untraced reads only to establish a distribution; report first/cold-ish and later/warm observations separately. OS page cache, Cassandra caches, background compaction, and container scheduling can all change between attempts.

4. Compare physical rewrite with model improvement

bash · controlled compaction; lab only
docker exec atlasmart-cass-1 nodetool compact atlasmart_readpath slow_read_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath slow_read_probedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.slow_read_probe

Compaction changes the physical file set. A query-first redesign changes the unit of data a request addresses. For an ever-growing time partition, an hour/day bucket can bound partition growth if the application can route to the right bucket and accept bounded fan-out.

CQL · bounded hourly alternative
CREATE TABLE IF NOT EXISTS atlasmart_readpath.slow_read_probe_hourly (  bucket_day date, bucket_hour tinyint, event_time timestamp, event_id uuid, payload text,  PRIMARY KEY ((bucket_day,bucket_hour),event_time,event_id)) WITH CLUSTERING ORDER BY (event_time DESC,event_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_readpath.slow_read_probe_hourly VALUES('2026-09-07',18,'2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'p2-new');SELECT * FROM atlasmart_readpath.slow_read_probe_hourlyWHERE bucket_day='2026-09-07' AND bucket_hour=18;
Bad incident response: “Disk is slow—double the cache.”

Prove disk rather than replica wait, partition size, SSTable overlap, tombstone scanning, compaction pressure, or query shape is dominant. Cache changes redistribute memory pressure and can improve a warm micro-fixture without fixing production p99.

5. Incident record and acceptance checks

Record Minimum context
Version/topology Cassandra patch, Java/runtime, SSTable format, node/DC/rack, RF/CL, driver
Query partition/clustering predicates, page size, result rows/bytes
Partition distribution rows/bytes p50/p95/max, skew/hot-key evidence
SSTable/compaction strategy, SSTables/read distribution, pending compactions/repaired state
Deletes/TTL tombstones/read, gc_grace, TTL/delete rate, repair cadence
Bloom/cache false positives/ratio, key/chunk cache state, warmup
Host disk latency/queue, CPU, memory/GC, network
Outcome before/after distributions, rollback, schema/operational change

Check your understanding

  1. Why not tune cache from one slow query?
  2. Why is the minimum repeated-read latency misleading?
  3. Which histogram directly helps quantify SSTable read amplification?
  4. Why compare compaction and bucketing separately?
  5. When is the incident actually fixed?
Review the answers

1. The dominant cost may be replica wait, partition size, SSTable amplification, tombstones, compaction, disk, GC, or query shape.

2. Later runs can benefit from a tiny warm working set and multiple cache layers unlike production steady state.

3. The SSTables-per-read column in nodetool tablehistograms, interpreted with tablestats and tracing.

4. Compaction rewrites physical files for the same model; bucketing changes partition boundaries and the read unit.

5. Representative latency distributions meet the SLO while partition/SSTable/tombstone/resource/failure metrics remain acceptable and rollback is understood.

Production judgment

Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.

Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Chapter 11 opens the SSTable itself and examines its Data, index/trie, filter, statistics, compression, TOC, and streaming details.

Summary and next bridge

Slow-read diagnosis should move from distributed request evidence to partition shape, SSTable/tombstone amplification, caches/Bloom filters, and host I/O. Chapter 11 now examines SSTable components and on-disk storage internals directly.

Authoritative references

Re-check these version-sensitive sources when regenerating the lesson.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.