Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation
Partition Index / Summary, Bloom Filters, Key Cache, Chunk Cache, and SSTable Lookup
Walk inside a replica through Bloom membership checks, BIG/BTI index structures, key/chunk caches, and SSTable data lookup.
Learning outcomes
AtlasMart now knows which replica served the read, but that replica can hold many immutable SSTables. This lesson explains how Cassandra rules out files cheaply, locates partitions, uses caches, and still preserves correctness when a Bloom filter is falsely positive.
Connect Bloom filter, partition index/summary, key cache, chunk cache, compression chunks, and Data component in one lookup pipeline.
Distinguish Cassandra 5.0 classic BIG SSTables from optional BTI trie-indexed SSTables.
Explain why Bloom false positives cost work but cannot hide existing data.
Inspect DESCRIBE/table statistics, cache configuration, and on-disk component names without mutating files.
Avoid cache-first tuning until SSTables/read and partition shape are measured.
The mandatory labs continue the established AtlasMart
disposable cluster: cassandra:5.0.9, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. Chapter 10 uses atlasmart_readpath with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. New tables
explicitly use UnifiedCompactionStrategy (UCS); table TTL
defaults to zero and gc_grace_seconds is not
changed. Authentication, client TLS, internode TLS, and remote
JMX remain disabled only inside the isolated learning network.
The Apache Cassandra Java Driver 4.19.3 is optional for
client-routing discussion; all mandatory evidence uses free
local cqlsh/nodetool. Capture
nodetool version, cqlsh --version,
and java -version on your machine rather than
treating a prose baseline as runtime proof.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms before tracing a read
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.
1. Bloom is a gate, not the answer
For each candidate SSTable, Cassandra uses a Bloom filter to ask whether the partition key might be present. “Definitely absent” lets the engine skip deeper lookup; “possibly present” requires an index/data check. A false positive therefore adds work but does not create incorrect rows. A false negative would be unacceptable because it would skip data that exists.
The table's bloom_filter_fp_chance trades filter
memory for false-positive probability. Lower is not
automatically better. Existing SSTables retain the filter
created when they were written; changing the option mainly
affects newly written/re-written files.
| Structure | Role | Important boundary |
|---|---|---|
| Bloom filter | skip definitely-absent SSTables | false positives cost lookup; correctness comes later |
| BIG partition index/summary | locate Data.db offsets | format-specific; summary is sampled, not row cache |
| Key cache | cache index positions for relevant formats | hit rate alone does not explain total latency |
| Chunk cache | optional off-heap cache of uncompressed SSTable chunks | documented disabled by default in current 5.0 config |
| OS page cache | kernel file-page reuse | separate from Cassandra caches |
| BTI trie indexes | Cassandra 5.0 optional trie-indexed lookup | component names and cache relevance differ |
2. Cassandra 5.0 SSTable format matters
Current Cassandra 5.0 configuration documents
big as the default SSTable format and offers
bti as an optional trie-indexed format. BIG uses
the classic partition-index/summary mental model. BTI replaces
classic index components with trie-based structures. The stable
mental model is “filter candidates → locate partition/rows →
read relevant chunks → merge versions,” but the exact
files/metrics must be interpreted for the selected format.
docker exec atlasmart-cass-1 sh -lc "grep -n -A8 -B2 '^sstable:' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "grep -n -E 'file_cache_enabled|file_cache_size|key_cache_size' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 nodetool info | grep -i -A4 -B1 cache || true
3. Create one SSTable and inspect evidence
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.order_reads_by_customer ( customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND caching = {'keys':'ALL','rows_per_partition':'NONE'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_readpath.order_reads_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-07T16:00:00Z',00000000-0000-0000-0000-000000000042,'PAID',129.90,'baseline');SELECT * FROM atlasmart_readpath.order_reads_by_customerWHERE customer_id='cust-42' AND order_month='2026-09-01';
docker exec atlasmart-cass-1 nodetool flush atlasmart_readpath order_reads_by_customerdocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.order_reads_by_customerdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath order_reads_by_customerdocker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_readpath -maxdepth 3 -type f | sort | head -80"
TRACING ON;SELECT * FROM atlasmart_readpath.order_reads_by_customerWHERE customer_id='cust-42' AND order_month='2026-09-01';SELECT * FROM atlasmart_readpath.order_reads_by_customerWHERE customer_id='missing' AND order_month='2026-09-01';TRACING OFF;
Record Bloom false positives/ratio and memory, SSTable count, table/read latency, partition-size histograms, and SSTables-per-read. Exact labels vary by version and selected format.
This spends memory before showing Bloom or cache misses are limiting. Chunk cache is off-heap and currently disabled by default; key cache relevance depends on format. Measure first.
4. Verification and checks
- Record BIG versus BTI instead of assuming component names.
- Capture Bloom statistics before tuning.
- Explain “possibly present” versus “definitely absent.”
- Keep Cassandra caches distinct from OS page cache.
Check your understanding
- Can a Bloom false positive return wrong data?
- Can Bloom false negatives be tolerated?
- Is Index.db universal in Cassandra 5.0?
- Is the SSTable chunk cache enabled by default?
- Why can high key-cache hit rate coexist with slow reads?
Review the answers
1. No. It only causes an extra candidate lookup that later storage checks resolve.
2. No; they would skip a file that actually contains the key.
3. No. BTI uses trie-indexed components rather than the classic BIG layout.
4. Current Cassandra 5.0 configuration documents file_cache_enabled=false by default.
5. Replica wait, many SSTables/read, tombstones, large partitions, disk/GC/network, or query shape can still dominate.
Production judgment
Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.
Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Lesson 3 deliberately creates multiple immutable versions and tombstones so SSTable read amplification becomes visible.
Summary and next bridge
Bloom filters and caches reduce cost; they do not decide correctness. SSTable format changes the exact index components. Next, create several physical generations for one logical row and watch Cassandra merge them.
Authoritative references
Re-check these version-sensitive sources when regenerating the lesson.