Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation

Partition Index / Summary, Bloom Filters, Key Cache, Chunk Cache, and SSTable Lookup

Walk inside a replica through Bloom membership checks, BIG/BTI index structures, key/chunk caches, and SSTable data lookup.

Intermediate110–150 minutesSSTable lookup + cache evidence labApache Cassandra 5.0.9 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart now knows which replica served the read, but that replica can hold many immutable SSTables. This lesson explains how Cassandra rules out files cheaply, locates partitions, uses caches, and still preserves correctness when a Bloom filter is falsely positive.

01

Connect Bloom filter, partition index/summary, key cache, chunk cache, compression chunks, and Data component in one lookup pipeline.

02

Distinguish Cassandra 5.0 classic BIG SSTables from optional BTI trie-indexed SSTables.

03

Explain why Bloom false positives cost work but cannot hide existing data.

04

Inspect DESCRIBE/table statistics, cache configuration, and on-disk component names without mutating files.

05

Avoid cache-first tuning until SSTables/read and partition shape are measured.

Chapter 10 lab baseline

The mandatory labs continue the established AtlasMart disposable cluster: cassandra:5.0.9, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 10 uses atlasmart_readpath with NetworkTopologyStrategy, replication factor (RF) 3, and normally LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS); table TTL defaults to zero and gc_grace_seconds is not changed. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside the isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional for client-routing discussion; all mandatory evidence uses free local cqlsh/nodetool. Capture nodetool version, cqlsh --version, and java -version on your machine rather than treating a prose baseline as runtime proof.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms before tracing a read

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.

1. Bloom is a gate, not the answer

For each candidate SSTable, Cassandra uses a Bloom filter to ask whether the partition key might be present. “Definitely absent” lets the engine skip deeper lookup; “possibly present” requires an index/data check. A false positive therefore adds work but does not create incorrect rows. A false negative would be unacceptable because it would skip data that exists.

The table's bloom_filter_fp_chance trades filter memory for false-positive probability. Lower is not automatically better. Existing SSTables retain the filter created when they were written; changing the option mainly affects newly written/re-written files.

Structure Role Important boundary
Bloom filter skip definitely-absent SSTables false positives cost lookup; correctness comes later
BIG partition index/summary locate Data.db offsets format-specific; summary is sampled, not row cache
Key cache cache index positions for relevant formats hit rate alone does not explain total latency
Chunk cache optional off-heap cache of uncompressed SSTable chunks documented disabled by default in current 5.0 config
OS page cache kernel file-page reuse separate from Cassandra caches
BTI trie indexes Cassandra 5.0 optional trie-indexed lookup component names and cache relevance differ

2. Cassandra 5.0 SSTable format matters

Current Cassandra 5.0 configuration documents big as the default SSTable format and offers bti as an optional trie-indexed format. BIG uses the classic partition-index/summary mental model. BTI replaces classic index components with trie-based structures. The stable mental model is “filter candidates → locate partition/rows → read relevant chunks → merge versions,” but the exact files/metrics must be interpreted for the selected format.

bash · inspect format and cache configuration
docker exec atlasmart-cass-1 sh -lc "grep -n -A8 -B2 '^sstable:' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "grep -n -E 'file_cache_enabled|file_cache_size|key_cache_size' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 nodetool info | grep -i -A4 -B1 cache || true

3. Create one SSTable and inspect evidence

bash · verify or recreate the disposable three-node cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CQL · create the bounded read-path fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.order_reads_by_customer (    customer_id text,    order_month date,    order_time timestamp,    order_id uuid,    status text,    total decimal,    note text,    PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND caching = {'keys':'ALL','rows_per_partition':'NONE'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_readpath.order_reads_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-07T16:00:00Z',00000000-0000-0000-0000-000000000042,'PAID',129.90,'baseline');SELECT * FROM atlasmart_readpath.order_reads_by_customerWHERE customer_id='cust-42' AND order_month='2026-09-01';
bash · flush only for this disposable observation
docker exec atlasmart-cass-1 nodetool flush atlasmart_readpath order_reads_by_customerdocker exec atlasmart-cass-1 nodetool tablestats atlasmart_readpath.order_reads_by_customerdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_readpath order_reads_by_customerdocker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_readpath -maxdepth 3 -type f | sort | head -80"
CQL · trace present and absent keys
TRACING ON;SELECT * FROM atlasmart_readpath.order_reads_by_customerWHERE customer_id='cust-42' AND order_month='2026-09-01';SELECT * FROM atlasmart_readpath.order_reads_by_customerWHERE customer_id='missing' AND order_month='2026-09-01';TRACING OFF;

Record Bloom false positives/ratio and memory, SSTable count, table/read latency, partition-size histograms, and SSTables-per-read. Exact labels vary by version and selected format.

Unsafe shortcut: “Set Bloom false-positive chance to zero and enable every cache.”

This spends memory before showing Bloom or cache misses are limiting. Chunk cache is off-heap and currently disabled by default; key cache relevance depends on format. Measure first.

4. Verification and checks

  • Record BIG versus BTI instead of assuming component names.
  • Capture Bloom statistics before tuning.
  • Explain “possibly present” versus “definitely absent.”
  • Keep Cassandra caches distinct from OS page cache.

Check your understanding

  1. Can a Bloom false positive return wrong data?
  2. Can Bloom false negatives be tolerated?
  3. Is Index.db universal in Cassandra 5.0?
  4. Is the SSTable chunk cache enabled by default?
  5. Why can high key-cache hit rate coexist with slow reads?
Review the answers

1. No. It only causes an extra candidate lookup that later storage checks resolve.

2. No; they would skip a file that actually contains the key.

3. No. BTI uses trie-indexed components rather than the classic BIG layout.

4. Current Cassandra 5.0 configuration documents file_cache_enabled=false by default.

5. Replica wait, many SSTables/read, tombstones, large partitions, disk/GC/network, or query shape can still dominate.

Production judgment

Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.

Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Lesson 3 deliberately creates multiple immutable versions and tombstones so SSTable read amplification becomes visible.

Summary and next bridge

Bloom filters and caches reduce cost; they do not decide correctness. SSTable format changes the exact index components. Next, create several physical generations for one logical row and watch Cassandra merge them.

Authoritative references

Re-check these version-sensitive sources when regenerating the lesson.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.