Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation
Read Repair / Reconciliation Concepts and Detecting Inconsistent Replica Data
Create controlled replica divergence, observe reconciliation/read repair, and distinguish it from scheduled anti-entropy repair.
Learning outcomes
AtlasMart recovered a node after it missed a write. A quorum read returns the newest value, but the team must know whether a stale replica participated, whether it was repaired, and whether scheduled repair is still required.
Separate reconciliation, read repair, hinted handoff, and operator-run anti-entropy repair.
Explain read_repair=BLOCKING versus NONE and the documented consistency tradeoff.
Create safe replica divergence with a paused disposable node and temporary hint suppression.
Use traced quorum/direct local reads to detect stale versus reconciled state.
Prove why read repair does not replace scheduled repair.
The mandatory labs continue the established AtlasMart
disposable cluster: cassandra:5.0.9, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. Chapter 10 uses atlasmart_readpath with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. New tables
explicitly use UnifiedCompactionStrategy (UCS); table TTL
defaults to zero and gc_grace_seconds is not
changed. Authentication, client TLS, internode TLS, and remote
JMX remain disabled only inside the isolated learning network.
The Apache Cassandra Java Driver 4.19.3 is optional for
client-routing discussion; all mandatory evidence uses free
local cqlsh/nodetool. Capture
nodetool version, cqlsh --version,
and java -version on your machine rather than
treating a prose baseline as runtime proof.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms before tracing a read
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.
1. Reconciliation versus read repair
Reconciliation is the coordinator constructing the newest
visible result from differing replica versions. Table-level
read_repair='BLOCKING' (the current default) can
additionally write reconciled state to stale replicas involved
in the request and block until the requested CL is satisfied by
those repair writes. The documented tradeoff is monotonic quorum
reads. With NONE, the coordinator still reconciles
the response but does not issue read-repair writes, preserving
the documented partition-level write-atomicity behavior instead.
Neither setting sweeps unread partitions or every replica. Scheduled repair remains the broader anti-entropy mechanism over token ranges.
| Mechanism | Scope | Purpose | Replaces repair? |
|---|---|---|---|
| Reconciliation | current read responses | newest visible result | No |
| BLOCKING read repair | stale replicas involved in read | foreground write-back + monotonic quorum behavior | No |
| read_repair=NONE | current read only | reconcile without write-back | No |
| nodetool repair | selected token ranges/tables | anti-entropy convergence | This is scheduled/operator repair |
2. Create controlled divergence
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.read_repair_probe ( k text PRIMARY KEY, v text) WITH compaction = {'class':'UnifiedCompactionStrategy'} AND read_repair = 'BLOCKING';CONSISTENCY ALL;INSERT INTO atlasmart_readpath.read_repair_probe (k,v) VALUES ('rr-42','baseline');
Temporarily suppress hinted handoff on the coordinator, pause
node 3, write with LOCAL_QUORUM, then re-enable
hints before unpausing node 3. This creates a short-lived stale
replica without changing clocks or host firewalls.
docker exec atlasmart-cass-1 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_readpath.read_repair_probe SET v='new-on-two-replicas' WHERE k='rr-42';
docker exec atlasmart-cass-1 nodetool enablehandoffdocker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
3. Observe request-scoped repair, then perform explicit anti-entropy cleanup
docker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_readpath.read_repair_probe WHERE k='rr-42';"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; TRACING ON; SELECT * FROM atlasmart_readpath.read_repair_probe WHERE k='rr-42'; TRACING OFF;"
A quorum read may be satisfied by two replicas that already agree and therefore never touch the stale third replica. If node 3 participates and mismatches, blocking read repair may update it. That variability is precisely why read repair is not a whole-cluster convergence plan.
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool repair --full atlasmart_readpath read_repair_probedonefor n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_readpath.read_repair_probe WHERE k='rr-42';"done
The read touches only its selected data and replicas. Unread data may remain divergent. Hints are best effort; read repair is request-scoped; scheduled repair is still required according to the production repair plan.
4. Verification and checks
- All nodes are
UN. - Hinted handoff is re-enabled.
- All direct LOCAL_ONE reads agree after explicit repair.
- The learner can explain BLOCKING versus NONE without calling either universally best.
Check your understanding
- Does reconciliation imply write-back?
- What does BLOCKING read repair provide?
- Why can a stale replica survive a quorum read?
- Why suppress hints briefly?
- What provides broad anti-entropy convergence?
Review the answers
1. No. It constructs the newest result; read repair is the optional write-back mechanism.
2. When triggered, it blocks on repair writes needed for CL and provides documented monotonic quorum-read behavior.
3. The request can be satisfied by replicas that already agree, so the stale replica may not participate.
4. To keep the missed write visible long enough to study read-time behavior in a controlled lab.
5. Operator-run repair over the appropriate token ranges/tables/keyspaces.
Production judgment
Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.
Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Lesson 5 combines request, partition, SSTable, tombstone, Bloom/cache, and host evidence into one slow-read diagnostic workflow.
Summary and next bridge
Reconciliation decides the result; read repair may update stale replicas involved in that request; scheduled repair remains the comprehensive anti-entropy mechanism. Next, use these distinctions during a slow-read incident.
Authoritative references
Re-check these version-sensitive sources when regenerating the lesson.