Chapter 10 · The Read Path: Bloom Filters, Indexes, SSTables, Caches, Merging, and Reconciliation

Read Repair / Reconciliation Concepts and Detecting Inconsistent Replica Data

Create controlled replica divergence, observe reconciliation/read repair, and distinguish it from scheduled anti-entropy repair.

Intermediate110–150 minutesReplica divergence + read-repair labApache Cassandra 5.0.9 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart recovered a node after it missed a write. A quorum read returns the newest value, but the team must know whether a stale replica participated, whether it was repaired, and whether scheduled repair is still required.

01

Separate reconciliation, read repair, hinted handoff, and operator-run anti-entropy repair.

02

Explain read_repair=BLOCKING versus NONE and the documented consistency tradeoff.

03

Create safe replica divergence with a paused disposable node and temporary hint suppression.

04

Use traced quorum/direct local reads to detect stale versus reconciled state.

05

Prove why read repair does not replace scheduled repair.

Chapter 10 lab baseline

The mandatory labs continue the established AtlasMart disposable cluster: cassandra:5.0.9, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 10 uses atlasmart_readpath with NetworkTopologyStrategy, replication factor (RF) 3, and normally LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS); table TTL defaults to zero and gc_grace_seconds is not changed. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside the isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional for client-routing discussion; all mandatory evidence uses free local cqlsh/nodetool. Capture nodetool version, cqlsh --version, and java -version on your machine rather than treating a prose baseline as runtime proof.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms before tracing a read

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores a copy of a partition according to replication placement. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, which participates in replica ownership. A consistency level (CL) specifies how many appropriately scoped replica responses a coordinator needs for the operation. An SSTable (Sorted String Table) is immutable on-disk storage. A tombstone is a distributed deletion/expiration marker. A Bloom filter is a probabilistic membership structure that can say “definitely absent” or “possibly present”; for its intended membership check it can produce false positives but not false negatives. Reconciliation merges cell versions and tombstones to construct the newest visible result. Read repair may write that reconciled result back to stale replicas involved in a request; it does not replace scheduled anti-entropy repair.

1. Reconciliation versus read repair

Reconciliation is the coordinator constructing the newest visible result from differing replica versions. Table-level read_repair='BLOCKING' (the current default) can additionally write reconciled state to stale replicas involved in the request and block until the requested CL is satisfied by those repair writes. The documented tradeoff is monotonic quorum reads. With NONE, the coordinator still reconciles the response but does not issue read-repair writes, preserving the documented partition-level write-atomicity behavior instead.

Neither setting sweeps unread partitions or every replica. Scheduled repair remains the broader anti-entropy mechanism over token ranges.

Mechanism Scope Purpose Replaces repair?
Reconciliation current read responses newest visible result No
BLOCKING read repair stale replicas involved in read foreground write-back + monotonic quorum behavior No
read_repair=NONE current read only reconcile without write-back No
nodetool repair selected token ranges/tables anti-entropy convergence This is scheduled/operator repair

2. Create controlled divergence

bash · verify or recreate the disposable three-node cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the shared course lab does not already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CQL · explicit repair probe
CREATE KEYSPACE IF NOT EXISTS atlasmart_readpathWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_readpath.read_repair_probe (  k text PRIMARY KEY, v text) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND read_repair = 'BLOCKING';CONSISTENCY ALL;INSERT INTO atlasmart_readpath.read_repair_probe (k,v) VALUES ('rr-42','baseline');

Temporarily suppress hinted handoff on the coordinator, pause node 3, write with LOCAL_QUORUM, then re-enable hints before unpausing node 3. This creates a short-lived stale replica without changing clocks or host firewalls.

bash · lab-only hint suppression and pause
docker exec atlasmart-cass-1 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
CQL · write while node 3 cannot receive it
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_readpath.read_repair_probe SET v='new-on-two-replicas' WHERE k='rr-42';
bash · restore normal mechanisms
docker exec atlasmart-cass-1 nodetool enablehandoffdocker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status

3. Observe request-scoped repair, then perform explicit anti-entropy cleanup

bash · compare a direct stale-node read and a traced quorum read
docker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_readpath.read_repair_probe WHERE k='rr-42';"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; TRACING ON; SELECT * FROM atlasmart_readpath.read_repair_probe WHERE k='rr-42'; TRACING OFF;"

A quorum read may be satisfied by two replicas that already agree and therefore never touch the stale third replica. If node 3 participates and mismatches, blocking read repair may update it. That variability is precisely why read repair is not a whole-cluster convergence plan.

bash · authoritative cleanup for this tiny lab table
for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool repair --full atlasmart_readpath read_repair_probedonefor n in 1 2 3; do  docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_readpath.read_repair_probe WHERE k='rr-42';"done
Wrong claim: “A quorum read fixed the cluster.”

The read touches only its selected data and replicas. Unread data may remain divergent. Hints are best effort; read repair is request-scoped; scheduled repair is still required according to the production repair plan.

4. Verification and checks

  • All nodes are UN.
  • Hinted handoff is re-enabled.
  • All direct LOCAL_ONE reads agree after explicit repair.
  • The learner can explain BLOCKING versus NONE without calling either universally best.

Check your understanding

  1. Does reconciliation imply write-back?
  2. What does BLOCKING read repair provide?
  3. Why can a stale replica survive a quorum read?
  4. Why suppress hints briefly?
  5. What provides broad anti-entropy convergence?
Review the answers

1. No. It constructs the newest result; read repair is the optional write-back mechanism.

2. When triggered, it blocks on repair writes needed for CL and provides documented monotonic quorum-read behavior.

3. The request can be satisfied by replicas that already agree, so the stale replica may not participate.

4. To keep the missed write visible long enough to study read-time behavior in a controlled lab.

5. Operator-run repair over the appropriate token ranges/tables/keyspaces.

Production judgment

Read latency is not a single disk or cache number. Evaluate partition rows/bytes and skew, clustering slices, RF and CL, replica locality/health, p95/p99 latency, SSTables touched per logical read, tombstones scanned, compaction state, Bloom false positives, key/chunk/page-cache state, disk throughput/queueing, JVM/GC, driver timeout/retry/speculation, repair state, and failure-domain health. Large partitions or overlapping SSTables can dominate an otherwise healthy cache. Conversely, aggressive cache allocation can steal memory from the OS page cache or other off-heap structures.

Do not generalize a warm laptop result. Record the Cassandra patch, SSTable format, topology, RF/CL, dataset and partition distribution, payload/result size, TTL/delete rate, compaction state, concurrent load, disk/network, warmup, and failure injection. Forced flush/compaction is a controlled lab/maintenance action, not routine performance tuning. Lesson 5 combines request, partition, SSTable, tombstone, Bloom/cache, and host evidence into one slow-read diagnostic workflow.

Summary and next bridge

Reconciliation decides the result; read repair may update stale replicas involved in that request; scheduled repair remains the comprehensive anti-entropy mechanism. Next, use these distinctions during a slow-read incident.

Authoritative references

Re-check these version-sensitive sources when regenerating the lesson.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.