Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors
Read Repair / Reconciliation and the Relationship to Eventual Consistency
Create controlled replica divergence, force reconciliation/read repair, and distinguish request-scoped healing from scheduled anti-entropy repair.
Learning outcomes
AtlasMart restores a replica after a missed order update. A quorum read returns the newest order state, but the platform team cannot infer from that result whether every replica is now identical. This lesson separates the client-visible act of reconciliation from optional read-time write-back and from scheduled anti-entropy repair.
Distinguish reconciliation, BLOCKING read repair, read_repair=NONE, and scheduled repair.
Create a stale replica without hinted handoff immediately repairing it.
Use an ALL read plus tracing to force all three replicas into the comparison path.
Verify whether read repair changed the stale replica, then run repair as authoritative convergence cleanup.
Explain eventual consistency as a system property built from multiple mechanisms, not one read path feature.
The mandatory labs continue the established local AtlasMart
cluster with Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 Docker image, Java 17 inside the
image, cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and explicit per-request consistency levels (CLs). New
Chapter 15 tables use UnifiedCompactionStrategy (UCS), no
default Time To Live (TTL), and Cassandra's default
gc_grace_seconds unless a lesson intentionally
changes a setting. Authentication, client Transport Layer
Security (TLS), internode TLS, and remote Java Management
Extensions (JMX) are disabled only inside this isolated
single-host learning network. The optional application
examples use Apache Cassandra Java Driver 4.19.3.
Windows learners should run the Linux containers through
Docker Desktop/WSL rather than infer native Windows production
support. No Storage-Attached Index (SAI) or vector index is
required in this chapter, so failure-handling evidence is not
confounded by index rebuild/query behavior.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Failure-handling vocabulary
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.
1. Reconciliation answers the current read; read repair changes replica state
When multiple replicas return different versions, the
coordinator reconciles cells and tombstones using Cassandra's
timestamp/deletion semantics and constructs the newest visible
result. That is necessary even if table-level
read_repair is NONE. With
read_repair='BLOCKING', Cassandra may also write
the reconciled data to stale replicas involved in the read and
wait for the required repair acknowledgments before completing
the request. With NONE, reconciliation still occurs
but the read does not perform this repair write-back.
Neither mode traverses every partition or every replica. A quorum read might contact replicas that already agree and leave another replica stale. Therefore read repair is request-scoped. Operator-run repair compares replicas over token ranges and streams mismatches; that is why Cassandra documentation treats hints/read repair as best effort and repair as the mechanism that guarantees eventual replica synchronization.
| Concept | Returns correct current read? | Writes stale replica? | Covers unread data? |
|---|---|---|---|
| Reconciliation | Yes, for replicas/responses involved | Not by itself | No |
| read_repair=BLOCKING | Yes | Can, when mismatch is detected | No |
| read_repair=NONE | Yes | No read-repair write-back | No |
| nodetool repair | Not a foreground query feature | Streams detected differences | Yes, for selected ranges/table scope |
2. Build a controlled stale replica without hints
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id ( order_id text PRIMARY KEY, status text, version int, note text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'} AND read_repair = 'BLOCKING' AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
# Node 1 will coordinate the write below.docker exec atlasmart-cass-1 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_failure.order_state_by_idSET status='PACKED', version=2, note='read-repair-fixture', updated_at='2026-09-08T03:50:00Z'WHERE order_id='order-1501';
# No hint was created while handoff was disabled.docker exec atlasmart-cass-1 nodetool enablehandoffdocker unpause atlasmart-cass-3# Wait until all nodes are UN.docker exec atlasmart-cass-1 nodetool status# Direct evidence before read repair:docker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"
Expected state: node 3 can still show version 0/1 depending on prior lab state while nodes 1 and 2 have version 2. Recreate/drop the keyspace first if you need a perfectly clean baseline. The important evidence is version inequality, not a memorized output string.
3. Force the stale replica into the read and observe reconciliation
A normal LOCAL_QUORUM read needs only two replicas,
so it is not deterministic that the stale third replica
participates. For this small controlled lab, use
ALL while all three nodes are healthy. That
requires responses from all RF=3 replicas, making divergence
observable to the coordinator. Tracing is diagnostic evidence,
but exact event text/timing is implementation and runtime
dependent.
CONSISTENCY ALL;TRACING ON;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';TRACING OFF;
docker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"# Inspect the table contract itself.docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE TABLE atlasmart_failure.order_state_by_id"
If you alter a disposable copy of the table to
read_repair='NONE', an ALL read can still
reconcile and return the newest visible value, but Cassandra
will not perform the same read-repair write-back. This is a
consistency/atomicity tradeoff documented by Cassandra, not a
reason to disable repair operations.
docker exec atlasmart-cass-1 nodetool repair --full atlasmart_failure order_state_by_idfor n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"done
4. Wrong recovery shortcut: “just read every key”
Using application reads to “repair” a dataset is incomplete and expensive: unread partitions remain untouched, read repair can add unpredictable foreground latency/write load, and a read may not include every stale replica unless its CL/routing does so. A repair schedule is explicit, measurable, and token-range aware. Treat read repair as opportunistic healing along a real read path, not an anti-entropy scheduler.
Check your understanding
- What is reconciliation?
- Does read_repair=NONE mean stale data is returned when replicas disagree?
- Why use ALL instead of LOCAL_QUORUM in this teaching fixture?
- Why can a successful quorum read leave some data inconsistent?
- Which mechanism is intended for comprehensive anti-entropy convergence?
Review the answers
1. The coordinator selects the newest visible cell/tombstone versions from the replica responses used for the current read.
2. No. Reconciliation still determines the visible result; NONE changes read-time repair write-back behavior.
3. ALL forces responses from all RF=3 replicas, so the intentionally stale replica cannot be skipped by a two-response quorum.
4. The request touches only its partition/slice and may not involve every replica; unread data and uninvolved replicas are outside its scope.
5. Operator-run repair over the appropriate token ranges/keyspaces/tables.
Production judgment
Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.
Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Lesson 3 moves from repairing divergence to reducing tail latency, while showing how server rapid-read protection and driver speculative execution can create duplicate work if enabled without evidence.
Summary and next bridge
Reconciliation makes the current read correct; read repair can opportunistically update stale replicas involved in that read; scheduled repair covers the wider dataset. Next, study speculation as a latency optimization with explicit duplicate-work and idempotency costs.
Authoritative references
Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.
- Apache Cassandra downloads and current 5.0 release
- Apache Cassandra hinted handoff
- Apache Cassandra replica synchronization / consistency architecture
- Apache Cassandra repair
- CQL CREATE TABLE: read_repair and speculative_retry options
- Apache Cassandra native protocol error codes
- Apache Cassandra Java Driver changelog
- Java Driver retries
- Java Driver idempotence
- Java Driver speculative execution