Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors
Failure Drill: Stop Replicas, Vary Consistency Levels, Recover Nodes, and Verify Convergence
Run an end-to-end replica failure drill with CL changes, pending hints, stale replicas, read repair, explicit repair, recovery, and final convergence checks.
Learning outcomes
AtlasMart's release checklist now needs a failure drill that proves what happens—not just what architecture diagrams promise—when replicas stop, writes continue, reads use different CLs, hints accumulate, nodes recover, replicas diverge, and repair runs. This lesson turns Chapters 14–15 into a repeatable incident exercise with explicit acceptance criteria.
Execute a timed RF=3 failure drill with reversible replica pauses and explicit CL changes.
Observe successful LOCAL_QUORUM operations alongside deterministic ALL/LOCAL_QUORUM unavailability cases.
Prove hint-based convergence, then separately create divergence that requires read repair/repair.
Classify every client-visible error before deciding whether a retry would be safe.
Finish with post-repair equality and operational reset checks rather than declaring recovery when nodes merely return UN.
The mandatory labs continue the established local AtlasMart
cluster with Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 Docker image, Java 17 inside the
image, cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and explicit per-request consistency levels (CLs). New
Chapter 15 tables use UnifiedCompactionStrategy (UCS), no
default Time To Live (TTL), and Cassandra's default
gc_grace_seconds unless a lesson intentionally
changes a setting. Authentication, client Transport Layer
Security (TLS), internode TLS, and remote Java Management
Extensions (JMX) are disabled only inside this isolated
single-host learning network. The optional application
examples use Apache Cassandra Java Driver 4.19.3.
Windows learners should run the Linux containers through
Docker Desktop/WSL rather than infer native Windows production
support. No Storage-Attached Index (SAI) or vector index is
required in this chapter, so failure-handling evidence is not
confounded by index rebuild/query behavior.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Failure-handling vocabulary
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.
1. Drill contract and safety envelope
Use only the disposable Docker cluster. Do not run these
commands against production, managed Cassandra, or an unrelated
local cluster. docker pause/unpause
isolates the experiment without changing host firewall rules or
clocks. Record timestamps for each step,
nodetool status, requested CL, client result/error,
pending hints, direct-replica values, and repair output. Exact
failure-detection time, hint replay time, and error wording are
runtime dependent.
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id ( order_id text PRIMARY KEY, status text, version int, note text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'} AND read_repair = 'BLOCKING' AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool statushandoff docker exec atlasmart-cass-$n nodetool versiondonedocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool listpendinghints
2. Phase A — one replica down: quorum availability plus pending hints
docker exec atlasmart-cass-1 nodetool pausehandoffdocker pause atlasmart-cass-3# Wait for failure detector state before interpreting errors.docker exec atlasmart-cass-1 nodetool status
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_failure.order_state_by_idSET status='PAID', version=10, note='phase-a', updated_at='2026-09-08T04:20:00Z'WHERE order_id='order-1501';SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';CONSISTENCY ALL;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
Interpretation: the LOCAL_QUORUM write/read need two local responses and can succeed with one replica down. The ALL read requires all three and should become unavailable once failure detection knows node 3 is down. A longer timeout would not satisfy ALL. Meanwhile node 1 should hold a hint because its delivery process is paused.
docker exec atlasmart-cass-1 nodetool listpendinghintsdocker unpause atlasmart-cass-3# Wait for node 3 to be UN, but node 1 is still not delivering hints.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"
docker exec atlasmart-cass-1 nodetool resumehandoff# Poll until the relevant hints drain.docker exec atlasmart-cass-1 nodetool listpendinghintsfor n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"done
3. Phase B — create divergence without hints, then expose repair behavior
Phase A demonstrates the fast best-effort path. Phase B intentionally prevents hint creation on node 1 for one write so that the stale replica remains visible after recovery. This proves why node status alone is not data convergence.
docker exec atlasmart-cass-1 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_failure.order_state_by_idSET status='PACKED', version=11, note='phase-b-no-hint', updated_at='2026-09-08T04:25:00Z'WHERE order_id='order-1501';
docker exec atlasmart-cass-1 nodetool enablehandoffdocker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"
Node 3 can now be UN and still contain the older
version. This is the central operational distinction between
membership/health and data convergence. Use an
ALL read to force the three replicas into reconciliation on this
table, which is configured with
read_repair='BLOCKING'.
CONSISTENCY ALL;TRACING ON;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';TRACING OFF;
for n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"donedocker exec atlasmart-cass-1 nodetool repair --full atlasmart_failure order_state_by_idfor n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"done
4. Phase C — two replicas down: classify before retrying
docker pause atlasmart-cass-2docker pause atlasmart-cass-3# Wait until node 1 marks both peers down.docker exec atlasmart-cass-1 nodetool status
CONSISTENCY LOCAL_ONE;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';CONSISTENCY LOCAL_QUORUM;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
The same topology can satisfy LOCAL_ONE but not LOCAL_QUORUM. This does not mean LOCAL_ONE is the “better” setting: it is a different business contract with weaker freshness/overlap guarantees. An application that promised quorum semantics must surface/fail the request rather than silently downgrade it. Restore the nodes and let failure detection converge before the final acceptance check.
docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool enablehandoff; donefor n in 1 2 3; do docker exec atlasmart-cass-$n nodetool resumehandoff || true; done# Wait until all nodes are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool repair --full atlasmart_failure order_state_by_idfor n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"donedocker exec atlasmart-cass-1 nodetool listpendinghints
5. Acceptance record
| Checkpoint | Evidence to retain | Pass condition |
|---|---|---|
| One replica down | status + LOCAL_QUORUM result + ALL error | quorum succeeds; ALL unavailable after detector convergence |
| Hint creation | listpendinghints / read-only hints directory | pending hint appears while delivery paused |
| Hint replay | direct LOCAL_ONE values | recovered replica converges after replay |
| No-hint divergence | direct node 3 read | node can be UN yet stale before reconciliation/repair |
| Read repair | ALL trace + post-read direct values | stale value is reconciled/updated when mismatch participates |
| Scheduled repair | nodetool repair + three direct reads | all replicas equal after repair |
| Two replicas down | LOCAL_ONE vs LOCAL_QUORUM | different CL availability matches RF arithmetic |
| Reset | statushandoff/status/pending hints | all nodes UN; handoff normal; no unexplained backlog |
If the client receives UNAVAILABLE, retrying the
same CL immediately while topology is unchanged is usually
wasted work. If a mutation times out, treat the outcome as
ambiguous and require idempotency/postcondition logic. If the
system reports overload, reduce offered load and investigate
capacity rather than launching simultaneous
retries/speculations.
Check your understanding
- Why does node status UN not prove replica equality?
- What did Phase A prove that Phase B did not?
- Why does LOCAL_ONE succeed with one live RF=3 replica while LOCAL_QUORUM fails?
- Why run repair after a successful read-repair demonstration?
- What is the retry rule for an ambiguous write timeout?
Review the answers
1. Membership/failure detection says the node is reachable/alive; it does not compare every partition version across replicas.
2. Phase A isolated best-effort hinted handoff; Phase B deliberately removed hints so reconciliation/read repair and explicit repair could be observed separately.
3. LOCAL_ONE requires one local response; LOCAL_QUORUM requires a majority, which is two for RF=3.
4. Read repair is request-scoped and does not prove all other partitions/ranges are consistent; repair is the anti-entropy mechanism.
5. Do not infer failure. Retry only if the operation remains safe when the first attempt may already have applied, preferably with an idempotency/postcondition design.
Production judgment
Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.
Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Chapter 16 now focuses on the special case where an application truly needs a conditional invariant: Cassandra lightweight transactions use Paxos-backed compare-and-set rather than generic retries or ordinary quorum writes.
Summary and next bridge
A production failure drill should end with data equality, handoff state, repair evidence, and client error classification—not merely green node status. You can now distinguish hints, read repair, repair, speculation, retries, timeout, unavailable, and overload as separate layers. Chapter 16 builds on that discipline with Paxos-backed lightweight transactions for true conditional invariants.
Authoritative references
Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.
- Apache Cassandra downloads and current 5.0 release
- Apache Cassandra hinted handoff
- Apache Cassandra replica synchronization / consistency architecture
- Apache Cassandra repair
- CQL CREATE TABLE: read_repair and speculative_retry options
- Apache Cassandra native protocol error codes
- Apache Cassandra Java Driver changelog
- Java Driver retries
- Java Driver idempotence
- Java Driver speculative execution