Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors
Timeout vs Unavailable vs Overloaded / Failure Responses and Correct Client Interpretation
Classify unavailable, server read/write timeout, overload/failure, and client deadline so retry behavior preserves correctness instead of amplifying incidents.
Learning outcomes
AtlasMart's client currently catches every Cassandra exception
and sleeps 200 ms before retrying three times. During a replica
outage this turns UNAVAILABLE into more useless
traffic; during a write timeout it can duplicate a mutation
whose outcome is already unknown; during overload it can amplify
the server queue. This lesson replaces “retry on exception” with
evidence-based classification.
Distinguish UNAVAILABLE, READ_TIMEOUT, WRITE_TIMEOUT, overloaded, read/write failure, connection abort, and client request timeout.
Explain why a write timeout has an ambiguous commit outcome and why longer timeout does not solve unavailable.
Use controlled replica loss to make UNAVAILABLE deterministic and capture required/alive evidence.
Classify driver exceptions and gate automated retry on both error semantics and statement idempotency.
Design postcondition checks and application-level retry/backoff instead of infinite driver retries.
The mandatory labs continue the established local AtlasMart
cluster with Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 Docker image, Java 17 inside the
image, cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and explicit per-request consistency levels (CLs). New
Chapter 15 tables use UnifiedCompactionStrategy (UCS), no
default Time To Live (TTL), and Cassandra's default
gc_grace_seconds unless a lesson intentionally
changes a setting. Authentication, client Transport Layer
Security (TLS), internode TLS, and remote Java Management
Extensions (JMX) are disabled only inside this isolated
single-host learning network. The optional application
examples use Apache Cassandra Java Driver 4.19.3.
Windows learners should run the Linux containers through
Docker Desktop/WSL rather than infer native Windows production
support. No Storage-Attached Index (SAI) or vector index is
required in this chapter, so failure-handling evidence is not
confounded by index rebuild/query behavior.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Failure-handling vocabulary
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.
1. Error classes are observations about different failure stages
The native protocol carries structured errors. UNAVAILABLE includes the requested CL plus the number of replicas required and known alive; it is decided before Cassandra can satisfy the operation at that CL. A write timeout means the coordinator did not receive enough acknowledgments before the server-side write deadline; some replicas may already have applied the mutation. A read timeout means the read failed to gather the required responses/data before its deadline. Overloaded means the coordinator is refusing/deferring work because resources/backlog are exhausted. ReadFailure/WriteFailure report failures from replicas rather than simply slow responses. Separately, a driver can hit its own request deadline or lose a connection without receiving a definitive coordinator response.
| Signal | What Cassandra/driver knows | Unsafe interpretation | Better first action |
|---|---|---|---|
| UNAVAILABLE | alive replicas < required for CL | “wait longer” | restore availability or intentionally choose a different business contract |
| WRITE_TIMEOUT | insufficient acks before server deadline | “the write did not happen” | treat outcome as ambiguous; verify/idempotency-aware retry |
| READ_TIMEOUT | insufficient timely read responses/data | “data is missing” | inspect replica latency/load/tombstones/CL; retry only by policy |
| OVERLOADED | coordinator cannot accept/process current work | “retry immediately everywhere” | reduce pressure/back off/fail over only if safe |
| READ/WRITE_FAILURE | replica-side failures were reported | “same as timeout” | inspect server failure cause; default driver generally rethrows |
| Driver timeout/connection abort | client did not receive a definitive result | “server definitely rejected it” | classify statement idempotency and verify outcome |
2. Reproduce UNAVAILABLE instead of waiting for a timeout
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id ( order_id text PRIMARY KEY, status text, version int, note text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'} AND read_repair = 'BLOCKING' AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
docker pause atlasmart-cass-2docker pause atlasmart-cass-3# Wait until failure detection on node 1 marks them down; this is not instantaneous.docker exec atlasmart-cass-1 nodetool status
CONSISTENCY LOCAL_QUORUM;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';UPDATE atlasmart_failure.order_state_by_idSET status='SHIPPED', version=3, updated_at='2026-09-08T04:10:00Z'WHERE order_id='order-1501';
Once the coordinator knows only one replica is alive,
LOCAL_QUORUM needs two and returns unavailable
without waiting for a longer read/write timeout. Increasing a
client timeout cannot manufacture a second live replica. Restore
both nodes before continuing.
docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3# Wait until all nodes return UN.docker exec atlasmart-cass-1 nodetool status
3. A timeout is ambiguous; a retry must be safe even if the first attempt succeeded
For writes, the core safety question is not “did the client see success?” but “what happens if this mutation is applied twice?” A server write timeout can arrive after one or more replicas persisted the mutation. A client-side deadline can expire while the coordinator is still finishing work. Therefore a blind second execution of a counter increment, list append, generated-ID insert, or external side effect can corrupt the business outcome.
catch-all exception → retry erases the
distinction among unavailable, ambiguous timeout, overload,
validation errors, authentication errors, and non-idempotent
mutations. It can also synchronize thousands of clients into a
retry storm. Correctness requires operation classification,
bounded attempts, jitter/backoff where appropriate,
circuit/throttle behavior, and postcondition or
idempotency-key design.
import com.datastax.oss.driver.api.core.DriverTimeoutException;import com.datastax.oss.driver.api.core.servererrors.*;try { session.execute(statement);} catch (UnavailableException e) { // Required replicas are not alive for this CL. Longer timeout is not the fix. record("unavailable", e.getConsistencyLevel(), e.getRequired(), e.getAlive()); throw e;} catch (WriteTimeoutException e) { // Outcome may be partially/applied. Retry only if the business operation is proven idempotent. record("write_timeout", e.getConsistencyLevel(), e.getWriteType()); throw e;} catch (ReadTimeoutException e) { record("read_timeout", e.getConsistencyLevel()); throw e;} catch (OverloadedException e) { // Capacity/backpressure signal; do not create an immediate retry storm. record("overloaded"); throw e;} catch (ReadFailureException | WriteFailureException e) { record("replica_failure"); throw e;} catch (DriverTimeoutException e) { // Client deadline expired; server outcome may be unknown for mutations. record("driver_timeout"); throw e;}
The Apache Java Driver's default retry policy is intentionally conservative and retry-policy callbacks are only used for eligible/idempotent requests in ambiguous cases. The documentation warns that consistency-downgrading retry can break application invariants and datacenter locality; do not use it as an availability shortcut without a deliberate business contract.
4. Error response versus capacity response
Overload is often better handled by reducing offered load than by multiplying attempts: driver throttling, bounded queues, admission control, backoff/jitter, and capacity repair can protect the system. A retry on another node can sometimes succeed, but if the whole cluster is saturated, retries simply redistribute the same pressure. Likewise, increasing timeouts can reduce visible timeout counts while making queues and user latency worse.
Check your understanding
- Why does UNAVAILABLE not improve merely by increasing a timeout?
- Does WRITE_TIMEOUT prove the mutation was not applied?
- What does OVERLOADED communicate?
- Why must idempotency be checked before retrying an ambiguous write?
- Why is a consistency-downgrading retry policy dangerous?
Review the answers
1. The coordinator already knows fewer replicas are alive than the requested CL requires; time is not the missing resource.
2. No. Some replicas may already have persisted it, so the outcome can be ambiguous.
3. The coordinator cannot currently process the request because its resources/backlog are exhausted; it is a capacity/backpressure signal.
4. The first attempt may already have applied, so a repeated non-idempotent mutation can change state twice.
5. It can silently violate invariants and locality assumptions that were based on the originally requested consistency level.
Production judgment
Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.
Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Lesson 5 combines hints, CL changes, unavailable errors, stale replicas, read repair, driver-safe retry reasoning, recovery, and explicit anti-entropy repair into one timed failure drill.
Summary and next bridge
Unavailable, timeout, overload, replica failure, and client deadline are different evidence. Correct clients classify them, preserve the requested business invariant, and retry only when the operation remains safe under an ambiguous first outcome. The final lesson rehearses the full recovery timeline.
Authoritative references
Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.
- Apache Cassandra downloads and current 5.0 release
- Apache Cassandra hinted handoff
- Apache Cassandra replica synchronization / consistency architecture
- Apache Cassandra repair
- CQL CREATE TABLE: read_repair and speculative_retry options
- Apache Cassandra native protocol error codes
- Apache Cassandra Java Driver changelog
- Java Driver retries
- Java Driver idempotence
- Java Driver speculative execution