Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors
Speculative Retry: Tail-Latency Reduction vs Duplicate Replica Work
Separate Cassandra table rapid-read protection from Java-driver speculative execution, then gate duplicate work on measured tail latency and idempotency.
Learning outcomes
AtlasMart's p50 order lookup is healthy but rare coordinator/GC/disk pauses push p99.9 above the checkout SLO. The team proposes “speculative retry,” but Cassandra exposes one server-side table mechanism and the Java driver exposes another client-side mechanism. Turning both on blindly can turn one slow request into several concurrent requests.
Distinguish Cassandra table speculative_retry (rapid read protection) from Java-driver speculative execution.
Explain why redundant requests can improve tail latency while increasing replica/coordinator/network work.
Use a disposable table with speculative_retry=ALWAYS and tracing to make server-side duplicate reads observable.
Configure driver speculation only for statements explicitly classified as idempotent.
Measure speculation counts/latency distributions before choosing a threshold and capacity margin.
The mandatory labs continue the established local AtlasMart
cluster with Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 Docker image, Java 17 inside the
image, cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and explicit per-request consistency levels (CLs). New
Chapter 15 tables use UnifiedCompactionStrategy (UCS), no
default Time To Live (TTL), and Cassandra's default
gc_grace_seconds unless a lesson intentionally
changes a setting. Authentication, client Transport Layer
Security (TLS), internode TLS, and remote Java Management
Extensions (JMX) are disabled only inside this isolated
single-host learning network. The optional application
examples use Apache Cassandra Java Driver 4.19.3.
Windows learners should run the Linux containers through
Docker Desktop/WSL rather than infer native Windows production
support. No Storage-Attached Index (SAI) or vector index is
required in this chapter, so failure-handling evidence is not
confounded by index rebuild/query behavior.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Failure-handling vocabulary
A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.
1. Two speculation layers, different owners
The CQL table option speculative_retry controls
Cassandra's server-side rapid read protection. The coordinator
normally sends read commands to enough replicas to satisfy CL;
this option can cause it to send redundant read requests to
other replicas after a configured trigger, or always/never. It
is about replica reads inside one coordinator request.
The Java driver's speculative execution policy is client-side. If an execution against one coordinator is slow, the driver can start another execution against the next coordinator in its query plan. The first response wins from the application's perspective; the earlier in-flight request is not magically cancelled inside Cassandra—its later response is discarded. The driver disables speculation by default and only schedules speculative executions for statements it considers idempotent. These two layers can stack: one client request can hit two coordinators, and each coordinator can itself send extra replica reads.
| Layer | Trigger owner | Duplicates | Idempotency gate | Primary evidence |
|---|---|---|---|---|
| Table speculative_retry | Cassandra coordinator/table config | replica read commands | read-only mechanism; not driver idempotency flag | CQL schema + tracing/server metrics |
| Driver speculative execution | client driver policy | whole request to another coordinator | Yes: driver only speculates idempotent statements | driver speculative-executions/request metrics + request tracker |
| Driver retry | retry policy after failure/response | another execution after an error | Only safe paths / idempotent gate | driver retry metrics + error classification |
2. Make server-side rapid read protection visible
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id ( order_id text PRIMARY KEY, status text, version int, note text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'} AND read_repair = 'BLOCKING' AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
DESCRIBE TABLE atlasmart_failure.order_state_by_id;ALTER TABLE atlasmart_failure.order_state_by_id WITH speculative_retry = 'ALWAYS';CONSISTENCY LOCAL_ONE;TRACING ON;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';TRACING OFF;-- Restore the intentionally conservative lab setting.ALTER TABLE atlasmart_failure.order_state_by_id WITH speculative_retry = 'NONE';
With ALWAYS, the coordinator is instructed to send
extra read requests to other replicas for each read of this
table. Trace event wording and exact ordering are version/timing
dependent, so the learner should count participating endpoints
rather than copy an example trace literally. This tiny local
cluster cannot establish a production latency threshold; it only
demonstrates the mechanism.
Redundant reads consume coordinator CPU, network, replica read capacity, disk/page-cache bandwidth, and protocol streams. Under overload they can worsen the queue that caused the tail latency. Use a representative benchmark and monitor both latency distribution and additional work.
3. Java Driver 4.19.3: speculation requires idempotency
In the driver, idempotency is a semantic promise from the application: executing the statement more than once leaves the intended database state the same as one execution. A partition-key lookup is naturally idempotent. An append to a list, incrementing a counter, a mutation that creates a new random/time UUID on each attempt, or a business side effect outside Cassandra is not automatically safe. The driver default idempotence is false; set it deliberately per statement or through carefully scoped execution profiles.
datastax-java-driver { basic.load-balancing-policy.local-datacenter = dc1 advanced.speculative-execution-policy { class = ConstantSpeculativeExecutionPolicy max-executions = 2 delay = 40 milliseconds } advanced.metrics.session.enabled = [cql-requests] advanced.metrics.node.enabled = [cql-messages]}# 40 ms is deliberately illustrative. Derive a threshold from healthy# workload percentiles and cluster capacity; do not copy this value.
import com.datastax.oss.driver.api.core.CqlSession;import com.datastax.oss.driver.api.core.cql.SimpleStatement;SimpleStatement read = SimpleStatement.builder( "SELECT status, version FROM atlasmart_failure.order_state_by_id WHERE order_id=?") .addPositionalValue("order-1501") .setIdempotence(true) .build();try (CqlSession session = CqlSession.builder() .withLocalDatacenter("dc1") .build()) { session.execute(read);}// Do not mark an operation idempotent merely to make retries/speculation happen.// Analyze the CQL mutation AND the surrounding business side effects.
Driver manual pages document the
speculative-executions metric and warn that
excessive speculation creates more traffic and can consume
protocol stream IDs. In a production test, correlate this metric
with request p99/p99.9, in-flight requests, pool
available/orphaned streams, server read latency, overload
responses, and host CPU/network/disk rather than watching
latency alone.
4. Idempotency decision examples
| Operation | Usually idempotent? | Why / caveat |
|---|---|---|
| SELECT by full partition key | Yes | repeating the read does not mutate state |
| SET status='PAID' for fixed key/value | Often yes | same intended value; timestamps/LWW and external side effects still need analysis |
| counter = counter + 1 | No | repeating increments twice |
| list_col = list_col + [x] | No | repeating appends another element |
| INSERT with a newly generated UUID/timeuuid per attempt | No | each execution can create a distinct row |
| LWT conditional update | Do not infer from CQL alone | conditional outcome and application retry contract must be designed explicitly |
Check your understanding
- What is Cassandra table speculative_retry?
- What is Java-driver speculative execution?
- Can the two mechanisms operate at the same time?
- Why does the driver gate speculation on idempotency?
- What should determine the speculation delay?
Review the answers
1. A server-side rapid-read-protection option that can make a coordinator send redundant read commands to additional replicas.
2. A client-side policy that starts another execution against a different coordinator when the original is slow.
3. Yes, which can multiply duplicate work and must be considered in capacity/testing.
4. Both executions might ultimately reach/apply the request; non-idempotent mutations could change state more than once.
5. Representative healthy-platform latency distributions and available cluster/client capacity, validated under failures—not a copied blog value.
Production judgment
Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.
Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Lesson 4 classifies the errors that drive retry decisions: unavailable, server read/write timeout, overloaded/server failure, and client-side request timeout are different evidence and must not collapse into one “retry” branch.
Summary and next bridge
Speculation trades duplicate work for lower tail latency. Cassandra's table option and the driver's client-side policy are separate mechanisms, and idempotency is the safety gate for client speculation. Next, classify failure responses before deciding whether to retry at all.
Authoritative references
Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.
- Apache Cassandra downloads and current 5.0 release
- Apache Cassandra hinted handoff
- Apache Cassandra replica synchronization / consistency architecture
- Apache Cassandra repair
- CQL CREATE TABLE: read_repair and speculative_retry options
- Apache Cassandra native protocol error codes
- Apache Cassandra Java Driver changelog
- Java Driver retries
- Java Driver idempotence
- Java Driver speculative execution