Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors

Speculative Retry: Tail-Latency Reduction vs Duplicate Replica Work

Separate Cassandra table rapid-read protection from Java-driver speculative execution, then gate duplicate work on measured tail latency and idempotency.

Intermediate110–150 minutesServer/driver speculation labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's p50 order lookup is healthy but rare coordinator/GC/disk pauses push p99.9 above the checkout SLO. The team proposes “speculative retry,” but Cassandra exposes one server-side table mechanism and the Java driver exposes another client-side mechanism. Turning both on blindly can turn one slow request into several concurrent requests.

01

Distinguish Cassandra table speculative_retry (rapid read protection) from Java-driver speculative execution.

02

Explain why redundant requests can improve tail latency while increasing replica/coordinator/network work.

03

Use a disposable table with speculative_retry=ALWAYS and tracing to make server-side duplicate reads observable.

04

Configure driver speculation only for statements explicitly classified as idempotent.

05

Measure speculation counts/latency distributions before choosing a threshold and capacity margin.

Chapter 15 lab baseline

The mandatory labs continue the established local AtlasMart cluster with Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 Docker image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy, replication factor (RF) 3, and explicit per-request consistency levels (CLs). New Chapter 15 tables use UnifiedCompactionStrategy (UCS), no default Time To Live (TTL), and Cassandra's default gc_grace_seconds unless a lesson intentionally changes a setting. Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network. The optional application examples use Apache Cassandra Java Driver 4.19.3. Windows learners should run the Linux containers through Docker Desktop/WSL rather than infer native Windows production support. No Storage-Attached Index (SAI) or vector index is required in this chapter, so failure-handling evidence is not confounded by index rebuild/query behavior.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Failure-handling vocabulary

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.

1. Two speculation layers, different owners

The CQL table option speculative_retry controls Cassandra's server-side rapid read protection. The coordinator normally sends read commands to enough replicas to satisfy CL; this option can cause it to send redundant read requests to other replicas after a configured trigger, or always/never. It is about replica reads inside one coordinator request.

The Java driver's speculative execution policy is client-side. If an execution against one coordinator is slow, the driver can start another execution against the next coordinator in its query plan. The first response wins from the application's perspective; the earlier in-flight request is not magically cancelled inside Cassandra—its later response is discarded. The driver disables speculation by default and only schedules speculative executions for statements it considers idempotent. These two layers can stack: one client request can hit two coordinators, and each coordinator can itself send extra replica reads.

Layer Trigger owner Duplicates Idempotency gate Primary evidence
Table speculative_retry Cassandra coordinator/table config replica read commands read-only mechanism; not driver idempotency flag CQL schema + tracing/server metrics
Driver speculative execution client driver policy whole request to another coordinator Yes: driver only speculates idempotent statements driver speculative-executions/request metrics + request tracker
Driver retry retry policy after failure/response another execution after an error Only safe paths / idempotent gate driver retry metrics + error classification

2. Make server-side rapid read protection visible

bash · verify or recreate the three-node dc1 course cluster
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CQL · Chapter 15 RF=3 failure fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id (    order_id text PRIMARY KEY,    status text,    version int,    note text,    updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND read_repair = 'BLOCKING'  AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
CQL · compare NONE with ALWAYS on the disposable table
DESCRIBE TABLE atlasmart_failure.order_state_by_id;ALTER TABLE atlasmart_failure.order_state_by_id WITH speculative_retry = 'ALWAYS';CONSISTENCY LOCAL_ONE;TRACING ON;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';TRACING OFF;-- Restore the intentionally conservative lab setting.ALTER TABLE atlasmart_failure.order_state_by_id WITH speculative_retry = 'NONE';

With ALWAYS, the coordinator is instructed to send extra read requests to other replicas for each read of this table. Trace event wording and exact ordering are version/timing dependent, so the learner should count participating endpoints rather than copy an example trace literally. This tiny local cluster cannot establish a production latency threshold; it only demonstrates the mechanism.

Wrong conclusion: “ALWAYS improved one request, so keep it.”

Redundant reads consume coordinator CPU, network, replica read capacity, disk/page-cache bandwidth, and protocol streams. Under overload they can worsen the queue that caused the tail latency. Use a representative benchmark and monitor both latency distribution and additional work.

3. Java Driver 4.19.3: speculation requires idempotency

In the driver, idempotency is a semantic promise from the application: executing the statement more than once leaves the intended database state the same as one execution. A partition-key lookup is naturally idempotent. An append to a list, incrementing a counter, a mutation that creates a new random/time UUID on each attempt, or a business side effect outside Cassandra is not automatically safe. The driver default idempotence is false; set it deliberately per statement or through carefully scoped execution profiles.

HOCON · illustrative driver profile, not a universal threshold
datastax-java-driver {  basic.load-balancing-policy.local-datacenter = dc1  advanced.speculative-execution-policy {    class = ConstantSpeculativeExecutionPolicy    max-executions = 2    delay = 40 milliseconds  }  advanced.metrics.session.enabled = [cql-requests]  advanced.metrics.node.enabled = [cql-messages]}# 40 ms is deliberately illustrative. Derive a threshold from healthy# workload percentiles and cluster capacity; do not copy this value.
Java · mark only a proven-idempotent lookup eligible
import com.datastax.oss.driver.api.core.CqlSession;import com.datastax.oss.driver.api.core.cql.SimpleStatement;SimpleStatement read = SimpleStatement.builder(        "SELECT status, version FROM atlasmart_failure.order_state_by_id WHERE order_id=?")    .addPositionalValue("order-1501")    .setIdempotence(true)    .build();try (CqlSession session = CqlSession.builder()        .withLocalDatacenter("dc1")        .build()) {    session.execute(read);}// Do not mark an operation idempotent merely to make retries/speculation happen.// Analyze the CQL mutation AND the surrounding business side effects.

Driver manual pages document the speculative-executions metric and warn that excessive speculation creates more traffic and can consume protocol stream IDs. In a production test, correlate this metric with request p99/p99.9, in-flight requests, pool available/orphaned streams, server read latency, overload responses, and host CPU/network/disk rather than watching latency alone.

4. Idempotency decision examples

Operation Usually idempotent? Why / caveat
SELECT by full partition key Yes repeating the read does not mutate state
SET status='PAID' for fixed key/value Often yes same intended value; timestamps/LWW and external side effects still need analysis
counter = counter + 1 No repeating increments twice
list_col = list_col + [x] No repeating appends another element
INSERT with a newly generated UUID/timeuuid per attempt No each execution can create a distinct row
LWT conditional update Do not infer from CQL alone conditional outcome and application retry contract must be designed explicitly

Check your understanding

  1. What is Cassandra table speculative_retry?
  2. What is Java-driver speculative execution?
  3. Can the two mechanisms operate at the same time?
  4. Why does the driver gate speculation on idempotency?
  5. What should determine the speculation delay?
Review the answers

1. A server-side rapid-read-protection option that can make a coordinator send redundant read commands to additional replicas.

2. A client-side policy that starts another execution against a different coordinator when the original is slow.

3. Yes, which can multiply duplicate work and must be considered in capacity/testing.

4. Both executions might ultimately reach/apply the request; non-idempotent mutations could change state more than once.

5. Representative healthy-platform latency distributions and available cluster/client capacity, validated under failures—not a copied blog value.

Production judgment

Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.

Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Lesson 4 classifies the errors that drive retry decisions: unavailable, server read/write timeout, overloaded/server failure, and client-side request timeout are different evidence and must not collapse into one “retry” branch.

Summary and next bridge

Speculation trades duplicate work for lower tail latency. Cassandra's table option and the driver's client-side policy are separate mechanisms, and idempotency is the safety gate for client speculation. Next, classify failure responses before deciding whether to retry at all.

Authoritative references

Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.