Chapter 17 · Batches, Counters, Idempotency, and Write Coordination

Idempotent vs Non-Idempotent Statements and Driver Retry Decisions

Classify statement idempotency before Java-driver retries or speculation and reconcile ambiguous write outcomes safely.

Intermediate → Advanced110–150 minutesDriver idempotency + duplicate replay labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 · RF=3 dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart sees intermittent write timeouts. The platform team wants the Java driver to retry everything once. That policy is safe only if executing the exact statement twice cannot create a different final state or duplicate an external side effect. This lesson makes idempotency an explicit property of each request.

01

Define idempotency by database effect rather than HTTP verb, CQL keyword, or developer intent.

02

Classify deterministic upserts, deletes, collection operations, counters, LWT, and batches for retry safety.

03

Use Java Driver 4.19.3 statement idempotence metadata correctly, including immutable setter behavior.

04

Explain which driver retry decisions are bypassed for non-idempotent requests and why write timeouts are ambiguous.

05

Build a duplicate-write test that proves whether a chosen mutation is retry-safe.

Chapter 17 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, NetworkTopologyStrategy with replication factor (RF) 3, and consistency level (CL) LOCAL_QUORUM unless an experiment explicitly changes it. New normal tables use UnifiedCompactionStrategy (UCS), gc_grace_seconds = 864000, and no default Time To Live (TTL). Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside this isolated local learning network. Application examples use Apache Cassandra Java Driver 4.19.3. Verify your actual runtime with nodetool version, cqlsh --version, and java -version; do not infer host Java from the container runtime.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms for this chapter

Apache Cassandra is a peer-to-peer distributed database. CQL is the Cassandra Query Language. A coordinator is the node handling one client request; it routes mutations to the replicas that own the target partition. A partition groups rows sharing a partition key, and the partitioner's token mapping determines which replica set owns it. A batch is one native-protocol request carrying multiple CQL mutations. A logged batch uses Cassandra's distributed batch log to support the documented atomic batch guarantee across the mutations; an unlogged batch skips that batch log and therefore can be partially applied on failure. Isolation means readers do not observe intermediate changes within the documented same-partition scope. A counter is a special 64-bit distributed value changed only by increment/decrement operations. Idempotent means repeating an operation produces the same final database state as executing it once. A retry re-executes a request after certain failures; because write outcomes can be ambiguous, retry safety depends on idempotency and error classification. A request ID is an application-supplied stable identifier used to recognize duplicate attempts. Reconciliation is application or Cassandra logic that converges divergent/partial state after failures; it is not the same as pretending every write happened exactly once.

1. Idempotency is a property of the effect

“INSERT” is not automatically idempotent and “UPDATE” is not automatically unsafe. An ordinary Cassandra upsert to a deterministic primary key and deterministic values is often idempotent: repeating it leaves the same final row. DELETE of the same cell/row is generally idempotent. In contrast, counter increments, list prepend/append operations, application-generated random keys created anew on each attempt, and side effects such as sending a payment/webhook can duplicate work.

The Java driver cannot infer every application invariant. Its default idempotence setting is false. You may override it per statement, but marking a request idempotent is a correctness assertion. Retries and speculative executions are considered only for idempotent requests because a connection drop or write timeout can leave the driver unable to know whether Cassandra already applied the mutation.

Mutation shape Usually idempotent? Why / caveat
SET status = fixed value on deterministic PK yes repeat converges to same database state; timestamp conflicts still require discipline
DELETE same row/cell yes repeat leaves deleted state; tombstone timestamp interactions still matter
counter = counter + 1 no each execution increments again
list = [x] + list no replay prepends duplicate element
set = set + {x} often yes set membership deduplicates; validate whole statement and timestamps
INSERT with new random UUID generated per attempt no at workflow level each attempt writes a different row
payment charge + Cassandra upsert not automatically database idempotency does not deduplicate the external payment

2. Driver 4.19.3: make the assertion explicit

bash · verify or recreate the disposable three-node lab
# Verify the existing course cluster first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# If the shared course cluster does not exist, recreate the same local topology.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CQL · create the Chapter 17 normal-write fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_writecoordWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.orders_by_customer (    customer_id text,    order_month date,    order_id uuid,    status text,    total decimal,    request_id uuid,    updated_at timestamp,    PRIMARY KEY ((customer_id,order_month),order_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.order_events_by_order (    order_id uuid,    event_id uuid,    event_type text,    detail text,    event_time timestamp,    PRIMARY KEY (order_id,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;
Java · statement-specific idempotence and immutable setters
PreparedStatement ps = session.prepare(  "UPDATE atlasmart_writecoord.orders_by_customer " +  "SET status=?, request_id=?, updated_at=? " +  "WHERE customer_id=? AND order_month=? AND order_id=?");BoundStatement safe = ps.bind("SHIPPED", requestId, Instant.now(), customerId, month, orderId)    .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_QUORUM)    .setIdempotent(true); // assertion: replay leaves the intended row state unchanged// Driver statements are immutable. This does NOT modify 'safe' unless reassigned:safe.setTimeout(Duration.ofSeconds(2));// Correct:safe = safe.setTimeout(Duration.ofSeconds(2));session.execute(safe);

Do not mark a statement true just because the default retry policy is conservative. The flag also affects whether speculative executions are allowed if enabled. A wrong idempotence flag can therefore create duplicate work through more than one execution path.

Java · counter/list operations must stay non-idempotent
SimpleStatement counter = SimpleStatement.builder(    "UPDATE atlasmart_writecoord.activity_counters SET value=value+1 " +    "WHERE tenant_id=? AND metric_day=? AND metric=?")  .addPositionalValues("atlasmart", LocalDate.of(2026,9,8), "clicks")  .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_QUORUM)  .setIdempotence(false)  .build();// Also review list prepend/append, random-ID inserts, and external side effects.// Do not force retries by lying about idempotence.

3. Ambiguous failure: timeout is not “not applied”

A write timeout can arrive after some replicas accepted the mutation. A driver/request timeout can happen even after the coordinator committed enough work but the response did not reach the client. Therefore the safe algorithm is not “timeout → retry.” It is “classify operation → classify error → reconcile when outcome is ambiguous → retry only when repeating is safe.” The default driver retry policy is intentionally limited and invoked only for idempotent requests in these ambiguous categories.

Java · error classification sketch
try {  session.execute(statement);} catch (WriteTimeoutException e) {  // Ambiguous mutation outcome. Retry only if the statement is truly idempotent  // and the configured retry policy chooses a safe verdict.  reconcileByRequestIdOrPrimaryKey();} catch (UnavailableException e) {  // alive < required: a longer client timeout cannot create missing replicas.  surfaceAvailabilityIncident(e);} catch (OverloadedException e) {  // Backpressure/resource signal. Avoid an immediate retry storm.  shedOrQueueWithBudget();} catch (DriverTimeoutException e) {  // Client-side outcome is unknown; reconcile before repeating non-idempotent work.  reconcileByRequestIdOrPrimaryKey();}

4. Duplicate-write proof, not a code-review guess

CQL · safe deterministic replay
INSERT INTO atlasmart_writecoord.order_events_by_order(order_id,event_id,event_type,detail,event_time)VALUES (17171717-1717-1717-1717-171717171717,aaaaaaaa-1717-1717-1717-171717171717,'FULFILLMENT_REQUESTED','stable request-id event','2026-09-08T07:00:00Z');-- Repeat the exact same primary key and values:INSERT INTO atlasmart_writecoord.order_events_by_order(order_id,event_id,event_type,detail,event_time)VALUES (17171717-1717-1717-1717-171717171717,aaaaaaaa-1717-1717-1717-171717171717,'FULFILLMENT_REQUESTED','stable request-id event','2026-09-08T07:00:00Z');SELECT * FROM atlasmart_writecoord.order_events_by_orderWHERE order_id=17171717-1717-1717-1717-171717171717;

The row set remains one logical event because the event ID is stable. Now compare that with Lesson 3's counter increment or an insert that calls now()/uuid() independently on every retry; those attempts need not converge to one logical effect.

Wrong approach: “Set default-idempotence=true globally; the retry policy is smart.”

The driver cannot know whether a counter, list mutation, random-ID write, or application side effect is safe to repeat. Prefer the default false and mark reviewed statements/profiles true only when the complete operation has a stable replay contract.

Check your understanding

  1. What does idempotent mean here?
  2. Why does the driver default idempotence to false?
  3. Does a write timeout prove the write failed?
  4. Why can a deterministic UUID make an event insert retry-safe?
  5. Does database idempotency make a payment API call idempotent?
Review the answers

1. Executing the request multiple times leaves the database in the same intended final state as executing it once.

2. It cannot safely infer application semantics; a false negative costs retry opportunity, while a false positive can corrupt behavior through duplicate effects.

3. No. Some or enough replicas may already have applied it; the client outcome can be ambiguous.

4. Replays target the same primary key instead of creating a new logical event row on every attempt.

5. No. External side effects need their own idempotency key/protocol.

Production judgment

Choose batch/counter/retry behavior from the business invariant and partition model, not from a generic “fewer requests is faster” rule. Record RF/CL, partition cardinality and bytes, mutations per request, partitions per batch, coordinator locality, batchlog warnings/timeouts, counter contention, p50/p95/p99 latency, write timeout/failure type, retry/speculation counts, duplicate-attempt rate, request-ID reconciliation backlog, SSTable/compaction/tombstone pressure, disk/network/JVM headroom, and downstream side effects. A logged batch does not turn independent partitions into a relational transaction; a counter does not become an exact ledger; an idempotent database mutation does not automatically make an email/payment/webhook side effect idempotent.

For SAI/vector tables later in the course, remember that every extra mutation also updates index structures; large batches can concentrate write/index pressure. Security and tenancy boundaries still require authentication/authorization/network controls—same partition or LOCAL_QUORUM is not isolation. Managed Cassandra services can cap batch size, hide JMX/internal tables, or expose different metrics, so preserve the semantic tests even when the observability surface changes. Every rollout needs a rollback path: remove unsafe driver idempotence flags, stop duplicate retries, split cross-partition batches into independent async writes, or migrate counters to an event/reconciliation model when exactness requirements change. Lesson 5 combines stable request IDs, narrowly scoped conditional claims, deterministic projections, and reconciliation into a complete retry-safe write workflow.

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Design Retry-Safe Writes with Request IDs, Conditional Logic, and Eventual Reconciliation.

Authoritative references

Use these version-sensitive sources as the contract. Re-check them when regenerating the course rather than freezing this lesson's dated snapshot.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.