Chapter 17 · Batches, Counters, Idempotency, and Write Coordination

Why Batches Are Not a Bulk-Load Performance Tool and When Single-Partition Batches Help

Prove why unrelated cross-partition batches are not a bulk-loading shortcut and when same-partition grouping is actually useful.

Intermediate → Advanced105–145 minutesBulk-write shape benchmark labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 · RF=3 dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart needs to backfill millions of catalog projections. Someone proposes 5,000-row logged batches because “one request is faster than 5,000.” Cassandra can indeed save some client round trips, but a giant cross-partition batch concentrates coordination, memory, network, batchlog, and replica work on one coordinator. This lesson distinguishes a correctness batch from throughput concurrency.

01

Explain why batching unrelated partitions does not turn the coordinator into a bulk-loader.

02

Compare same-partition batches, cross-partition batches, and bounded asynchronous individual writes.

03

Observe batch-size warnings and latency distributions without raising guardrails to hide a bad workload.

04

Explain when a single-partition batch can help both correctness and transport efficiency.

05

Design a throughput test that reports p50/p95/p99, partitions/request, concurrency, payload, and errors.

Chapter 17 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, NetworkTopologyStrategy with replication factor (RF) 3, and consistency level (CL) LOCAL_QUORUM unless an experiment explicitly changes it. New normal tables use UnifiedCompactionStrategy (UCS), gc_grace_seconds = 864000, and no default Time To Live (TTL). Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside this isolated local learning network. Application examples use Apache Cassandra Java Driver 4.19.3. Verify your actual runtime with nodetool version, cqlsh --version, and java -version; do not infer host Java from the container runtime.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms for this chapter

Apache Cassandra is a peer-to-peer distributed database. CQL is the Cassandra Query Language. A coordinator is the node handling one client request; it routes mutations to the replicas that own the target partition. A partition groups rows sharing a partition key, and the partitioner's token mapping determines which replica set owns it. A batch is one native-protocol request carrying multiple CQL mutations. A logged batch uses Cassandra's distributed batch log to support the documented atomic batch guarantee across the mutations; an unlogged batch skips that batch log and therefore can be partially applied on failure. Isolation means readers do not observe intermediate changes within the documented same-partition scope. A counter is a special 64-bit distributed value changed only by increment/decrement operations. Idempotent means repeating an operation produces the same final database state as executing it once. A retry re-executes a request after certain failures; because write outcomes can be ambiguous, retry safety depends on idempotency and error classification. A request ID is an application-supplied stable identifier used to recognize duplicate attempts. Reconciliation is application or Cassandra logic that converges divergent/partial state after failures; it is not the same as pretending every write happened exactly once.

1. One request can still fan out into many distributed writes

Client-side request count and server-side work are different quantities. A 1,000-partition logged batch may save client/coordinator network round trips, but the coordinator must decode and retain the request, persist batchlog state, route hundreds or thousands of partition mutations, track acknowledgments at the requested CL, and replay incomplete work after failure. This creates a single-request burst and often worse tail latency than a bounded number of independent asynchronous writes distributed across coordinators by a token-aware driver.

A same-partition batch is different: the mutations share one replica set, the batchlog is optimized away, and the application may actually need same-partition atomic/isolation semantics. Even there, “more statements” is not free—payload size, tombstones, SAI/vector index maintenance, and replica CPU/disk work still matter.

Pattern Coordinator shape Correctness value Throughput guidance
Huge logged cross-partition batch one coordinator fans out widely cross-partition atomic completion usually wrong for bulk loading
UNLOGGED cross-partition batch one coordinator fans out, no batchlog no all-or-none guarantee only if fewer client round trips outweigh hotspot and partial-success risk
Small same-partition batch one replica set same-partition atomic/isolation reasonable when the domain operation is one partition
Prepared async individual writes many coordinators via token-aware routing per-write semantics normal bulk/backfill pattern with bounded concurrency

2. Build a same-partition comparison

bash · verify or recreate the disposable three-node lab
# Verify the existing course cluster first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# If the shared course cluster does not exist, recreate the same local topology.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CQL · create the Chapter 17 normal-write fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_writecoordWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.orders_by_customer (    customer_id text,    order_month date,    order_id uuid,    status text,    total decimal,    request_id uuid,    updated_at timestamp,    PRIMARY KEY ((customer_id,order_month),order_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.order_events_by_order (    order_id uuid,    event_id uuid,    event_type text,    detail text,    event_time timestamp,    PRIMARY KEY (order_id,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;
CQL · same-partition batch: three order rows in one customer-month partition
CONSISTENCY LOCAL_QUORUM;TRACING ON;BEGIN BATCH  INSERT INTO atlasmart_writecoord.orders_by_customer  (customer_id,order_month,order_id,status,total,request_id,updated_at)  VALUES ('cust-bulk','2026-09-01',00000000-0000-0000-0000-000000000171,'CREATED',10.00,10000000-0000-0000-0000-000000000171,toTimestamp(now()));  INSERT INTO atlasmart_writecoord.orders_by_customer  (customer_id,order_month,order_id,status,total,request_id,updated_at)  VALUES ('cust-bulk','2026-09-01',00000000-0000-0000-0000-000000000172,'CREATED',20.00,10000000-0000-0000-0000-000000000172,toTimestamp(now()));  INSERT INTO atlasmart_writecoord.orders_by_customer  (customer_id,order_month,order_id,status,total,request_id,updated_at)  VALUES ('cust-bulk','2026-09-01',00000000-0000-0000-0000-000000000173,'CREATED',30.00,10000000-0000-0000-0000-000000000173,toTimestamp(now()));APPLY BATCH;TRACING OFF;SELECT order_id,status,total FROM atlasmart_writecoord.orders_by_customerWHERE customer_id='cust-bulk' AND order_month='2026-09-01';

Record the trace and compare it with Lesson 1's cross-partition trace. Do not expect identical event text. The useful question is whether all mutations share one partition/replica set and whether the logged-batch optimization removes separate batchlog work.

3. Wrong benchmark: one huge logged batch

Cassandra has server-side batch-size warning/failure thresholds and other guardrails because oversized batches are operationally dangerous, but the thresholds are safeguards—not target sizes. Never tune them upward just to silence warnings. The correct experiment holds data and write count constant while comparing multiple request shapes.

Java · bounded asynchronous write sketch with Driver 4.19.3
// Pseudocode-level lab: use the same prepared INSERT for each row.PreparedStatement ps = session.prepare(    "INSERT INTO atlasmart_writecoord.orders_by_customer " +    "(customer_id,order_month,order_id,status,total,request_id,updated_at) " +    "VALUES (?,?,?,?,?,?,?)");Semaphore inFlight = new Semaphore(64); // example test input, NOT a universal tuning valueList<CompletionStage<AsyncResultSet>> futures = new ArrayList<>();for (Order row : rows) {  inFlight.acquire();  BoundStatement bs = ps.bind(/* row fields */)      .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_QUORUM)      .setIdempotent(true);  CompletionStage<AsyncResultSet> f = session.executeAsync(bs)      .whenComplete((ok,err) -> inFlight.release());  futures.add(f);}// Measure p50/p95/p99 and error/retry counts for each tested concurrency level.
PowerShell · record repeatable environment evidence
$nodes = 'atlasmart-cass-1','atlasmart-cass-2','atlasmart-cass-3'docker stats --no-stream $nodesdocker exec atlasmart-cass-1 nodetool tpstatsdocker exec atlasmart-cass-1 nodetool proxyhistogramsdocker exec atlasmart-cass-1 nodetool tablestats atlasmart_writecoord.orders_by_customer# Also record rows/request, partitions/request, payload bytes, CL, concurrency, warmup, and failures.
Wrong approach: “Raise batch_size_fail_threshold so my 5 MB batch works.”

A guardrail warning is evidence about request shape, not an obstacle to disable reflexively. Split unrelated partitions into bounded async requests, preserve per-row idempotency, and let token-aware routing distribute coordinator work. If a true invariant requires a logged cross-partition batch, size and load-test that specific invariant independently.

4. Verification and benchmark acceptance

Record for every test Why it matters
Cassandra/driver version, RF/CL, topology changes routing, acknowledgments, and protocol behavior
rows + partitions per request separates same-partition grouping from fan-out
payload bytes and column/index cost large values/SAI can dominate request count
client concurrency + retries determines queueing and duplicate work
p50/p95/p99 + errors tail latency matters more than one average
CPU/network/disk/JVM + compaction shows whether throughput simply moved pressure elsewhere

Check your understanding

  1. Why can 1,000 writes in one batch be slower than many async writes?
  2. When is a single-partition batch most defensible?
  3. Should batch warning thresholds be used as recommended batch sizes?
  4. What is the normal bulk-loading pattern?
  5. Why is average latency insufficient?
Review the answers

1. The single coordinator must fan out and track all mutations, possibly including batchlog work, causing concentrated memory/network/replica pressure and tail latency.

2. When multiple mutations in one partition belong to one domain operation and same-partition atomic/isolation behavior is useful.

3. No. They are safeguards; workload tests and partition semantics determine appropriate request shape.

4. Prepared/token-aware individual writes with bounded asynchronous concurrency, or a purpose-built loader, measured under representative load.

5. Large batches may produce tail spikes and queueing hidden by an average; p95/p99 and error/retry rates expose that behavior.

Production judgment

Choose batch/counter/retry behavior from the business invariant and partition model, not from a generic “fewer requests is faster” rule. Record RF/CL, partition cardinality and bytes, mutations per request, partitions per batch, coordinator locality, batchlog warnings/timeouts, counter contention, p50/p95/p99 latency, write timeout/failure type, retry/speculation counts, duplicate-attempt rate, request-ID reconciliation backlog, SSTable/compaction/tombstone pressure, disk/network/JVM headroom, and downstream side effects. A logged batch does not turn independent partitions into a relational transaction; a counter does not become an exact ledger; an idempotent database mutation does not automatically make an email/payment/webhook side effect idempotent.

For SAI/vector tables later in the course, remember that every extra mutation also updates index structures; large batches can concentrate write/index pressure. Security and tenancy boundaries still require authentication/authorization/network controls—same partition or LOCAL_QUORUM is not isolation. Managed Cassandra services can cap batch size, hide JMX/internal tables, or expose different metrics, so preserve the semantic tests even when the observability surface changes. Every rollout needs a rollback path: remove unsafe driver idempotence flags, stop duplicate retries, split cross-partition batches into independent async writes, or migrate counters to an event/reconciliation model when exactness requirements change. Lesson 3 introduces counters, whose increment/decrement semantics break the idempotent retry assumptions used by ordinary deterministic upserts.

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Counter Tables, Distributed Counter Semantics, Restrictions, and Workload Suitability.

Authoritative references

Use these version-sensitive sources as the contract. Re-check them when regenerating the course rather than freezing this lesson's dated snapshot.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.