Chapter 17 · Batches, Counters, Idempotency, and Write Coordination
Counter Tables, Distributed Counter Semantics, Restrictions, and Workload Suitability
Use Cassandra counters correctly, expose their retry/non-idempotency risks, and compare them with event-derived exact totals.
Learning outcomes
AtlasMart wants a fast “views today” number and a financial “captured amount” total. Both look like numbers, but only the first can tolerate approximate/eventual distributed counter behavior. This lesson shows why Cassandra counters are specialized state, why retries can double increments, and why exact ledgers need a different model.
Create and update a legal Cassandra counter table and explain its type restrictions.
Explain why counter increments/decrements are non-idempotent and how ambiguous retries can overcount.
Use COUNTER BATCH only with counter mutations and distinguish it from normal logged/unlogged batches.
Run concurrent counter increments and compare the observed final value with an event-derived exact total.
Decide whether a workload needs a counter, an immutable event table, or an LWT-protected invariant.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
NetworkTopologyStrategy with replication factor
(RF) 3, and consistency level (CL)
LOCAL_QUORUM unless an experiment explicitly
changes it. New normal tables use UnifiedCompactionStrategy
(UCS), gc_grace_seconds = 864000, and no default
Time To Live (TTL). Authentication, client Transport Layer
Security (TLS), internode TLS, and remote Java Management
Extensions (JMX) are disabled only inside this isolated local
learning network. Application examples use Apache Cassandra
Java Driver 4.19.3. Verify your actual runtime
with nodetool version,
cqlsh --version, and java -version;
do not infer host Java from the container runtime.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms for this chapter
Apache Cassandra is a peer-to-peer distributed database. CQL is the Cassandra Query Language. A coordinator is the node handling one client request; it routes mutations to the replicas that own the target partition. A partition groups rows sharing a partition key, and the partitioner's token mapping determines which replica set owns it. A batch is one native-protocol request carrying multiple CQL mutations. A logged batch uses Cassandra's distributed batch log to support the documented atomic batch guarantee across the mutations; an unlogged batch skips that batch log and therefore can be partially applied on failure. Isolation means readers do not observe intermediate changes within the documented same-partition scope. A counter is a special 64-bit distributed value changed only by increment/decrement operations. Idempotent means repeating an operation produces the same final database state as executing it once. A retry re-executes a request after certain failures; because write outcomes can be ambiguous, retry safety depends on idempotency and error classification. A request ID is an application-supplied stable identifier used to recognize duplicate attempts. Reconciliation is application or Cassandra logic that converges divergent/partial state after failures; it is not the same as pretending every write happened exactly once.
1. Counter semantics are deliberately different
A Cassandra counter is a special 64-bit signed
value that can only be incremented or decremented. You do not
INSERT a chosen counter value; you
UPDATE ... SET count = count + n. Counter columns
cannot be part of the primary key. Aside from primary-key
columns, a counter table contains counter columns rather than
ordinary mutable value columns. Counter updates reject user TTL
and timestamp options, and counter mutations cannot be mixed
with non-counter mutations in a normal batch; use
BEGIN COUNTER BATCH when batching counters.
The crucial retry property is that
count = count + 1 is not idempotent. If a client
times out after Cassandra applied the increment, repeating the
request can increment twice. This is why counters are a bad
substrate for exact money, inventory decrements that must never
double, or auditable ledgers.
| Workload | Counter fit | Reason |
|---|---|---|
| Page/view/activity approximation | possible | duplicate retry error may be acceptable if explicitly budgeted |
| Payment captured cents | poor | duplicate increment violates financial invariant |
| Inventory exact available quantity | poor | retry ambiguity can oversell/undersell |
| Operational monotonic-ish metric | possible | eventual distributed value may be acceptable |
| Auditable total | prefer events + derived projection | source events provide replay/reconciliation evidence |
2. Reproducible counter lab
# Verify the existing course cluster first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# If the shared course cluster does not exist, recreate the same local topology.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_writecoordWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.activity_counters ( tenant_id text, metric_day date, metric text, value counter, PRIMARY KEY ((tenant_id,metric_day),metric));CREATE TABLE IF NOT EXISTS atlasmart_writecoord.activity_events ( tenant_id text, metric_day date, event_id uuid, metric text, delta int, PRIMARY KEY ((tenant_id,metric_day),event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_writecoord.activity_countersSET value = value + 1WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view';BEGIN COUNTER BATCH UPDATE atlasmart_writecoord.activity_counters SET value = value + 3 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view'; UPDATE atlasmart_writecoord.activity_counters SET value = value + 1 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='checkout_started';APPLY BATCH;SELECT * FROM atlasmart_writecoord.activity_countersWHERE tenant_id='atlasmart' AND metric_day='2026-09-08';
-- These are learning failures; run one at a time and capture the server error.-- INSERT INTO atlasmart_writecoord.activity_counters (...) VALUES (...); -- counters are updated, not inserted-- UPDATE atlasmart_writecoord.activity_counters USING TTL 60-- SET value = value + 1 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view';-- UPDATE atlasmart_writecoord.activity_counters USING TIMESTAMP 1-- SET value = value + 1 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view';-- Do not mix normal INSERT/UPDATE statements into BEGIN COUNTER BATCH.
3. Duplicate-attempt simulation: the operation, not the transport, is the problem
You do not need a flaky network to prove non-idempotence. Execute the same increment twice with the same conceptual request identifier. The final counter rises twice because Cassandra counters do not deduplicate by application request ID.
UPDATE atlasmart_writecoord.activity_counters SET value = value + 10WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='retry_demo';-- Simulate an application retry after an ambiguous timeout:UPDATE atlasmart_writecoord.activity_counters SET value = value + 10WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='retry_demo';SELECT value FROM atlasmart_writecoord.activity_countersWHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='retry_demo';
The expected semantic result is 20, not 10. That is not a
Cassandra bug; the application submitted two increments. By
contrast, an event table can give each logical increment a
deterministic event_id. Replaying the same
INSERT with the same complete primary key is an
ordinary idempotent upsert.
INSERT INTO atlasmart_writecoord.activity_events(tenant_id,metric_day,event_id,metric,delta)VALUES ('atlasmart','2026-09-08',17000000-0000-0000-0000-000000000001,'retry_demo',10);-- Same logical request, same primary key: final row set is unchanged.INSERT INTO atlasmart_writecoord.activity_events(tenant_id,metric_day,event_id,metric,delta)VALUES ('atlasmart','2026-09-08',17000000-0000-0000-0000-000000000001,'retry_demo',10);SELECT * FROM atlasmart_writecoord.activity_eventsWHERE tenant_id='atlasmart' AND metric_day='2026-09-08';
Blind retry creates additional increments. For exact totals, write immutable/deterministic events and derive/reconcile a projection. If you keep counters for low-cost approximate metrics, document the duplicate-attempt error budget and never mark counter statements idempotent in the driver.
4. Concurrent updates and reconciliation
$jobs = 1..20 | ForEach-Object { Start-Job -ScriptBlock { docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_writecoord.activity_counters SET value = value + 1 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='concurrency_demo';" }}$jobs | Wait-Job | Receive-Job$jobs | Remove-Jobdocker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_writecoord.activity_counters WHERE tenant_id='atlasmart' AND metric_day='2026-09-08';"
A clean local run may show the expected number of increments, but this is not proof of exactly-once behavior under ambiguous failures. The non-idempotency problem appears when an attempt's outcome is unknown and the client duplicates it.
Check your understanding
- Why is a counter increment non-idempotent?
- Can a counter UPDATE use TTL or a client timestamp?
- Can normal and counter mutations be mixed in one batch?
- Why can a local concurrency test still end at the expected value?
- What is safer for exact auditable totals?
Review the answers
1. Executing +1 twice changes the state twice; repeating the request is not equivalent to executing it once.
2. No; Cassandra rejects TTL/TIMESTAMP options for counter updates.
3. No. Counter updates use counter-specific semantics and COUNTER BATCH.
4. No ambiguous retries may have occurred; that does not establish exactly-once behavior under failures.
5. Immutable/deterministic source events plus a derived/reconciled total, or another transaction system if the invariant requires it.
Production judgment
Choose batch/counter/retry behavior from the business invariant and partition model, not from a generic “fewer requests is faster” rule. Record RF/CL, partition cardinality and bytes, mutations per request, partitions per batch, coordinator locality, batchlog warnings/timeouts, counter contention, p50/p95/p99 latency, write timeout/failure type, retry/speculation counts, duplicate-attempt rate, request-ID reconciliation backlog, SSTable/compaction/tombstone pressure, disk/network/JVM headroom, and downstream side effects. A logged batch does not turn independent partitions into a relational transaction; a counter does not become an exact ledger; an idempotent database mutation does not automatically make an email/payment/webhook side effect idempotent.
For SAI/vector tables later in the course, remember that every extra mutation also updates index structures; large batches can concentrate write/index pressure. Security and tenancy boundaries still require authentication/authorization/network controls—same partition or LOCAL_QUORUM is not isolation. Managed Cassandra services can cap batch size, hide JMX/internal tables, or expose different metrics, so preserve the semantic tests even when the observability surface changes. Every rollout needs a rollback path: remove unsafe driver idempotence flags, stop duplicate retries, split cross-partition batches into independent async writes, or migrate counters to an event/reconciliation model when exactness requirements change. Lesson 4 generalizes the counter lesson into a driver rule: classify the statement itself as idempotent or non-idempotent before allowing automatic retries or speculation.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Idempotent vs Non-Idempotent Statements and Driver Retry Decisions.
Authoritative references
Use these version-sensitive sources as the contract. Re-check them when regenerating the course rather than freezing this lesson's dated snapshot.