Chapter 17 · Batches, Counters, Idempotency, and Write Coordination

Counter Tables, Distributed Counter Semantics, Restrictions, and Workload Suitability

Use Cassandra counters correctly, expose their retry/non-idempotency risks, and compare them with event-derived exact totals.

Intermediate → Advanced105–145 minutesCounter retry/concurrency labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 · RF=3 dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart wants a fast “views today” number and a financial “captured amount” total. Both look like numbers, but only the first can tolerate approximate/eventual distributed counter behavior. This lesson shows why Cassandra counters are specialized state, why retries can double increments, and why exact ledgers need a different model.

01

Create and update a legal Cassandra counter table and explain its type restrictions.

02

Explain why counter increments/decrements are non-idempotent and how ambiguous retries can overcount.

03

Use COUNTER BATCH only with counter mutations and distinguish it from normal logged/unlogged batches.

04

Run concurrent counter increments and compare the observed final value with an event-derived exact total.

05

Decide whether a workload needs a counter, an immutable event table, or an LWT-protected invariant.

Chapter 17 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, NetworkTopologyStrategy with replication factor (RF) 3, and consistency level (CL) LOCAL_QUORUM unless an experiment explicitly changes it. New normal tables use UnifiedCompactionStrategy (UCS), gc_grace_seconds = 864000, and no default Time To Live (TTL). Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside this isolated local learning network. Application examples use Apache Cassandra Java Driver 4.19.3. Verify your actual runtime with nodetool version, cqlsh --version, and java -version; do not infer host Java from the container runtime.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms for this chapter

Apache Cassandra is a peer-to-peer distributed database. CQL is the Cassandra Query Language. A coordinator is the node handling one client request; it routes mutations to the replicas that own the target partition. A partition groups rows sharing a partition key, and the partitioner's token mapping determines which replica set owns it. A batch is one native-protocol request carrying multiple CQL mutations. A logged batch uses Cassandra's distributed batch log to support the documented atomic batch guarantee across the mutations; an unlogged batch skips that batch log and therefore can be partially applied on failure. Isolation means readers do not observe intermediate changes within the documented same-partition scope. A counter is a special 64-bit distributed value changed only by increment/decrement operations. Idempotent means repeating an operation produces the same final database state as executing it once. A retry re-executes a request after certain failures; because write outcomes can be ambiguous, retry safety depends on idempotency and error classification. A request ID is an application-supplied stable identifier used to recognize duplicate attempts. Reconciliation is application or Cassandra logic that converges divergent/partial state after failures; it is not the same as pretending every write happened exactly once.

1. Counter semantics are deliberately different

A Cassandra counter is a special 64-bit signed value that can only be incremented or decremented. You do not INSERT a chosen counter value; you UPDATE ... SET count = count + n. Counter columns cannot be part of the primary key. Aside from primary-key columns, a counter table contains counter columns rather than ordinary mutable value columns. Counter updates reject user TTL and timestamp options, and counter mutations cannot be mixed with non-counter mutations in a normal batch; use BEGIN COUNTER BATCH when batching counters.

The crucial retry property is that count = count + 1 is not idempotent. If a client times out after Cassandra applied the increment, repeating the request can increment twice. This is why counters are a bad substrate for exact money, inventory decrements that must never double, or auditable ledgers.

Workload Counter fit Reason
Page/view/activity approximation possible duplicate retry error may be acceptable if explicitly budgeted
Payment captured cents poor duplicate increment violates financial invariant
Inventory exact available quantity poor retry ambiguity can oversell/undersell
Operational monotonic-ish metric possible eventual distributed value may be acceptable
Auditable total prefer events + derived projection source events provide replay/reconciliation evidence

2. Reproducible counter lab

bash · verify or recreate the disposable three-node lab
# Verify the existing course cluster first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# If the shared course cluster does not exist, recreate the same local topology.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CQL · create a legal counter table and normal event table
CREATE KEYSPACE IF NOT EXISTS atlasmart_writecoordWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.activity_counters (  tenant_id text,  metric_day date,  metric text,  value counter,  PRIMARY KEY ((tenant_id,metric_day),metric));CREATE TABLE IF NOT EXISTS atlasmart_writecoord.activity_events (  tenant_id text,  metric_day date,  event_id uuid,  metric text,  delta int,  PRIMARY KEY ((tenant_id,metric_day),event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_writecoord.activity_countersSET value = value + 1WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view';BEGIN COUNTER BATCH  UPDATE atlasmart_writecoord.activity_counters SET value = value + 3  WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view';  UPDATE atlasmart_writecoord.activity_counters SET value = value + 1  WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='checkout_started';APPLY BATCH;SELECT * FROM atlasmart_writecoord.activity_countersWHERE tenant_id='atlasmart' AND metric_day='2026-09-08';
CQL · deliberately invalid counter operations
-- These are learning failures; run one at a time and capture the server error.-- INSERT INTO atlasmart_writecoord.activity_counters (...) VALUES (...);  -- counters are updated, not inserted-- UPDATE atlasmart_writecoord.activity_counters USING TTL 60--   SET value = value + 1 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view';-- UPDATE atlasmart_writecoord.activity_counters USING TIMESTAMP 1--   SET value = value + 1 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='product_view';-- Do not mix normal INSERT/UPDATE statements into BEGIN COUNTER BATCH.

3. Duplicate-attempt simulation: the operation, not the transport, is the problem

You do not need a flaky network to prove non-idempotence. Execute the same increment twice with the same conceptual request identifier. The final counter rises twice because Cassandra counters do not deduplicate by application request ID.

CQL · deterministic duplicate simulation
UPDATE atlasmart_writecoord.activity_counters SET value = value + 10WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='retry_demo';-- Simulate an application retry after an ambiguous timeout:UPDATE atlasmart_writecoord.activity_counters SET value = value + 10WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='retry_demo';SELECT value FROM atlasmart_writecoord.activity_countersWHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='retry_demo';

The expected semantic result is 20, not 10. That is not a Cassandra bug; the application submitted two increments. By contrast, an event table can give each logical increment a deterministic event_id. Replaying the same INSERT with the same complete primary key is an ordinary idempotent upsert.

CQL · retry-safe event alternative with deterministic event ID
INSERT INTO atlasmart_writecoord.activity_events(tenant_id,metric_day,event_id,metric,delta)VALUES ('atlasmart','2026-09-08',17000000-0000-0000-0000-000000000001,'retry_demo',10);-- Same logical request, same primary key: final row set is unchanged.INSERT INTO atlasmart_writecoord.activity_events(tenant_id,metric_day,event_id,metric,delta)VALUES ('atlasmart','2026-09-08',17000000-0000-0000-0000-000000000001,'retry_demo',10);SELECT * FROM atlasmart_writecoord.activity_eventsWHERE tenant_id='atlasmart' AND metric_day='2026-09-08';
Wrong approach: “Counters are eventually consistent, so retry until the total looks right.”

Blind retry creates additional increments. For exact totals, write immutable/deterministic events and derive/reconcile a projection. If you keep counters for low-cost approximate metrics, document the duplicate-attempt error budget and never mark counter statements idempotent in the driver.

4. Concurrent updates and reconciliation

PowerShell · issue concurrent counter increments in the disposable lab
$jobs = 1..20 | ForEach-Object {  Start-Job -ScriptBlock {    docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_writecoord.activity_counters SET value = value + 1 WHERE tenant_id='atlasmart' AND metric_day='2026-09-08' AND metric='concurrency_demo';"  }}$jobs | Wait-Job | Receive-Job$jobs | Remove-Jobdocker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_writecoord.activity_counters WHERE tenant_id='atlasmart' AND metric_day='2026-09-08';"

A clean local run may show the expected number of increments, but this is not proof of exactly-once behavior under ambiguous failures. The non-idempotency problem appears when an attempt's outcome is unknown and the client duplicates it.

Check your understanding

  1. Why is a counter increment non-idempotent?
  2. Can a counter UPDATE use TTL or a client timestamp?
  3. Can normal and counter mutations be mixed in one batch?
  4. Why can a local concurrency test still end at the expected value?
  5. What is safer for exact auditable totals?
Review the answers

1. Executing +1 twice changes the state twice; repeating the request is not equivalent to executing it once.

2. No; Cassandra rejects TTL/TIMESTAMP options for counter updates.

3. No. Counter updates use counter-specific semantics and COUNTER BATCH.

4. No ambiguous retries may have occurred; that does not establish exactly-once behavior under failures.

5. Immutable/deterministic source events plus a derived/reconciled total, or another transaction system if the invariant requires it.

Production judgment

Choose batch/counter/retry behavior from the business invariant and partition model, not from a generic “fewer requests is faster” rule. Record RF/CL, partition cardinality and bytes, mutations per request, partitions per batch, coordinator locality, batchlog warnings/timeouts, counter contention, p50/p95/p99 latency, write timeout/failure type, retry/speculation counts, duplicate-attempt rate, request-ID reconciliation backlog, SSTable/compaction/tombstone pressure, disk/network/JVM headroom, and downstream side effects. A logged batch does not turn independent partitions into a relational transaction; a counter does not become an exact ledger; an idempotent database mutation does not automatically make an email/payment/webhook side effect idempotent.

For SAI/vector tables later in the course, remember that every extra mutation also updates index structures; large batches can concentrate write/index pressure. Security and tenancy boundaries still require authentication/authorization/network controls—same partition or LOCAL_QUORUM is not isolation. Managed Cassandra services can cap batch size, hide JMX/internal tables, or expose different metrics, so preserve the semantic tests even when the observability surface changes. Every rollout needs a rollback path: remove unsafe driver idempotence flags, stop duplicate retries, split cross-partition batches into independent async writes, or migrate counters to an event/reconciliation model when exactness requirements change. Lesson 4 generalizes the counter lesson into a driver rule: classify the statement itself as idempotent or non-idempotent before allowing automatic retries or speculation.

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Idempotent vs Non-Idempotent Statements and Driver Retry Decisions.

Authoritative references

Use these version-sensitive sources as the contract. Re-check them when regenerating the course rather than freezing this lesson's dated snapshot.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.