Chapter 17 · Batches, Counters, Idempotency, and Write Coordination
LOGGED vs UNLOGGED BATCH: Atomicity Scope, Batch Log, and Multi-Partition Costs
Separate logged atomic completion from unlogged partial-failure behavior and make cross-partition batchlog cost observable.
Learning outcomes
AtlasMart's checkout service wants to update an order summary,
write an order event, and mark a cart as converted in one
request. A developer proposes one large
BEGIN BATCH because “batch means transaction.” This
lesson separates the documented batch atomicity guarantee from
isolation, throughput, and cross-partition coordination cost.
Explain LOGGED, UNLOGGED, and same-partition batch behavior without mapping Cassandra batches to SQL transactions.
Explain what the batch log protects and why a multi-partition logged batch adds coordinator/replica work.
Use tracing and the system batch table as best-effort evidence without depending on internal rows remaining visible.
Create a controlled RF=1 failure sandbox that can expose partial application of an UNLOGGED batch.
Choose independent asynchronous writes, same-partition grouping, or logged batch according to correctness—not folklore.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
NetworkTopologyStrategy with replication factor
(RF) 3, and consistency level (CL)
LOCAL_QUORUM unless an experiment explicitly
changes it. New normal tables use UnifiedCompactionStrategy
(UCS), gc_grace_seconds = 864000, and no default
Time To Live (TTL). Authentication, client Transport Layer
Security (TLS), internode TLS, and remote Java Management
Extensions (JMX) are disabled only inside this isolated local
learning network. Application examples use Apache Cassandra
Java Driver 4.19.3. Verify your actual runtime
with nodetool version,
cqlsh --version, and java -version;
do not infer host Java from the container runtime.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms for this chapter
Apache Cassandra is a peer-to-peer distributed database. CQL is the Cassandra Query Language. A coordinator is the node handling one client request; it routes mutations to the replicas that own the target partition. A partition groups rows sharing a partition key, and the partitioner's token mapping determines which replica set owns it. A batch is one native-protocol request carrying multiple CQL mutations. A logged batch uses Cassandra's distributed batch log to support the documented atomic batch guarantee across the mutations; an unlogged batch skips that batch log and therefore can be partially applied on failure. Isolation means readers do not observe intermediate changes within the documented same-partition scope. A counter is a special 64-bit distributed value changed only by increment/decrement operations. Idempotent means repeating an operation produces the same final database state as executing it once. A retry re-executes a request after certain failures; because write outcomes can be ambiguous, retry safety depends on idempotency and error classification. A request ID is an application-supplied stable identifier used to recognize duplicate attempts. Reconciliation is application or Cassandra logic that converges divergent/partial state after failures; it is not the same as pretending every write happened exactly once.
1. Atomicity, isolation, and transport are three different questions
A CQL batch is one native-protocol request containing several mutations. Cassandra's default batch is logged. For a batch spanning multiple partitions, the coordinator first persists batchlog state and then sends the constituent mutations to their owning replicas. The batch log exists so Cassandra can replay outstanding mutations if the coordinator fails after the batch was accepted. This is the source of the documented “eventually all or none” atomic batch behavior. It does not create SQL-style snapshot isolation across unrelated partitions. The documentation only promises isolation for updates in the batch that belong to the same partition key.
An UNLOGGED batch skips the batch log. That removes the cross-partition batchlog overhead, but a failure can leave only part of the batch applied. If all mutations target one partition, a normal logged batch is optimized to an unlogged batch because Cassandra already provides atomic/isolation semantics within that partition. This optimization is one reason “logged versus unlogged” cannot be interpreted purely from the CQL keyword without also asking how many partitions are involved.
| Shape | Batch log? | Atomicity / isolation scope | Typical reason |
|---|---|---|---|
| One partition, normal BEGIN BATCH | optimized away | same-partition atomic/isolation semantics | group mutations that must be observed together in one partition |
| Multiple partitions, LOGGED | yes | atomic completion semantics; isolation remains partition-scoped | rare cross-partition atomicity requirement |
| Multiple partitions, UNLOGGED | no | partial application is possible | reduce network round trips only when partial success is acceptable |
| Many unrelated partitions for ingestion | wrong abstraction | coordination hotspot / warning risk | use concurrency-aware async writes or a loader instead |
2. Reproduce logged and unlogged requests with tracing
# Verify the existing course cluster first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# If the shared course cluster does not exist, recreate the same local topology.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_writecoordWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.orders_by_customer ( customer_id text, order_month date, order_id uuid, status text, total decimal, request_id uuid, updated_at timestamp, PRIMARY KEY ((customer_id,order_month),order_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_writecoord.order_events_by_order ( order_id uuid, event_id uuid, event_type text, detail text, event_time timestamp, PRIMARY KEY (order_id,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;TRACING ON;BEGIN BATCH INSERT INTO atlasmart_writecoord.orders_by_customer (customer_id,order_month,order_id,status,total,request_id,updated_at) VALUES ('cust-17','2026-09-01',11111111-1111-1111-1111-111111111117,'CREATED',44.20,aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaa17,toTimestamp(now())); INSERT INTO atlasmart_writecoord.order_events_by_order (order_id,event_id,event_type,detail,event_time) VALUES (11111111-1111-1111-1111-111111111117,bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbb17,'ORDER_CREATED','logged cross-partition batch',toTimestamp(now()));APPLY BATCH;TRACING OFF;TRACING ON;BEGIN UNLOGGED BATCH UPDATE atlasmart_writecoord.orders_by_customer SET status='PAID',updated_at=toTimestamp(now()) WHERE customer_id='cust-17' AND order_month='2026-09-01' AND order_id=11111111-1111-1111-1111-111111111117; INSERT INTO atlasmart_writecoord.order_events_by_order (order_id,event_id,event_type,detail,event_time) VALUES (11111111-1111-1111-1111-111111111117,cccccccc-cccc-cccc-cccc-cccccccccc17,'PAYMENT_CAPTURED','unlogged comparison',toTimestamp(now()));APPLY BATCH;TRACING OFF;
Capture the trace rather than copying the example wording into an incident report. The exact trace events, endpoints, and timing vary with replica placement, cache state, and version. The important evidence is that a multi-partition logged batch performs additional batch coordination while an unlogged batch skips it.
-- system.batches is an internal implementation table. Rows are intentionally short-lived.-- A successful small batch may leave nothing visible by the time you query it.SELECT id, version FROM system.batches;-- Treat “zero rows” as inconclusive, not as proof that batch logging never happened.-- Prefer tracing + error write type + server metrics/logs for durable operational evidence.
3. Controlled edge case: expose partial UNLOGGED application
The shared RF=3 keyspace is intentionally resilient, which makes
partial failure harder to force. For this one experiment create
an RF=1 sandbox. RF=1 is deliberately unsafe
for production; here it lets different partition keys have one
owning replica each. Find three keys owned by different nodes,
stop one owning replica, then submit an unlogged batch at
ONE. Mutations whose owners remain live can succeed
while the mutation for the down owner fails. This demonstrates
possibility, not a promise that every run will split at exactly
the same statement boundary.
CREATE KEYSPACE IF NOT EXISTS atlasmart_batch_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':1};CREATE TABLE IF NOT EXISTS atlasmart_batch_failure.by_key ( k text PRIMARY KEY, value text) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY ONE;
# Try candidate keys until you have examples whose single endpoint differs.for k in k1 k2 k3 k4 k5 k6 k7 k8 k9 k10; do echo "=== $k ===" docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_batch_failure by_key "$k"done# Record one key per node; do not assume k1/k2/k3 map to different replicas.
After you record one key whose endpoint is node 3 and at least one key owned by a live node, pause only node 3. Substitute your measured keys below.
docker pause atlasmart-cass-3# In cqlsh on a live node, substitute LIVE_KEY and DOWN_KEY:# CONSISTENCY ONE;# BEGIN UNLOGGED BATCH# INSERT INTO atlasmart_batch_failure.by_key (k,value) VALUES ('LIVE_KEY','may-apply');# INSERT INTO atlasmart_batch_failure.by_key (k,value) VALUES ('DOWN_KEY','cannot-reach-owner');# APPLY BATCH;# Then read LIVE_KEY and DOWN_KEY separately and capture the actual result/error.docker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
Skipping the batch log removes the atomic-completion mechanism. Use UNLOGGED only when the application can tolerate partial success and reconcile it. If cross-partition all-or-none completion is truly required, LOGGED is the Cassandra batch tool—but first ask whether the data model can put the invariant in one partition or avoid the distributed batch altogether.
4. Verification and reset
-
All three course nodes are
UNafter the failure drill. - The normal keyspace still has RF=3; the RF=1 sandbox is clearly labeled disposable.
- You captured traces for both LOGGED and UNLOGGED requests and did not infer timing from example output.
- You can explain why same-partition logged batches are optimized.
- You did not use the RF=1 sandbox for any durability claim.
DROP KEYSPACE IF EXISTS atlasmart_batch_failure;-- Keep atlasmart_writecoord for the remaining Chapter 17 lessons.DESCRIBE KEYSPACE atlasmart_writecoord;
Check your understanding
- Does a LOGGED batch provide SQL-style transaction isolation across independent partitions?
- Why can Cassandra optimize a logged single-partition batch?
- What can happen when an UNLOGGED cross-partition batch fails?
- Does an empty system.batches query prove a batch was unlogged?
- Why is the RF=1 experiment acceptable only as a lab?
Review the answers
1. No. The batchlog supports documented atomic batch completion; isolation is documented within a partition.
2. Same-partition mutations already have atomic/isolation semantics, so the extra distributed batchlog is unnecessary.
3. Some mutations can be applied while others are not; the application must tolerate or reconcile that state.
4. No. Batchlog rows are internal and short-lived; tracing/errors/metrics are better evidence.
5. It deliberately removes replica redundancy to make partial failure observable and is not a production resilience design.
Production judgment
Choose batch/counter/retry behavior from the business invariant and partition model, not from a generic “fewer requests is faster” rule. Record RF/CL, partition cardinality and bytes, mutations per request, partitions per batch, coordinator locality, batchlog warnings/timeouts, counter contention, p50/p95/p99 latency, write timeout/failure type, retry/speculation counts, duplicate-attempt rate, request-ID reconciliation backlog, SSTable/compaction/tombstone pressure, disk/network/JVM headroom, and downstream side effects. A logged batch does not turn independent partitions into a relational transaction; a counter does not become an exact ledger; an idempotent database mutation does not automatically make an email/payment/webhook side effect idempotent.
For SAI/vector tables later in the course, remember that every extra mutation also updates index structures; large batches can concentrate write/index pressure. Security and tenancy boundaries still require authentication/authorization/network controls—same partition or LOCAL_QUORUM is not isolation. Managed Cassandra services can cap batch size, hide JMX/internal tables, or expose different metrics, so preserve the semantic tests even when the observability surface changes. Every rollout needs a rollback path: remove unsafe driver idempotence flags, stop duplicate retries, split cross-partition batches into independent async writes, or migrate counters to an event/reconciliation model when exactness requirements change. Lesson 2 keeps atomicity separate from throughput and compares a same-partition batch with many independent writes so bulk ingestion is modeled correctly.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Why Batches Are Not a Bulk-Load Performance Tool and When Single-Partition Batches Help.
Authoritative references
Use these version-sensitive sources as the contract. Re-check them when regenerating the course rather than freezing this lesson's dated snapshot.