Chapter 16 · Lightweight Transactions, Paxos, CAS, and Linearizable Conditional Updates
Contention, Hot Partitions, Retries, and Modeling to Minimize LWT Frequency
Create a hot CAS key, measure contention, reconcile ambiguous outcomes, and redesign the invariant to reduce LWT frequency.
Learning outcomes
AtlasMart launches a flash sale and routes every reservation through one global CAS row. Correctness holds, but p99 explodes because hundreds of contenders serialize on the same Paxos state. The fix is not “retry harder”; it is to reduce contention and shrink the set of operations that truly need LWT.
Explain how concurrent ballots on one Paxos key create contention and repeated coordination.
Generate a safe hot-key contention fixture and observe [applied] losers plus CAS latency/ContentionHistogram evidence.
Distinguish a rejected business condition from an infrastructure failure or ambiguous timeout.
Design retry behavior that reconciles a timed-out LWT before deciding whether to reissue it.
Remodel global/hot invariants into partitioned ownership keys, preallocation, or ordinary idempotent writes where possible.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
NetworkTopologyStrategy with replication factor
(RF) 3, and regular consistency level (CL)
LOCAL_QUORUM unless an experiment says otherwise.
New tables use UnifiedCompactionStrategy (UCS),
gc_grace_seconds = 864000 unless explicitly
isolated for an exercise, and no default TTL. Authentication,
client TLS, internode TLS, and remote JMX are disabled only
inside this isolated local learning network. The optional
application examples use Apache Cassandra Java Driver
4.19.3. Verify your actual runtime with
nodetool version, cqlsh --version,
and java -version.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms for this chapter
A lightweight transaction (LWT) is Cassandra's
conditional mutation mechanism. It uses the
Paxos consensus protocol so competing
operations on the same logical Paxos scope can agree on one
ordered outcome. Compare-and-set (CAS) means
“apply this mutation only if the current value satisfies the
condition.” A conditional mutation is CQL such
as INSERT ... IF NOT EXISTS or
UPDATE ... IF column = value.
Linearizable means successful operations appear
to occur in a single real-time-compatible order within the
guarantee's documented scope; it is stronger than ordinary
eventual consistency.
A coordinator is the Cassandra node handling one request. A replica is a node that stores the partition according to the keyspace replication strategy. A partition groups rows by partition key and maps through the partitioner to a token; the token determines natural replicas. RF is replication factor. A regular CL controls the data/learn phase of the operation. SERIAL and LOCAL_SERIAL are serial consistency levels that control the Paxos phase. A ballot is the proposal identity/order used by Paxos rounds. Contention occurs when concurrent conditional operations compete for the same Paxos state. A hot partition is a partition receiving disproportionately high request volume. A retry repeats an operation after an error or timeout; for LWT, an ambiguous outcome must be reconciled before blind retry. SSTables are immutable on-disk table files, while compaction rewrites SSTables. repair is Cassandra's anti-entropy process and is separate from Paxos agreement. Storage-Attached Indexing (SAI) and vector search can help locate rows but do not create uniqueness or cross-row transactional invariants.
1. Correct can still be operationally wrong
Paxos makes competing CAS operations agree, but agreement on one
hot key necessarily coordinates contenders around the same
state. As concurrency rises, ballots can collide, conditions can
become stale, and the coordinator may perform more work before
an operation wins or is rejected. The application should treat
[applied]=False as a normal concurrency/business
result when the condition simply lost; only
timeouts/failures/unavailability are infrastructure-style
errors.
| Signal | Meaning | Action |
|---|---|---|
[applied]=False |
condition was not true at compare point | read returned current values; choose business path, usually no retry loop |
| CAS contention histogram rises | multiple Paxos operations compete | find hot keys/invariants; reduce shared contention |
| CAS timeout | outcome may be ambiguous | reconcile/read current state before reissuing |
| Unavailable | serial/regular quorum cannot be reached | restore required replica scope or degrade invariant intentionally at application level—not blind retry |
2. Safe hot-key contention lab
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name, data_center, rack, release_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer, data_center, rack, release_version FROM system.peers_v2;"# Continue only when all three nodes are UN in dc1.# If the course cluster is absent, recreate it with the same Chapter 01 conventions# and pinned cassandra:5.0.9 image before running this chapter.
CREATE TABLE IF NOT EXISTS atlasmart_lwt.global_slot ( slot_id text PRIMARY KEY, owner text, revision int) WITH compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_lwt.global_slot (slot_id, owner, revision)VALUES ('global-flash-slot','OPEN',0);
docker exec atlasmart-cass-1 nodetool proxyhistograms
$jobs = 1..20 | ForEach-Object { $n = $_ Start-Job -ArgumentList $n -ScriptBlock { param($n) docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SERIAL CONSISTENCY LOCAL_SERIAL; UPDATE atlasmart_lwt.global_slot SET owner='order-$n', revision=1 WHERE slot_id='global-flash-slot' IF owner='OPEN' AND revision=0;" }}$jobs | Wait-Job | Receive-Job$jobs | Remove-Jobdocker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_lwt.global_slot WHERE slot_id='global-flash-slot';"docker exec atlasmart-cass-1 nodetool proxyhistograms
Expect one logical winner from an untouched row, but do not
expect a particular order number. Record how many requests were
rejected by condition versus failed by infrastructure. For
production metrics, collect
ClientRequest.CASWrite.ContentionHistogram,
ConditionNotMet, latency, timeouts, failures, and
unavailables from JMX/your metrics pipeline. The course
containers intentionally do not expose remote JMX.
3. Wrong approach: blind retry of a timed-out LWT
A timeout means the client did not receive a definitive result before its deadline. The Paxos operation might still have been chosen or learned. Blindly retrying can produce confusing user experience, repeated contention, or duplicate side effects outside Cassandra even if the CAS itself protects one row. Instead attach a business request/claim identifier, read/reconcile the invariant key after an ambiguous timeout, and only issue another conditional operation if current state proves the original intent did not win.
try { Row row = session.execute(claimStatement).one(); boolean applied = row != null && row.getBoolean("[applied]"); // handle applied or condition-not-met normally} catch (DriverTimeoutException ambiguous) { // Do NOT blindly resubmit a high-value CAS. Row current = session.execute( SimpleStatement.builder( "SELECT owner_id FROM atlasmart_lwt.unique_claim WHERE claim_type=? AND claim_value=?") .addPositionalValues("promo", "WELCOME-42") .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_QUORUM) .setSerialConsistencyLevel(DefaultConsistencyLevel.LOCAL_SERIAL) .build()).one(); // Compare current owner/request identity with the original intent, then decide.}
4. Model to reduce LWT frequency
Suppose AtlasMart needs one reservation per SKU per order. A single global row for “all inventory” creates unnecessary contention. Better choices include partitioning claims by the actual invariant key, preallocating independently owned inventory buckets, using normal idempotent writes for append-only events/projections, or accepting occasional duplicates and reconciling them when the business invariant permits it. Do not hash-shard a uniqueness constraint randomly unless the application can deterministically route every contender for the same invariant to the same shard.
| Requirement | LWT candidate? | Lower-contention alternative |
|---|---|---|
| unique promo code claim | Yes, key by the code/invariant | one CAS key per promo code, not one global promo partition |
| append order event | Usually no | normal idempotent insert keyed by order/event ID |
| analytics counter | No generic LWT default | counter/event aggregation depending semantics |
| global flash-sale capacity | Maybe, but high contention risk | preallocate capacity buckets/tokens with deterministic ownership |
| search/vector nearest-neighbor result | No uniqueness guarantee from index result | resolve invariant by direct primary-key CAS after search |
CREATE TABLE IF NOT EXISTS atlasmart_lwt.inventory_claim_by_sku ( sku text, reservation_id text, owner_order text, PRIMARY KEY ((sku, reservation_id))) WITH compaction = {'class':'UnifiedCompactionStrategy'};-- If reservation_id itself must be unique, CAS this exact key.INSERT INTO atlasmart_lwt.inventory_claim_by_sku (sku,reservation_id,owner_order)VALUES ('sku-42','res-9001','order-9001') IF NOT EXISTS;
5. Verification checklist
- Capture before/after CAS percentile evidence.
- Separate condition-not-met outcomes from request failures.
- Document the hot key and contender count.
- Show an outcome-reconciliation path for timeout before retry.
- Produce a redesigned invariant key that distributes independent decisions without weakening the invariant.
Check your understanding
- Why can a perfectly correct LWT design still have poor p99?
- Should [applied]=False be automatically retried?
- Why is a timed-out LWT ambiguous?
- How do you reduce LWT contention without weakening uniqueness?
- Can SAI or vector search itself enforce a uniqueness invariant?
Review the answers
1. Many concurrent operations may contend on the same Paxos state, causing extra coordination/retries and queueing.
2. No. It usually means the business condition lost; retrying the same stale condition can just create more contention.
3. The client deadline can expire even if the Paxos value was chosen or later learned; the client must reconcile current state.
4. Partition by the actual invariant key so only contenders for the same invariant serialize; do not route independent invariants through one global key.
5. No. Use the index to locate candidates, then enforce the invariant on a direct primary-key conditional mutation if needed.
Production judgment
LWT is an invariant tool, not a “strong consistency” checkbox
for every write. Before using it, define the exact invariant,
its partition/key scope, contender cardinality, expected
concurrency, RF and datacenter placement, regular and serial CL,
failure behavior, acceptable p95/p99 latency,
retry/reconciliation rules, and how a timed-out outcome will be
discovered. Measure CASRead/CASWrite
latency, timeouts, failures, unavailables, condition-not-met
counts, and contention histograms alongside ordinary request
latency, CPU, JVM garbage collection, network, disk, compaction,
repair state, and hot-key distribution.
Do not infer a universal LWT throughput number from this laptop
lab. SAI/vector indexes do not make a search result safe as a
uniqueness lock; security/tenant boundaries require
authorization and data-model controls, not
LOCAL_SERIAL. Managed Cassandra services can
constrain JMX, Paxos variants, or topology settings, so
translate the same invariant and evidence model to the service's
supported telemetry. Migration and rollback must account for
application semantics: replacing a normal write with LWT can
change latency and availability, while removing LWT can silently
weaken an invariant. Lesson 5 turns these mechanics into an
invariant decision framework and shows where Cassandra LWT ends
and relational multi-row transaction expectations begin.
Summary and next bridge
Contention is a modeling signal. Paxos can serialize a hot invariant, but it cannot make the hot key cheap. Safe systems reduce the number of decisions requiring LWT, reconcile ambiguous outcomes, and avoid blind retry loops. The final lesson defines when LWT is justified at all.
Authoritative references
Use these current official sources as the version-sensitive source of truth. Re-check them when regenerating this chapter because Paxos variants, driver behavior, metrics, and operational recommendations can evolve.
- Apache Cassandra downloads — 5.0.9 and Java Driver 4.19.3
- Cassandra guarantees — LWT and linearizable consistency
- cqlsh consistency and serial consistency
- cassandra.yaml — paxos_variant
- Cassandra monitoring metrics — CASRead/CASWrite
- nodetool proxyhistograms — CAS latency distributions
- Java Driver statement attributes — regular and serial consistency
- Java Driver retries and idempotence
- Java Driver query timestamps — LWT restriction