Chapter 16 · Lightweight Transactions, Paxos, CAS, and Linearizable Conditional Updates
Paxos Phases and Why LWT Requires More Coordination Than Regular Writes
Trace the Paxos agreement model, inspect Cassandra 5.0 Paxos variants, and compare regular request latency with CAS read/write distributions.
Learning outcomes
AtlasMart's developers understand that CAS works, but a latency review shows conditional writes are materially slower than ordinary mutations. The useful question is not “is LWT slow?” but “which extra coordination does Paxos perform, which Cassandra 5.0 Paxos variant is actually configured, and what do CAS latency/contended metrics show?”
Explain the prepare/promise, propose/accept, commit/learn mental model without pretending every implementation variant uses identical packet sequences.
Inspect the configured Cassandra 5.0 Paxos variant and distinguish the default v1 from recommended v2 guidance.
Compare ordinary write latency evidence with CAS write/read latency distributions using nodetool proxyhistograms.
Use tracing and CAS metrics to identify extra coordination, rejected conditions, and contention rather than timing one request.
Explain why changing paxos_variant is an operational migration requiring version/repair/restart prerequisites, not an ad-hoc query tweak.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
NetworkTopologyStrategy with replication factor
(RF) 3, and regular consistency level (CL)
LOCAL_QUORUM unless an experiment says otherwise.
New tables use UnifiedCompactionStrategy (UCS),
gc_grace_seconds = 864000 unless explicitly
isolated for an exercise, and no default TTL. Authentication,
client TLS, internode TLS, and remote JMX are disabled only
inside this isolated local learning network. The optional
application examples use Apache Cassandra Java Driver
4.19.3. Verify your actual runtime with
nodetool version, cqlsh --version,
and java -version. Cassandra 5.0 documents
multiple Paxos variants. The default remains v1,
while v2 is documented as recommended; do not
assume which one your cluster uses.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms for this chapter
A lightweight transaction (LWT) is Cassandra's
conditional mutation mechanism. It uses the
Paxos consensus protocol so competing
operations on the same logical Paxos scope can agree on one
ordered outcome. Compare-and-set (CAS) means
“apply this mutation only if the current value satisfies the
condition.” A conditional mutation is CQL such
as INSERT ... IF NOT EXISTS or
UPDATE ... IF column = value.
Linearizable means successful operations appear
to occur in a single real-time-compatible order within the
guarantee's documented scope; it is stronger than ordinary
eventual consistency.
A coordinator is the Cassandra node handling one request. A replica is a node that stores the partition according to the keyspace replication strategy. A partition groups rows by partition key and maps through the partitioner to a token; the token determines natural replicas. RF is replication factor. A regular CL controls the data/learn phase of the operation. SERIAL and LOCAL_SERIAL are serial consistency levels that control the Paxos phase. A ballot is the proposal identity/order used by Paxos rounds. Contention occurs when concurrent conditional operations compete for the same Paxos state. A hot partition is a partition receiving disproportionately high request volume. A retry repeats an operation after an error or timeout; for LWT, an ambiguous outcome must be reconciled before blind retry. SSTables are immutable on-disk table files, while compaction rewrites SSTables. repair is Cassandra's anti-entropy process and is separate from Paxos agreement. Storage-Attached Indexing (SAI) and vector search can help locate rows but do not create uniqueness or cross-row transactional invariants.
1. Paxos turns one write into an agreement protocol
A normal Cassandra write can be coordinated directly to the replicas required by the requested regular CL. An LWT first needs agreement about whether the condition can be accepted and which ballot/value wins. A useful conceptual sequence is prepare/promise → propose/accept → commit/learn. Cassandra optimizations can collapse or alter network round trips, so treat that as a reasoning model rather than a packet-for-packet API contract. The invariant comes from quorum intersection and Paxos state, not from a permanent leader.
| Conceptual phase | Question answered | Failure/contention implication |
|---|---|---|
| Prepare / promise | Can this ballot proceed, and is there earlier Paxos state to finish? | another ballot may force retry/reconciliation |
| Propose / accept | Will enough replicas accept this proposed value? | concurrent proposers can contend |
| Commit / learn | Publish the chosen value at the regular consistency requirement | regular CL affects the learned mutation/visibility phase |
| Subsequent linearizable read/conditional op | Is unfinished state present and what is current value? | may complete unfinished work before returning |
2. Cassandra 5.0: inspect the actual Paxos variant
Current Cassandra 5.0 configuration documents v1 as
the default: roughly 4 round trips for a write and 3 for a
linearizable read. It documents v2 as recommended:
roughly 2 round trips for a write and 1–2 for a read. Those are
protocol-level expectations, not end-to-end millisecond
guarantees. Upgrading from v1 to v2 requires cluster-version
prerequisites, a full primary-range repair on every node,
changing configuration, and rolling restarts; the chapter does
not mutate this setting automatically.
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name, data_center, rack, release_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer, data_center, rack, release_version FROM system.peers_v2;"# Continue only when all three nodes are UN in dc1.# If the course cluster is absent, recreate it with the same Chapter 01 conventions# and pinned cassandra:5.0.9 image before running this chapter.
docker exec atlasmart-cass-1 nodetool version# system_views.settings shows the running configuration, including defaults.for n in 1 2 3; do echo "=== atlasmart-cass-$n ===" docker exec atlasmart-cass-$n cqlsh -e "SELECT name, value FROM system_views.settings WHERE name='paxos_variant';"done# The YAML file can still be useful to see whether the value was explicitly set or left commented.docker exec atlasmart-cass-1 sh -lc "grep -n -A4 -B2 'paxos_variant' /etc/cassandra/cassandra.yaml | head -20 || true"
The documented migration includes version alignment, full primary-range repair, configuration change, and rolling restart. Treat it as an operational change with rollback and validation, not a per-table tuning option.
3. Measure distributions, not one stopwatch sample
CREATE KEYSPACE IF NOT EXISTS atlasmart_lwtWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_lwt.inventory_guard ( sku text PRIMARY KEY, available int, reservation_owner text, revision int) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_lwt.unique_claim ( claim_type text, claim_value text, owner_id text, created_at timestamp, PRIMARY KEY ((claim_type, claim_value))) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_lwt.order_projection ( order_id text PRIMARY KEY, customer_id text, state text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;SERIAL CONSISTENCY LOCAL_SERIAL;INSERT INTO atlasmart_lwt.inventory_guard(sku, available, reservation_owner, revision)VALUES ('sku-42', 10, 'NONE', 0);
docker exec atlasmart-cass-1 nodetool proxyhistograms
CONSISTENCY LOCAL_QUORUM;SERIAL CONSISTENCY LOCAL_SERIAL;CREATE TABLE IF NOT EXISTS atlasmart_lwt.latency_probe ( k text PRIMARY KEY, v int, guard int) WITH compaction = {'class':'UnifiedCompactionStrategy'};DELETE FROM atlasmart_lwt.latency_probe WHERE k='regular-1';DELETE FROM atlasmart_lwt.latency_probe WHERE k='cas-1';-- Regular mutations: normal write path.INSERT INTO atlasmart_lwt.latency_probe (k,v,guard) VALUES ('regular-1',1,0);UPDATE atlasmart_lwt.latency_probe SET v=2 WHERE k='regular-1';-- Conditional mutations: Paxos/CAS path.INSERT INTO atlasmart_lwt.latency_probe (k,v,guard) VALUES ('cas-1',1,0) IF NOT EXISTS;UPDATE atlasmart_lwt.latency_probe SET v=2, guard=1 WHERE k='cas-1' IF guard=0;UPDATE atlasmart_lwt.latency_probe SET v=3, guard=2 WHERE k='cas-1' IF guard=0;
docker exec atlasmart-cass-1 nodetool proxyhistograms# proxyhistograms exposes Read, Write, CAS Read, and CAS Write percentile columns.# Counters/histograms are cumulative for the node process; annotate warmup and sample volume.
The CAS Read Latency and
CAS Write Latency columns expose the
coordinator-level transactional distributions. Capture the
conditional result too: a chosen mutation should show
[applied]=True, while a stale compare should show
[applied]=False. With only a few lab requests,
percentiles are statistically weak; generate a larger synthetic
sample before comparing p95/p99, and disclose that the histogram
includes other requests since process start.
4. Trace one condition to understand shape, not benchmark it
TRACING ON;CONSISTENCY LOCAL_QUORUM;SERIAL CONSISTENCY LOCAL_SERIAL;UPDATE atlasmart_lwt.latency_probeSET v=4, guard=3WHERE k='cas-1'IF guard=1;TRACING OFF;
Tracing can reveal extra replica/coordinator activity and timing, but trace overhead makes it unsuitable as a latency benchmark. Use it to answer “what path did this request take?” and use production metrics/load tests to answer “what is the p99 cost under representative concurrency?”
# No remote JMX is exposed by the course containers.# Record the metric names you would collect in a production-safe JMX/metrics pipeline:# org.apache.cassandra.metrics.ClientRequest.Latency.CASRead# org.apache.cassandra.metrics.ClientRequest.Latency.CASWrite# ... ConditionNotMet, ContentionHistogram, Timeouts, Failures, Unavailables# Do not expose unauthenticated remote JMX just to complete this lesson.
5. Verification checklist
- Record the actual Paxos variant on every node.
-
Capture
proxyhistogramsbefore/after and preserve CAS Read/CAS Write p95/p99 values with sample caveats. - Capture one successful and one rejected condition.
- Explain why trace output is mechanism evidence, not a benchmark.
-
Do not change
paxos_variantin this mandatory lab.
Check your understanding
- Why does LWT need more coordination than a regular write?
- Is Cassandra 5.0 paxos_variant=v2 the default?
- Can you compare one regular write and one LWT stopwatch result as a production benchmark?
- What does CASWrite ContentionHistogram help reveal?
- Why is changing paxos_variant operationally significant?
Review the answers
1. It must first reach Paxos agreement about the conditional value/order before the learned mutation is completed at the regular consistency requirement.
2. No. Current documentation lists v1 as the default and v2 as recommended.
3. No. Use distributions under representative load and disclose topology, RF/CL, contention, warmup, and background state.
4. How much conditional-write contention the coordinator encountered, which is critical when hot keys cause repeated Paxos rounds/retries.
5. The documented migration has cluster-version, repair, configuration, and rolling-restart prerequisites; it is not a statement-level toggle.
Production judgment
LWT is an invariant tool, not a “strong consistency” checkbox
for every write. Before using it, define the exact invariant,
its partition/key scope, contender cardinality, expected
concurrency, RF and datacenter placement, regular and serial CL,
failure behavior, acceptable p95/p99 latency,
retry/reconciliation rules, and how a timed-out outcome will be
discovered. Measure CASRead/CASWrite
latency, timeouts, failures, unavailables, condition-not-met
counts, and contention histograms alongside ordinary request
latency, CPU, JVM garbage collection, network, disk, compaction,
repair state, and hot-key distribution.
Do not infer a universal LWT throughput number from this laptop
lab. SAI/vector indexes do not make a search result safe as a
uniqueness lock; security/tenant boundaries require
authorization and data-model controls, not
LOCAL_SERIAL. Managed Cassandra services can
constrain JMX, Paxos variants, or topology settings, so
translate the same invariant and evidence model to the service's
supported telemetry. Migration and rollback must account for
application semantics: replacing a normal write with LWT can
change latency and availability, while removing LWT can silently
weaken an invariant. Lesson 3 separates SERIAL from LOCAL_SERIAL
in a multi-datacenter model and shows how the serial phase and
normal phase have different replica scopes.
Summary and next bridge
Paxos adds an agreement protocol before Cassandra can safely
apply a conditional mutation. Cassandra 5.0's exact round-trip
behavior depends on the configured Paxos variant, and CAS
metrics make the cost observable. Next, distinguish
SERIAL and LOCAL_SERIAL from the
regular consistency of the learned write.
Authoritative references
Use these current official sources as the version-sensitive source of truth. Re-check them when regenerating this chapter because Paxos variants, driver behavior, metrics, and operational recommendations can evolve.
- Apache Cassandra downloads — 5.0.9 and Java Driver 4.19.3
- Cassandra guarantees — LWT and linearizable consistency
- cqlsh consistency and serial consistency
- cassandra.yaml — paxos_variant
- Cassandra monitoring metrics — CASRead/CASWrite
- nodetool proxyhistograms — CAS latency distributions
- Java Driver statement attributes — regular and serial consistency
- Java Driver retries and idempotence
- Java Driver query timestamps — LWT restriction