Chapter 14 · Consistency Levels, Quorums, Availability, and Client Guarantees
Write vs Read Consistency Choices: Latency, Availability, Staleness, and Failure Behavior
Pair read and write consistency deliberately, then observe a controlled stale-read window and repair convergence.
Learning outcomes
AtlasMart accepts an order update at a write CL chosen for availability, then immediately serves the order page from a weak read. During a replica outage, the page occasionally shows the old state. The write did not “vanish”; the read simply did not require an intersecting replica set.
Compare read and write CL choices by latency, availability, and stale-read risk.
Use quorum-intersection reasoning carefully rather than as an unconditional freshness theorem.
Create a controlled stale replica by disabling hints, pausing one replica, and writing at LOCAL_QUORUM.
Observe a possible stale LOCAL_ONE/ONE read and then repair the replica safely.
Distinguish unavailable errors, timeouts, successful weak reads, and application-level freshness promises.
The mandatory single-DC labs use Apache Cassandra
5.0.9 in the pinned Docker image
cassandra:5.0.9, Java 17 inside the image,
cluster atlasmart-course, network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
NetworkTopologyStrategy, replication factor (RF)
3, and explicit per-request consistency levels (CLs). New
tables use UnifiedCompactionStrategy (UCS), no default Time To
Live (TTL), and the Cassandra default
gc_grace_seconds unless a lesson states
otherwise. Authentication, client Transport Layer Security
(TLS), internode TLS, and remote Java Management Extensions
(JMX) are disabled only inside the isolated local learning
network. The optional application example uses Apache
Cassandra Java Driver 4.19.3. Windows learners
should use Docker Desktop/WSL-style Linux containers; commands
run inside containers unless labeled host-side.
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three nodes are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CREATE KEYSPACE IF NOT EXISTS atlasmart_consistencyWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_consistency.order_status_by_id ( order_id text PRIMARY KEY, status text, version int, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_consistency.inventory_reservation_by_product ( product_id text, reservation_id text, customer_id text, state text, created_at timestamp, PRIMARY KEY (product_id, reservation_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_consistency.order_status_by_id(order_id,status,version,updated_at)VALUES ('order-42','PAID',1,toTimestamp(now()));DESCRIBE KEYSPACE atlasmart_consistency;
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Core terms used throughout this chapter
Replication factor (RF) is the number of
replicas Cassandra is configured to keep for each partition in a
datacenter. A replica stores a copy of a
partition; the coordinator is the node handling
one client request and is not a permanent leader. A
consistency level (CL) is the per-request rule
that says how many and which replicas must acknowledge a write
or satisfy a read before the coordinator can return success. A
datacenter (DC) is Cassandra's logical
locality/failure-domain grouping; a rack is a
smaller placement grouping within a DC. A
quorum is a majority, calculated as
floor(RF/2)+1 for the relevant replica scope.
CQL means Cassandra Query Language;
cqlsh is its shell. A
Lightweight Transaction (LWT) is a conditional
CQL operation using Paxos consensus; it has a serial phase and a
regular learn/data phase. Availability here
means whether enough replicas are alive/responding for the
requested CL—not whether the cluster has any node alive.
Staleness means a read can return an older
replica version because the requested read CL did not
necessarily intersect the replicas that acknowledged a prior
write.
1. Read and write consistency are independent knobs
With RF=3, a successful LOCAL_QUORUM write required
two acknowledgments. A later LOCAL_QUORUM read also
requires two responses; any two-of-three read set must intersect
any two-of-three write-ack set by at least one replica.
Cassandra then reconciles returned versions. By contrast, a
LOCAL_ONE read can land on the third replica that
missed the write and return stale data until hints/read
repair/scheduled repair converge it.
| Write CL | Read CL | Availability tendency | Freshness/intersection note |
|---|---|---|---|
| LOCAL_QUORUM | LOCAL_QUORUM | tolerates one local replica failure at RF=3 | quorum sets intersect in the local DC |
| ONE | ONE | high availability/low wait | no guaranteed intersection; stale reads are possible |
| LOCAL_QUORUM | LOCAL_ONE | write is quorum-acknowledged; read is weak | read can select the non-acknowledging stale replica |
| ALL | ONE | write unavailable on any replica failure | successful ALL write gives all replicas the acknowledged version at that moment, but later failures/updates still matter |
| ANY | ONE | very high write availability | hint-only success can precede any replica storing queryable data |
The intersection model is not a substitute for understanding timeouts and failures. A timed-out write may have reached some replicas even though the client did not receive success; retrying a non-idempotent operation blindly can therefore duplicate effects at the application level.
2. Build a stale replica deliberately
Disable hinted handoff briefly on the two live nodes so the missed mutation is not immediately delivered, then pause node 3, write at LOCAL_QUORUM, re-enable hints before node 3 returns, and unpause it. This is safe only because the table/cluster are disposable.
docker exec atlasmart-cass-1 nodetool disablehandoffdocker exec atlasmart-cass-2 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_consistency.order_status_by_idSET status='SHIPPED',version=20,updated_at=toTimestamp(now())WHERE order_id='order-42';SELECT * FROM atlasmart_consistency.order_status_by_id WHERE order_id='order-42';
docker exec atlasmart-cass-1 nodetool enablehandoffdocker exec atlasmart-cass-2 nodetool enablehandoffdocker unpause atlasmart-cass-3# Wait until node 3 is UN.docker exec atlasmart-cass-1 nodetool status
3. Observe weak read behavior, then converge
Connect directly to node 3 and request
LOCAL_ONE with tracing. Because all three nodes are
natural replicas at RF=3, the contacted node is a candidate for
the direct read, but dynamic replica selection can vary; use
tracing to verify the actual replica that served the request. If
node 3 serves it before repair, the old value may appear. Repeat
rather than claiming a deterministic stale result if tracing
selects another replica.
docker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; TRACING ON; SELECT * FROM atlasmart_consistency.order_status_by_id WHERE order_id='order-42'; TRACING OFF;"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_consistency.order_status_by_id WHERE order_id='order-42';"
Now repair the table and verify from every node. Repair is the authoritative anti-entropy cleanup; read CL is not a replacement for repair.
docker exec atlasmart-cass-1 nodetool repair --full atlasmart_consistency order_status_by_idfor n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_consistency.order_status_by_id WHERE order_id='order-42';"done
Quorum acknowledgment does not force a subsequent ONE/LOCAL_ONE read to include an acknowledging replica. If the application promises read-after-write freshness, the read CL, routing/session design, LWT requirements, failure behavior, and retry policy must be selected and tested for that invariant.
4. Verification
- Hinted handoff is enabled again.
- All three nodes are UN.
- After full repair, direct weak reads show the same value on all replicas.
- The learner can explain why timeout ≠ definite failure and unavailable ≠ latency.
- The business freshness promise is stated separately from the raw CL names.
Check your understanding
- Why can LOCAL_ONE return stale after a successful LOCAL_QUORUM write?
- Why does LOCAL_QUORUM read after LOCAL_QUORUM write at RF=3 have an intersection?
- Does a write timeout prove the mutation was not applied?
- What repairs the deliberately stale replica comprehensively?
- What should an API specification say instead of “uses quorum”?
Review the answers
1. The one-replica read set can be the replica that missed the write; it does not necessarily intersect the two replicas that acknowledged it.
2. Any two-of-three replica set intersects any other two-of-three set by at least one replica, allowing reconciliation to include an acknowledged version.
3. No. Some or even enough replicas may have applied it while the coordinator/client timed out; retry safety depends on idempotency and operation semantics.
4. Anti-entropy repair over the relevant range/table, followed by verification.
5. State the concrete invariant: which reads must see which writes, topology/DC scope, allowed stale window, and behavior during replica/DC failures.
Production judgment
Choose consistency from a business invariant, topology, RF, and failure budget—not from a cluster-wide slogan. Record which reads must observe which writes, whether the invariant is local to one DC or global, whether conditional uniqueness/compare-and-set is required, and what latency/availability degradation is acceptable when replicas or a whole DC are unavailable. Then test that contract with the real driver, routing policy, request timeout, retry/speculative-execution policy, idempotency classification, and representative network latency.
Consistency level does not replace durable commit-log/storage
design, repair, backup/restore, security isolation,
schema/partition design, compaction/tombstone management, or
application-level idempotency. ALL is not
“permanent durability”; LOCAL_ONE is not tenant
isolation; a timeout is not proof that a write failed; and a
successful weak write does not promise a subsequent weak read
will be fresh. Storage-Attached Indexing (SAI) or vector search
can add read work but do not redefine RF/CL arithmetic. Managed
Cassandra services may restrict topology visibility or CL
choices; verify provider semantics rather than assuming Apache
Cassandra behavior is exposed unchanged. Lesson 4 introduces the
second consistency dimension: SERIAL/LOCAL_SERIAL for the Paxos
phase of LWT, separate from ordinary read/write consistency.
Summary and next bridge
Read and write CLs must be chosen as a pair against a freshness/availability contract. Quorum intersection is useful, but weak reads can still expose stale replicas and timeouts create ambiguous outcomes. Next, separate ordinary consistency from the serial consistency used by conditional LWT operations.
Authoritative references
Re-check these version-sensitive sources when regenerating this course.