Chapter 14 · Consistency Levels, Quorums, Availability, and Client Guarantees

Define Per-Query Consistency from Business Invariants Instead of One Cluster-Wide Habit

Translate business invariants into per-query regular/serial consistency, failure tests, retry rules, and Java-driver settings.

Intermediate110–155 minutesPer-query invariant + driver labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 dc1 baselineLast reviewed: September 2026

Learning outcomes

AtlasMart has product browsing, order-state reads, payment/idempotency records, inventory reservations, analytics counters, and operational control data. Giving every request the same consistency level either wastes availability/latency or fails important invariants. This lesson builds a per-query contract and encodes it in the application driver.

01

Translate business invariants into regular and serial consistency requirements.

02

Build a consistency matrix that includes RF, DC scope, allowed staleness, and failure behavior.

03

Use Apache Cassandra Java Driver 4.19.3 to set consistency per statement/execution profile.

04

Classify retry/idempotency behavior alongside CL rather than as an afterthought.

05

Design invariant tests that fail replicas/DCs and verify client-visible guarantees.

Chapter 14 lab baseline

The mandatory single-DC labs use Apache Cassandra 5.0.9 in the pinned Docker image cassandra:5.0.9, Java 17 inside the image, cluster atlasmart-course, network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, NetworkTopologyStrategy, replication factor (RF) 3, and explicit per-request consistency levels (CLs). New tables use UnifiedCompactionStrategy (UCS), no default Time To Live (TTL), and the Cassandra default gc_grace_seconds unless a lesson states otherwise. Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside the isolated local learning network. The optional application example uses Apache Cassandra Java Driver 4.19.3. Windows learners should use Docker Desktop/WSL-style Linux containers; commands run inside containers unless labeled host-side. Driver examples assume Java Driver 4.19.3 and a local DC name of dc1. The mandatory CQL reasoning remains usable without a Java build environment.

bash · verify or recreate the three-node dc1 lab
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three nodes are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CQL · Chapter 14 single-DC fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_consistencyWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_consistency.order_status_by_id (    order_id text PRIMARY KEY,    status text,    version int,    updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_consistency.inventory_reservation_by_product (    product_id text,    reservation_id text,    customer_id text,    state text,    created_at timestamp,    PRIMARY KEY (product_id, reservation_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_consistency.order_status_by_id(order_id,status,version,updated_at)VALUES ('order-42','PAID',1,toTimestamp(now()));DESCRIBE KEYSPACE atlasmart_consistency;
Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Core terms used throughout this chapter

Replication factor (RF) is the number of replicas Cassandra is configured to keep for each partition in a datacenter. A replica stores a copy of a partition; the coordinator is the node handling one client request and is not a permanent leader. A consistency level (CL) is the per-request rule that says how many and which replicas must acknowledge a write or satisfy a read before the coordinator can return success. A datacenter (DC) is Cassandra's logical locality/failure-domain grouping; a rack is a smaller placement grouping within a DC. A quorum is a majority, calculated as floor(RF/2)+1 for the relevant replica scope. CQL means Cassandra Query Language; cqlsh is its shell. A Lightweight Transaction (LWT) is a conditional CQL operation using Paxos consensus; it has a serial phase and a regular learn/data phase. Availability here means whether enough replicas are alive/responding for the requested CL—not whether the cluster has any node alive. Staleness means a read can return an older replica version because the requested read CL did not necessarily intersect the replicas that acknowledged a prior write.

1. Start with the invariant, not the consistency label

AtlasMart query Business invariant Candidate regular CL Serial CL Failure stance
product catalog read seconds of staleness acceptable LOCAL_ONE none serve locally through one replica failure; cache allowed
order status after checkout read-after-write freshness in local region LOCAL_QUORUM write + LOCAL_QUORUM read none fail/queue if local quorum unavailable
inventory reservation only one reservation wins LOCAL_QUORUM LOCAL_SERIAL LWT; reject/resolve ambiguous timeout safely
payment idempotency key duplicate external charge must be prevented LOCAL_QUORUM plus LWT/application idempotency LOCAL_SERIAL never blind-retry non-idempotent external side effect
analytics/event read bounded staleness acceptable LOCAL_ONE none favor availability; repair/retention still required
global security/config invariant all DCs may need coordinated visibility derive from actual requirement; maybe EACH_QUORUM/QUORUM possibly SERIAL accept WAN latency only if invariant truly global

These are candidate designs, not universal prescriptions. A different topology, RF, region failover model, regulator requirement, or user-visible freshness promise changes the answer.

2. Encode per-query CL in the Java driver

Apache Cassandra Java Driver 4.19.3 exposes normal and serial consistency on each statement. Statements are immutable: the setter returns a new statement instance. In a production application, define named execution profiles for common request classes so consistency, timeout, retry, speculative execution, and idempotency are reviewed together.

java · per-statement consistency contracts
import com.datastax.oss.driver.api.core.CqlSession;import com.datastax.oss.driver.api.core.DefaultConsistencyLevel;import com.datastax.oss.driver.api.core.cql.SimpleStatement;// Catalog: availability-first, bounded staleness accepted.SimpleStatement catalogRead = SimpleStatement.builder(    "SELECT * FROM atlasmart_consistency.order_status_by_id WHERE order_id=?")    .addPositionalValue("order-42")    .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_ONE)    .setIdempotent(true)    .build();// Order workflow: local quorum intersection for the stated freshness contract.SimpleStatement orderRead = catalogRead    .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_QUORUM);// Conditional reservation: ordinary learn phase + Paxos serial phase.SimpleStatement reserve = SimpleStatement.builder(    "INSERT INTO atlasmart_consistency.inventory_reservation_by_product " +    "(product_id,reservation_id,customer_id,state,created_at) " +    "VALUES (?,?,?,?,toTimestamp(now())) IF NOT EXISTS")    .addPositionalValues("sku-9","slot-2","customer-C","HELD")    .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_QUORUM)    .setSerialConsistencyLevel(DefaultConsistencyLevel.LOCAL_SERIAL)    .setIdempotent(false) // review semantics before enabling retries/speculation    .build();

3. Test invariants under failure, not just happy-path correctness

For every query class, write an acceptance test with the failure conditions the API promises to tolerate. At minimum: one replica down, stale replica after a missed write, local quorum unavailable, remote DC unavailable (if multi-DC), timeout/ambiguous write outcome, and concurrent LWT contention. Record the actual exception class, received/applied status, trace/metrics, and post-repair state.

text · invariant-test matrix
CASE catalog_read:  RF: dc1=3  CL: LOCAL_ONE  promise: may be stale; should survive up to two replica losses if one local replica remains  test: pause two replicas -> read; restore -> repair -> verify equalityCASE order_status_read_after_write:  RF: dc1=3  write CL: LOCAL_QUORUM  read CL: LOCAL_QUORUM  promise: local quorum intersection under one replica loss  test: pause one replica -> write -> read -> restore -> repairCASE inventory_reservation:  RF: dc1=3  regular CL: LOCAL_QUORUM  serial CL: LOCAL_SERIAL  promise: only one concurrent IF NOT EXISTS wins  test: race two writers; assert exactly one applied=trueCASE remote_dc_outage (if dc2 exists):  compare LOCAL_QUORUM vs QUORUM/EACH_QUORUM  promise: only the explicitly local query class should remain available
Wrong architecture: one CL in a global driver config and no per-query review.

A cluster-wide default is useful as a safe baseline, but different query classes have different freshness, uniqueness, latency, locality, and failure requirements. Encode reviewed profiles, make exceptional stronger/weaker levels visible in code, and test them under failure. Also review retries/speculative execution: a CL does not make a non-idempotent operation safe to duplicate.

4. Decision worksheet

Question Evidence required before choosing CL
What must a successful write mean? number/scope of acknowledgments, durability settings, ambiguity on timeout
What must the next read observe? read-after-write/staleness contract and allowed replica/DC scope
Is compare-and-set required? if yes, LWT and SERIAL/LOCAL_SERIAL design
Which failures must remain available? replica, rack, DC, WAN partition scenarios and RF math
What is acceptable p95/p99 latency? real topology measurements with routing/retry/speculation disclosed
Can the operation be safely retried? idempotency and external-side-effect analysis
How will convergence be restored? hints/read repair/scheduled repair plan, not CL alone
How will the contract be monitored? driver error classes, CL, latency, unavailable/timeout rate, replica health

Check your understanding

  1. Why is a single cluster-wide CL often inadequate?
  2. What two CLs must an LWT design review?
  3. Does LOCAL_QUORUM make a retry safe?
  4. What should a failure test assert besides “request succeeded”?
  5. When is EACH_QUORUM justified?
Review the answers

1. Different queries have different staleness, uniqueness, locality, latency, and failure invariants.

2. The regular consistency level and the serial consistency level.

3. No. Retry safety depends on operation idempotency and whether the prior attempt may already have applied.

4. The client-visible value/invariant, exact failure class where relevant, topology state, and convergence after recovery/repair.

5. When the business invariant truly requires a quorum in every replicated DC and the application accepts the resulting WAN latency/availability coupling.

Production judgment

Choose consistency from a business invariant, topology, RF, and failure budget—not from a cluster-wide slogan. Record which reads must observe which writes, whether the invariant is local to one DC or global, whether conditional uniqueness/compare-and-set is required, and what latency/availability degradation is acceptable when replicas or a whole DC are unavailable. Then test that contract with the real driver, routing policy, request timeout, retry/speculative-execution policy, idempotency classification, and representative network latency.

Consistency level does not replace durable commit-log/storage design, repair, backup/restore, security isolation, schema/partition design, compaction/tombstone management, or application-level idempotency. ALL is not “permanent durability”; LOCAL_ONE is not tenant isolation; a timeout is not proof that a write failed; and a successful weak write does not promise a subsequent weak read will be fresh. Storage-Attached Indexing (SAI) or vector search can add read work but do not redefine RF/CL arithmetic. Managed Cassandra services may restrict topology visibility or CL choices; verify provider semantics rather than assuming Apache Cassandra behavior is exposed unchanged. Chapter 15 now focuses on what happens when requests do not complete cleanly: hints, read repair, speculative retry, timeouts, unavailable errors, overload, and safe client retry behavior.

Summary and next bridge

Tunable consistency becomes engineering only when each request has an explicit invariant, RF/DC arithmetic, latency/availability target, retry/idempotency rule, and failure test. Chapter 15 continues with failure handling, where timeout, unavailable, hints, reconciliation, repair, and speculative execution must be kept distinct.

Authoritative references

Re-check these version-sensitive sources when regenerating this course.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.