Chapter 04 · Keyspaces, Replication Strategies, Replication Factor, and Placement
Replication Factor, Failure Tolerance, Storage Cost, and Consistency-Level Math
Derive RF/consistency-level failure behavior from replica counts and test controlled replica loss safely.
Learning outcomes
AtlasMart now has topology-aware replicas, but request availability still depends on how many replicas each operation requires. This lesson derives behavior from RF/CL math instead of relying on generic “high availability” language.
Calculate ONE, TWO, QUORUM, LOCAL_QUORUM, and ALL requirements for RF=3.
Separate intended replica count from per-request response/acknowledgment count.
Test one-node and two-node loss safely.
Explain quorum overlap without overclaiming linearizability.
Connect RF to storage, repair, streaming, and cost.
Apache Cassandra 5.0.9 is the current GA 5.0 patch as of 7
September 2026. The pinned Docker Official Image
cassandra:5.0.9 uses Java 17. The course cluster
is atlasmart-course, one DC (dc1),
three rack labels, 16 vnodes per node, no TLS/auth in this
isolated chapter lab, and no host-published CQL/JMX ports.
Docker and Cassandra are unavailable in this generation environment. Commands and result shapes were checked against current official Cassandra documentation, but outputs are expected shapes rather than fabricated captured runs. Record exact values on your own disposable cluster.
Reusable Chapter 04 lab topology
Continue the Chapter 01–03 naming and failure-domain conventions. The rack labels let Cassandra exercise rack-aware placement, but all containers still share one physical host, so the lab demonstrates placement semantics rather than true rack/AZ isolation. Require all three nodes to be Up/Normal before interpreting replica evidence.
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
Capture
SELECT cluster_name, data_center, rack, release_version FROM
system.local;,
SELECT peer, data_center, rack, release_version FROM
system.peers_v2;, and nodetool status. If those disagree with the
intended topology, stop and fix the lab rather than explaining
placement from bad metadata.
1. Derive the counts first
For RF=N, QUORUM requires floor(N/2)+1. At RF=3,
QUORUM and LOCAL_QUORUM in this one-DC lab each require two
replicas. ALL requires three; ONE requires one. In multi-DC
deployments the scope of LOCAL_QUORUM differs from global QUORUM
even when a local count looks similar.
| RF=3 CL | Required | 1 replica down | 2 replicas down |
|---|---|---|---|
| ONE / LOCAL_ONE | 1 | Can succeed | Can succeed if remaining replica is reachable |
| TWO | 2 | Can succeed | Cannot meet CL |
| QUORUM / LOCAL_QUORUM | 2 | Can succeed | Cannot meet CL |
| ALL | 3 | Cannot meet CL | Cannot meet CL |
2. Create an RF=3 fixture
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_consistency; CREATE KEYSPACE atlasmart_consistency WITH replication = {'class':'NetworkTopologyStrategy','dc1':3}; CREATE TABLE atlasmart_consistency.stock_by_sku (sku text PRIMARY KEY, quantity int); INSERT INTO atlasmart_consistency.stock_by_sku (sku,quantity) VALUES ('sku-1001',42);"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_consistency stock_by_sku sku-1001docker exec atlasmart-cass-1 nodetool status atlasmart_consistency
Because there are three nodes and RF=3, all three are replicas for this partition. That makes the failure boundary easy to observe.
3. One replica down: LOCAL_QUORUM survives, ALL does not
docker stop atlasmart-cass-3# Wait until a live node reflects the failure.docker exec atlasmart-cass-1 nodetool status atlasmart_consistencydocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_consistency.stock_by_sku SET quantity=41 WHERE sku='sku-1001';"# Expected to fail because ALL needs all 3 replicas.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY ALL; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"
Capture the actual exception rather than hard-coding one message. Failure detection timing can affect whether the client sees an unavailable versus timeout-style outcome.
4. Two replicas down: cross the quorum boundary
docker stop atlasmart-cass-2docker exec atlasmart-cass-1 nodetool status atlasmart_consistency# Expected to fail: LOCAL_QUORUM needs two RF=3 replicas.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"docker start atlasmart-cass-2 atlasmart-cass-3# Wait for Up/Normal before accepting recovery.docker exec atlasmart-cass-1 nodetool status atlasmart_consistencydocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"
The lesson is not that lower CL is “better.” Lower CL can preserve availability with fewer replicas, but it changes participation/freshness guarantees and can expose more ambiguous failure/retry behavior.
5. Quorum overlap and cost
The condition R + W > RF guarantees overlap
between the read and successful write replica sets. At RF=3,
QUORUM+QUORUM gives 2+2>3. This is useful reasoning, but it
does not make every non-LWT Cassandra workload globally
linearizable: concurrent timestamps, topology scope, retries,
outages, and repair still matter.
Higher RF also means more stored data, write fan-out, streaming/repair work, backup volume and capacity headroom. Choose RF/CL from tolerated failures, latency SLOs, RPO/RTO and operating cost.
Verification checklist
- You derived the counts before testing.
- One node down preserved LOCAL_QUORUM but not ALL.
- Two nodes down prevented LOCAL_QUORUM.
- Both nodes were restored and returned Up/Normal.
- You did not equate RF with backup or CL with universal correctness.
Check your understanding
- How many replicas does QUORUM require at RF=3?
- Why can ALL fail while other traffic works?
- Does RF=5 improve availability for ALL?
- What does R+W>RF prove?
- What later mechanism repairs divergence?
Review the answers
1. Two.
2. ALL is an operation-specific requirement for every replica.
3. Not necessarily; ALL would then depend on all five replicas.
4. Replica-set overlap, not every concurrency or linearizability property.
5. Anti-entropy repair, with hints/reconciliation serving separate roles.
Cleanup/reset
Remove only the disposable course-owned containers, volumes, and network. Never copy these cleanup commands into production.
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Alter Replication Safely: Repair Requirements, Data Movement, and Validation.
Authoritative references
- Apache Cassandra downloads — current GA version evidence.
- CQL data definition — CREATE/ALTER KEYSPACE, NTS, SimpleStrategy, RF, durable_writes.
- Repair — incremental/full repair, preview, and RF increase implications.
- Cassandra FAQ — RF increase/decrease operational guidance.
- nodetool — status, getendpoints, repair, cleanup and related operator commands.
- Docker Official Image 5.0 Dockerfile — Cassandra 5.0.9 + Java 17 image baseline.