Chapter 04 · Keyspaces, Replication Strategies, Replication Factor, and Placement
CREATE KEYSPACE and Replication Configuration: Durable Writes and Namespaces
Build a precise Cassandra keyspace model: namespace, replication policy, replication factor, durable writes, and observable natural replicas.
Learning outcomes
AtlasMart needs a Cassandra namespace for inventory data, but a keyspace is more than a database-like name: it is also the replication and durable-write contract for every table inside it.
Explain keyspace, replication strategy, RF, CL, durable_writes, and backup as distinct mechanisms.
Create and inspect a NetworkTopologyStrategy keyspace.
Resolve a concrete partition key to its natural replica endpoints.
Explain why schema agreement does not prove physical data convergence.
Diagnose unsafe durable_writes and RF assumptions.
Apache Cassandra 5.0.9 is the current GA 5.0 patch as of 7
September 2026. The pinned Docker Official Image
cassandra:5.0.9 uses Java 17. The course cluster
is atlasmart-course, one DC (dc1),
three rack labels, 16 vnodes per node, no TLS/auth in this
isolated chapter lab, and no host-published CQL/JMX ports.
Docker and Cassandra are unavailable in this generation environment. Commands and result shapes were checked against current official Cassandra documentation, but outputs are expected shapes rather than fabricated captured runs. Record exact values on your own disposable cluster.
Reusable Chapter 04 lab topology
Continue the Chapter 01–03 naming and failure-domain conventions. The rack labels let Cassandra exercise rack-aware placement, but all containers still share one physical host, so the lab demonstrates placement semantics rather than true rack/AZ isolation. Require all three nodes to be Up/Normal before interpreting replica evidence.
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
Capture
SELECT cluster_name, data_center, rack, release_version FROM
system.local;,
SELECT peer, data_center, rack, release_version FROM
system.peers_v2;, and nodetool status. If those disagree with the
intended topology, stop and fix the lab rather than explaining
placement from bad metadata.
1. Keyspace = namespace + replication contract
A keyspace groups tables and carries the
replication policy that determines which nodes should hold
replicas of each token range.
Replication factor (RF) is the configured
number of copies; for NetworkTopologyStrategy it is stated per
datacenter. Consistency level (CL) is chosen
per operation and determines how many replicas must participate.
durable_writes controls the commit-log durability
path for normal writes. A backup is a separate recoverability
artifact. Treating any two of these as synonyms produces bad
operational decisions.
| Control | Scope | Question answered |
|---|---|---|
| RF | Keyspace placement | How many replicas should exist in each DC? |
| CL | Per request | How many/location of replicas must respond? |
| durable_writes | Write path | Should mutations use the commit log? |
| Backup | Recovery | Can we restore after destructive or correlated failure? |
2. Create RF=2 and inspect the exact schema
Use a chapter-specific keyspace so replication experiments do not mutate earlier AtlasMart fixtures.
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_replication WITH replication = {'class':'NetworkTopologyStrategy','dc1':2} AND durable_writes = true;"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_replication.inventory_by_sku (sku text PRIMARY KEY, quantity int, warehouse text);"docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_replication.inventory_by_sku (sku,quantity,warehouse) VALUES ('sku-1001',42,'tehran-01');"docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE KEYSPACE atlasmart_replication"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, durable_writes, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_replication';"
Expected evidence: NTS, dc1:2, and
durable_writes=true.
DESCRIBE reconstructs schema for humans;
system_schema.keyspaces is queryable metadata.
Neither shows whether a later RF change has already moved
historical rows.
3. Map one AtlasMart key to replicas
docker exec atlasmart-cass-1 cqlsh -e "SELECT sku, token(sku) AS token_value FROM atlasmart_replication.inventory_by_sku WHERE sku='sku-1001';"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_replication inventory_by_sku sku-1001docker exec atlasmart-cass-1 nodetool status atlasmart_replication
At RF=2, the command should return two natural endpoints. Record the actual endpoint addresses and map them to rack labels; do not hard-code which two racks you expect because token allocation determines the starting point.
4. Deliberately wrong: disable durable writes as a “performance tweak”
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE atlasmart_no_commitlog_demo WITH replication = {'class':'NetworkTopologyStrategy','dc1':2} AND durable_writes = false;"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, durable_writes, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_no_commitlog_demo';"docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE atlasmart_no_commitlog_demo;"
This proves the risky policy exists; it does not manufacture a crash or fake an exact loss window. The correction is to keep durable writes enabled unless a deliberately analyzed workload justifies otherwise, while still treating replication and backup separately.
5. Production judgment
Choose keyspace boundaries from shared replication/recovery needs. Record exact DC names, RF per DC, intended CLs, repair cadence, storage/network cost, backup policy, and failure assumptions. RF=2 is useful here for teaching but is not a universal production recommendation.
Verification checklist
- Three nodes are Up/Normal.
-
atlasmart_replicationis NTS with RF=2 in dc1. - You captured both DESCRIBE and system_schema evidence.
- You resolved
sku-1001to two endpoints. - You can distinguish RF, CL, durable writes, and backup.
Check your understanding
- What does a keyspace define beyond naming?
- Does RF=3 tell you a request used three acknowledgments?
- What does schema agreement prove?
- Is replication a backup?
- What does Lesson 2 add?
Review the answers
1. Replication strategy/factor and durable-write policy.
2. No; CL defines the per-request requirement.
3. Nodes agree on schema metadata, not that historical data has converged.
4. No; replicas can share the same logical corruption or destructive operation.
5. How NTS uses DC/rack topology to place replicas.
Cleanup/reset
Remove only the disposable course-owned containers, volumes, and network. Never copy these cleanup commands into production.
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to NetworkTopologyStrategy: Replication Factor per Datacenter and Rack-Aware Placement.
Authoritative references
- Apache Cassandra downloads — current GA version evidence.
- CQL data definition — CREATE/ALTER KEYSPACE, NTS, SimpleStrategy, RF, durable_writes.
- Repair — incremental/full repair, preview, and RF increase implications.
- Cassandra FAQ — RF increase/decrease operational guidance.
- nodetool — status, getendpoints, repair, cleanup and related operator commands.
- Docker Official Image 5.0 Dockerfile — Cassandra 5.0.9 + Java 17 image baseline.