Chapter 01 · Cassandra Foundations, Version 5.0, Architecture, and Lab Setup
Build a Safe Course Lab with Multiple Nodes, Sample Keyspaces, Metrics, and Repeatable Failure Tests
Scale the same lab conventions to three peers, then make RF and QUORUM observable by crossing the exact failure boundary one node at a time and restoring the cluster safely.
Learning outcomes
AtlasMart can now start and inspect one node, but the chapter goal is a distributed mental model. The capstone lab creates three peers, uses explicit racks and RF=3, proves membership from multiple nodes, writes a deterministic AtlasMart row at QUORUM, removes replicas one at a time, observes the exact consistency boundary, restores the cluster, and cleans up only course-owned resources.
Build a three-node, one-datacenter Cassandra 5.0.9 simulation with stable node/rack/volume names and seed discovery.
Create a NetworkTopologyStrategy keyspace with RF=3 and verify schema/replica-related topology evidence before testing failure.
Use QUORUM operations to connect consistency requirements to the number of reachable replicas rather than to node count alone.
Inject reversible node failures and distinguish expected availability behavior from startup/failure-detector timing artifacts.
Produce a repeatable evidence manifest and explain why a single-host multi-container cluster is not equivalent to independent production failure domains.
The current generally available Apache Cassandra line is 5.0
and the current patch verified for this chapter is
5.0.9. Reproducible container examples pin
cassandra:5.0.9 instead of using a moving
latest tag. Cassandra 5.0 binary releases support
documented Java 11/17 runtime paths; the lesson always asks
you to record the Java runtime actually present in your
installation or image rather than inferring it from the
Cassandra version.
The generation environment used to build this chapter does not
contain Docker, Cassandra, cqlsh, or
nodetool. Commands were checked against current
official documentation but were not executed here. Expected
output is therefore described by shape and invariant, never
presented as captured output. The lab uses an isolated Docker
network and course-specific containers/volumes. Do not point
any cleanup, failure, configuration, or schema command at
unrelated or production Cassandra data.
1. Resource and blast-radius contract
Three Cassandra JVMs are materially heavier than a single database container. Use a development machine with enough free memory/CPU/disk that the host does not thrash; a practical learning target is roughly 6–8 GiB of free RAM, multiple CPU cores, and several gigabytes of free disk for a small empty cluster, but treat that as a planning range rather than a Cassandra sizing rule. If the host cannot sustain three nodes, use the single-node lessons for execution and read this lab as a deterministic topology exercise rather than forcing unsafe heap hacks.
Every command below names only
atlasmart-cass-* containers,
atlasmart-cass-*-data volumes, and the
atlasmart-cassandra network. Do not reuse
production seeds, cluster names, data directories, firewall
rules, system clocks, or host kernel settings. A stopped
container is the only failure mechanism required.
2. Start three peers deliberately
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 answers cqlsh before starting peers.docker exec atlasmart-cass-1 cqlsh -e "SELECT release_version FROM system.local;"docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait until all nodes are ready, then inspect:docker exec atlasmart-cass-1 nodetool status
Do not proceed until all three nodes are Up/Normal. If a peer fails to join, inspect its logs and verify cluster name, seed hostname, network membership and rack/DC configuration. Starting all three simultaneously can work, but sequencing node 1 first reduces ambiguity in a teaching lab.
docker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name, host_id, data_center, rack, listen_address, release_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer, data_center, rack, release_version FROM system.peers_v2;"docker exec atlasmart-cass-2 cqlsh atlasmart-cass-2 9042 -e "SELECT cluster_name, host_id, data_center, rack FROM system.local;"docker exec atlasmart-cass-3 cqlsh atlasmart-cass-3 9042 -e "SELECT cluster_name, host_id, data_center, rack FROM system.local;"docker exec atlasmart-cass-1 nodetool describecluster
Expected invariants: three distinct node identities, all in
atlasmart-course/dc1, with racks 1–3,
same Cassandra release, and converged schema/topology state.
Exact IPs, host IDs, tokens and schema UUIDs vary.
3. Create RF=3 data and establish a QUORUM baseline
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_course WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_course.order_status_by_order (order_id text, changed_at timestamp, status text, source text, PRIMARY KEY ((order_id), changed_at)) WITH CLUSTERING ORDER BY (changed_at DESC);"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; INSERT INTO atlasmart_course.order_status_by_order (order_id,changed_at,status,source) VALUES ('o-9001','2026-09-07T11:00:00Z','PLACED','checkout');"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; INSERT INTO atlasmart_course.order_status_by_order (order_id,changed_at,status,source) VALUES ('o-9001','2026-09-07T11:03:00Z','PAID','payments');"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_course.order_status_by_order WHERE order_id='o-9001';"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_course';"
In one datacenter with RF=3, QUORUM requires a
majority of replicas: two. That arithmetic describes replica
responses for the operation, not “two of any three processes in
the universe.” The coordinator must locate the replicas for the
partition and reach enough of them. Chapter 08 later formalizes
CL math across more topologies and LOCAL_QUORUM.
Run the same SELECT through node 2 and node 3 to show that different peers can coordinate access to the replicated partition. Do not infer exactly-once delivery from the inserts; retry safety and timestamps are separate concerns.
4. Failure test A: one replica down, QUORUM remains possible
docker stop atlasmart-cass-3# Wait until the remaining node's failure detector marks node 3 down.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_course.order_status_by_order WHERE order_id='o-9001';"
After failure detection converges, two replicas remain reachable, so a QUORUM read for RF=3 can normally complete. If you test immediately after stopping the container, the observed delay/error can reflect failure detection, transport timeout or coordinator state rather than the steady-state CL boundary. Record the timeline instead of hiding it.
Now write a third event while node 3 is down:
docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; INSERT INTO atlasmart_course.order_status_by_order (order_id,changed_at,status,source) VALUES ('o-9001','2026-09-07T11:08:00Z','PACKING','warehouse');"
A successful QUORUM acknowledgement establishes the requested replica participation at that moment; it does not prove an off-host backup, repair completeness, or that every replica contains the newest state immediately. Cassandra may use hints and later reconciliation mechanisms, while scheduled repair remains an explicit operational responsibility.
5. Failure test B: two replicas down, QUORUM cannot be satisfied
docker stop atlasmart-cass-2# Wait until nodetool on node 1 reflects both peers as down.docker exec atlasmart-cass-1 nodetool status# This query is expected to fail once Cassandra knows only one RF=3 replica is reachable.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_course.order_status_by_order WHERE order_id='o-9001';"
Do not freeze the exact exception text as the lesson's truth. Depending on timing and failure knowledge, the client may observe an unavailable/timeout-style failure. The invariant is that one reachable replica cannot satisfy a consistency requirement of two responses. A timeout also does not always prove that a write was not applied anywhere; application retry policy must treat ambiguous outcomes carefully.
Recover and reconcile the lab
docker start atlasmart-cass-2 atlasmart-cass-3# Wait for both nodes to become queryable and Up/Normal.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-3 cqlsh atlasmart-cass-3 9042 -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_course.order_status_by_order WHERE order_id='o-9001';"
For this short lab, hints may help a returning replica catch
missed writes, but a production repair strategy cannot be
replaced by “we restarted it and the row appeared.” Chapter 18
treats hints, anti-entropy repair, incremental repair and
gc_grace safety explicitly.
6. Wrong conclusions to reject
| Misleading conclusion | Why it is wrong | Safer evidence |
|---|---|---|
| “Three containers means three failure domains.” | They share one host/Docker daemon/storage/network. | Use real rack/zone/DC placement for production resilience tests. |
| “Seed node 1 is the leader.” | It aided discovery; healthy peers coordinate requests. | Query through multiple peers and inspect local identities. |
| “QUORUM success means every replica is current.” | It means the required acknowledgements/responses were obtained. | Observe replica/recovery/repair state separately. |
| “Replication is backup.” | Replicas can copy logical mistakes and share failure domains. | Independent, tested backup/restore with RPO/RTO. |
| “Restarted nodes are repaired.” | Restart, hint replay and anti-entropy repair are different mechanisms. | Use the documented repair plan and validate data convergence. |
7. Verification, reset, and production handoff
Verification checklist
- All three nodes report the same cluster name/release and distinct host identities.
-
Topology metadata records
dc1and racks 1–3. -
The AtlasMart keyspace reports
NetworkTopologyStrategywith RF=3 indc1. - A QUORUM baseline read succeeds with all nodes healthy.
- After one node is classified down, QUORUM remains satisfiable with two replicas.
- After two replica nodes are classified down, the QUORUM operation fails rather than silently lowering consistency.
- Restarted nodes return to Up/Normal and the application-visible timeline is re-readable.
- No host firewall, system clock, production endpoint, or unrelated volume was modified.
Cleanup/reset
# Optional logical cleanup while cluster is healthy:docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_course;"# Destroy only course-owned disposable containers and volumes:docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
Production judgment
This lab establishes the central Cassandra reasoning loop: topology + partition placement + RF + requested CL + failure state → client-visible outcome. Production adds independent racks/zones, multiple datacenters, real drivers and routing policies, TLS/auth, monitored JVM/disk/network capacity, compaction, tombstones, repair, backups, guardrails and tested upgrades. The lab's purpose is to make those later layers observable without pretending a laptop is a production cluster.
Chapter 02 begins by formalizing nodes, datacenters, racks, endpoint snitches, gossip and failure detection so the topology labels used here become operational placement and failure-domain decisions.
Check your understanding
- For RF=3 in one datacenter, how many replicas does QUORUM require?
- Why might a query immediately after docker stop behave differently from the same query after nodetool reports the node down?
- Does a successful QUORUM write prove all three replicas stored the write?
- Why is the three-node Docker lab not a disaster-recovery test?
- What is the most important invariant carried into Chapter 02?
Review the answers
1. A majority of three replicas: two. The requirement concerns replicas for the partition, not simply any two cluster processes.
2. Failure detection and transport state need time to converge. Early behavior can include timeout/connection effects before steady-state unavailability is known.
3. No. It proves enough replicas responded for the requested consistency level. Other replicas can be temporarily behind and require hints/reconciliation/repair.
4. All nodes share one physical failure domain, and the lab does not create independent backups, regions, networks, or recovery infrastructure.
5. Client-visible availability depends on topology and replica placement plus RF/CL and current failures; labels such as “peer-to-peer” are not sufficient evidence by themselves.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Cluster Topology: Nodes, Datacenters, Racks, Endpoints, and Failure Domains.
Authoritative references
- Apache Cassandra 5.0 documentation — Current official documentation entry point for the 5.0 line.
- Apache Cassandra downloads — Official release page used to verify the current 5.0 patch.
- Cassandra architecture overview — Official architecture and wide-column/distributed design framing.
- Cassandra quickstart — Official Docker-oriented learning workflow and isolated-network approach.
- Cassandra configuration reference — Official cassandra.yaml semantics, including native transport and security-related settings.
- CQL querying and cqlsh — Official CQL/cqlsh connection and system.local examples.
- Java support for Cassandra 5.0 — Official Java build/runtime compatibility notes for Cassandra 5.0.
- Docker Official Image: Cassandra — Container-image usage and tag information for the Docker Official Image.
- Cassandra consistency guarantees — Official consistency/availability semantics and tunable-consistency context.