Chapter 01 · Cassandra Foundations, Version 5.0, Architecture, and Lab Setup
What Cassandra Is: Distributed Wide-Column Storage, Availability Goals, and Workload Fit
Build the distributed mental model before choosing tables: understand what Cassandra distributes, what one coordinator does for a request, how partitions and replicas shape availability, and which workloads fit the trade.
Learning outcomes
AtlasMart has outgrown a single-region event store for order and fulfillment timelines. During promotions it receives a sustained stream of writes from many application instances, and product teams want regional availability without routing every operation through one permanent database primary. The useful question is not “is Cassandra fast?” It is whether Cassandra's partitioned, replicated, peer-to-peer storage model matches AtlasMart's access patterns, failure assumptions, and operational skills.
Explain Apache Cassandra as a distributed, partitioned wide-column database without reducing it to “a key-value store” or “SQL without joins.”
Distinguish node, cluster, coordinator, replica, keyspace, table, row, partition, token, replication factor, and consistency level before using those terms operationally.
Trace a single CQL request from a client connection to a coordinator and replicas, while separating request coordination from permanent leadership.
Identify workloads that fit Cassandra well and workloads that conflict with query-first, bounded-partition modeling.
Prove which Cassandra process and cluster you reached by inspecting system tables and nodetool evidence rather than trusting a container name.
The current generally available Apache Cassandra line is 5.0
and the current patch verified for this chapter is
5.0.9. Reproducible container examples pin
cassandra:5.0.9 instead of using a moving
latest tag. Cassandra 5.0 binary releases support
documented Java 11/17 runtime paths; the lesson always asks
you to record the Java runtime actually present in your
installation or image rather than inferring it from the
Cassandra version.
The generation environment used to build this chapter does not
contain Docker, Cassandra, cqlsh, or
nodetool. Commands were checked against current
official documentation but were not executed here. Expected
output is therefore described by shape and invariant, never
presented as captured output. The lab uses an isolated Docker
network and course-specific containers/volumes. Do not point
any cleanup, failure, configuration, or schema command at
unrelated or production Cassandra data.
1. Cassandra is a distributed database, not a storage format
Apache Cassandra is an open-source distributed NoSQL database that uses a partitioned wide-column storage model. “Wide-column” describes how rows can be organized under a partition key with clustering columns and typed non-key columns; it does not mean a table should contain an arbitrary unbounded number of unrelated columns. CQL, the Cassandra Query Language, gives the model an SQL-like syntax, but the execution assumptions are different: tables are normally designed around known query patterns rather than around normalized entities that will later be joined freely.
A Cassandra cluster is a set of peer nodes that share a logical cluster identity and data-placement metadata. A node is one Cassandra server process. There is no permanent database primary for ordinary reads and writes. Instead, the client contacts a node and that node becomes the coordinator for that request. The coordinator determines which replica nodes own the requested partition, sends work to them as needed, and returns success or failure according to the requested consistency level (CL). Coordination is therefore a per-request role, not a permanent leader role.
A keyspace is the top-level CQL namespace and carries replication configuration. A table's partition key determines which logical partition a row belongs to. Cassandra's partitioner hashes that key to a token, and token ownership plus the keyspace replication strategy determines replica placement. The replication factor (RF) is the number of replicas requested per datacenter by the replication strategy. RF describes copies; CL describes how many relevant replicas must participate in a particular operation. They are related but not interchangeable.
| Term | AtlasMart Chapter 01 meaning | Not the same as |
|---|---|---|
| Coordinator | The node handling one client request | A permanent primary/leader |
| Replica | A node responsible for a copy of a partition | A passive read-only follower |
| Partition | Rows grouped by one partition-key value | A whole table or a filesystem partition |
| RF | Configured number of data replicas in a DC | The number of acknowledgements required by every query |
| CL | Per-operation replica participation requirement | Disk durability or backup policy |
2. Workload fit begins with access patterns and partitions
Cassandra is attractive when the application can name its important queries in advance, route them through well-designed partition keys, keep partitions bounded, and benefit from horizontal scale and replicated availability. AtlasMart examples include append-heavy order timelines by order ID, customer activity buckets by customer and time window, or regional inventory projections keyed so a request reaches a predictable bounded partition. These shapes can be denormalized deliberately: multiple tables may hold the same business fact in forms optimized for different reads.
The same architecture is a poor default for ad-hoc relational exploration, cross-partition joins, arbitrary filtering over fields that were never modeled for access, strict multi-row invariants spanning many partitions, or datasets dominated by a few extreme hot partition keys. Adding nodes increases capacity only when the data model and traffic distribute work. A single celebrity product or tenant mapped to one unbounded hot partition can remain a bottleneck in a very large cluster.
| Workload property | Cassandra fit | Reasoning question |
|---|---|---|
| High sustained write volume with known keys | Often strong | Can writes distribute across bounded partitions? |
| Multi-region replicas and local traffic | Often strong | Are RF, CL, repair, and failure policies explicit? |
| Unknown future joins/ad-hoc analytics | Usually weak as the primary store | Would every new question require a new table or scan? |
| One globally serialized invariant | Requires careful design | Is a Cassandra conditional operation's scope sufficient? |
| Few unbounded hot keys | Risky | Can the partition key/bucketing model spread load? |
3. Follow one request: client → coordinator → replicas
Suppose an AtlasMart service issues a CQL read for one partition. A driver first chooses a reachable Cassandra node according to its routing/load-balancing policy. That contacted node coordinates the operation. It uses cluster metadata to identify replicas for the partition and obtains enough replica responses to satisfy the requested CL. A later request can be coordinated by a different peer. This is why “peer-to-peer” must not be translated into “every node performs every operation” or “there is no coordination.”
The partition key matters before storage details such as SSTables or compaction. It chooses the partition, whose token determines placement. Later chapters will unpack the token ring, replication strategies, write path, read path, compaction, tombstones, and repair; here the acceptance criterion is simpler: you can state which component made each decision and which guarantees are not established by a successful query.
A node can be reachable while a requested CL cannot be satisfied. Conversely, a node failure does not necessarily make the service unavailable if enough replicas remain reachable. The outcome depends on topology, RF, CL, replica health, routing, and the exact operation. Cassandra's availability goals are architectural capabilities that must be configured and tested.
4. Minimal observable lab: prove what answered
This first lab intentionally stays small. It creates one disposable node so the evidence vocabulary is concrete; it does not claim to demonstrate replication or high availability. Docker Desktop on Windows/macOS or Docker Engine on Linux can use the same one-line commands. The private Docker network means port 9042 need not be published on the host.
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9docker logs --tail 80 atlasmart-cass-1
Wait until the node reports that native transport is ready
before treating a failed cqlsh connection as a
configuration problem. Then run the tools inside the container:
docker exec atlasmart-cass-1 cassandra -vdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"docker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name, release_version, cql_version, data_center, rack, listen_address, partitioner FROM system.local;"docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool info
Your exact addresses, host ID, token counts, load numbers,
schema IDs, uptime, and Java patch can differ. The stable
evidence you want is: Cassandra reports the pinned 5.0.9 server
release; system.local reports
atlasmart-course, dc1, and
rack1; nodetool status shows one node
in an Up/Normal state after startup; and the partitioner is the
configured/default partitioner for this release. A one-node
status does not prove RF>1, replica failure tolerance, or
production security.
Smallest AtlasMart state
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_lab WITH replication = {'class':'NetworkTopologyStrategy','dc1':1};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_lab.products_by_id (product_id text PRIMARY KEY, name text, category text, price decimal);"docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_lab.products_by_id (product_id,name,category,price) VALUES ('p-1001','Trail Camera','cameras',129.90);"docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_lab.products_by_id WHERE product_id='p-1001';"
The table name products_by_id states the intended
lookup. That naming convention is deliberate: Cassandra modeling
should make the query shape obvious. One row proves CQL state
exists; it does not yet prove good partition sizing,
replication, repair, or failure behavior.
5. Deliberately wrong approach: “one node worked, so Cassandra is highly available”
A common mistake is to run one container, see a successful
query, and infer Cassandra's distributed guarantees. The
evidence actually proves only that one server accepted the
request. RF=1 means only one replica for this keyspace in
dc1. If that single node disappears, no alternate
replica can satisfy the read.
The repair is conceptual before operational: record topology, RF and CL next to every availability claim. In Lesson 5, AtlasMart will create three peers with RF=3, perform a QUORUM exercise, remove peers in a controlled order, and observe the boundary where the operation can no longer satisfy its consistency requirement.
Cleanup
docker rm -f atlasmart-cass-1docker volume rm atlasmart-cass-1-datadocker network rm atlasmart-cassandra
Removing the named volume intentionally destroys this disposable lab. Never substitute a production or unrelated volume name.
Production judgment
Choose Cassandra because measured workload characteristics match its architecture: bounded query-first partitions, distributed write/read demand, explicit replication and consistency requirements, and a team prepared to operate compaction, tombstones, repair, backups, security, JVM resources, upgrades, and multi-datacenter behavior. Do not choose it merely because “NoSQL scales.” A future migration must also account for denormalized tables, dual writes/backfills, client retry semantics, and a rollback window; those costs are part of the architecture.
The next lesson isolates the peer-to-peer part of this model and explains why seeds, coordinators, replicas, and client contact points are four different concepts.
Check your understanding
- Why is a Cassandra coordinator not the same thing as a primary database node?
- What does the partition key decide before Cassandra reads an SSTable?
- Why does RF=1 not demonstrate Cassandra high availability?
- What does a successful SELECT from system.local prove?
- Name one AtlasMart workload that is a poor Cassandra fit without redesign.
Review the answers
1. A coordinator is chosen per request: it is the contacted peer that routes work to replicas and applies the requested consistency rule. It is not a permanent writer/leader for the cluster.
2. It identifies the logical partition; the partitioner hashes it to a token, which feeds replica placement. Storage-engine lookup happens after that routing context exists.
3. There is only one replica of each partition in that datacenter, so losing the owning node can make that partition unavailable regardless of Cassandra’s peer-to-peer architecture.
4. It proves the identity/configuration metadata reported by the specific node you reached. It does not prove replica health, backups, repair, security, or multi-node failure tolerance.
5. An ad-hoc workload requiring arbitrary joins or filters across many unrelated entities, or a workload concentrated on one unbounded hot key, conflicts with Cassandra’s query-first bounded-partition model.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Peer-to-Peer Architecture vs Primary/Replica Systems: Coordination Without a Permanent Leader.
Authoritative references
- Apache Cassandra 5.0 documentation — Current official documentation entry point for the 5.0 line.
- Apache Cassandra downloads — Official release page used to verify the current 5.0 patch.
- Cassandra architecture overview — Official architecture and wide-column/distributed design framing.
- Cassandra quickstart — Official Docker-oriented learning workflow and isolated-network approach.
- Cassandra configuration reference — Official cassandra.yaml semantics, including native transport and security-related settings.
- CQL querying and cqlsh — Official CQL/cqlsh connection and system.local examples.
- Java support for Cassandra 5.0 — Official Java build/runtime compatibility notes for Cassandra 5.0.
- Docker Official Image: Cassandra — Container-image usage and tag information for the Docker Official Image.