Chapter 01 · Cassandra Foundations, Version 5.0, Architecture, and Lab Setup

Start a Single-Node or Container Lab, Connect with cqlsh, and Inspect Cluster Metadata

Build the chapter’s single-node reference environment and prove each layer—container, Cassandra readiness, CQL, JMX management, configuration, storage, schema and application-visible state—before scaling out.

Intermediate125–150 minutesSingle-node cqlsh + nodetool + AtlasMart state labApache Cassandra 5.0.9 · pinned Docker Official ImageLast reviewed: September 2026

Learning outcomes

AtlasMart now has a version contract; the next risk is operational ambiguity. A container can be “running” while Cassandra is still initializing, port 9042 can be unreachable from the host because it was intentionally not published, and nodetool can fail for a JMX reason while CQL is healthy. This lesson builds a single-node lab whose process, configuration, storage, network and metadata are all observable.

01

Start a pinned single-node Cassandra 5.0.9 lab on a private Docker network with persistent course-scoped data.

02

Distinguish Docker process state, Cassandra readiness, CQL native transport and local JMX management connectivity.

03

Use cqlsh, system.local, nodetool status/info and logs to prove the exact node identity and topology metadata.

04

Create a minimal query-shaped AtlasMart keyspace/table and verify before/after state deterministically.

05

Diagnose wrong host/port, startup-not-ready, accidental host exposure and ephemeral-storage mistakes without unsafe host changes.

Chapter baseline reviewed 7 September 2026

The current generally available Apache Cassandra line is 5.0 and the current patch verified for this chapter is 5.0.9. Reproducible container examples pin cassandra:5.0.9 instead of using a moving latest tag. Cassandra 5.0 binary releases support documented Java 11/17 runtime paths; the lesson always asks you to record the Java runtime actually present in your installation or image rather than inferring it from the Cassandra version.

Execution and safety note

The generation environment used to build this chapter does not contain Docker, Cassandra, cqlsh, or nodetool. Commands were checked against current official documentation but were not executed here. Expected output is therefore described by shape and invariant, never presented as captured output. The lab uses an isolated Docker network and course-specific containers/volumes. Do not point any cleanup, failure, configuration, or schema command at unrelated or production Cassandra data.

1. Create the boundary before starting the database

The safest default for this course is an isolated Docker network. Running cqlsh and nodetool inside the container means neither CQL port 9042 nor JMX needs to be published to the host. A named data volume makes the persistence boundary visible. These choices reduce the blast radius while preserving the same Cassandra server semantics needed for the lesson.

setup · network, volume, pinned node
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9docker ps --filter name=atlasmart-cass-1docker inspect atlasmart-cass-1 --format '{{json .NetworkSettings.Networks}}'docker inspect atlasmart-cass-1 --format '{{json .Mounts}}' 

docker ps proves the container process is running. It does not prove Cassandra has finished bootstrapping or that native transport is accepting CQL. Cassandra startup is asynchronous and can take materially longer on a constrained laptop.

Readiness, not guessing

readiness · CQL evidence plus logs
# Repeat until it succeeds; do not treat the first startup refusal as a permanent fault.docker exec atlasmart-cass-1 cqlsh -e "SELECT release_version FROM system.local;"# Inspect the most recent server messages if it does not succeed yet.docker logs --tail 120 atlasmart-cass-1

On Bash you may wrap the first command in a retry loop; in PowerShell use Start-Sleep and check $LASTEXITCODE. The course avoids a fixed “sleep 20 seconds” because startup time depends on CPU, storage and memory.

2. Inspect server, CQL and management identity

evidence · identity, topology and tool surfaces
docker exec atlasmart-cass-1 cassandra -vdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE CLUSTER"docker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name, host_id, release_version, cql_version, data_center, rack, listen_address, partitioner FROM system.local;"docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool infodocker exec atlasmart-cass-1 nodetool describecluster

Record the actual output. Expected invariants are: the release identifies the pinned Cassandra patch; cluster metadata says atlasmart-course; topology says dc1/rack1; nodetool status reports one Up/Normal node after startup; and system.local exposes one node-local host ID. Token counts, load, addresses, schema version, uptime and heap values are machine-specific.

cqlsh exercises the CQL native transport. nodetool exercises management/JMX paths. One can fail while the other succeeds. That distinction matters when diagnosing “Cassandra is down”: a JMX authentication or port problem is not the same failure as a native-transport connection refusal.

3. Inspect configuration and storage without changing them

inspection · config paths, data paths, listening sockets
docker exec atlasmart-cass-1 sh -lc 'echo "$CASSANDRA_CONF"; grep -nE "^(cluster_name|listen_address|rpc_address|endpoint_snitch|authenticator|authorizer|native_transport_port|commitlog_directory|saved_caches_directory):" "$CASSANDRA_CONF/cassandra.yaml"'docker exec atlasmart-cass-1 sh -lc 'cat "$CASSANDRA_CONF/cassandra-rackdc.properties"'docker exec atlasmart-cass-1 sh -lc 'find /var/lib/cassandra -maxdepth 2 -type d | sort | head -60'docker exec atlasmart-cass-1 sh -lc 'ss -lnt 2>/dev/null | grep -E ":(7000|7001|7199|9042)\b" || true' 

The listening sockets are inside the container namespace. Docker's image metadata may expose ports without publishing them on the host. Verify host publication separately with docker port atlasmart-cass-1; the expected safe result for this lab is no host mapping. If you later need a host-native driver, publish only 127.0.0.1:9042:9042 for that exercise and keep JMX private unless the lesson explicitly secures and needs it.

Default auth is a lab boundary, not a recommendation.

The standard development image can start with permissive authentication/authorization defaults. The lab is acceptable only because the database is isolated on a course-only Docker network and not exposed publicly. Production authentication, authorization, TLS, certificates, JMX security, secrets rotation and network policy are taught separately in Chapter 26.

4. Create AtlasMart state and prove every transition

AtlasMart lab · schema + before/after query evidence
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_lab WITH replication = {'class':'NetworkTopologyStrategy','dc1':1};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_lab.inventory_by_product (product_id text, warehouse_id text, quantity int, updated_at timestamp, PRIMARY KEY ((product_id), warehouse_id));"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_lab';"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, table_name FROM system_schema.tables WHERE keyspace_name='atlasmart_lab';"docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_lab.inventory_by_product (product_id,warehouse_id,quantity,updated_at) VALUES ('p-1001','wh-baku',12,'2026-09-07T10:00:00Z');"docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_lab.inventory_by_product WHERE product_id='p-1001';"docker exec atlasmart-cass-1 cqlsh -e "UPDATE atlasmart_lab.inventory_by_product SET quantity=9, updated_at='2026-09-07T10:05:00Z' WHERE product_id='p-1001' AND warehouse_id='wh-baku';"docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_lab.inventory_by_product WHERE product_id='p-1001';"

The first query proves keyspace replication metadata, the second proves table existence, and the row queries prove application-visible state transitions. Cassandra writes are upserts: an UPDATE can create state even if the row did not previously exist. Chapter 06 treats timestamps, TTL, null/unset and mutation semantics in depth; do not import SQL row-existence assumptions into this lesson.

The partition key here is product_id, with warehouse_id clustering rows inside that product partition. This is intentionally small and pedagogical. It is not yet a defended production model: AtlasMart must later bound warehouses/cardinality, model required queries, measure hot products, and decide whether a different partition key or bucketing scheme is needed.

5. Diagnose three realistic failures

Failure A — Cassandra is still starting

A connection-refused error immediately after docker run does not prove an address error. Inspect logs and retry CQL readiness. Do not “fix” the situation by disabling security or changing random ports.

Failure B — host cqlsh cannot reach 9042

That is expected if you intentionally did not publish a host port. Use docker exec for the course lab. If a host-native driver is necessary, recreate the disposable node with a loopback-only mapping such as -p 127.0.0.1:9042:9042. Never use broad -p 9042:9042 as a reflex on an untrusted network.

Failure C — data vanishes after container replacement

If the node was started without the named volume, its writable container layer was the persistence boundary. The repair is to use the explicit course volume and verify the mount with docker inspect. Do not conclude that Cassandra persistence failed; the container-storage design failed.

Verification checklist

  • Container is running and CQL readiness succeeds.
  • system.local reports the intended cluster/DC/rack and exact Cassandra release.
  • nodetool status/info can query management state.
  • No CQL or JMX port is published on the host in the default lab.
  • The named data volume is mounted at the Cassandra data boundary.
  • AtlasMart keyspace/table/row state can be re-read deterministically.

Cleanup/reset

cleanup · schema first, then disposable infrastructure
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_lab;"docker rm -f atlasmart-cass-1# Remove the volume only when you intentionally want a full data reset.docker volume rm atlasmart-cass-1-datadocker network rm atlasmart-cassandra

Production judgment

A healthy single node is a useful syntax and storage-path laboratory, not a Cassandra production architecture. It has no replica failure tolerance, shares one failure domain, and leaves repair/failover behavior untested. Its value is observability: every later multi-node claim can now be compared with known server, CQL, JMX, storage and topology evidence.

The final lesson scales this exact lab—not a new naming scheme—to three nodes, RF=3, explicit racks, QUORUM operations, and reversible node failures.

Check your understanding

  1. Why is docker ps insufficient as Cassandra readiness evidence?
  2. What different surfaces do cqlsh and nodetool exercise?
  3. Why is no host mapping for 9042 a feature in this lab?
  4. What does mounting /var/lib/cassandra to a named volume change?
  5. Why is inventory_by_product not automatically a good production table?
Review the answers

1. It proves the container process exists, not that Cassandra has completed startup and is accepting native-protocol CQL requests.

2. cqlsh uses the Cassandra native protocol/CQL service; nodetool uses the management/JMX surface. Failures can therefore have different causes.

3. The course can run cqlsh inside the private Docker network, reducing accidental exposure of an unauthenticated development database.

4. It makes local data survive container replacement and makes the persistence boundary explicit, but it still does not create an independent backup.

5. Its partition cardinality, query set, hot-key distribution, retention and growth have not been measured or defended yet; Chapter 07 modeling work is still required.

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Build a Safe Course Lab with Multiple Nodes, Sample Keyspaces, Metrics, and Repeatable Failure Tests.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.