Chapter 22 · Configuration, cqlsh, nodetool, Dynamic Settings, and Operational Tooling
nodetool status / info / describecluster / ring Awareness and Reading Cluster Health
Interpret nodetool status, info, describecluster and ring as scoped operational evidence, including vnode/keyspace ownership and observer-dependent failure detection.
Learning outcomes
AtlasMart sees three green UN rows in
nodetool status, yet p99 latency is rising and one
replica has not repaired recently. A status screen can tell you
membership/state/load/topology from one node's perspective; it
cannot prove application health, replica convergence, disk
latency or backup recoverability.
Interpret nodetool status state codes, load, tokens, host IDs, racks and keyspace-scoped effective ownership.
Use nodetool info and describecluster to collect process/partitioner/snitch/schema evidence without overclaiming health.
Explain why nodetool ring becomes verbose with vnodes and why ownership must be interpreted with a keyspace.
Compare status from multiple nodes because failure detection is local and may temporarily disagree.
Build a layered health checklist that includes query latency, repair, compaction, disk and client evidence beyond nodetool status.
The mandatory labs continue the established free/local
AtlasMart cluster: Docker Official Image
cassandra:5.0.9 (latest GA 5.0 patch at
generation time), Java 17 inside that image, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and named disposable data volumes. Keyspace
atlasmart_ops uses
NetworkTopologyStrategy with replication factor
(RF) 3; ordinary reads/writes use LOCAL_QUORUM.
New tables explicitly use UnifiedCompactionStrategy (UCS), no
table default time-to-live (TTL), and Cassandra's normal
gc_grace_seconds. Authentication,
client/internode Transport Layer Security (TLS), and remote
Java Management Extensions (JMX) are disabled only on this
isolated single-host lab. JMX remains local to each container;
do not publish port 7199 to an untrusted network. Apache
Cassandra Java Driver 4.19.3 is the course application
baseline but is optional in this operations chapter. Exact
IPs, host IDs, tokens, configuration rows, logs, metrics,
guardrail output, and timings are learner-captured runtime
evidence.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and configuration-provenance mental model
Configuration provenance is the chain from an
intended baseline to the value a Cassandra process is actually
running. cassandra.yaml holds most server settings;
cassandra-env.sh, jvm*.options,
environment variables, Java system properties, container
orchestration and package defaults can also influence startup. A
startup-only/static setting requires process
restart to take effect. A
runtime/dynamic setting can be changed through
a supported JMX or nodetool surface, but that
change may be node-local and non-persistent unless the
deployment baseline is also updated.
A seed is a discovery/contact point used during
startup and gossip; it is not a leader, primary, quorum
authority or special replica.
listen_address identifies the interface/address
used for internode traffic, while
rpc_address is the bind address for native
client transport; broadcast_address and
broadcast_rpc_address are addresses advertised
to peers/drivers when binding differs from reachability.
JMX is the Java management plane used by
nodetool; it has a different security boundary from
CQL. A virtual table, such as
system_views.settings, exposes node-local runtime
information through CQL but is not replicated and ignores
consistency level. A guardrail warns about or
rejects risky operations/values.
Configuration drift means nodes or deployment
artifacts no longer share the intended settings.
Rolling change means applying a compatible
change one node/failure domain at a time while verifying service
and rollback between steps.
1. status is one observer's topology snapshot
docker exec atlasmart-cass-1 nodetool version -vdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh --versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool statusbinarydocker exec atlasmart-cass-1 nodetool getseeds# Repeat from another node because nodetool/JMX observations are node-local.docker exec atlasmart-cass-2 nodetool status
for n in 1 2 3; do echo "===== observer atlasmart-cass-$n =====" docker exec atlasmart-cass-$n nodetool status atlasmart_ops docker exec atlasmart-cass-$n nodetool info docker exec atlasmart-cass-$n nodetool describeclusterdone
U/D is the observer's current up/down view;
N/L/J/M describes normal/leaving/joining/moving
token state. Load is not “query load”; it is
storage data size reported by the node. Host ID is node
identity. Rack/DC labels are placement metadata. With a keyspace
argument, effective ownership accounts for its replication
strategy; without the right keyspace context, ownership
percentages are easy to misread.
2. ring is a token diagnostic, not the preferred health dashboard
# With 16 vnodes per node this is intentionally verbose.docker exec atlasmart-cass-1 nodetool ring | head -80# Specify the keyspace when reasoning about effective ownership/replication.docker exec atlasmart-cass-1 nodetool ring atlasmart_ops | head -80docker exec atlasmart-cass-1 nodetool describering atlasmart_ops
nodetool ring and declare capacity
balanced.
Vnodes produce many token rows; ownership is
topology/replication dependent. The ring command
itself tells operators to provide a keyspace for accurate
ownership information. Capacity decisions also need actual
bytes, partition distribution, hot keys, compaction debt, disk
performance and workload throughput—not just token math.
3. Controlled disagreement: failure detection is node-local
CREATE KEYSPACE IF NOT EXISTS atlasmart_opsWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_ops.config_probe ( probe_id text PRIMARY KEY, status text, owner text, note text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-1','READY','platform','chapter22',toTimestamp(now()));INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-2','READY','orders','copy-fixture',toTimestamp(now()));SELECT * FROM atlasmart_ops.config_probe WHERE probe_id='probe-1';
docker pause atlasmart-cass-3# Immediately sample; observers may not agree yet.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status# Repeat after failure detection has had time to update; exact timing is environment dependent.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool statusdocker unpause atlasmart-cass-3# Verify recovery from at least two observers.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status
Temporary disagreement is not automatically a split-brain event; Cassandra's phi-accrual failure detector is local to each node. Correlate with driver node state, internode metrics, logs and request outcomes before deciding what failed.
4. A layered health evidence matrix
| Layer | Useful evidence | What a green result still does not prove |
|---|---|---|
| membership/topology | status, gossip, host IDs, DC/rack |
replica convergence or query SLO |
| process/JVM | info, uptime, heap, GC/JVM metrics |
disk health or correctness |
| schema/partitioner | describecluster, schema tables |
application-compatible migrations complete everywhere |
| storage | tablestats, compactionstats, disk metrics | repair freshness or client routing |
| repair/consistency | repair history/state, failure drills | backup restore RTO/RPO |
| application | driver p50/p95/p99, errors, retries, coordinator traces | future failure capacity |
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_ops.config_probedocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM system_views.thread_pools;" | head -80docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM system_views.coordinator_read_latency WHERE keyspace_name='atlasmart_ops' ALLOW FILTERING;"
5. Verification
- All nodes return to UN after the failure drill.
-
Effective ownership is interpreted with
atlasmart_ops, not a context-free ring percentage. -
The learner explains why
statuscan differ by observer. - A health assessment includes storage/repair/client evidence beyond topology state.
- No destructive topology command is used.
Check your understanding
- Does UN mean every replica is consistent?
- Why pass a keyspace to status/ring when discussing ownership?
- What does nodetool info load represent?
- Why run status from more than one node during an incident?
- What should accompany nodetool output before declaring the cluster healthy?
Review the answers
1. No. It means the observer sees the node Up and Normal; it says nothing by itself about repair freshness or replica equality.
2. Effective ownership depends on that keyspace replication strategy and topology.
3. Node storage/data load, not request throughput or CPU utilization.
4. Up/down is determined independently by each failure detector, so observers can temporarily disagree.
5. Application latency/errors, storage/compaction, repair state, disk/JVM/network metrics and recovery evidence appropriate to the SLO.
Production judgment
Operational configuration is part of the database design. Record
the Cassandra patch/JDK/container or package, topology and
failure domains, RF/consistency levels, compaction, repair and
backup schedules, TTL/gc_grace_seconds, SAI/vector
dependencies, disk/memory/network/JVM limits, auth/TLS/JMX
boundaries, native transport exposure, guardrails, driver
retry/idempotency behavior, observability endpoints, and
managed-service overrides. A syntactically valid setting can
still be wrong for workload cardinality, retention, tail
latency, disk headroom or failure behavior. Likewise, a
healthy-looking nodetool status does not prove
repair freshness, query SLOs, disk latency, compaction health,
application correctness or backup recoverability.
Make every change with an owner, compatibility check, blast-radius estimate, pre/post evidence, rollback path and version-controlled intent. Avoid copying tuning values from another cluster without workload evidence. Security-sensitive files, passwords and keystores belong in secret-management systems, not source control. Lesson 3 turns to cqlsh: useful for interactive CQL, tracing, schema inspection and small COPY/scripting tasks, but not a production bulk-migration or application-client substitute.
Summary and next bridge
nodetool narrows operational questions; it does not
collapse them into one health score. Read topology, process and
ownership output in context, from multiple observers where
necessary, then correlate with storage, repair and client
evidence. Next, use cqlsh precisely within its
interactive/scripting boundaries.
Authoritative references
Re-check these version-sensitive references when regenerating the course. Tool availability, defaults, guardrails and configuration names evolve across Cassandra releases and managed services.
- Apache Cassandra downloads / current 5.0 patch
- cassandra.yaml configuration reference
- Unit-aware cassandra.yaml parameters
- Virtual tables and system_views.settings
- nodetool command reference
- cqlsh special commands and COPY
- Cassandra security / JMX access
- Cassandra FAQ / seed semantics
- Docker Official Cassandra 5.0 image
- nodetool status
- nodetool ring
- nodetool describecluster