Chapter 22 · Configuration, cqlsh, nodetool, Dynamic Settings, and Operational Tooling
cassandra.yaml, Environment / JVM Options, Seeds, Listen / RPC Addresses, and Unit-Aware Settings
Trace Cassandra configuration from Docker/environment and cassandra.yaml through JVM options and node-local running settings; distinguish seeds from leaders and client/internode/JMX addresses.
Learning outcomes
AtlasMart's platform team finds that node 2 advertises a different client address than expected after a container rebuild, while an engineer insists the seed node must remain available because it is “the Cassandra master.” The incident is really about configuration provenance: which source supplied each value, when Cassandra read it, what it advertised, and whether all nodes agree.
Locate cassandra.yaml, cassandra-env.sh, JVM option files, container environment and running virtual-table values.
Explain seeds as discovery contacts rather than leaders and prove the cluster can operate after a seed is unavailable.
Distinguish listen/broadcast internode addresses from rpc/broadcast-rpc native-client addresses and native port 9042.
Read unit-aware duration/storage/rate settings without confusing legacy unit-suffixed names or implicit units.
Separate startup-only configuration from runtime-adjustable settings and identify which source must be version controlled.
The mandatory labs continue the established free/local
AtlasMart cluster: Docker Official Image
cassandra:5.0.9 (latest GA 5.0 patch at
generation time), Java 17 inside that image, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and named disposable data volumes. Keyspace
atlasmart_ops uses
NetworkTopologyStrategy with replication factor
(RF) 3; ordinary reads/writes use LOCAL_QUORUM.
New tables explicitly use UnifiedCompactionStrategy (UCS), no
table default time-to-live (TTL), and Cassandra's normal
gc_grace_seconds. Authentication,
client/internode Transport Layer Security (TLS), and remote
Java Management Extensions (JMX) are disabled only on this
isolated single-host lab. JMX remains local to each container;
do not publish port 7199 to an untrusted network. Apache
Cassandra Java Driver 4.19.3 is the course application
baseline but is optional in this operations chapter. Exact
IPs, host IDs, tokens, configuration rows, logs, metrics,
guardrail output, and timings are learner-captured runtime
evidence.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and configuration-provenance mental model
Configuration provenance is the chain from an
intended baseline to the value a Cassandra process is actually
running. cassandra.yaml holds most server settings;
cassandra-env.sh, jvm*.options,
environment variables, Java system properties, container
orchestration and package defaults can also influence startup. A
startup-only/static setting requires process
restart to take effect. A
runtime/dynamic setting can be changed through
a supported JMX or nodetool surface, but that
change may be node-local and non-persistent unless the
deployment baseline is also updated.
A seed is a discovery/contact point used during
startup and gossip; it is not a leader, primary, quorum
authority or special replica.
listen_address identifies the interface/address
used for internode traffic, while
rpc_address is the bind address for native
client transport; broadcast_address and
broadcast_rpc_address are addresses advertised
to peers/drivers when binding differs from reachability.
JMX is the Java management plane used by
nodetool; it has a different security boundary from
CQL. A virtual table, such as
system_views.settings, exposes node-local runtime
information through CQL but is not replicated and ignores
consistency level. A guardrail warns about or
rejects risky operations/values.
Configuration drift means nodes or deployment
artifacts no longer share the intended settings.
Rolling change means applying a compatible
change one node/failure domain at a time while verifying service
and rollback between steps.
1. Trace intent → rendered file → running process
In the Docker Official Image, environment variables are
transformed into Cassandra configuration before the process
starts. That makes the container environment part of the
provenance chain; the file inside the container is the rendered
artifact, while system_views.settings is evidence
of what the running node believes. These three layers can
diverge after manual edits or runtime JMX changes.
docker exec atlasmart-cass-1 nodetool version -vdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh --versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool statusbinarydocker exec atlasmart-cass-1 nodetool getseeds# Repeat from another node because nodetool/JMX observations are node-local.docker exec atlasmart-cass-2 nodetool status
# Layer 1: container inputs (do not print secrets in production logs).docker inspect atlasmart-cass-1 --format '{{json .Config.Env}}'# Layer 2: rendered files and JVM option inventory.docker exec atlasmart-cass-1 sh -lc "grep -nE '^(cluster_name|num_tokens|listen_address|broadcast_address|rpc_address|broadcast_rpc_address|native_transport_port|endpoint_snitch|seed_provider):' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "find /etc/cassandra -maxdepth 1 -name 'jvm*.options' -print | sort"docker exec atlasmart-cass-1 sh -lc "grep -nE '^-Xms|^-Xmx|^-XX:MaxDirectMemorySize' /etc/cassandra/jvm*.options 2>/dev/null || true"# Layer 3: node-local running values.docker exec atlasmart-cass-1 cqlsh -e "SELECT name,value FROM system_views.settings WHERE name='cluster_name';"docker exec atlasmart-cass-1 cqlsh -e "SELECT name,value FROM system_views.settings WHERE name='rpc_address';"docker exec atlasmart-cass-1 cqlsh -e "SELECT name,value FROM system_views.settings WHERE name='native_transport_port';"
2. Addresses answer different networking questions
| Setting | Role | Important boundary |
|---|---|---|
listen_address |
bind/advertise internode communication | 0.0.0.0 is explicitly wrong |
broadcast_address |
peer-reachable internode address when different | must match routable topology |
rpc_address |
bind address for native CQL transport |
may be 0.0.0.0 only with a valid broadcast
RPC address
|
broadcast_rpc_address |
address advertised to drivers | cannot be 0.0.0.0 |
native_transport_port |
CQL/native protocol TCP port | 9042 by default; do not expose indiscriminately |
| JMX port 7199 | management plane for nodetool | separate auth/TLS boundary; local-only by default |
Native CQL and JMX are different security planes. Cassandra
documentation recommends keeping JMX local by default; if
remote JMX is enabled, use authentication/authorization and
SSL and validate nodetool configuration. Keep
internode/native ports behind deliberate network controls as
well.
3. Seeds are discovery contacts, not leaders
CREATE KEYSPACE IF NOT EXISTS atlasmart_opsWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_ops.config_probe ( probe_id text PRIMARY KEY, status text, owner text, note text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-1','READY','platform','chapter22',toTimestamp(now()));INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-2','READY','orders','copy-fixture',toTimestamp(now()));SELECT * FROM atlasmart_ops.config_probe WHERE probe_id='probe-1';
# Node 2 should report the configured seed contact used for discovery.docker exec atlasmart-cass-2 nodetool getseeds# Safe local failure injection: pause the disposable seed/contact node.docker pause atlasmart-cass-1# Wait until node 2's failure detector reports node 1 down; repeat rather than assuming a fixed delay.docker exec atlasmart-cass-2 nodetool status# Connect directly to a non-seed peer. With RF=3 and two live replicas, LOCAL_QUORUM can still be satisfied.docker exec atlasmart-cass-2 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_ops.config_probe SET note='seed-not-leader',updated_at=toTimestamp(now()) WHERE probe_id='probe-1'; SELECT * FROM atlasmart_ops.config_probe WHERE probe_id='probe-1';"docker unpause atlasmart-cass-1# Wait for all nodes UN, then make convergence authoritative.docker exec atlasmart-cass-2 nodetool statusdocker exec atlasmart-cass-2 nodetool repair --full atlasmart_ops config_probe
This proves only that an already-formed cluster does not elect
or depend on a permanent seed leader for ordinary coordination.
New/bootstrap nodes still need usable discovery information, and
seed-list changes should be managed deliberately (Cassandra 5.0
also exposes nodetool reloadseeds).
4. Unit-aware configuration prevents hidden-unit mistakes
Modern Cassandra configuration accepts explicit units for
durations, data sizes and data rates—for example
60000ms, 16MiB or
50MiB/s. Some parameters impose a minimum accepted
unit because Cassandra's internal representation cannot preserve
finer precision. Legacy names with unit suffixes may remain for
compatibility, but new baselines should use current names and
explicit units where supported.
docker exec atlasmart-cass-1 cqlsh -e "SELECT name,value FROM system_views.settings WHERE name='native_transport_max_frame_size';"docker exec atlasmart-cass-1 cqlsh -e "SELECT name,value FROM system_views.settings WHERE name='request_timeout';"docker exec atlasmart-cass-1 sh -lc "grep -nE 'native_transport_max_frame_size|request_timeout|stream_throughput' /etc/cassandra/cassandra.yaml | head -20"
The exact names/defaults shown by your 5.0.9 image are the evidence. Do not copy an older Cassandra 3.x/4.0 unit-suffixed setting into a 5.0 baseline without checking the current configuration reference and startup logs for deprecation/compatibility messages.
5. Verification and reset
-
All three nodes are
UNafter the seed failure drill. -
system_views.settingsand the rendered file are compared on each node instead of assuming they match. - The learner can state why seeds are not leaders and why address bind/advertise roles differ.
- JMX remains local-only in the mandatory lab.
- No persistent configuration file was modified in this lesson.
Check your understanding
- Why is CASSANDRA_SEEDS not a list of masters?
- Which value is better evidence of the running node after someone edits cassandra.yaml without restarting?
- Why are rpc_address and listen_address not interchangeable?
- Why prefer explicit units?
- What does pausing a seed node prove?
Review the answers
1. Seeds accelerate startup/discovery and gossip. Once joined, Cassandra nodes are peers and any suitable node can coordinate a request.
2. The node-local running configuration, for example system_views.settings, because a file edit alone does not retroactively change startup-only state.
3. rpc_address binds native client transport; listen_address is for internode communication. They serve different network planes and broadcast settings may also differ.
4. They make duration/storage/rate intent auditable and avoid implicit-unit mistakes; the current parameter still defines which units are accepted.
5. Only that a formed peer-to-peer cluster does not depend on a permanent seed leader; it does not remove the need for discovery/bootstrap design.
Production judgment
Operational configuration is part of the database design. Record
the Cassandra patch/JDK/container or package, topology and
failure domains, RF/consistency levels, compaction, repair and
backup schedules, TTL/gc_grace_seconds, SAI/vector
dependencies, disk/memory/network/JVM limits, auth/TLS/JMX
boundaries, native transport exposure, guardrails, driver
retry/idempotency behavior, observability endpoints, and
managed-service overrides. A syntactically valid setting can
still be wrong for workload cardinality, retention, tail
latency, disk headroom or failure behavior. Likewise, a
healthy-looking nodetool status does not prove
repair freshness, query SLOs, disk latency, compaction health,
application correctness or backup recoverability.
Make every change with an owner, compatibility check, blast-radius estimate, pre/post evidence, rollback path and version-controlled intent. Avoid copying tuning values from another cluster without workload evidence. Security-sensitive files, passwords and keystores belong in secret-management systems, not source control. Lesson 2 moves from configuration provenance to operational observation: nodetool status/info/describecluster/ring are useful evidence, but none is a complete health verdict.
Summary and next bridge
Configuration is a provenance chain, not just a YAML file. Seeds
are discovery contacts, network addresses have distinct
bind/advertise roles, units are part of the value, and
node-local running state is the final evidence. Next, interpret
nodetool output without turning one screen into a
false “cluster healthy” conclusion.
Authoritative references
Re-check these version-sensitive references when regenerating the course. Tool availability, defaults, guardrails and configuration names evolve across Cassandra releases and managed services.
- Apache Cassandra downloads / current 5.0 patch
- cassandra.yaml configuration reference
- Unit-aware cassandra.yaml parameters
- Virtual tables and system_views.settings
- nodetool command reference
- cqlsh special commands and COPY
- Cassandra security / JMX access
- Cassandra FAQ / seed semantics
- Docker Official Cassandra 5.0 image
- nodetool reloadseeds