Chapter 22 · Configuration, cqlsh, nodetool, Dynamic Settings, and Operational Tooling
Build a Configuration Baseline and Version-Control Operational Settings Safely
Create a version-controlled Cassandra operational baseline with drift evidence, secrets boundaries, rolling-change gates, rollback metadata and node-lifecycle readiness.
Learning outcomes
AtlasMart can now inspect and change Cassandra correctly, but six months later nobody can answer why node 3 has a different timeout, which seed list is intended, or whether a security-sensitive file was ever committed. The final chapter task is to turn configuration into a versioned operational contract rather than tribal knowledge.
Define a baseline that records Cassandra/JDK/image, topology, networking, storage, security, guardrails and operational policies without committing secrets.
Capture rendered files and node-local running settings for all peers and produce deterministic drift diffs.
Distinguish declarative intent, environment-specific values, secrets and runtime evidence in source-control layout.
Attach change/rollback/testing metadata to rolling configuration updates.
Use the baseline as the handoff into node-lifecycle work, where stale configuration can make bootstrap/replacement unsafe.
The mandatory labs continue the established free/local
AtlasMart cluster: Docker Official Image
cassandra:5.0.9 (latest GA 5.0 patch at
generation time), Java 17 inside that image, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and named disposable data volumes. Keyspace
atlasmart_ops uses
NetworkTopologyStrategy with replication factor
(RF) 3; ordinary reads/writes use LOCAL_QUORUM.
New tables explicitly use UnifiedCompactionStrategy (UCS), no
table default time-to-live (TTL), and Cassandra's normal
gc_grace_seconds. Authentication,
client/internode Transport Layer Security (TLS), and remote
Java Management Extensions (JMX) are disabled only on this
isolated single-host lab. JMX remains local to each container;
do not publish port 7199 to an untrusted network. Apache
Cassandra Java Driver 4.19.3 is the course application
baseline but is optional in this operations chapter. Exact
IPs, host IDs, tokens, configuration rows, logs, metrics,
guardrail output, and timings are learner-captured runtime
evidence.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and configuration-provenance mental model
Configuration provenance is the chain from an
intended baseline to the value a Cassandra process is actually
running. cassandra.yaml holds most server settings;
cassandra-env.sh, jvm*.options,
environment variables, Java system properties, container
orchestration and package defaults can also influence startup. A
startup-only/static setting requires process
restart to take effect. A
runtime/dynamic setting can be changed through
a supported JMX or nodetool surface, but that
change may be node-local and non-persistent unless the
deployment baseline is also updated.
A seed is a discovery/contact point used during
startup and gossip; it is not a leader, primary, quorum
authority or special replica.
listen_address identifies the interface/address
used for internode traffic, while
rpc_address is the bind address for native
client transport; broadcast_address and
broadcast_rpc_address are addresses advertised
to peers/drivers when binding differs from reachability.
JMX is the Java management plane used by
nodetool; it has a different security boundary from
CQL. A virtual table, such as
system_views.settings, exposes node-local runtime
information through CQL but is not replicated and ignores
consistency level. A guardrail warns about or
rejects risky operations/values.
Configuration drift means nodes or deployment
artifacts no longer share the intended settings.
Rolling change means applying a compatible
change one node/failure domain at a time while verifying service
and rollback between steps.
1. What belongs in the baseline
| Domain | Version-controlled intent | Runtime evidence |
|---|---|---|
| software | Cassandra 5.0.9 tag/digest, JDK family, package/image source |
nodetool version -v,
java -version
|
| topology | cluster/DC/rack conventions, vnode count, seed-provider policy | status, getseeds, system.local/peers_v2 |
| network | listen/broadcast/native addresses/ports, firewall intent | settings virtual table, clients/internode virtual tables |
| data | directories, RF defaults/policies, compaction baseline, TTL/gc_grace standards | datapaths, DESCRIBE, tablestats |
| security | auth/TLS/JMX policy and secret references | client SSL/auth state, JMX reachability, role/audit evidence |
| guardrails | approved warning/fail/feature policies | getguardrailsconfig per node |
| operations | repair/backup/snapshot/upgrade/change procedures | repair/snapshot metrics and run logs |
Store secret references and certificate identity/rotation metadata, not plaintext passwords, private keys, truststore passwords or tokens. Pin mutable container tags with a digest in production if reproducibility requirements demand it.
2. Capture rendered and running state without editing nodes
docker exec atlasmart-cass-1 nodetool version -vdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh --versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool statusbinarydocker exec atlasmart-cass-1 nodetool getseeds# Repeat from another node because nodetool/JMX observations are node-local.docker exec atlasmart-cass-2 nodetool status
mkdir -p chapter22-baseline/node1 chapter22-baseline/node2 chapter22-baseline/node3for n in 1 2 3; do docker cp atlasmart-cass-$n:/etc/cassandra/cassandra.yaml chapter22-baseline/node$n/cassandra.yaml docker cp atlasmart-cass-$n:/etc/cassandra/cassandra-env.sh chapter22-baseline/node$n/cassandra-env.sh docker exec atlasmart-cass-$n sh -lc "cat /etc/cassandra/jvm*.options" > chapter22-baseline/node$n/jvm-options.txt docker exec atlasmart-cass-$n cqlsh -e "SELECT * FROM system_views.settings;" > chapter22-baseline/node$n/running-settings.txt docker exec atlasmart-cass-$n nodetool getguardrailsconfig > chapter22-baseline/node$n/guardrails.txt docker exec atlasmart-cass-$n nodetool status > chapter22-baseline/node$n/status.txtdone# Unix/Git Bash comparison example:diff -u chapter22-baseline/node1/cassandra.yaml chapter22-baseline/node2/cassandra.yaml || true# PowerShell equivalent:# Compare-Object (Get-Content chapter22-baseline/node1/cassandra.yaml) (Get-Content chapter22-baseline/node2/cassandra.yaml)
Expected differences are not automatically defects: rack labels and some addresses are intentionally per-node. The baseline needs a normalization/allowlist so drift tooling distinguishes declared variance from accidental variance.
3. Build an auditable repository layout
ops/cassandra/ README.md # ownership, change workflow, restore/rollback links versions.yaml # Cassandra/JDK/image digest/driver compatibility common/cassandra.yaml # common intended settings common/jvm.options # reviewed JVM intent environments/prod-dc1.yaml # allowed environment/DC overrides nodes/rack-map.yaml # explicit failure-domain mapping, if needed guardrails.yaml # intended warn/fail/feature policies runbooks/ rolling-config-change.md repair.md backup-restore.md node-replacement.md evidence/README.md # how to capture current state; not long-lived secrets secrets/README.md # secret-manager paths only; NO secret values
cassandra: release: "5.0.9" image: "cassandra:5.0.9" java_major: 17 cluster_name: atlasmart-course num_tokens: 16 topology: dc1: [rack1, rack2, rack3] native_transport_port: 9042 jmx: remote_enabled: false keyspace_policy: replication: NetworkTopologyStrategy default_application_rf_dc1: 3 compaction_new_tables: UnifiedCompactionStrategy guardrails: allow_filtering_enabled: truesecurity: secret_values_in_git: falsechange_policy: rolling: true require_pre_post_evidence: true require_rollback: true
4. Detect deliberate and accidental drift
find chapter22-baseline -type f -maxdepth 2 -print0 | sort -z | xargs -0 sha256sum > chapter22-baseline/SHA256SUMS# Compare specific settings whose values should be identical.for n in 1 2 3; do echo "node$n" grep -E '^ *cluster_name:|^ *num_tokens:|^ *native_transport_port:' chapter22-baseline/node$n/cassandra.yamldone# Rack/address differences should be explicitly allowed by policy, not ignored wholesale.
/etc/cassandra directory as “the backup.”
That can capture credentials, keystore/truststore paths or environment-specific state and still fail to document the intended deployment. Version control should store reviewed intent plus secret references; evidence bundles should be access-controlled, redacted and retention-managed.
5. Change record and acceptance contract
| Before change | During rolling change | After change |
|---|---|---|
| release docs + compatibility | one node/failure domain at a time | running settings equal intent |
| current evidence + diff | watch client/server error + p95/p99 | representative read/write + failure test |
| capacity/RF/CL headroom | do not overlap topology/repair surprises | schema/repair/backup assumptions revalidated |
| rollback artifact/procedure | stop rollout on acceptance failure | commit change reason/evidence reference |
# Guardrail restored on every node.for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool getguardrailsconfig allow_filtering_enabled; done# All peers normal from more than one observer.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status# Representative data path still works.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_ops.config_probe WHERE probe_id='probe-1';"# Remove local evidence bundle whenever your lab policy no longer needs it; it is not part of the course ZIP.
Check your understanding
- Why store both intended config and runtime evidence?
- Should node-specific rack/address values make every diff irrelevant?
- What must never be committed just because it exists under /etc/cassandra?
- Why pin software/JDK/image identity in the baseline?
- How does this baseline help Chapter 23 node lifecycle work?
Review the answers
1. Intent tells you what should run; runtime evidence tells you what actually runs. Drift exists when they differ without an approved reason.
2. No. Declare allowed variance explicitly, then flag unexpected differences in settings that should be common.
3. Passwords, private keys, keystore/truststore secrets and other credentials; source control should store secure references and policy.
4. Defaults, syntax and operational behavior are version dependent, so reproducibility and rollback require exact software provenance.
5. Bootstrap/replacement/decommission depend on correct node identity, addresses, topology, seed/discovery, tokens, security and resource settings; drift can turn a topology operation into an outage.
Production judgment
Operational configuration is part of the database design. Record
the Cassandra patch/JDK/container or package, topology and
failure domains, RF/consistency levels, compaction, repair and
backup schedules, TTL/gc_grace_seconds, SAI/vector
dependencies, disk/memory/network/JVM limits, auth/TLS/JMX
boundaries, native transport exposure, guardrails, driver
retry/idempotency behavior, observability endpoints, and
managed-service overrides. A syntactically valid setting can
still be wrong for workload cardinality, retention, tail
latency, disk headroom or failure behavior. Likewise, a
healthy-looking nodetool status does not prove
repair freshness, query SLOs, disk latency, compaction health,
application correctness or backup recoverability.
Make every change with an owner, compatibility check, blast-radius estimate, pre/post evidence, rollback path and version-controlled intent. Avoid copying tuning values from another cluster without workload evidence. Security-sensitive files, passwords and keystores belong in secret-management systems, not source control. Chapter 23 uses this baseline during bootstrap, replacement, decommission and cleanup. Those operations change token ownership and stream data, so configuration identity and failure-domain correctness become prerequisites rather than documentation niceties.
Summary and next bridge
A production Cassandra configuration is an auditable contract
spanning declarative intent, secrets management, rendered files,
running node-local values, runtime policies and change evidence.
With that contract in place, Chapter 23 can safely change
cluster membership and ownership instead of treating
nodetool topology commands as one-line recipes.
Authoritative references
Re-check these version-sensitive references when regenerating the course. Tool availability, defaults, guardrails and configuration names evolve across Cassandra releases and managed services.
- Apache Cassandra downloads / current 5.0 patch
- cassandra.yaml configuration reference
- Unit-aware cassandra.yaml parameters
- Virtual tables and system_views.settings
- nodetool command reference
- cqlsh special commands and COPY
- Cassandra security / JMX access
- Cassandra FAQ / seed semantics
- Docker Official Cassandra 5.0 image