Chapter 22 · Configuration, cqlsh, nodetool, Dynamic Settings, and Operational Tooling
Runtime / Guardrail Configuration, Rolling Changes, and Avoiding Configuration Drift
Change a Cassandra 5.0 runtime guardrail safely, expose one-node configuration drift, roll the policy across peers, and restore the documented baseline.
Learning outcomes
AtlasMart wants to forbid accidental
ALLOW FILTERING in production. An engineer changes
a runtime guardrail on node 1, sees success, and closes the
ticket. Requests coordinated by nodes 2 and 3 can still behave
differently. This lesson makes the drift visible, rolls the
change across nodes, and restores the lab safely.
Distinguish runtime JMX/nodetool configuration from persisted startup configuration.
Inspect guardrail configuration and use an isolated allow_filtering_enabled toggle as a reversible dynamic-setting experiment.
Demonstrate node-local behavior differences when only one coordinator receives the runtime change.
Apply and verify a rolling runtime change across all nodes, then restore the documented lab baseline.
Explain why runtime success must be followed by persistent configuration/deployment changes and restart/upgrade compatibility review.
The mandatory labs continue the established free/local
AtlasMart cluster: Docker Official Image
cassandra:5.0.9 (latest GA 5.0 patch at
generation time), Java 17 inside that image, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and named disposable data volumes. Keyspace
atlasmart_ops uses
NetworkTopologyStrategy with replication factor
(RF) 3; ordinary reads/writes use LOCAL_QUORUM.
New tables explicitly use UnifiedCompactionStrategy (UCS), no
table default time-to-live (TTL), and Cassandra's normal
gc_grace_seconds. Authentication,
client/internode Transport Layer Security (TLS), and remote
Java Management Extensions (JMX) are disabled only on this
isolated single-host lab. JMX remains local to each container;
do not publish port 7199 to an untrusted network. Apache
Cassandra Java Driver 4.19.3 is the course application
baseline but is optional in this operations chapter. Exact
IPs, host IDs, tokens, configuration rows, logs, metrics,
guardrail output, and timings are learner-captured runtime
evidence.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and configuration-provenance mental model
Configuration provenance is the chain from an
intended baseline to the value a Cassandra process is actually
running. cassandra.yaml holds most server settings;
cassandra-env.sh, jvm*.options,
environment variables, Java system properties, container
orchestration and package defaults can also influence startup. A
startup-only/static setting requires process
restart to take effect. A
runtime/dynamic setting can be changed through
a supported JMX or nodetool surface, but that
change may be node-local and non-persistent unless the
deployment baseline is also updated.
A seed is a discovery/contact point used during
startup and gossip; it is not a leader, primary, quorum
authority or special replica.
listen_address identifies the interface/address
used for internode traffic, while
rpc_address is the bind address for native
client transport; broadcast_address and
broadcast_rpc_address are addresses advertised
to peers/drivers when binding differs from reachability.
JMX is the Java management plane used by
nodetool; it has a different security boundary from
CQL. A virtual table, such as
system_views.settings, exposes node-local runtime
information through CQL but is not replicated and ignores
consistency level. A guardrail warns about or
rejects risky operations/values.
Configuration drift means nodes or deployment
artifacts no longer share the intended settings.
Rolling change means applying a compatible
change one node/failure domain at a time while verifying service
and rollback between steps.
1. Guardrails are policy at the coordinator—not magic cluster state
Cassandra 5.0 exposes guardrail configuration through
nodetool getguardrailsconfig and
setguardrailsconfig. The command uses JMX and
therefore targets one node. A runtime change can be immediately
useful for incident mitigation/testing, but it does not
automatically modify every peer or your version-controlled
cassandra.yaml. Some settings are
runtime-adjustable; others remain startup-only. Never infer
dynamism merely because a similarly named nodetool command
exists.
docker exec atlasmart-cass-1 nodetool version -vdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh --versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool statusbinarydocker exec atlasmart-cass-1 nodetool getseeds# Repeat from another node because nodetool/JMX observations are node-local.docker exec atlasmart-cass-2 nodetool status
CREATE KEYSPACE IF NOT EXISTS atlasmart_opsWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_ops.config_probe ( probe_id text PRIMARY KEY, status text, owner text, note text, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-1','READY','platform','chapter22',toTimestamp(now()));INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-2','READY','orders','copy-fixture',toTimestamp(now()));SELECT * FROM atlasmart_ops.config_probe WHERE probe_id='probe-1';
for n in 1 2 3; do echo "===== node $n =====" docker exec atlasmart-cass-$n nodetool getguardrailsconfig allow_filtering_enabled docker exec atlasmart-cass-$n cqlsh -e "SELECT name,value FROM system_views.settings WHERE name='allow_filtering_enabled';"done# See the supported runtime surface on this exact 5.0.9 build.docker exec atlasmart-cass-1 nodetool help setguardrailsconfig
2. Build a query that requires ALLOW FILTERING
CREATE TABLE IF NOT EXISTS atlasmart_ops.filter_guardrail_probe ( id text PRIMARY KEY, status text, amount int) WITH compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_ops.filter_guardrail_probe (id,status,amount) VALUES ('a','PAID',10);INSERT INTO atlasmart_ops.filter_guardrail_probe (id,status,amount) VALUES ('b','OPEN',20);SELECT * FROM atlasmart_ops.filter_guardrail_probe WHERE status='PAID' ALLOW FILTERING;
The query is deliberately poor Cassandra modeling and small enough to be safe. It exists only to expose the guardrail. Production should normally create a query-first table or appropriate SAI index rather than relying on broad filtering.
3. Create real configuration drift, observe it, then roll
# Cassandra 5.0 source/tests expose this exact runtime setter; default is true.docker exec atlasmart-cass-1 nodetool setguardrailsconfig allow_filtering_enabled falsedocker exec atlasmart-cass-1 nodetool getguardrailsconfig allow_filtering_enabled# Coordinator node 1 should reject the ALLOW FILTERING query.docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_ops.filter_guardrail_probe WHERE status='PAID' ALLOW FILTERING;"# Node 2 still has its old runtime value, demonstrating drift.docker exec atlasmart-cass-2 nodetool getguardrailsconfig allow_filtering_enableddocker exec atlasmart-cass-2 cqlsh -e "SELECT * FROM atlasmart_ops.filter_guardrail_probe WHERE status='PAID' ALLOW FILTERING;"
The precise error text is version-dependent; the important evidence is that the same CQL is accepted/rejected based on the coordinator's runtime guardrail. That is why a change ticket must name scope and verify every node.
docker exec atlasmart-cass-2 nodetool setguardrailsconfig allow_filtering_enabled falsedocker exec atlasmart-cass-3 nodetool setguardrailsconfig allow_filtering_enabled falsefor n in 1 2 3; do docker exec atlasmart-cass-$n nodetool getguardrailsconfig allow_filtering_enabled; done
4. Restore first, then discuss persistence
# Restore the established lab default on every node.for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool setguardrailsconfig allow_filtering_enabled truedonefor n in 1 2 3; do docker exec atlasmart-cass-$n nodetool getguardrailsconfig allow_filtering_enabled; done# Optional cleanup table only; keep config_probe for later lessons.docker exec atlasmart-cass-1 cqlsh -e "DROP TABLE IF EXISTS atlasmart_ops.filter_guardrail_probe;"
A runtime JMX change is operational state. If the policy must survive restart/replacement, update the version-controlled deployment configuration using the exact release's supported YAML/management surface, roll it under a maintenance/change plan, and verify after restart. Managed Cassandra services may expose different policy APIs and may prohibit direct JMX/YAML access.
5. Rolling-change acceptance gates
| Gate | Evidence before next node |
|---|---|
| compatibility | release docs/deprecations/upgrade notes checked |
| availability | RF/CL still tolerates one-node work in the current failure domain |
| runtime | target setting matches intent on changed node |
| service | native transport + representative read/write succeed |
| observability | errors/latency/GC/compaction not regressing |
| rollback | previous value/config artifact ready and tested |
Check your understanding
- Why did the same ALLOW FILTERING query behave differently on node 1 and node 2?
- Does setguardrailsconfig prove the change will survive restart?
- Why roll one node/failure domain at a time?
- Can every cassandra.yaml option be changed dynamically?
- What is the safe final state of this lab?
Review the answers
1. The guardrail was changed through node-local JMX on only one coordinator, creating runtime configuration drift.
2. No. Treat it as runtime state unless the deployment/persistent configuration is also updated and verified.
3. It limits blast radius and lets you verify compatibility, availability and rollback before propagating the change.
4. No. Some have supported runtime controls; many are startup-only. Check the exact release documentation/tool surface.
5. allow_filtering_enabled is restored to true on all three nodes and the temporary filtering table is removed.
Production judgment
Operational configuration is part of the database design. Record
the Cassandra patch/JDK/container or package, topology and
failure domains, RF/consistency levels, compaction, repair and
backup schedules, TTL/gc_grace_seconds, SAI/vector
dependencies, disk/memory/network/JVM limits, auth/TLS/JMX
boundaries, native transport exposure, guardrails, driver
retry/idempotency behavior, observability endpoints, and
managed-service overrides. A syntactically valid setting can
still be wrong for workload cardinality, retention, tail
latency, disk headroom or failure behavior. Likewise, a
healthy-looking nodetool status does not prove
repair freshness, query SLOs, disk latency, compaction health,
application correctness or backup recoverability.
Make every change with an owner, compatibility check, blast-radius estimate, pre/post evidence, rollback path and version-controlled intent. Avoid copying tuning values from another cluster without workload evidence. Security-sensitive files, passwords and keystores belong in secret-management systems, not source control. Lesson 5 converts these practices into a version-controlled baseline: intended settings, rendered/running evidence, secrets boundaries, drift checks, change metadata and rollback become one auditable operational artifact.
Summary and next bridge
Dynamic configuration is powerful precisely because it can create invisible node-local state. A safe change is scoped, rolled, verified, restored or persisted, and linked to a baseline. The final lesson builds that baseline and a repeatable drift-review workflow.
Authoritative references
Re-check these version-sensitive references when regenerating the course. Tool availability, defaults, guardrails and configuration names evolve across Cassandra releases and managed services.
- Apache Cassandra downloads / current 5.0 patch
- cassandra.yaml configuration reference
- Unit-aware cassandra.yaml parameters
- Virtual tables and system_views.settings
- nodetool command reference
- cqlsh special commands and COPY
- Cassandra security / JMX access
- Cassandra FAQ / seed semantics
- Docker Official Cassandra 5.0 image
- nodetool setguardrailsconfig