Chapter 22 · Configuration, cqlsh, nodetool, Dynamic Settings, and Operational Tooling

Runtime / Guardrail Configuration, Rolling Changes, and Avoiding Configuration Drift

Change a Cassandra 5.0 runtime guardrail safely, expose one-node configuration drift, roll the policy across peers, and restore the documented baseline.

Intermediate → Advanced115–155 minutesDynamic guardrail + drift labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · RF=3 · LOCAL_QUORUM · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart wants to forbid accidental ALLOW FILTERING in production. An engineer changes a runtime guardrail on node 1, sees success, and closes the ticket. Requests coordinated by nodes 2 and 3 can still behave differently. This lesson makes the drift visible, rolls the change across nodes, and restores the lab safely.

01

Distinguish runtime JMX/nodetool configuration from persisted startup configuration.

02

Inspect guardrail configuration and use an isolated allow_filtering_enabled toggle as a reversible dynamic-setting experiment.

03

Demonstrate node-local behavior differences when only one coordinator receives the runtime change.

04

Apply and verify a rolling runtime change across all nodes, then restore the documented lab baseline.

05

Explain why runtime success must be followed by persistent configuration/deployment changes and restart/upgrade compatibility review.

Chapter 22 lab baseline

The mandatory labs continue the established free/local AtlasMart cluster: Docker Official Image cassandra:5.0.9 (latest GA 5.0 patch at generation time), Java 17 inside that image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, and named disposable data volumes. Keyspace atlasmart_ops uses NetworkTopologyStrategy with replication factor (RF) 3; ordinary reads/writes use LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no table default time-to-live (TTL), and Cassandra's normal gc_grace_seconds. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only on this isolated single-host lab. JMX remains local to each container; do not publish port 7199 to an untrusted network. Apache Cassandra Java Driver 4.19.3 is the course application baseline but is optional in this operations chapter. Exact IPs, host IDs, tokens, configuration rows, logs, metrics, guardrail output, and timings are learner-captured runtime evidence.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and configuration-provenance mental model

Configuration provenance is the chain from an intended baseline to the value a Cassandra process is actually running. cassandra.yaml holds most server settings; cassandra-env.sh, jvm*.options, environment variables, Java system properties, container orchestration and package defaults can also influence startup. A startup-only/static setting requires process restart to take effect. A runtime/dynamic setting can be changed through a supported JMX or nodetool surface, but that change may be node-local and non-persistent unless the deployment baseline is also updated.

A seed is a discovery/contact point used during startup and gossip; it is not a leader, primary, quorum authority or special replica. listen_address identifies the interface/address used for internode traffic, while rpc_address is the bind address for native client transport; broadcast_address and broadcast_rpc_address are addresses advertised to peers/drivers when binding differs from reachability. JMX is the Java management plane used by nodetool; it has a different security boundary from CQL. A virtual table, such as system_views.settings, exposes node-local runtime information through CQL but is not replicated and ignores consistency level. A guardrail warns about or rejects risky operations/values. Configuration drift means nodes or deployment artifacts no longer share the intended settings. Rolling change means applying a compatible change one node/failure domain at a time while verifying service and rollback between steps.

1. Guardrails are policy at the coordinator—not magic cluster state

Cassandra 5.0 exposes guardrail configuration through nodetool getguardrailsconfig and setguardrailsconfig. The command uses JMX and therefore targets one node. A runtime change can be immediately useful for incident mitigation/testing, but it does not automatically modify every peer or your version-controlled cassandra.yaml. Some settings are runtime-adjustable; others remain startup-only. Never infer dynamism merely because a similarly named nodetool command exists.

Docker · verify versions, topology, JMX-local tooling and native transport
docker exec atlasmart-cass-1 nodetool version -vdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh --versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool statusbinarydocker exec atlasmart-cass-1 nodetool getseeds# Repeat from another node because nodetool/JMX observations are node-local.docker exec atlasmart-cass-2 nodetool status
CQL · create the bounded operations fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_opsWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_ops.config_probe (  probe_id text PRIMARY KEY,  status text,  owner text,  note text,  updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-1','READY','platform','chapter22',toTimestamp(now()));INSERT INTO atlasmart_ops.config_probe (probe_id,status,owner,note,updated_at)VALUES ('probe-2','READY','orders','copy-fixture',toTimestamp(now()));SELECT * FROM atlasmart_ops.config_probe WHERE probe_id='probe-1';
Docker · inspect the current guardrail and runtime settings on all nodes
for n in 1 2 3; do  echo "===== node $n ====="  docker exec atlasmart-cass-$n nodetool getguardrailsconfig allow_filtering_enabled  docker exec atlasmart-cass-$n cqlsh -e "SELECT name,value FROM system_views.settings WHERE name='allow_filtering_enabled';"done# See the supported runtime surface on this exact 5.0.9 build.docker exec atlasmart-cass-1 nodetool help setguardrailsconfig

2. Build a query that requires ALLOW FILTERING

CQL · small diagnostic-only filtering fixture
CREATE TABLE IF NOT EXISTS atlasmart_ops.filter_guardrail_probe (  id text PRIMARY KEY,  status text,  amount int) WITH compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_ops.filter_guardrail_probe (id,status,amount) VALUES ('a','PAID',10);INSERT INTO atlasmart_ops.filter_guardrail_probe (id,status,amount) VALUES ('b','OPEN',20);SELECT * FROM atlasmart_ops.filter_guardrail_probe WHERE status='PAID' ALLOW FILTERING;

The query is deliberately poor Cassandra modeling and small enough to be safe. It exists only to expose the guardrail. Production should normally create a query-first table or appropriate SAI index rather than relying on broad filtering.

3. Create real configuration drift, observe it, then roll

Docker · node-local guardrail change and evidence
# Cassandra 5.0 source/tests expose this exact runtime setter; default is true.docker exec atlasmart-cass-1 nodetool setguardrailsconfig allow_filtering_enabled falsedocker exec atlasmart-cass-1 nodetool getguardrailsconfig allow_filtering_enabled# Coordinator node 1 should reject the ALLOW FILTERING query.docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_ops.filter_guardrail_probe WHERE status='PAID' ALLOW FILTERING;"# Node 2 still has its old runtime value, demonstrating drift.docker exec atlasmart-cass-2 nodetool getguardrailsconfig allow_filtering_enableddocker exec atlasmart-cass-2 cqlsh -e "SELECT * FROM atlasmart_ops.filter_guardrail_probe WHERE status='PAID' ALLOW FILTERING;"

The precise error text is version-dependent; the important evidence is that the same CQL is accepted/rejected based on the coordinator's runtime guardrail. That is why a change ticket must name scope and verify every node.

Docker · roll the same value to peers and verify
docker exec atlasmart-cass-2 nodetool setguardrailsconfig allow_filtering_enabled falsedocker exec atlasmart-cass-3 nodetool setguardrailsconfig allow_filtering_enabled falsefor n in 1 2 3; do docker exec atlasmart-cass-$n nodetool getguardrailsconfig allow_filtering_enabled; done

4. Restore first, then discuss persistence

Docker · mandatory cleanup/reset
# Restore the established lab default on every node.for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool setguardrailsconfig allow_filtering_enabled truedonefor n in 1 2 3; do docker exec atlasmart-cass-$n nodetool getguardrailsconfig allow_filtering_enabled; done# Optional cleanup table only; keep config_probe for later lessons.docker exec atlasmart-cass-1 cqlsh -e "DROP TABLE IF EXISTS atlasmart_ops.filter_guardrail_probe;"
Wrong approach: “nodetool returned success, therefore the cluster baseline is changed.”

A runtime JMX change is operational state. If the policy must survive restart/replacement, update the version-controlled deployment configuration using the exact release's supported YAML/management surface, roll it under a maintenance/change plan, and verify after restart. Managed Cassandra services may expose different policy APIs and may prohibit direct JMX/YAML access.

5. Rolling-change acceptance gates

Gate Evidence before next node
compatibility release docs/deprecations/upgrade notes checked
availability RF/CL still tolerates one-node work in the current failure domain
runtime target setting matches intent on changed node
service native transport + representative read/write succeed
observability errors/latency/GC/compaction not regressing
rollback previous value/config artifact ready and tested

Check your understanding

  1. Why did the same ALLOW FILTERING query behave differently on node 1 and node 2?
  2. Does setguardrailsconfig prove the change will survive restart?
  3. Why roll one node/failure domain at a time?
  4. Can every cassandra.yaml option be changed dynamically?
  5. What is the safe final state of this lab?
Review the answers

1. The guardrail was changed through node-local JMX on only one coordinator, creating runtime configuration drift.

2. No. Treat it as runtime state unless the deployment/persistent configuration is also updated and verified.

3. It limits blast radius and lets you verify compatibility, availability and rollback before propagating the change.

4. No. Some have supported runtime controls; many are startup-only. Check the exact release documentation/tool surface.

5. allow_filtering_enabled is restored to true on all three nodes and the temporary filtering table is removed.

Production judgment

Operational configuration is part of the database design. Record the Cassandra patch/JDK/container or package, topology and failure domains, RF/consistency levels, compaction, repair and backup schedules, TTL/gc_grace_seconds, SAI/vector dependencies, disk/memory/network/JVM limits, auth/TLS/JMX boundaries, native transport exposure, guardrails, driver retry/idempotency behavior, observability endpoints, and managed-service overrides. A syntactically valid setting can still be wrong for workload cardinality, retention, tail latency, disk headroom or failure behavior. Likewise, a healthy-looking nodetool status does not prove repair freshness, query SLOs, disk latency, compaction health, application correctness or backup recoverability.

Make every change with an owner, compatibility check, blast-radius estimate, pre/post evidence, rollback path and version-controlled intent. Avoid copying tuning values from another cluster without workload evidence. Security-sensitive files, passwords and keystores belong in secret-management systems, not source control. Lesson 5 converts these practices into a version-controlled baseline: intended settings, rendered/running evidence, secrets boundaries, drift checks, change metadata and rollback become one auditable operational artifact.

Summary and next bridge

Dynamic configuration is powerful precisely because it can create invisible node-local state. A safe change is scoped, rolled, verified, restored or persisted, and linked to a baseline. The final lesson builds that baseline and a repeatable drift-review workflow.

Authoritative references

Re-check these version-sensitive references when regenerating the course. Tool availability, defaults, guardrails and configuration names evolve across Cassandra releases and managed services.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.