Chapter 27 · Observability, Performance, Capacity, Guardrails, Upgrades, Multi-DC Resilience, and Capstone
Guardrails and Performance Controls: Prevent Dangerous Schema / Query Patterns Before They Become Incidents
Use Cassandra guardrails and layered performance controls to reject dangerous patterns before they become production incidents.
Learning outcomes
AtlasMart's incident channel fills with
ALLOW FILTERING warnings and oversized
batch/partition alerts. A responder proposes disabling the
guardrails because “they are creating noise.” The warnings
disappear, but the dangerous queries continue and p99 gets
worse. Guardrails are policy boundaries and early-warning
signals, not performance fixes.
Explain warning versus failure guardrails and their node-local/runtime configuration behavior.
Use getguardrailsconfig/setguardrailsconfig to inspect and safely change a reversible guardrail in the disposable cluster.
Connect guardrails to partition size, query filtering, batch size, collection/vector/schema limits and operational risk.
Distinguish guardrails from driver/application timeouts, rate limits, circuit breakers and capacity controls.
Build a warn→investigate→redesign→verify workflow instead of silencing the signal.
The mandatory single-datacenter labs use Apache Cassandra
5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the
official image, cqlsh/nodetool from
that same image, and Apache Cassandra Java Driver
4.19.3 where client behavior matters. Use Docker
network atlasmart-cassandra-capstone, cluster
atlasmart-capstone, nodes
atlasmart-cap-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and application reads/writes at
LOCAL_QUORUM unless the lesson deliberately
changes consistency level (CL). Tables explicitly use
UnifiedCompactionStrategy (UCS),
default_time_to_live=0, and
gc_grace_seconds=864000.
For repeatability, mandatory Lessons 1–3 keep authentication/client TLS/internode TLS disabled only inside this isolated Docker network; no Cassandra or JMX port is published on the host, and JMX remains local-only inside containers. Chapter 26 remains the production security baseline. Lesson 4 creates separate disposable upgrade/multi-DC clusters, and Lesson 5 makes the security state an explicit capstone acceptance gate. A full RF=3 two-DC game day requires six Cassandra containers and roughly 10–14 GiB of available RAM depending on container limits/JVM ergonomics; the lesson also provides a reduced four-node RF=2 simulation for constrained laptops and labels the semantic difference.
Exact token values, latencies, GC pauses, SSTable counts, compaction/repair bytes, disk throughput, network rates, failure-detection timing, guardrail messages and benchmark throughput are runtime evidence. The lesson never treats example numbers as results from your machine. Before every benchmark/failure drill capture host CPU/RAM/disk, Docker limits, Cassandra/Java/driver versions, topology, RF/CL, schema/compaction, dataset shape, concurrency, warmup and security state.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms used across the capstone
A coordinator is the Cassandra node handling
one client request; a replica stores a copy of
the requested partition according to the keyspace replication
strategy. A partition is the rows sharing a
partition key; the partitioner hashes that key to a
token, and virtual nodes
(vnodes) give each physical node multiple token
ranges. A datacenter (DC) and
rack model failure/locality domains.
RF (replication factor) is the number of
replicas per DC configured by
NetworkTopologyStrategy; a
CL (consistency level) controls how many
appropriately scoped replica responses are required.
An SSTable (Sorted String Table) is an immutable on-disk data-file set. Compaction rewrites SSTables to merge versions/tombstones and manage read/space amplification. Repair is anti-entropy comparison/streaming between replicas. heap is Java-managed memory; off-heap covers native/direct structures outside the Java heap; the operating-system page cache is another memory consumer. GC (garbage collection) reclaims Java heap objects and can introduce pauses.
A histogram is a distribution, not an average. p50/p95/p99 are percentiles: p99 means 99% of observed values are at or below that value. tail latency is the high-percentile response time users often feel during saturation/failure. throughput is completed work per time; error rate is failed work divided by attempted work. A Service Level Objective (SLO) is a target such as p99 latency/availability; RPO (Recovery Point Objective) is tolerated data-loss time; RTO (Recovery Time Objective) is tolerated service-recovery time.
JMX (Java Management Extensions) exposes
node-local Cassandra metrics/management operations;
nodetool is itself a JMX client. An
exported metric is forwarded to an external
time-series system; Cassandra metrics are node-local until an
operator aggregates them. A guardrail warns or
rejects dangerous schema/query/operational patterns. A
rolling upgrade changes one node at a time
while the cluster remains available.
schema agreement means nodes report the same
current schema version. A driver is the client
library implementing native protocol, topology discovery, load
balancing, timeouts, retries, idempotency and speculative
execution. SAI expands to Storage-Attached
Indexing; a vector index supports approximate
nearest-neighbor retrieval and has separate
memory/disk/build/recall costs.
1. Guardrails are prevention/feedback, not a substitute for modeling
Cassandra 5.0 includes guardrails for dangerous or surprising
schema/query patterns. Some are boolean feature gates (for
example allow_filtering_enabled or drop/truncate
controls); others have warning/failure thresholds. A warning
allows work but emits evidence. A failure rejects the operation
before it executes. This changes user-visible behavior, so
guardrail settings are part of the application contract and
rolling configuration baseline.
| Guardrail family | Risk it addresses | What a trigger means | What it does not mean |
|---|---|---|---|
| ALLOW FILTERING | unbounded/unpredictable reads | query uses a prohibited filtered access path | the correct threshold can rescue a bad model |
| partition size | very large partitions | estimated/observed size crosses policy | all smaller partitions meet SLO |
| batch size/partitions | coordinator/batchlog/mutation pressure | batch shape exceeds configured policy | batches are a bulk-load tool |
| collection size/items | unbounded collection mutations/reads | collection breaches policy | collection is safe merely below threshold |
| tables/columns/indexes | schema sprawl/resource overhead | schema object count breaches policy | current schema/query fit is good |
| vector dimensions | index/memory/query cost | vector shape exceeds policy | lower dimensions guarantee recall/latency |
| DROP/TRUNCATE | destructive DDL | dangerous operation is disabled | backup/restore/authorization is unnecessary |
Guardrails are evaluated by the node coordinating the operation.
Runtime changes through
nodetool setguardrailsconfig are per-node; changing
only one node creates coordinator-dependent behavior until
rolled consistently. Persistence across restarts depends on the
corresponding deployment configuration—not merely the live JMX
state.
2. Inspect the running guardrail baseline on every node
docker network inspect atlasmart-cassandra-capstone >/dev/null 2>&1 || docker network create atlasmart-cassandra-capstonefor n in 1 2 3; do docker volume create atlasmart-cap-$n-data; donedocker inspect atlasmart-cap-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-1 --hostname atlasmart-cap-1 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-cap-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-2 --hostname atlasmart-cap-2 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cap-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-3 --hostname atlasmart-cap-3 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cap-1 nodetool versiondocker exec atlasmart-cap-1 java -versiondocker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-1 --format '{{json .HostConfig.PortBindings}}'
CREATE KEYSPACE IF NOT EXISTS atlasmart_capstoneWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_customer_month ( customer_id text, order_month date, order_time timestamp, order_id uuid, region text, status text, total decimal, payload text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_region_day ( region text, order_day date, bucket tinyint, order_time timestamp, order_id uuid, customer_id text, status text, total decimal, PRIMARY KEY ((region,order_day,bucket),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;
for n in 1 2 3; do echo "=== atlasmart-cap-$n ===" docker exec atlasmart-cap-$n nodetool getguardrailsconfig --expand > "guardrails-node$n.before.txt" docker exec atlasmart-cap-$n nodetool getguardrailsconfig -- allow_filtering_enabled docker exec atlasmart-cap-$n nodetool getguardrailsconfig -- drop_truncate_table_enableddone
The exact list/defaults are Cassandra-version dependent. Store the full expanded baseline under configuration management, but avoid treating undocumented/internal names as stable across major versions. Compare the running values on all nodes, not just the file you intended to deploy.
3. Reproduce a dangerous query, then enforce the policy
-- This table's primary key does not support a global status lookup.-- Without an SAI index or purpose-built query table, this is a bad access path.SELECT customer_id,order_month,order_time,status,totalFROM atlasmart_capstone.orders_by_customer_monthWHERE status='PAID'ALLOW FILTERING;
for n in 1 2 3; do docker exec atlasmart-cap-$n nodetool setguardrailsconfig -- allow_filtering_enabled false docker exec atlasmart-cap-$n nodetool getguardrailsconfig -- allow_filtering_enableddone# This should now be rejected regardless of which node is coordinator.for n in 1 2 3; do docker exec atlasmart-cap-$n cqlsh -e \ "SELECT customer_id FROM atlasmart_capstone.orders_by_customer_month WHERE status='PAID' ALLOW FILTERING;" || truedone
The corrected design depends on the product query. If AtlasMart
needs “orders by status” with bounded time/region, create a
purpose-built query table with a bounded partition key or
evaluate Storage-Attached Indexing (SAI) with
selectivity/fanout/index-cost evidence. Removing
ALLOW FILTERING from the CQL without changing the
access path is not a fix.
This removes prevention/visibility while leaving the underlying query/schema/capacity risk. The repair is to preserve or strengthen the guardrail, identify the offending query/owner from logs/audit/driver telemetry, redesign or rate-limit the operation, then prove the corrected workload meets p99/error/resource SLOs.
4. Guardrails and performance controls live at different layers
| Control | Layer | Purpose | Failure if confused |
|---|---|---|---|
| Cassandra guardrail | server coordinator/schema | warn/reject dangerous operations | cannot shape every business workload or traffic burst |
| driver timeout | client | bound client wait time | does not cancel/rollback all server work; timeout can be ambiguous |
| driver retry | client | retry selected failures | unsafe for non-idempotent operations or overload storms |
| speculative execution | client/server read mechanisms | reduce tail by duplicate work | can multiply load during saturation |
| rate/concurrency limit | application/gateway | bound offered load | does not repair bad partition model |
| circuit breaker/load shed | application/platform | protect dependencies during failure | can reduce availability if thresholds are wrong |
| capacity/headroom | architecture | survive steady + maintenance/failure load | cannot prevent one pathological query |
For example, increasing a driver timeout can convert a timeout into a 5-second user wait while the same unbounded query consumes more cluster resources. A guardrail rejection can be healthier if the access pattern is not production-safe.
5. Build a guardrail event workflow
1. CAPTURE: node/DC/coordinator, role/service, query/schema operation, guardrail name/value.2. CLASSIFY: warning vs failure; one request vs repeated deployment/workload pattern.3. CORRELATE: p95/p99/errors, partition/SSTable/tombstone metrics, GC/disk/network, deploy/version.4. CONTAIN: rate-limit/disable offending feature/client if needed; DO NOT globally disable the guardrail.5. REPAIR DESIGN: query table/bucketing/SAI/vector dimension/batch/application workflow change.6. TEST: representative load + N-1/compaction/repair; prove SLO and resource headroom.7. ROLL: canary then all nodes; verify runtime settings are identical.8. PERSIST: configuration-as-code, alert/runbook, owner and rollback.
for n in 1 2 3; do docker exec atlasmart-cap-$n nodetool setguardrailsconfig -- allow_filtering_enabled true docker exec atlasmart-cap-$n nodetool getguardrailsconfig -- allow_filtering_enableddonediff -u guardrails-node1.before.txt guardrails-node2.before.txt || truediff -u guardrails-node1.before.txt guardrails-node3.before.txt || true
The lesson resets allow_filtering_enabled to the
course default only because this isolated lab began there. A
production baseline may intentionally disable it. Never copy the
lab's boolean into production without reviewing your deployed
policy.
6. Verification checklist
- All three nodes have a captured expanded guardrail baseline.
- The same bad query is accepted/rejected according to the deliberate before/after policy, with the exact server message captured.
- Coordinator-dependent drift is demonstrated or explicitly avoided by changing all nodes.
- The corrected query/model path is specified; the guardrail itself is not called the performance fix.
- Driver timeout/retry/speculation/rate limiting are documented as separate controls.
- The runtime value is restored/persisted according to the intended configuration baseline.
Check your understanding
- What is the difference between a warn and fail guardrail?
- Why is changing one node dangerous?
- Does disabling ALLOW FILTERING fix the application query?
- Why is raising a driver timeout not equivalent to fixing a guardrail event?
- What should happen after a guardrail incident is fixed?
Review the answers
1. A warn allows the operation while emitting evidence; a fail guardrail rejects it before execution.
2. Requests coordinated by different nodes can see different guardrail behavior until the setting is rolled consistently.
3. No. It prevents that unsafe access path; the application still needs a query-first table or suitable index/design.
4. It only changes how long the client waits and can increase resource occupancy; it does not change the dangerous query/storage mechanism.
5. Representative load/failure verification plus persistent config/runbook/owner/rollback updates.
Production judgment
Do not promote one local run into a universal Cassandra tuning rule. Record workload fit and non-goals; partition cardinality/rows/bytes and retention; read/write mix and p50/p95/p99/max latency; RF/CL and coordinator/replica failure behavior; JVM heap/GC, off-heap/page-cache, disk capacity/latency/IOPS/throughput and network bandwidth/packet loss; SSTable/read amplification, compaction backlog, tombstones and repair state; SAI/vector build/query/write and recall costs where used; authentication/authorization/TLS/JMX/secret/tenant boundaries; driver local-DC routing, timeout/retry/idempotency/speculation behavior; metrics/log/tracing coverage; backup RPO/RTO/restore proof; and the operator skill/runbook needed to perform repair, topology, upgrade and recovery safely.
Managed Cassandra services may hide disks, JMX, repair, backup, upgrade sequencing or metric names, and their quotas/cost model can change capacity decisions. Translate the same evidence questions into provider-native signals; do not assume the provider removes application data-model, driver, consistency, SLO, security, migration or rollback responsibility. Lesson 4 now changes the software/topology itself: rolling patch upgrades, major-version compatibility concepts, schema agreement, driver compatibility and a two-DC failover/repair game day.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Rolling Upgrades, Schema Agreement, Driver Compatibility, Multi-DC Failover/Repair, and Disaster-Game-Day Planning.
Authoritative references
Re-check these version-sensitive sources before a real upgrade, capacity commitment, security change or game day.
- Apache Cassandra 5.0 release/download baseline
- Monitoring metrics
- Troubleshooting with nodetool histograms
- nodetool tablestats
- nodetool tablehistograms
- nodetool proxyhistograms
- cassandra-stress user mode
- cassandra.yaml including guardrails/storage compatibility
- nodetool setguardrailsconfig
- Repair operations
- Backups
- Security
- Java Driver core documentation