Chapter 27 · Observability, Performance, Capacity, Guardrails, Upgrades, Multi-DC Resilience, and Capstone
Metrics, Logs, Tracing, nodetool tablestats / tablehistograms / proxyhistograms, and Alerting Signals
Build an evidence hierarchy from coordinator/table histograms, logs, tracing, JMX-backed metrics and representative workload distributions.
Learning outcomes
AtlasMart receives a customer-facing alert: checkout reads look “fine” at 4 ms average latency, but support reports intermittent 200–500 ms pauses and a few timeouts. A single average cannot tell whether the problem lives at the coordinator, local table read path, one hot partition, GC, compaction, repair, disk or network. This lesson builds the observability hierarchy before anyone tunes a knob.
Separate coordinator latency from local table latency and explain which nodetool histogram answers which question.
Use tablestats, tablehistograms, proxyhistograms, logs, tracing, JMX-backed metrics and host/container signals together.
Run a representative user-mode workload rather than a uniform tiny-key microbenchmark and report distributions, throughput and errors.
Define alert signals with context/sustained windows instead of one magic Cassandra threshold.
Build an evidence bundle that later capacity, guardrail, upgrade and disaster lessons can reuse.
The mandatory single-datacenter labs use Apache Cassandra
5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the
official image, cqlsh/nodetool from
that same image, and Apache Cassandra Java Driver
4.19.3 where client behavior matters. Use Docker
network atlasmart-cassandra-capstone, cluster
atlasmart-capstone, nodes
atlasmart-cap-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and application reads/writes at
LOCAL_QUORUM unless the lesson deliberately
changes consistency level (CL). Tables explicitly use
UnifiedCompactionStrategy (UCS),
default_time_to_live=0, and
gc_grace_seconds=864000.
For repeatability, mandatory Lessons 1–3 keep authentication/client TLS/internode TLS disabled only inside this isolated Docker network; no Cassandra or JMX port is published on the host, and JMX remains local-only inside containers. Chapter 26 remains the production security baseline. Lesson 4 creates separate disposable upgrade/multi-DC clusters, and Lesson 5 makes the security state an explicit capstone acceptance gate. A full RF=3 two-DC game day requires six Cassandra containers and roughly 10–14 GiB of available RAM depending on container limits/JVM ergonomics; the lesson also provides a reduced four-node RF=2 simulation for constrained laptops and labels the semantic difference.
Exact token values, latencies, GC pauses, SSTable counts, compaction/repair bytes, disk throughput, network rates, failure-detection timing, guardrail messages and benchmark throughput are runtime evidence. The lesson never treats example numbers as results from your machine. Before every benchmark/failure drill capture host CPU/RAM/disk, Docker limits, Cassandra/Java/driver versions, topology, RF/CL, schema/compaction, dataset shape, concurrency, warmup and security state.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms used across the capstone
A coordinator is the Cassandra node handling
one client request; a replica stores a copy of
the requested partition according to the keyspace replication
strategy. A partition is the rows sharing a
partition key; the partitioner hashes that key to a
token, and virtual nodes
(vnodes) give each physical node multiple token
ranges. A datacenter (DC) and
rack model failure/locality domains.
RF (replication factor) is the number of
replicas per DC configured by
NetworkTopologyStrategy; a
CL (consistency level) controls how many
appropriately scoped replica responses are required.
An SSTable (Sorted String Table) is an immutable on-disk data-file set. Compaction rewrites SSTables to merge versions/tombstones and manage read/space amplification. Repair is anti-entropy comparison/streaming between replicas. heap is Java-managed memory; off-heap covers native/direct structures outside the Java heap; the operating-system page cache is another memory consumer. GC (garbage collection) reclaims Java heap objects and can introduce pauses.
A histogram is a distribution, not an average. p50/p95/p99 are percentiles: p99 means 99% of observed values are at or below that value. tail latency is the high-percentile response time users often feel during saturation/failure. throughput is completed work per time; error rate is failed work divided by attempted work. A Service Level Objective (SLO) is a target such as p99 latency/availability; RPO (Recovery Point Objective) is tolerated data-loss time; RTO (Recovery Time Objective) is tolerated service-recovery time.
JMX (Java Management Extensions) exposes
node-local Cassandra metrics/management operations;
nodetool is itself a JMX client. An
exported metric is forwarded to an external
time-series system; Cassandra metrics are node-local until an
operator aggregates them. A guardrail warns or
rejects dangerous schema/query/operational patterns. A
rolling upgrade changes one node at a time
while the cluster remains available.
schema agreement means nodes report the same
current schema version. A driver is the client
library implementing native protocol, topology discovery, load
balancing, timeouts, retries, idempotency and speculative
execution. SAI expands to Storage-Attached
Indexing; a vector index supports approximate
nearest-neighbor retrieval and has separate
memory/disk/build/recall costs.
1. Start with an observation stack, not a dashboard screenshot
Cassandra exposes many metrics per node.
nodetool obtains many of them over local Java
Management Extensions (JMX). Because metrics are node-local, an
external monitoring system must aggregate per-node time series
if you want a cluster/DC view. Aggregation must preserve
dimensions such as node, DC, rack, keyspace/table and operation;
summing latency percentiles across nodes is mathematically
wrong.
| Evidence | Scope | Best question | Important non-proof |
|---|---|---|---|
| proxyhistograms | coordinator node | what distribution did client-facing coordinator operations see? | does not isolate local storage time |
| tablehistograms | one table on one node | local read/write distribution, SSTables/read, partition/cell size | does not include all coordinator/network waiting |
| tablestats | one/more tables on one node | SSTable, Bloom, tombstone, read/write, repaired-state counters | cumulative counters need rate/delta context |
| TRACING | one request | which coordinator/replicas/stages contributed time? | tracing adds overhead and is not continuous monitoring |
| system.log/debug.log | node events | GC, dropped messages, compaction/repair/security/guardrail clues | absence of a log line is not health proof |
| JMX/exported metrics | node → external TSDB | time correlation across throughput/errors/GC/disk/network | external aggregation/config can drop or mislabel series |
| host/container metrics | OS/container | CPU, RSS, page cache, disk queue, network saturation | high utilization alone is not causality |
The first boundary case is coordinator versus replica-local
time. A high proxyhistograms p99 with modest
tablehistograms local read p99 can point toward
network/replica coordination, overload, cross-DC routing, GC on
coordinators or client/request behavior. High local table p99
plus high SSTables-per-read/tombstones/partition bytes can point
toward the storage path. Both can be high simultaneously.
2. Create the AtlasMart schema and a representative stress profile
docker network inspect atlasmart-cassandra-capstone >/dev/null 2>&1 || docker network create atlasmart-cassandra-capstonefor n in 1 2 3; do docker volume create atlasmart-cap-$n-data; donedocker inspect atlasmart-cap-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-1 --hostname atlasmart-cap-1 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-cap-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-2 --hostname atlasmart-cap-2 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cap-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-3 --hostname atlasmart-cap-3 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cap-1 nodetool versiondocker exec atlasmart-cap-1 java -versiondocker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-1 --format '{{json .HostConfig.PortBindings}}'
CREATE KEYSPACE IF NOT EXISTS atlasmart_capstoneWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_customer_month ( customer_id text, order_month date, order_time timestamp, order_id uuid, region text, status text, total decimal, payload text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_region_day ( region text, order_day date, bucket tinyint, order_time timestamp, order_id uuid, customer_id text, status text, total decimal, PRIMARY KEY ((region,order_day,bucket),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;
specname: atlasmart_orderskeyspace: atlasmart_capstonekeyspace_definition: | CREATE KEYSPACE IF NOT EXISTS atlasmart_capstone WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};table: orders_by_customer_monthtable_definition: | CREATE TABLE IF NOT EXISTS orders_by_customer_month ( customer_id text, order_month date, order_time timestamp, order_id uuid, region text, status text, total decimal, payload text, PRIMARY KEY ((customer_id,order_month),order_time,order_id) ) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'};columnspec: - name: customer_id size: fixed(12) population: gaussian(1..50000,10000) - name: order_month cluster: fixed(1) - name: order_time cluster: gaussian(10..500,80) - name: payload size: gaussian(40..1200,180)insert: partitions: fixed(1) select: fixed(1)/500 batchtype: UNLOGGEDqueries: latest_orders: cql: SELECT order_time,order_id,status,total FROM orders_by_customer_month WHERE customer_id=? AND order_month=? LIMIT 20 fields: samerow
# Save the YAML above as ./atlasmart-stress.yaml on the host.docker cp ./atlasmart-stress.yaml atlasmart-cap-1:/tmp/atlasmart-stress.yaml# Warmup is intentionally separate from the measured run.docker exec atlasmart-cap-1 cassandra-stress user \ profile=/tmp/atlasmart-stress.yaml duration=30s \ "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \ -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 \ -rate threads=8# Measured short lab run. Capture the complete output; do not copy its example# numbers into a production capacity document.docker exec atlasmart-cap-1 cassandra-stress user \ profile=/tmp/atlasmart-stress.yaml duration=2m \ "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \ -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 \ -rate threads=16 | tee /tmp/atlasmart-stress-run.txt
User mode exercises the actual query table and can generate clustered rows with non-uniform population/size distributions. The exact distribution functions are workload models, not claims about AtlasMart production. Export the run metadata beside the result: workload mix, duration, concurrency, partition population, payload size, CL, coordinator nodes, warmup, Docker CPU/RAM limits and host storage.
It suppresses the hot/large-partition tail and can make a cluster appear safe until production traffic hits skew, wide partitions, compaction/repair or failure. Repair the benchmark by matching representative cardinality/partition size/read-write mix, include hot-key/large-value cases, warmup explicitly, report p50/p95/p99/max plus throughput/errors, and repeat under compaction/repair and N-1 conditions.
3. Capture the three nodetool views before tracing
docker exec atlasmart-cap-1 nodetool proxyhistogramsdocker exec atlasmart-cap-1 nodetool tablehistograms atlasmart_capstone orders_by_customer_monthdocker exec atlasmart-cap-1 nodetool tablestats -H atlasmart_capstone.orders_by_customer_month# Capture all nodes because metrics are node-local.for n in 1 2 3; do echo "=== node $n ===" docker exec atlasmart-cap-$n nodetool tablehistograms atlasmart_capstone orders_by_customer_month docker exec atlasmart-cap-$n nodetool tablestats -F json atlasmart_capstone.orders_by_customer_month > "tablestats-node$n.json"done
proxyhistograms exposes coordinator
read/write/range/CAS/view latency distributions.
tablehistograms exposes local table percentiles
including SSTables touched per read, local read/write latency,
partition size and cell count. tablestats adds
table counters/state such as live disk space, SSTables,
Bloom-filter and tombstone statistics,
compaction-related/repaired information exposed by the running
version. The human-readable and JSON forms serve different uses:
operators read the former; automation should parse the latter
instead of scraping aligned text.
4. Add logs, tracing and JVM/host signals only when they answer a hypothesis
CONSISTENCY LOCAL_QUORUM;TRACING ON;SELECT order_time,order_id,status,totalFROM atlasmart_capstone.orders_by_customer_monthWHERE customer_id='REPLACE_WITH_A_GENERATED_KEY' AND order_month='2026-09-01'LIMIT 20;TRACING OFF;
docker logs --since 10m atlasmart-cap-1 2>&1 | tail -200docker exec atlasmart-cap-1 sh -lc 'tail -200 /var/log/cassandra/system.log 2>/dev/null || true'docker exec atlasmart-cap-1 nodetool gcstatsdocker exec atlasmart-cap-1 nodetool infodocker exec atlasmart-cap-1 nodetool tpstatsdocker exec atlasmart-cap-1 nodetool compactionstatsdocker exec atlasmart-cap-1 nodetool netstatsdocker stats --no-stream atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3
Tracing can show coordinator/replica messaging, local read
stages and mismatch/repair events for that query, but it adds
work and the exact event text/timing is version/topology
dependent. Use it for sampled diagnosis, not every request.
gcstats/info give JVM-level clues,
while container/OS tooling is needed for
CPU/RSS/page-cache/disk/network. If Docker Desktop hides
physical disk queue/latency, record that observability
limitation instead of inventing it.
For long-lived production monitoring, export Cassandra JMX
metrics to a time-series system using a reviewed JMX/reporting
integration and aggregate by node/DC/rack/table. The mandatory
lab remains free/local and uses nodetool as the JMX
client; a Prometheus/JMX exporter or managed observability
product is optional, not required to learn the metric semantics.
5. Alert on symptoms + saturation + context
| Signal family | Useful alert shape | Correlate with | Avoid |
|---|---|---|---|
| client/coordinator SLO | sustained p99 + timeout/failure rate | proxyhistograms, driver metrics, CL/topology | average latency only |
| local read path | local p99 + SSTables/read/tombstones | tablehistograms/tablestats/compaction | one high sample |
| write path | coordinator p99/errors + pending mutations/dropped work | tpstats, commitlog/memtable/compaction | throughput alone |
| JVM | GC pause/time + heap pressure | allocation/load, off-heap/page cache | heap % without GC context |
| disk | free-space + latency/queue/throughput | compaction/repair/snapshot/streaming | % used without temp-headroom model |
| network | throughput/loss/errors + stream traffic | netstats, repair/bootstrap, cross-DC | bandwidth % alone |
| repair | age/failure/pending repair + unrepaired bytes | gc_grace/delete rate | “last command succeeded” only |
| guardrails/security | warn/fail/auth/audit events | query identity/deployment change | disabling the signal |
Alerts should be actionable: identify the affected SLO/resource, a sustained window or burn-rate, scope (node/DC/table) and a runbook link. Exact thresholds must come from workload SLOs and saturation tests. The same p99 may be acceptable for an offline report and unacceptable for checkout.
6. Verification checklist and reset
- All three nodes are UN and version/topology/CL/security assumptions are recorded.
- The workload profile uses the real query table and non-uniform partition/payload distributions.
- You captured p50/p95/p99/max and throughput/errors rather than one average.
-
proxyhistogramsandtablehistogramsare interpreted at their correct coordinator/local scopes. - At least one traced request is correlated with logs/JVM/compaction/network state.
- Any external exporter is labeled optional and its aggregation semantics are documented.
docker exec atlasmart-cap-1 cqlsh -e "TRUNCATE atlasmart_capstone.orders_by_customer_month;"# Keep the cluster/schema for Lesson 2, or remove it deliberately at the end:# docker rm -f atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3# docker volume rm atlasmart-cap-1-data atlasmart-cap-2-data atlasmart-cap-3-data# docker network rm atlasmart-cassandra-capstone
Check your understanding
- Why can proxyhistograms and tablehistograms disagree?
- Why are Cassandra metrics not automatically cluster metrics?
- Why is p99 more useful than average for an intermittent checkout complaint?
- What does tracing prove?
- Why is a uniform tiny-key benchmark dangerous?
Review the answers
1. They measure different scopes: coordinator/client-facing operation time versus local table work on one node.
2. Metrics are emitted per node; an external system must aggregate while preserving meaningful dimensions.
3. A small fraction of very slow requests can dominate user pain while barely moving the average.
4. The stages/endpoints observed for that traced request; it does not prove normal workload distribution or continuous health.
5. It hides partition-size/cardinality skew and may miss compaction/repair/failure costs that dominate production tails.
Production judgment
Do not promote one local run into a universal Cassandra tuning rule. Record workload fit and non-goals; partition cardinality/rows/bytes and retention; read/write mix and p50/p95/p99/max latency; RF/CL and coordinator/replica failure behavior; JVM heap/GC, off-heap/page-cache, disk capacity/latency/IOPS/throughput and network bandwidth/packet loss; SSTable/read amplification, compaction backlog, tombstones and repair state; SAI/vector build/query/write and recall costs where used; authentication/authorization/TLS/JMX/secret/tenant boundaries; driver local-DC routing, timeout/retry/idempotency/speculation behavior; metrics/log/tracing coverage; backup RPO/RTO/restore proof; and the operator skill/runbook needed to perform repair, topology, upgrade and recovery safely.
Managed Cassandra services may hide disks, JMX, repair, backup, upgrade sequencing or metric names, and their quotas/cost model can change capacity decisions. Translate the same evidence questions into provider-native signals; do not assume the provider removes application data-model, driver, consistency, SLO, security, migration or rollback responsibility. Lesson 2 converts these measured distributions and resource signals into partition/SSTable/compaction/repair/memory/disk/network capacity and explicit N-1 headroom.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Capacity Planning for Partitions, SSTables, Compaction, Repair, Heap/Off-Heap, Disk, Network, and Failure Headroom.
Authoritative references
Re-check these version-sensitive sources before a real upgrade, capacity commitment, security change or game day.
- Apache Cassandra 5.0 release/download baseline
- Monitoring metrics
- Troubleshooting with nodetool histograms
- nodetool tablestats
- nodetool tablehistograms
- nodetool proxyhistograms
- cassandra-stress user mode
- cassandra.yaml including guardrails/storage compatibility
- nodetool setguardrailsconfig
- Repair operations
- Backups
- Security
- Java Driver core documentation