Chapter 27 · Observability, Performance, Capacity, Guardrails, Upgrades, Multi-DC Resilience, and Capstone

Metrics, Logs, Tracing, nodetool tablestats / tablehistograms / proxyhistograms, and Alerting Signals

Build an evidence hierarchy from coordinator/table histograms, logs, tracing, JMX-backed metrics and representative workload distributions.

Advanced · Production capstone120–170 minutesObservability + workload labApache Cassandra 5.0.9 · Java 17 · Java Driver 4.19.3 · cqlsh/nodetool/cassandra-stress · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart receives a customer-facing alert: checkout reads look “fine” at 4 ms average latency, but support reports intermittent 200–500 ms pauses and a few timeouts. A single average cannot tell whether the problem lives at the coordinator, local table read path, one hot partition, GC, compaction, repair, disk or network. This lesson builds the observability hierarchy before anyone tunes a knob.

01

Separate coordinator latency from local table latency and explain which nodetool histogram answers which question.

02

Use tablestats, tablehistograms, proxyhistograms, logs, tracing, JMX-backed metrics and host/container signals together.

03

Run a representative user-mode workload rather than a uniform tiny-key microbenchmark and report distributions, throughput and errors.

04

Define alert signals with context/sustained windows instead of one magic Cassandra threshold.

05

Build an evidence bundle that later capacity, guardrail, upgrade and disaster lessons can reuse.

Chapter 27 capstone lab baseline

The mandatory single-datacenter labs use Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the official image, cqlsh/nodetool from that same image, and Apache Cassandra Java Driver 4.19.3 where client behavior matters. Use Docker network atlasmart-cassandra-capstone, cluster atlasmart-capstone, nodes atlasmart-cap-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy, replication factor (RF) 3, and application reads/writes at LOCAL_QUORUM unless the lesson deliberately changes consistency level (CL). Tables explicitly use UnifiedCompactionStrategy (UCS), default_time_to_live=0, and gc_grace_seconds=864000.

For repeatability, mandatory Lessons 1–3 keep authentication/client TLS/internode TLS disabled only inside this isolated Docker network; no Cassandra or JMX port is published on the host, and JMX remains local-only inside containers. Chapter 26 remains the production security baseline. Lesson 4 creates separate disposable upgrade/multi-DC clusters, and Lesson 5 makes the security state an explicit capstone acceptance gate. A full RF=3 two-DC game day requires six Cassandra containers and roughly 10–14 GiB of available RAM depending on container limits/JVM ergonomics; the lesson also provides a reduced four-node RF=2 simulation for constrained laptops and labels the semantic difference.

Exact token values, latencies, GC pauses, SSTable counts, compaction/repair bytes, disk throughput, network rates, failure-detection timing, guardrail messages and benchmark throughput are runtime evidence. The lesson never treats example numbers as results from your machine. Before every benchmark/failure drill capture host CPU/RAM/disk, Docker limits, Cassandra/Java/driver versions, topology, RF/CL, schema/compaction, dataset shape, concurrency, warmup and security state.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms used across the capstone

A coordinator is the Cassandra node handling one client request; a replica stores a copy of the requested partition according to the keyspace replication strategy. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, and virtual nodes (vnodes) give each physical node multiple token ranges. A datacenter (DC) and rack model failure/locality domains. RF (replication factor) is the number of replicas per DC configured by NetworkTopologyStrategy; a CL (consistency level) controls how many appropriately scoped replica responses are required.

An SSTable (Sorted String Table) is an immutable on-disk data-file set. Compaction rewrites SSTables to merge versions/tombstones and manage read/space amplification. Repair is anti-entropy comparison/streaming between replicas. heap is Java-managed memory; off-heap covers native/direct structures outside the Java heap; the operating-system page cache is another memory consumer. GC (garbage collection) reclaims Java heap objects and can introduce pauses.

A histogram is a distribution, not an average. p50/p95/p99 are percentiles: p99 means 99% of observed values are at or below that value. tail latency is the high-percentile response time users often feel during saturation/failure. throughput is completed work per time; error rate is failed work divided by attempted work. A Service Level Objective (SLO) is a target such as p99 latency/availability; RPO (Recovery Point Objective) is tolerated data-loss time; RTO (Recovery Time Objective) is tolerated service-recovery time.

JMX (Java Management Extensions) exposes node-local Cassandra metrics/management operations; nodetool is itself a JMX client. An exported metric is forwarded to an external time-series system; Cassandra metrics are node-local until an operator aggregates them. A guardrail warns or rejects dangerous schema/query/operational patterns. A rolling upgrade changes one node at a time while the cluster remains available. schema agreement means nodes report the same current schema version. A driver is the client library implementing native protocol, topology discovery, load balancing, timeouts, retries, idempotency and speculative execution. SAI expands to Storage-Attached Indexing; a vector index supports approximate nearest-neighbor retrieval and has separate memory/disk/build/recall costs.

1. Start with an observation stack, not a dashboard screenshot

Cassandra exposes many metrics per node. nodetool obtains many of them over local Java Management Extensions (JMX). Because metrics are node-local, an external monitoring system must aggregate per-node time series if you want a cluster/DC view. Aggregation must preserve dimensions such as node, DC, rack, keyspace/table and operation; summing latency percentiles across nodes is mathematically wrong.

Evidence Scope Best question Important non-proof
proxyhistograms coordinator node what distribution did client-facing coordinator operations see? does not isolate local storage time
tablehistograms one table on one node local read/write distribution, SSTables/read, partition/cell size does not include all coordinator/network waiting
tablestats one/more tables on one node SSTable, Bloom, tombstone, read/write, repaired-state counters cumulative counters need rate/delta context
TRACING one request which coordinator/replicas/stages contributed time? tracing adds overhead and is not continuous monitoring
system.log/debug.log node events GC, dropped messages, compaction/repair/security/guardrail clues absence of a log line is not health proof
JMX/exported metrics node → external TSDB time correlation across throughput/errors/GC/disk/network external aggregation/config can drop or mislabel series
host/container metrics OS/container CPU, RSS, page cache, disk queue, network saturation high utilization alone is not causality

The first boundary case is coordinator versus replica-local time. A high proxyhistograms p99 with modest tablehistograms local read p99 can point toward network/replica coordination, overload, cross-DC routing, GC on coordinators or client/request behavior. High local table p99 plus high SSTables-per-read/tombstones/partition bytes can point toward the storage path. Both can be high simultaneously.

2. Create the AtlasMart schema and a representative stress profile

bash · create or verify the three-node capstone cluster
docker network inspect atlasmart-cassandra-capstone >/dev/null 2>&1 || docker network create atlasmart-cassandra-capstonefor n in 1 2 3; do docker volume create atlasmart-cap-$n-data; donedocker inspect atlasmart-cap-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-1 --hostname atlasmart-cap-1 \  --network atlasmart-cassandra-capstone \  -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-cap-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-2 --hostname atlasmart-cap-2 \  --network atlasmart-cassandra-capstone \  -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cap-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-3 --hostname atlasmart-cap-3 \  --network atlasmart-cassandra-capstone \  -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cap-1 nodetool versiondocker exec atlasmart-cap-1 java -versiondocker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-1 --format '{{json .HostConfig.PortBindings}}'
CQL · AtlasMart query-first capstone schema
CREATE KEYSPACE IF NOT EXISTS atlasmart_capstoneWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_customer_month (    customer_id text,    order_month date,    order_time timestamp,    order_id uuid,    region text,    status text,    total decimal,    payload text,    PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND default_time_to_live = 0  AND gc_grace_seconds = 864000;CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_region_day (    region text,    order_day date,    bucket tinyint,    order_time timestamp,    order_id uuid,    customer_id text,    status text,    total decimal,    PRIMARY KEY ((region,order_day,bucket),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND default_time_to_live = 0  AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;
yaml · cassandra-stress user profile for the real query table
specname: atlasmart_orderskeyspace: atlasmart_capstonekeyspace_definition: |  CREATE KEYSPACE IF NOT EXISTS atlasmart_capstone WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};table: orders_by_customer_monthtable_definition: |  CREATE TABLE IF NOT EXISTS orders_by_customer_month (    customer_id text,    order_month date,    order_time timestamp,    order_id uuid,    region text,    status text,    total decimal,    payload text,    PRIMARY KEY ((customer_id,order_month),order_time,order_id)  ) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)    AND compaction = {'class':'UnifiedCompactionStrategy'};columnspec:  - name: customer_id    size: fixed(12)    population: gaussian(1..50000,10000)  - name: order_month    cluster: fixed(1)  - name: order_time    cluster: gaussian(10..500,80)  - name: payload    size: gaussian(40..1200,180)insert:  partitions: fixed(1)  select: fixed(1)/500  batchtype: UNLOGGEDqueries:  latest_orders:    cql: SELECT order_time,order_id,status,total FROM orders_by_customer_month WHERE customer_id=? AND order_month=? LIMIT 20    fields: samerow
bash · copy the profile and run warmup plus a measured mixed workload
# Save the YAML above as ./atlasmart-stress.yaml on the host.docker cp ./atlasmart-stress.yaml atlasmart-cap-1:/tmp/atlasmart-stress.yaml# Warmup is intentionally separate from the measured run.docker exec atlasmart-cap-1 cassandra-stress user \  profile=/tmp/atlasmart-stress.yaml duration=30s \  "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \  -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 \  -rate threads=8# Measured short lab run. Capture the complete output; do not copy its example# numbers into a production capacity document.docker exec atlasmart-cap-1 cassandra-stress user \  profile=/tmp/atlasmart-stress.yaml duration=2m \  "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \  -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 \  -rate threads=16 | tee /tmp/atlasmart-stress-run.txt

User mode exercises the actual query table and can generate clustered rows with non-uniform population/size distributions. The exact distribution functions are workload models, not claims about AtlasMart production. Export the run metadata beside the result: workload mix, duration, concurrency, partition population, payload size, CL, coordinator nodes, warmup, Docker CPU/RAM limits and host storage.

Deliberately misleading benchmark: uniform tiny keys + average latency only.

It suppresses the hot/large-partition tail and can make a cluster appear safe until production traffic hits skew, wide partitions, compaction/repair or failure. Repair the benchmark by matching representative cardinality/partition size/read-write mix, include hot-key/large-value cases, warmup explicitly, report p50/p95/p99/max plus throughput/errors, and repeat under compaction/repair and N-1 conditions.

3. Capture the three nodetool views before tracing

bash · coordinator, table and storage evidence
docker exec atlasmart-cap-1 nodetool proxyhistogramsdocker exec atlasmart-cap-1 nodetool tablehistograms atlasmart_capstone orders_by_customer_monthdocker exec atlasmart-cap-1 nodetool tablestats -H atlasmart_capstone.orders_by_customer_month# Capture all nodes because metrics are node-local.for n in 1 2 3; do  echo "=== node $n ==="  docker exec atlasmart-cap-$n nodetool tablehistograms atlasmart_capstone orders_by_customer_month  docker exec atlasmart-cap-$n nodetool tablestats -F json atlasmart_capstone.orders_by_customer_month > "tablestats-node$n.json"done

proxyhistograms exposes coordinator read/write/range/CAS/view latency distributions. tablehistograms exposes local table percentiles including SSTables touched per read, local read/write latency, partition size and cell count. tablestats adds table counters/state such as live disk space, SSTables, Bloom-filter and tombstone statistics, compaction-related/repaired information exposed by the running version. The human-readable and JSON forms serve different uses: operators read the former; automation should parse the latter instead of scraping aligned text.

4. Add logs, tracing and JVM/host signals only when they answer a hypothesis

CQL · trace one bounded partition read
CONSISTENCY LOCAL_QUORUM;TRACING ON;SELECT order_time,order_id,status,totalFROM atlasmart_capstone.orders_by_customer_monthWHERE customer_id='REPLACE_WITH_A_GENERATED_KEY'  AND order_month='2026-09-01'LIMIT 20;TRACING OFF;
bash · correlate node logs, GC/JVM and container resources
docker logs --since 10m atlasmart-cap-1 2>&1 | tail -200docker exec atlasmart-cap-1 sh -lc 'tail -200 /var/log/cassandra/system.log 2>/dev/null || true'docker exec atlasmart-cap-1 nodetool gcstatsdocker exec atlasmart-cap-1 nodetool infodocker exec atlasmart-cap-1 nodetool tpstatsdocker exec atlasmart-cap-1 nodetool compactionstatsdocker exec atlasmart-cap-1 nodetool netstatsdocker stats --no-stream atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3

Tracing can show coordinator/replica messaging, local read stages and mismatch/repair events for that query, but it adds work and the exact event text/timing is version/topology dependent. Use it for sampled diagnosis, not every request. gcstats/info give JVM-level clues, while container/OS tooling is needed for CPU/RSS/page-cache/disk/network. If Docker Desktop hides physical disk queue/latency, record that observability limitation instead of inventing it.

For long-lived production monitoring, export Cassandra JMX metrics to a time-series system using a reviewed JMX/reporting integration and aggregate by node/DC/rack/table. The mandatory lab remains free/local and uses nodetool as the JMX client; a Prometheus/JMX exporter or managed observability product is optional, not required to learn the metric semantics.

5. Alert on symptoms + saturation + context

Signal family Useful alert shape Correlate with Avoid
client/coordinator SLO sustained p99 + timeout/failure rate proxyhistograms, driver metrics, CL/topology average latency only
local read path local p99 + SSTables/read/tombstones tablehistograms/tablestats/compaction one high sample
write path coordinator p99/errors + pending mutations/dropped work tpstats, commitlog/memtable/compaction throughput alone
JVM GC pause/time + heap pressure allocation/load, off-heap/page cache heap % without GC context
disk free-space + latency/queue/throughput compaction/repair/snapshot/streaming % used without temp-headroom model
network throughput/loss/errors + stream traffic netstats, repair/bootstrap, cross-DC bandwidth % alone
repair age/failure/pending repair + unrepaired bytes gc_grace/delete rate “last command succeeded” only
guardrails/security warn/fail/auth/audit events query identity/deployment change disabling the signal

Alerts should be actionable: identify the affected SLO/resource, a sustained window or burn-rate, scope (node/DC/table) and a runbook link. Exact thresholds must come from workload SLOs and saturation tests. The same p99 may be acceptable for an offline report and unacceptable for checkout.

6. Verification checklist and reset

  • All three nodes are UN and version/topology/CL/security assumptions are recorded.
  • The workload profile uses the real query table and non-uniform partition/payload distributions.
  • You captured p50/p95/p99/max and throughput/errors rather than one average.
  • proxyhistograms and tablehistograms are interpreted at their correct coordinator/local scopes.
  • At least one traced request is correlated with logs/JVM/compaction/network state.
  • Any external exporter is labeled optional and its aggregation semantics are documented.
bash · reset only the capstone workload data when needed
docker exec atlasmart-cap-1 cqlsh -e "TRUNCATE atlasmart_capstone.orders_by_customer_month;"# Keep the cluster/schema for Lesson 2, or remove it deliberately at the end:# docker rm -f atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3# docker volume rm atlasmart-cap-1-data atlasmart-cap-2-data atlasmart-cap-3-data# docker network rm atlasmart-cassandra-capstone

Check your understanding

  1. Why can proxyhistograms and tablehistograms disagree?
  2. Why are Cassandra metrics not automatically cluster metrics?
  3. Why is p99 more useful than average for an intermittent checkout complaint?
  4. What does tracing prove?
  5. Why is a uniform tiny-key benchmark dangerous?
Review the answers

1. They measure different scopes: coordinator/client-facing operation time versus local table work on one node.

2. Metrics are emitted per node; an external system must aggregate while preserving meaningful dimensions.

3. A small fraction of very slow requests can dominate user pain while barely moving the average.

4. The stages/endpoints observed for that traced request; it does not prove normal workload distribution or continuous health.

5. It hides partition-size/cardinality skew and may miss compaction/repair/failure costs that dominate production tails.

Production judgment

Do not promote one local run into a universal Cassandra tuning rule. Record workload fit and non-goals; partition cardinality/rows/bytes and retention; read/write mix and p50/p95/p99/max latency; RF/CL and coordinator/replica failure behavior; JVM heap/GC, off-heap/page-cache, disk capacity/latency/IOPS/throughput and network bandwidth/packet loss; SSTable/read amplification, compaction backlog, tombstones and repair state; SAI/vector build/query/write and recall costs where used; authentication/authorization/TLS/JMX/secret/tenant boundaries; driver local-DC routing, timeout/retry/idempotency/speculation behavior; metrics/log/tracing coverage; backup RPO/RTO/restore proof; and the operator skill/runbook needed to perform repair, topology, upgrade and recovery safely.

Managed Cassandra services may hide disks, JMX, repair, backup, upgrade sequencing or metric names, and their quotas/cost model can change capacity decisions. Translate the same evidence questions into provider-native signals; do not assume the provider removes application data-model, driver, consistency, SLO, security, migration or rollback responsibility. Lesson 2 converts these measured distributions and resource signals into partition/SSTable/compaction/repair/memory/disk/network capacity and explicit N-1 headroom.

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Capacity Planning for Partitions, SSTables, Compaction, Repair, Heap/Off-Heap, Disk, Network, and Failure Headroom.

Authoritative references

Re-check these version-sensitive sources before a real upgrade, capacity commitment, security change or game day.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.