Chapter 27 · Observability, Performance, Capacity, Guardrails, Upgrades, Multi-DC Resilience, and Capstone
Rolling Upgrades, Schema Agreement, Driver Compatibility, Multi-DC Failover / Repair, and Disaster-Game-Day Planning
Canary a rolling patch upgrade, verify schema/driver compatibility, and rehearse full multi-DC failover, repair and recovery.
Learning outcomes
AtlasMart schedules a Cassandra patch upgrade and a disaster-recovery game day in the same quarter. The unsafe plan is “replace images one by one, then if dc1 fails point clients at dc2.” That omits schema agreement, compatibility modes, driver/local-DC behavior, consistency semantics, repair after isolation, rollback windows and an evidence timeline.
Run a free local rolling patch upgrade from Cassandra 5.0.8 to 5.0.9 with canary/rollback evidence.
Explain storage_compatibility_mode for a major 4.x→5.0 transition without pretending a patch drill reproduces a major upgrade.
Verify schema agreement and Java Driver 4.19.3 connectivity/routing before, during and after the roll.
Build a six-node two-DC RF=3 topology and fail one DC while comparing LOCAL_QUORUM/EACH_QUORUM/global behavior.
Repair/converge the recovered DC and close the game day against SLO/RPO/RTO acceptance criteria.
The mandatory single-datacenter labs use Apache Cassandra
5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the
official image, cqlsh/nodetool from
that same image, and Apache Cassandra Java Driver
4.19.3 where client behavior matters. Use Docker
network atlasmart-cassandra-capstone, cluster
atlasmart-capstone, nodes
atlasmart-cap-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and application reads/writes at
LOCAL_QUORUM unless the lesson deliberately
changes consistency level (CL). Tables explicitly use
UnifiedCompactionStrategy (UCS),
default_time_to_live=0, and
gc_grace_seconds=864000.
For repeatability, mandatory Lessons 1–3 keep authentication/client TLS/internode TLS disabled only inside this isolated Docker network; no Cassandra or JMX port is published on the host, and JMX remains local-only inside containers. Chapter 26 remains the production security baseline. Lesson 4 creates separate disposable upgrade/multi-DC clusters, and Lesson 5 makes the security state an explicit capstone acceptance gate. A full RF=3 two-DC game day requires six Cassandra containers and roughly 10–14 GiB of available RAM depending on container limits/JVM ergonomics; the lesson also provides a reduced four-node RF=2 simulation for constrained laptops and labels the semantic difference.
Exact token values, latencies, GC pauses, SSTable counts, compaction/repair bytes, disk throughput, network rates, failure-detection timing, guardrail messages and benchmark throughput are runtime evidence. The lesson never treats example numbers as results from your machine. Before every benchmark/failure drill capture host CPU/RAM/disk, Docker limits, Cassandra/Java/driver versions, topology, RF/CL, schema/compaction, dataset shape, concurrency, warmup and security state.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms used across the capstone
A coordinator is the Cassandra node handling
one client request; a replica stores a copy of
the requested partition according to the keyspace replication
strategy. A partition is the rows sharing a
partition key; the partitioner hashes that key to a
token, and virtual nodes
(vnodes) give each physical node multiple token
ranges. A datacenter (DC) and
rack model failure/locality domains.
RF (replication factor) is the number of
replicas per DC configured by
NetworkTopologyStrategy; a
CL (consistency level) controls how many
appropriately scoped replica responses are required.
An SSTable (Sorted String Table) is an immutable on-disk data-file set. Compaction rewrites SSTables to merge versions/tombstones and manage read/space amplification. Repair is anti-entropy comparison/streaming between replicas. heap is Java-managed memory; off-heap covers native/direct structures outside the Java heap; the operating-system page cache is another memory consumer. GC (garbage collection) reclaims Java heap objects and can introduce pauses.
A histogram is a distribution, not an average. p50/p95/p99 are percentiles: p99 means 99% of observed values are at or below that value. tail latency is the high-percentile response time users often feel during saturation/failure. throughput is completed work per time; error rate is failed work divided by attempted work. A Service Level Objective (SLO) is a target such as p99 latency/availability; RPO (Recovery Point Objective) is tolerated data-loss time; RTO (Recovery Time Objective) is tolerated service-recovery time.
JMX (Java Management Extensions) exposes
node-local Cassandra metrics/management operations;
nodetool is itself a JMX client. An
exported metric is forwarded to an external
time-series system; Cassandra metrics are node-local until an
operator aggregates them. A guardrail warns or
rejects dangerous schema/query/operational patterns. A
rolling upgrade changes one node at a time
while the cluster remains available.
schema agreement means nodes report the same
current schema version. A driver is the client
library implementing native protocol, topology discovery, load
balancing, timeouts, retries, idempotency and speculative
execution. SAI expands to Storage-Attached
Indexing; a vector index supports approximate
nearest-neighbor retrieval and has separate
memory/disk/build/recall costs.
1. Separate patch upgrade mechanics from major-version compatibility
The mandatory upgrade drill uses cassandra:5.0.8 →
cassandra:5.0.9 so it is reproducible on one laptop
and teaches rolling mechanics without introducing every 4.1→5.0
format/protocol constraint. A real major upgrade must follow the
official upgrade path and release notes. Cassandra 5.0's
storage_compatibility_mode deliberately gates
format/features: CASSANDRA_4 remains compatible
with 4.x formats/features, UPGRADING monitors
cluster versions until all nodes are ready, and
NONE enables full 5.0 behavior after stability.
Changing that mode is a separate rolling-restart stage, not a
cosmetic setting.
| Upgrade gate | Evidence | Rollback implication |
|---|---|---|
| supported source/target path | official release/upgrade notes | unsupported jumps have no safe rollback promise |
| driver compatibility | 4.19.3 test suite + native-protocol connections | client rollout may need independent rollback |
| schema agreement | describecluster/system schema versions | do not layer schema changes on disagreement |
| one-node canary | p99/errors/logs/GC/stream/compaction | rollback image/config before more nodes |
| storage compatibility major upgrade | CASSANDRA_4→UPGRADING→NONE stages | format/feature activation narrows downgrade options |
| SSTable upgrade | automatic or deliberate rewrite after binary stability | disk/I/O/headroom event; snapshots need version care |
2. Rolling patch lab: 5.0.8 → 5.0.9
docker network create atlasmart-upgrade 2>/dev/null || truefor n in 1 2 3; do docker volume create atlasmart-up-$n-data; donedocker run -d --name atlasmart-up-1 --hostname atlasmart-up-1 --network atlasmart-upgrade \ -e CASSANDRA_CLUSTER_NAME=atlasmart-upgrade -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-up-1-data:/var/lib/cassandra cassandra:5.0.8docker exec atlasmart-up-1 nodetool statusfor n in 2 3; do docker run -d --name atlasmart-up-$n --hostname atlasmart-up-$n --network atlasmart-upgrade \ -e CASSANDRA_CLUSTER_NAME=atlasmart-upgrade -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack$n \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-up-1 -v atlasmart-up-$n-data:/var/lib/cassandra cassandra:5.0.8donedocker exec atlasmart-up-1 nodetool statusdocker exec atlasmart-up-1 nodetool describecluster
# Before each node: application SLO green, no abnormal repair/stream/compaction backlog,# schema agreement healthy, backup/rollback image+config available.for n in 1 2 3; do rack="rack$n" docker stop atlasmart-up-$n docker rm atlasmart-up-$n docker run -d --name atlasmart-up-$n --hostname atlasmart-up-$n --network atlasmart-upgrade \ -e CASSANDRA_CLUSTER_NAME=atlasmart-upgrade -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=$rack \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-up-1 -v atlasmart-up-$n-data:/var/lib/cassandra cassandra:5.0.9 # Wait for UN and verify before moving to the next node. docker exec atlasmart-up-$n nodetool version docker exec atlasmart-up-$n nodetool status docker exec atlasmart-up-$n nodetool describecluster docker logs --since 5m atlasmart-up-$n 2>&1 | tail -100done
Do not run this loop blindly in production—the loop is compact lab notation. The real runbook has a human/automation gate between nodes based on error rate, p99, dropped messages, GC, streaming/compaction, schema agreement, driver errors and rollback criteria. Patch downgrade behavior must still be checked against release notes/data-format changes; preserving the old image alone is not a rollback guarantee.
This converts a rolling compatibility problem into a full-cluster outage/data-format gamble. Freeze nonessential schema/topology changes, canary one node and client path, verify mixed-version behavior, preserve tested backup/config/image rollback, and only advance when metrics/logs/schema/driver SLOs are green.
3. Driver compatibility is an executable gate
<dependency> <groupId>org.apache.cassandra</groupId> <artifactId>java-driver-core</artifactId> <version>4.19.3</version></dependency>
try (CqlSession session = CqlSession.builder() .addContactPoint(new InetSocketAddress("atlasmart-up-1", 9042)) .withLocalDatacenter("dc1") .build()) { Row r = session.execute("SELECT release_version, schema_version FROM system.local").one(); System.out.println(r.getString("release_version") + " schema=" + r.getUuid("schema_version")); session.getMetadata().getNodes().values().forEach(n -> System.out.println(n.getEndPoint()+" dc="+n.getDatacenter()+" state="+n.getState()));}
A successful connection/select is only a smoke test. Run the application's prepared statements, paging, data types, LWT, SAI/vector paths, timeouts/retries/idempotency/speculation and error mapping against mixed and final versions. Record the Java/JVM/driver/native-protocol versions actually observed.
4. Build a two-DC game-day cluster
The full local scenario uses six nodes: three racks in
dc1 and three in dc2, with RF=3 in
each DC. That gives meaningful LOCAL_QUORUM in
either DC. If your machine cannot run six Cassandra JVMs, use
four nodes with two per DC and RF=2 per DC; label the reduced
failure math because local quorum becomes 2/2 and one node loss
removes local availability.
docker network create atlasmart-multidc 2>/dev/null || truefor n in 1 2 3 4 5 6; do docker volume create atlasmart-mdc-$n-data; done# dc1 seed/first nodedocker run -d --name atlasmart-mdc-1 --hostname atlasmart-mdc-1 --network atlasmart-multidc \ -e CASSANDRA_CLUSTER_NAME=atlasmart-multidc -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-mdc-1-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-mdc-1 nodetool statusfor n in 2 3; do docker run -d --name atlasmart-mdc-$n --hostname atlasmart-mdc-$n --network atlasmart-multidc \ -e CASSANDRA_CLUSTER_NAME=atlasmart-multidc -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack$n \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-mdc-1 -v atlasmart-mdc-$n-data:/var/lib/cassandra cassandra:5.0.9donefor n in 4 5 6; do rack=$((n-3)) docker run -d --name atlasmart-mdc-$n --hostname atlasmart-mdc-$n --network atlasmart-multidc \ -e CASSANDRA_CLUSTER_NAME=atlasmart-multidc -e CASSANDRA_DC=dc2 -e CASSANDRA_RACK=rack$rack \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-mdc-1 -v atlasmart-mdc-$n-data:/var/lib/cassandra cassandra:5.0.9donedocker exec atlasmart-mdc-1 nodetool status
CREATE KEYSPACE IF NOT EXISTS atlasmart_drWITH replication = {'class':'NetworkTopologyStrategy','dc1':3,'dc2':3};CREATE TABLE IF NOT EXISTS atlasmart_dr.checkout_state ( cart_id text PRIMARY KEY, state text, total decimal) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_dr.checkout_state (cart_id,state,total) VALUES ('cart-27','READY',42.00);
5. Fail dc1, continue locally in dc2, then converge
# T0: capture status, application p99/errors and last successful backup/repair time.docker exec atlasmart-mdc-4 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';"# T1: simulate complete dc1 loss.docker pause atlasmart-mdc-1 atlasmart-mdc-2 atlasmart-mdc-3docker exec atlasmart-mdc-4 nodetool status# LOCAL_QUORUM coordinated in dc2 can still use dc2's 2-of-3 majority.docker exec atlasmart-mdc-4 cqlsh -e \ "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_dr.checkout_state SET state='PAID' WHERE cart_id='cart-27'; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';"# EACH_QUORUM requires a quorum in every replicated DC and should be unavailable.docker exec atlasmart-mdc-4 cqlsh -e \ "CONSISTENCY EACH_QUORUM; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';" || true# T2: recover dc1.docker unpause atlasmart-mdc-1 atlasmart-mdc-2 atlasmart-mdc-3docker exec atlasmart-mdc-4 nodetool status# T3: anti-entropy convergence. Scope/order should follow your production repair plan.docker exec atlasmart-mdc-1 nodetool repair --full atlasmart_dr checkout_statefor n in 1 2 3 4 5 6; do docker exec atlasmart-mdc-$n cqlsh -e \ "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';"done
The game-day timeline should include failure detection, client local-DC reroute, error/p99 spike, first successful dc2 request, dc1 recovery, repair start/end, replica convergence and application normalization. RPO depends on which writes were acknowledged/replicated before isolation and your backup/PITR evidence; RTO ends only when the application is safely serving the intended region/DC, not when Cassandra nodes restart.
6. Acceptance and cleanup
| Gate | Pass condition |
|---|---|
| upgrade | one node at a time; mixed/final versions healthy; schema agreement; p99/errors within SLO |
| driver | 4.19.3 application query suite passes mixed/final versions and local-DC routing is correct |
| dc failure | dc2 LOCAL_QUORUM meets defined SLO while dc1 is unavailable |
| consistency | EACH_QUORUM/global choices fail or succeed according to documented RF/DC math |
| recovery | dc1 returns UN; repair completes per runbook; replicas/application state converge |
| RPO/RTO | measured from explicit timestamps/business checkpoints, not theoretical defaults |
| rollback | source image/config/backup/client rollback boundaries are documented and tested |
docker rm -f atlasmart-up-1 atlasmart-up-2 atlasmart-up-3 2>/dev/null || truedocker network rm atlasmart-upgrade 2>/dev/null || truefor n in 1 2 3; do docker volume rm atlasmart-up-$n-data 2>/dev/null || true; donedocker rm -f atlasmart-mdc-1 atlasmart-mdc-2 atlasmart-mdc-3 atlasmart-mdc-4 atlasmart-mdc-5 atlasmart-mdc-6 2>/dev/null || truedocker network rm atlasmart-multidc 2>/dev/null || truefor n in 1 2 3 4 5 6; do docker volume rm atlasmart-mdc-$n-data 2>/dev/null || true; done
Check your understanding
- Why does a 5.0.8→5.0.9 drill not prove a 4.1→5.0 major upgrade?
- What does schema agreement prove during a roll?
- Why can dc2 LOCAL_QUORUM survive total dc1 loss with RF3 per DC?
- What must happen after the isolated DC returns?
- When does disaster RTO end?
Review the answers
1. A major upgrade introduces additional compatibility/format/feature gates such as storage_compatibility_mode and must follow the major upgrade guide.
2. That nodes report the same current schema; it does not prove application/driver/data-path compatibility.
3. It needs a majority of local dc2 replicas only—2 of 3—whereas EACH_QUORUM needs a majority in every replicated DC.
4. Verify topology/health and run the documented repair/convergence process before declaring normal operation.
5. When the intended application service is verified and safely serving, not merely when Cassandra processes are running.
Production judgment
Do not promote one local run into a universal Cassandra tuning rule. Record workload fit and non-goals; partition cardinality/rows/bytes and retention; read/write mix and p50/p95/p99/max latency; RF/CL and coordinator/replica failure behavior; JVM heap/GC, off-heap/page-cache, disk capacity/latency/IOPS/throughput and network bandwidth/packet loss; SSTable/read amplification, compaction backlog, tombstones and repair state; SAI/vector build/query/write and recall costs where used; authentication/authorization/TLS/JMX/secret/tenant boundaries; driver local-DC routing, timeout/retry/idempotency/speculation behavior; metrics/log/tracing coverage; backup RPO/RTO/restore proof; and the operator skill/runbook needed to perform repair, topology, upgrade and recovery safely.
Managed Cassandra services may hide disks, JMX, repair, backup, upgrade sequencing or metric names, and their quotas/cost model can change capacity decisions. Translate the same evidence questions into provider-native signals; do not assume the provider removes application data-model, driver, consistency, SLO, security, migration or rollback responsibility. Lesson 5 combines every prior chapter into one production-defense acceptance run: model, deploy, load-test, fail, repair, restore, secure, tune and defend the architecture with evidence.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Capstone: Model, Deploy, Load-Test, Fail, Repair, Recover, Secure, Tune, and Defend a Production Cassandra Cluster.
Authoritative references
Re-check these version-sensitive sources before a real upgrade, capacity commitment, security change or game day.
- Apache Cassandra 5.0 release/download baseline
- Monitoring metrics
- Troubleshooting with nodetool histograms
- nodetool tablestats
- nodetool tablehistograms
- nodetool proxyhistograms
- cassandra-stress user mode
- cassandra.yaml including guardrails/storage compatibility
- nodetool setguardrailsconfig
- Repair operations
- Backups
- Security
- Java Driver core documentation
- Production recommendations
- SSTable upgrade tool
- Dynamo replication/consistency semantics