Chapter 27 · Observability, Performance, Capacity, Guardrails, Upgrades, Multi-DC Resilience, and Capstone

Rolling Upgrades, Schema Agreement, Driver Compatibility, Multi-DC Failover / Repair, and Disaster-Game-Day Planning

Canary a rolling patch upgrade, verify schema/driver compatibility, and rehearse full multi-DC failover, repair and recovery.

Advanced · Production capstone170–260 minutesUpgrade + multi-DC game dayApache Cassandra 5.0.9 · Java 17 · Java Driver 4.19.3 · cqlsh/nodetool/cassandra-stress · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart schedules a Cassandra patch upgrade and a disaster-recovery game day in the same quarter. The unsafe plan is “replace images one by one, then if dc1 fails point clients at dc2.” That omits schema agreement, compatibility modes, driver/local-DC behavior, consistency semantics, repair after isolation, rollback windows and an evidence timeline.

01

Run a free local rolling patch upgrade from Cassandra 5.0.8 to 5.0.9 with canary/rollback evidence.

02

Explain storage_compatibility_mode for a major 4.x→5.0 transition without pretending a patch drill reproduces a major upgrade.

03

Verify schema agreement and Java Driver 4.19.3 connectivity/routing before, during and after the roll.

04

Build a six-node two-DC RF=3 topology and fail one DC while comparing LOCAL_QUORUM/EACH_QUORUM/global behavior.

05

Repair/converge the recovered DC and close the game day against SLO/RPO/RTO acceptance criteria.

Chapter 27 capstone lab baseline

The mandatory single-datacenter labs use Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the official image, cqlsh/nodetool from that same image, and Apache Cassandra Java Driver 4.19.3 where client behavior matters. Use Docker network atlasmart-cassandra-capstone, cluster atlasmart-capstone, nodes atlasmart-cap-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy, replication factor (RF) 3, and application reads/writes at LOCAL_QUORUM unless the lesson deliberately changes consistency level (CL). Tables explicitly use UnifiedCompactionStrategy (UCS), default_time_to_live=0, and gc_grace_seconds=864000.

For repeatability, mandatory Lessons 1–3 keep authentication/client TLS/internode TLS disabled only inside this isolated Docker network; no Cassandra or JMX port is published on the host, and JMX remains local-only inside containers. Chapter 26 remains the production security baseline. Lesson 4 creates separate disposable upgrade/multi-DC clusters, and Lesson 5 makes the security state an explicit capstone acceptance gate. A full RF=3 two-DC game day requires six Cassandra containers and roughly 10–14 GiB of available RAM depending on container limits/JVM ergonomics; the lesson also provides a reduced four-node RF=2 simulation for constrained laptops and labels the semantic difference.

Exact token values, latencies, GC pauses, SSTable counts, compaction/repair bytes, disk throughput, network rates, failure-detection timing, guardrail messages and benchmark throughput are runtime evidence. The lesson never treats example numbers as results from your machine. Before every benchmark/failure drill capture host CPU/RAM/disk, Docker limits, Cassandra/Java/driver versions, topology, RF/CL, schema/compaction, dataset shape, concurrency, warmup and security state.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms used across the capstone

A coordinator is the Cassandra node handling one client request; a replica stores a copy of the requested partition according to the keyspace replication strategy. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, and virtual nodes (vnodes) give each physical node multiple token ranges. A datacenter (DC) and rack model failure/locality domains. RF (replication factor) is the number of replicas per DC configured by NetworkTopologyStrategy; a CL (consistency level) controls how many appropriately scoped replica responses are required.

An SSTable (Sorted String Table) is an immutable on-disk data-file set. Compaction rewrites SSTables to merge versions/tombstones and manage read/space amplification. Repair is anti-entropy comparison/streaming between replicas. heap is Java-managed memory; off-heap covers native/direct structures outside the Java heap; the operating-system page cache is another memory consumer. GC (garbage collection) reclaims Java heap objects and can introduce pauses.

A histogram is a distribution, not an average. p50/p95/p99 are percentiles: p99 means 99% of observed values are at or below that value. tail latency is the high-percentile response time users often feel during saturation/failure. throughput is completed work per time; error rate is failed work divided by attempted work. A Service Level Objective (SLO) is a target such as p99 latency/availability; RPO (Recovery Point Objective) is tolerated data-loss time; RTO (Recovery Time Objective) is tolerated service-recovery time.

JMX (Java Management Extensions) exposes node-local Cassandra metrics/management operations; nodetool is itself a JMX client. An exported metric is forwarded to an external time-series system; Cassandra metrics are node-local until an operator aggregates them. A guardrail warns or rejects dangerous schema/query/operational patterns. A rolling upgrade changes one node at a time while the cluster remains available. schema agreement means nodes report the same current schema version. A driver is the client library implementing native protocol, topology discovery, load balancing, timeouts, retries, idempotency and speculative execution. SAI expands to Storage-Attached Indexing; a vector index supports approximate nearest-neighbor retrieval and has separate memory/disk/build/recall costs.

1. Separate patch upgrade mechanics from major-version compatibility

The mandatory upgrade drill uses cassandra:5.0.8 → cassandra:5.0.9 so it is reproducible on one laptop and teaches rolling mechanics without introducing every 4.1→5.0 format/protocol constraint. A real major upgrade must follow the official upgrade path and release notes. Cassandra 5.0's storage_compatibility_mode deliberately gates format/features: CASSANDRA_4 remains compatible with 4.x formats/features, UPGRADING monitors cluster versions until all nodes are ready, and NONE enables full 5.0 behavior after stability. Changing that mode is a separate rolling-restart stage, not a cosmetic setting.

Upgrade gate Evidence Rollback implication
supported source/target path official release/upgrade notes unsupported jumps have no safe rollback promise
driver compatibility 4.19.3 test suite + native-protocol connections client rollout may need independent rollback
schema agreement describecluster/system schema versions do not layer schema changes on disagreement
one-node canary p99/errors/logs/GC/stream/compaction rollback image/config before more nodes
storage compatibility major upgrade CASSANDRA_4→UPGRADING→NONE stages format/feature activation narrows downgrade options
SSTable upgrade automatic or deliberate rewrite after binary stability disk/I/O/headroom event; snapshots need version care

2. Rolling patch lab: 5.0.8 → 5.0.9

bash · create a disposable three-node 5.0.8 upgrade cluster
docker network create atlasmart-upgrade 2>/dev/null || truefor n in 1 2 3; do docker volume create atlasmart-up-$n-data; donedocker run -d --name atlasmart-up-1 --hostname atlasmart-up-1 --network atlasmart-upgrade \  -e CASSANDRA_CLUSTER_NAME=atlasmart-upgrade -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-up-1-data:/var/lib/cassandra cassandra:5.0.8docker exec atlasmart-up-1 nodetool statusfor n in 2 3; do  docker run -d --name atlasmart-up-$n --hostname atlasmart-up-$n --network atlasmart-upgrade \    -e CASSANDRA_CLUSTER_NAME=atlasmart-upgrade -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack$n \    -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \    -e CASSANDRA_SEEDS=atlasmart-up-1 -v atlasmart-up-$n-data:/var/lib/cassandra cassandra:5.0.8donedocker exec atlasmart-up-1 nodetool statusdocker exec atlasmart-up-1 nodetool describecluster
bash · replace one container at a time while preserving its data volume
# Before each node: application SLO green, no abnormal repair/stream/compaction backlog,# schema agreement healthy, backup/rollback image+config available.for n in 1 2 3; do  rack="rack$n"  docker stop atlasmart-up-$n  docker rm atlasmart-up-$n  docker run -d --name atlasmart-up-$n --hostname atlasmart-up-$n --network atlasmart-upgrade \    -e CASSANDRA_CLUSTER_NAME=atlasmart-upgrade -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=$rack \    -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \    -e CASSANDRA_SEEDS=atlasmart-up-1 -v atlasmart-up-$n-data:/var/lib/cassandra cassandra:5.0.9  # Wait for UN and verify before moving to the next node.  docker exec atlasmart-up-$n nodetool version  docker exec atlasmart-up-$n nodetool status  docker exec atlasmart-up-$n nodetool describecluster  docker logs --since 5m atlasmart-up-$n 2>&1 | tail -100done

Do not run this loop blindly in production—the loop is compact lab notation. The real runbook has a human/automation gate between nodes based on error rate, p99, dropped messages, GC, streaming/compaction, schema agreement, driver errors and rollback criteria. Patch downgrade behavior must still be checked against release notes/data-format changes; preserving the old image alone is not a rollback guarantee.

Deliberately unsafe upgrade: all nodes at once, schema migrations during the roll, no driver canary.

This converts a rolling compatibility problem into a full-cluster outage/data-format gamble. Freeze nonessential schema/topology changes, canary one node and client path, verify mixed-version behavior, preserve tested backup/config/image rollback, and only advance when metrics/logs/schema/driver SLOs are green.

3. Driver compatibility is an executable gate

xml · Maven dependency for the current Apache Java Driver
<dependency>  <groupId>org.apache.cassandra</groupId>  <artifactId>java-driver-core</artifactId>  <version>4.19.3</version></dependency>
java · minimal compatibility/sanity probe
try (CqlSession session = CqlSession.builder()    .addContactPoint(new InetSocketAddress("atlasmart-up-1", 9042))    .withLocalDatacenter("dc1")    .build()) {  Row r = session.execute("SELECT release_version, schema_version FROM system.local").one();  System.out.println(r.getString("release_version") + " schema=" + r.getUuid("schema_version"));  session.getMetadata().getNodes().values().forEach(n ->      System.out.println(n.getEndPoint()+" dc="+n.getDatacenter()+" state="+n.getState()));}

A successful connection/select is only a smoke test. Run the application's prepared statements, paging, data types, LWT, SAI/vector paths, timeouts/retries/idempotency/speculation and error mapping against mixed and final versions. Record the Java/JVM/driver/native-protocol versions actually observed.

4. Build a two-DC game-day cluster

The full local scenario uses six nodes: three racks in dc1 and three in dc2, with RF=3 in each DC. That gives meaningful LOCAL_QUORUM in either DC. If your machine cannot run six Cassandra JVMs, use four nodes with two per DC and RF=2 per DC; label the reduced failure math because local quorum becomes 2/2 and one node loss removes local availability.

bash · six-node two-DC topology (full semantic lab)
docker network create atlasmart-multidc 2>/dev/null || truefor n in 1 2 3 4 5 6; do docker volume create atlasmart-mdc-$n-data; done# dc1 seed/first nodedocker run -d --name atlasmart-mdc-1 --hostname atlasmart-mdc-1 --network atlasmart-multidc \  -e CASSANDRA_CLUSTER_NAME=atlasmart-multidc -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-mdc-1-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-mdc-1 nodetool statusfor n in 2 3; do  docker run -d --name atlasmart-mdc-$n --hostname atlasmart-mdc-$n --network atlasmart-multidc \    -e CASSANDRA_CLUSTER_NAME=atlasmart-multidc -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack$n \    -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \    -e CASSANDRA_SEEDS=atlasmart-mdc-1 -v atlasmart-mdc-$n-data:/var/lib/cassandra cassandra:5.0.9donefor n in 4 5 6; do  rack=$((n-3))  docker run -d --name atlasmart-mdc-$n --hostname atlasmart-mdc-$n --network atlasmart-multidc \    -e CASSANDRA_CLUSTER_NAME=atlasmart-multidc -e CASSANDRA_DC=dc2 -e CASSANDRA_RACK=rack$rack \    -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \    -e CASSANDRA_SEEDS=atlasmart-mdc-1 -v atlasmart-mdc-$n-data:/var/lib/cassandra cassandra:5.0.9donedocker exec atlasmart-mdc-1 nodetool status
CQL · multi-DC RF and a deterministic failover row
CREATE KEYSPACE IF NOT EXISTS atlasmart_drWITH replication = {'class':'NetworkTopologyStrategy','dc1':3,'dc2':3};CREATE TABLE IF NOT EXISTS atlasmart_dr.checkout_state (  cart_id text PRIMARY KEY, state text, total decimal) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_dr.checkout_state (cart_id,state,total) VALUES ('cart-27','READY',42.00);

5. Fail dc1, continue locally in dc2, then converge

bash · reversible whole-DC failure injection and evidence timeline
# T0: capture status, application p99/errors and last successful backup/repair time.docker exec atlasmart-mdc-4 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';"# T1: simulate complete dc1 loss.docker pause atlasmart-mdc-1 atlasmart-mdc-2 atlasmart-mdc-3docker exec atlasmart-mdc-4 nodetool status# LOCAL_QUORUM coordinated in dc2 can still use dc2's 2-of-3 majority.docker exec atlasmart-mdc-4 cqlsh -e \  "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_dr.checkout_state SET state='PAID' WHERE cart_id='cart-27'; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';"# EACH_QUORUM requires a quorum in every replicated DC and should be unavailable.docker exec atlasmart-mdc-4 cqlsh -e \  "CONSISTENCY EACH_QUORUM; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';" || true# T2: recover dc1.docker unpause atlasmart-mdc-1 atlasmart-mdc-2 atlasmart-mdc-3docker exec atlasmart-mdc-4 nodetool status# T3: anti-entropy convergence. Scope/order should follow your production repair plan.docker exec atlasmart-mdc-1 nodetool repair --full atlasmart_dr checkout_statefor n in 1 2 3 4 5 6; do  docker exec atlasmart-mdc-$n cqlsh -e \    "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_dr.checkout_state WHERE cart_id='cart-27';"done

The game-day timeline should include failure detection, client local-DC reroute, error/p99 spike, first successful dc2 request, dc1 recovery, repair start/end, replica convergence and application normalization. RPO depends on which writes were acknowledged/replicated before isolation and your backup/PITR evidence; RTO ends only when the application is safely serving the intended region/DC, not when Cassandra nodes restart.

6. Acceptance and cleanup

Gate Pass condition
upgrade one node at a time; mixed/final versions healthy; schema agreement; p99/errors within SLO
driver 4.19.3 application query suite passes mixed/final versions and local-DC routing is correct
dc failure dc2 LOCAL_QUORUM meets defined SLO while dc1 is unavailable
consistency EACH_QUORUM/global choices fail or succeed according to documented RF/DC math
recovery dc1 returns UN; repair completes per runbook; replicas/application state converge
RPO/RTO measured from explicit timestamps/business checkpoints, not theoretical defaults
rollback source image/config/backup/client rollback boundaries are documented and tested
bash · remove only disposable upgrade/game-day clusters
docker rm -f atlasmart-up-1 atlasmart-up-2 atlasmart-up-3 2>/dev/null || truedocker network rm atlasmart-upgrade 2>/dev/null || truefor n in 1 2 3; do docker volume rm atlasmart-up-$n-data 2>/dev/null || true; donedocker rm -f atlasmart-mdc-1 atlasmart-mdc-2 atlasmart-mdc-3 atlasmart-mdc-4 atlasmart-mdc-5 atlasmart-mdc-6 2>/dev/null || truedocker network rm atlasmart-multidc 2>/dev/null || truefor n in 1 2 3 4 5 6; do docker volume rm atlasmart-mdc-$n-data 2>/dev/null || true; done

Check your understanding

  1. Why does a 5.0.8→5.0.9 drill not prove a 4.1→5.0 major upgrade?
  2. What does schema agreement prove during a roll?
  3. Why can dc2 LOCAL_QUORUM survive total dc1 loss with RF3 per DC?
  4. What must happen after the isolated DC returns?
  5. When does disaster RTO end?
Review the answers

1. A major upgrade introduces additional compatibility/format/feature gates such as storage_compatibility_mode and must follow the major upgrade guide.

2. That nodes report the same current schema; it does not prove application/driver/data-path compatibility.

3. It needs a majority of local dc2 replicas only—2 of 3—whereas EACH_QUORUM needs a majority in every replicated DC.

4. Verify topology/health and run the documented repair/convergence process before declaring normal operation.

5. When the intended application service is verified and safely serving, not merely when Cassandra processes are running.

Production judgment

Do not promote one local run into a universal Cassandra tuning rule. Record workload fit and non-goals; partition cardinality/rows/bytes and retention; read/write mix and p50/p95/p99/max latency; RF/CL and coordinator/replica failure behavior; JVM heap/GC, off-heap/page-cache, disk capacity/latency/IOPS/throughput and network bandwidth/packet loss; SSTable/read amplification, compaction backlog, tombstones and repair state; SAI/vector build/query/write and recall costs where used; authentication/authorization/TLS/JMX/secret/tenant boundaries; driver local-DC routing, timeout/retry/idempotency/speculation behavior; metrics/log/tracing coverage; backup RPO/RTO/restore proof; and the operator skill/runbook needed to perform repair, topology, upgrade and recovery safely.

Managed Cassandra services may hide disks, JMX, repair, backup, upgrade sequencing or metric names, and their quotas/cost model can change capacity decisions. Translate the same evidence questions into provider-native signals; do not assume the provider removes application data-model, driver, consistency, SLO, security, migration or rollback responsibility. Lesson 5 combines every prior chapter into one production-defense acceptance run: model, deploy, load-test, fail, repair, restore, secure, tune and defend the architecture with evidence.

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Capstone: Model, Deploy, Load-Test, Fail, Repair, Recover, Secure, Tune, and Defend a Production Cassandra Cluster.

Authoritative references

Re-check these version-sensitive sources before a real upgrade, capacity commitment, security change or game day.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.