Chapter 27 · Observability, Performance, Capacity, Guardrails, Upgrades, Multi-DC Resilience, and Capstone
Capstone: Model, Deploy, Load-Test, Fail, Repair, Recover, Secure, Tune, and Defend a Production Cassandra Cluster
Defend the complete AtlasMart Cassandra production design through measurable load, failure, repair, restore, security, tuning and change gates.
Learning outcomes
AtlasMart is ready for production review. The architecture diagram looks polished, but the approval board asks harder questions: Can the schema survive the real partition distribution? Can p99 stay inside the checkout SLO during compaction/repair and one-node loss? Can the team recover a tested backup within RTO, reject dangerous queries, rotate credentials/certificates, survive a DC loss, upgrade without schema/driver breakage, and explain every rollback? The capstone is accepted only when those claims are reproducible.
Turn product queries/retention/invariants into bounded Cassandra tables with an explicit RF/CL/driver contract.
Deploy and load-test representative distributions, report p50/p95/p99/max/throughput/errors and correlate them with Cassandra/JVM/host evidence.
Inject node/DC and bad-query failures, repair/converge, enforce guardrails and verify recovery rather than merely restarting processes.
Demonstrate a cataloged backup restore and security acceptance gates before application cutover.
Produce a production-defense scorecard/runbook with SLO/RPO/RTO, capacity headroom, upgrade/migration/rollback and ownership.
The mandatory single-datacenter labs use Apache Cassandra
5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the
official image, cqlsh/nodetool from
that same image, and Apache Cassandra Java Driver
4.19.3 where client behavior matters. Use Docker
network atlasmart-cassandra-capstone, cluster
atlasmart-capstone, nodes
atlasmart-cap-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and application reads/writes at
LOCAL_QUORUM unless the lesson deliberately
changes consistency level (CL). Tables explicitly use
UnifiedCompactionStrategy (UCS),
default_time_to_live=0, and
gc_grace_seconds=864000.
For repeatability, mandatory Lessons 1–3 keep authentication/client TLS/internode TLS disabled only inside this isolated Docker network; no Cassandra or JMX port is published on the host, and JMX remains local-only inside containers. Chapter 26 remains the production security baseline. Lesson 4 creates separate disposable upgrade/multi-DC clusters, and Lesson 5 makes the security state an explicit capstone acceptance gate. A full RF=3 two-DC game day requires six Cassandra containers and roughly 10–14 GiB of available RAM depending on container limits/JVM ergonomics; the lesson also provides a reduced four-node RF=2 simulation for constrained laptops and labels the semantic difference.
Exact token values, latencies, GC pauses, SSTable counts, compaction/repair bytes, disk throughput, network rates, failure-detection timing, guardrail messages and benchmark throughput are runtime evidence. The lesson never treats example numbers as results from your machine. Before every benchmark/failure drill capture host CPU/RAM/disk, Docker limits, Cassandra/Java/driver versions, topology, RF/CL, schema/compaction, dataset shape, concurrency, warmup and security state.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms used across the capstone
A coordinator is the Cassandra node handling
one client request; a replica stores a copy of
the requested partition according to the keyspace replication
strategy. A partition is the rows sharing a
partition key; the partitioner hashes that key to a
token, and virtual nodes
(vnodes) give each physical node multiple token
ranges. A datacenter (DC) and
rack model failure/locality domains.
RF (replication factor) is the number of
replicas per DC configured by
NetworkTopologyStrategy; a
CL (consistency level) controls how many
appropriately scoped replica responses are required.
An SSTable (Sorted String Table) is an immutable on-disk data-file set. Compaction rewrites SSTables to merge versions/tombstones and manage read/space amplification. Repair is anti-entropy comparison/streaming between replicas. heap is Java-managed memory; off-heap covers native/direct structures outside the Java heap; the operating-system page cache is another memory consumer. GC (garbage collection) reclaims Java heap objects and can introduce pauses.
A histogram is a distribution, not an average. p50/p95/p99 are percentiles: p99 means 99% of observed values are at or below that value. tail latency is the high-percentile response time users often feel during saturation/failure. throughput is completed work per time; error rate is failed work divided by attempted work. A Service Level Objective (SLO) is a target such as p99 latency/availability; RPO (Recovery Point Objective) is tolerated data-loss time; RTO (Recovery Time Objective) is tolerated service-recovery time.
JMX (Java Management Extensions) exposes
node-local Cassandra metrics/management operations;
nodetool is itself a JMX client. An
exported metric is forwarded to an external
time-series system; Cassandra metrics are node-local until an
operator aggregates them. A guardrail warns or
rejects dangerous schema/query/operational patterns. A
rolling upgrade changes one node at a time
while the cluster remains available.
schema agreement means nodes report the same
current schema version. A driver is the client
library implementing native protocol, topology discovery, load
balancing, timeouts, retries, idempotency and speculative
execution. SAI expands to Storage-Attached
Indexing; a vector index supports approximate
nearest-neighbor retrieval and has separate
memory/disk/build/recall costs.
1. Model and defend the application contract
The capstone begins with queries, not nodes. AtlasMart's core reads are: latest orders for one customer/month, recent orders for one region/day bucket, and a point lookup/update for a known order. The first two become separate denormalized query tables. The monthly/day bucket bounds retention work; a small synthetic bucket on the region feed prevents one hot region/day partition from growing without limit. Every duplicated projection has an application/streaming reconciliation owner.
CREATE KEYSPACE IF NOT EXISTS atlasmart_capstoneWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_customer_month ( customer_id text, order_month date, order_time timestamp, order_id uuid, region text, status text, total decimal, payload text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_region_day ( region text, order_day date, bucket tinyint, order_time timestamp, order_id uuid, customer_id text, status text, total decimal, PRIMARY KEY ((region,order_day,bucket),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND default_time_to_live = 0 AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;
| Business contract | Cassandra design | Acceptance evidence |
|---|---|---|
| customer order history | ((customer_id,order_month), order_time,order_id) | partition p95/max bounded; slice query p99 |
| regional recent feed | ((region,order_day,bucket), order_time,order_id) | known fan-out across bounded buckets; merged API limit |
| availability | RF=3 dc1; LOCAL_QUORUM normal | one replica loss still satisfies 2-of-3 local quorum |
| durability/convergence | commit log + replicas + repair cadence | repair evidence before tombstone purge policy |
| search/filter | purpose-built table first; SAI only if measured fit | index build/query/write/disk/fanout metrics |
| vector/RAG optional | fixed-dimension VECTOR + SAI ANN only for evaluated use case | recall@k + p99 + index capacity; not required for checkout |
Reject an unbounded “all orders by customer
forever” partition, cross-partition scans, blind
ALLOW FILTERING, random sharding without a
routing/fan-out plan, LWT for ordinary projections, giant
multi-partition batches and a single CL chosen by habit.
2. Deploy, seed, warm and load-test the exact schema
docker network inspect atlasmart-cassandra-capstone >/dev/null 2>&1 || docker network create atlasmart-cassandra-capstonefor n in 1 2 3; do docker volume create atlasmart-cap-$n-data; donedocker inspect atlasmart-cap-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-1 --hostname atlasmart-cap-1 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-cap-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-2 --hostname atlasmart-cap-2 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cap-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-3 --hostname atlasmart-cap-3 \ --network atlasmart-cassandra-capstone \ -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cap-1 nodetool versiondocker exec atlasmart-cap-1 java -versiondocker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-1 --format '{{json .HostConfig.PortBindings}}'
specname: atlasmart_orderskeyspace: atlasmart_capstonekeyspace_definition: | CREATE KEYSPACE IF NOT EXISTS atlasmart_capstone WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};table: orders_by_customer_monthtable_definition: | CREATE TABLE IF NOT EXISTS orders_by_customer_month ( customer_id text, order_month date, order_time timestamp, order_id uuid, region text, status text, total decimal, payload text, PRIMARY KEY ((customer_id,order_month),order_time,order_id) ) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'};columnspec: - name: customer_id size: fixed(12) population: gaussian(1..50000,10000) - name: order_month cluster: fixed(1) - name: order_time cluster: gaussian(10..500,80) - name: payload size: gaussian(40..1200,180)insert: partitions: fixed(1) select: fixed(1)/500 batchtype: UNLOGGEDqueries: latest_orders: cql: SELECT order_time,order_id,status,total FROM orders_by_customer_month WHERE customer_id=? AND order_month=? LIMIT 20 fields: samerow
docker cp ./atlasmart-stress.yaml atlasmart-cap-1:/tmp/atlasmart-stress.yaml# A. Warmup: do not include in reported latency SLO.docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \ duration=45s "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \ -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 -rate threads=8# B. Normal load.docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \ duration=2m "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \ -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 -rate threads=16# C. Maintenance load: run the same workload while a controlled full repair is active.docker exec atlasmart-cap-2 nodetool repair --full atlasmart_capstone orders_by_customer_month &repair_pid=$!docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \ duration=90s "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \ -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 -rate threads=16wait $repair_pid || true# D. N-1 load: one replica unavailable, same offered workload.docker pause atlasmart-cap-3docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \ duration=60s "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \ -node atlasmart-cap-1,atlasmart-cap-2 -rate threads=16docker unpause atlasmart-cap-3
Record the complete cassandra-stress output for each phase, but the production SLO comes from the real application/driver workload. Cassandra-stress is a schema/load tool, not a security-hardened production traffic generator, and its generated distributions must be compared with production telemetry. The same offered load across phases is what makes the failure/maintenance delta interpretable.
3. Capture one evidence bundle for every phase
mkdir -p capstone-evidencedate -u +%FT%TZ > capstone-evidence/captured_at_utc.txtdocker exec atlasmart-cap-1 nodetool status > capstone-evidence/status.txtdocker exec atlasmart-cap-1 nodetool describecluster > capstone-evidence/describecluster.txtdocker exec atlasmart-cap-1 nodetool proxyhistograms > capstone-evidence/proxyhistograms.txtdocker exec atlasmart-cap-1 nodetool tablehistograms atlasmart_capstone orders_by_customer_month > capstone-evidence/tablehistograms.txtdocker exec atlasmart-cap-1 nodetool tablestats -F json atlasmart_capstone.orders_by_customer_month > capstone-evidence/tablestats.jsondocker exec atlasmart-cap-1 nodetool compactionstats > capstone-evidence/compactionstats.txtdocker exec atlasmart-cap-1 nodetool netstats > capstone-evidence/netstats.txtdocker exec atlasmart-cap-1 nodetool gcstats > capstone-evidence/gcstats.txtdocker exec atlasmart-cap-1 nodetool tpstats > capstone-evidence/tpstats.txtdocker exec atlasmart-cap-1 nodetool getguardrailsconfig --expand > capstone-evidence/guardrails.txtdocker exec atlasmart-cap-1 sh -lc 'df -h /var/lib/cassandra' > capstone-evidence/disk.txtdocker stats --no-stream atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3 > capstone-evidence/docker-stats.txtdocker logs --since 15m atlasmart-cap-1 > capstone-evidence/node1.log 2>&1
Attach the workload metadata to this bundle: dataset size, partition p50/p95/max, retention, query mix, concurrency, payload sizes, CL, RF, driver version/local DC/retry/idempotency/speculation policy, security state, warmup, host/container limits and whether repair/compaction/failure was active. Otherwise the numbers cannot be compared later.
4. Fail, repair and prove convergence
# Create a known business checkpoint before failure.docker exec atlasmart-cap-1 cqlsh -e \"CONSISTENCY LOCAL_QUORUM; INSERT INTO atlasmart_capstone.orders_by_customer_month (customer_id,order_month,order_time,order_id,region,status,total,payload) VALUES ('capstone-cust','2026-09-01','2026-09-08T12:00:00Z',00000000-0000-0000-0000-000000000275,'EU','READY',27.00,'checkpoint');"docker pause atlasmart-cap-3docker exec atlasmart-cap-1 cqlsh -e \"CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_capstone.orders_by_customer_month SET status='PAID' WHERE customer_id='capstone-cust' AND order_month='2026-09-01' AND order_time='2026-09-08T12:00:00Z' AND order_id=00000000-0000-0000-0000-000000000275;"docker unpause atlasmart-cap-3docker exec atlasmart-cap-1 nodetool status# Hints/read repair may help, but anti-entropy repair is the authoritative convergence step.docker exec atlasmart-cap-1 nodetool repair --full atlasmart_capstone orders_by_customer_monthfor n in 1 2 3; do docker exec atlasmart-cap-$n cqlsh -e \ "CONSISTENCY LOCAL_ONE; SELECT status,total FROM atlasmart_capstone.orders_by_customer_month WHERE customer_id='capstone-cust' AND order_month='2026-09-01';"done
A successful LOCAL_QUORUM mutation during one
replica loss proves 2-of-3 local acknowledgments were available;
it does not prove the paused replica was current. The
post-repair direct checks make convergence observable. In a real
multi-DC deployment, repeat the Lesson 4 whole-DC scenario and
repair policy before production signoff.
5. Prevent the known dangerous path instead of accepting it
for n in 1 2 3; do docker exec atlasmart-cap-$n nodetool setguardrailsconfig -- allow_filtering_enabled falsedonefor n in 1 2 3; do docker exec atlasmart-cap-$n cqlsh -e \ "SELECT customer_id FROM atlasmart_capstone.orders_by_customer_month WHERE status='PAID' ALLOW FILTERING;" || truedone
The acceptance result is the rejection plus a documented supported alternative—not the rejection alone. If status discovery is a business query, model a bounded status query table or evaluate SAI with measured selectivity/index bytes/write amplification/query p99. If vector search is introduced later, add vector dimension/index build/recall@k/failure capacity guardrails to this scorecard.
6. Backup and restore is a release gate
The capstone does not re-teach Chapter 25, but production approval requires a fresh restore drill from the same schema/version/security configuration. A snapshot directory on the source node is not enough. Capture a named snapshot on every replica, copy/catalog/checksum it off the node failure domain, restore into a separate target topology with the schema recreated, stream/import SSTables, run repair/verification and execute the same business checkpoint query.
for n in 1 2 3; do docker exec atlasmart-cap-$n nodetool flush atlasmart_capstone orders_by_customer_month docker exec atlasmart-cap-$n nodetool snapshot -t capstone-acceptance -cf orders_by_customer_month atlasmart_capstone docker exec atlasmart-cap-$n nodetool listsnapshotsdone# Continue with the Chapter 25 catalog/checksum/off-host copy + separate# target-cluster sstableloader/repair workflow. Record timestamps:# T_backup, T_incident, T_target_ready, T_data_loaded, T_verified, T_cutover.
| Recovery gate | Evidence |
|---|---|
| RPO | latest mutation actually recoverable from tested snapshot/incremental/PITR chain |
| RTO | incident declaration → verified application cutover, including schema/security/repair/index readiness |
| integrity | file inventory/checksums before loading |
| data | known business IDs/totals/status + partition-aware samples/invariants |
| convergence | repair/replica checks after restore path |
| security | roles/TLS/JMX/secrets/backup key access restored before exposure |
| rollback | original source/target authority and write-freeze/reconciliation decision documented |
This fails the production-defense goal. A backup is accepted only after a separate restore target reaches the expected business state, security controls and application query path within the measured RPO/RTO, with rollback/authority explicit.
7. Secure the production variant
The performance lab intentionally kept auth/TLS off on an
isolated network. That state is a failing production gate. Apply
Chapter 26's staged security design to a disposable production
candidate: PasswordAuthenticator +
CassandraAuthorizer, replicated/repaired
system_auth, non-default break-glass
administration, inherited least-privilege application roles,
client TLS and internode TLS with rotation evidence, JMX
local-only or separately authenticated/TLS/network-limited,
protected secrets/backups and complete per-node audit
collection.
authentication anonymous CQL denied; intended roles authenticateauthorization API/support denial matrix passes; no application superusersystem_auth RF/topology/repair health recordedclient TLS active clients show TLS + validated trust/endpoint identityinternode TLS encrypted peer path survives rolling restart/repair/streamingJMX local-only OR dedicated authenticated+TLS management networksecrets owner/source/rotation; no Git/CLI/history leakagebackups encrypted/restricted/audited; restore keys availableaudit every intended coordinator node + durable off-node collectionpatch inventory Cassandra/JDK/image/driver maintained and digest recorded
Re-run the load/SLO tests after security is enabled because TLS/auth/audit can change CPU, connection and I/O behavior. Security is not exempt from performance/capacity evidence, and performance is not a justification for plaintext or superuser shortcuts.
8. Tune only the measured bottleneck, then repeat the same test
| Observed mechanism | Candidate action | Required before/after proof |
|---|---|---|
| large/hot partition | bucket/query-table redesign | partition histogram + same query p99 + fanout cost |
| SSTables/read + compaction debt | headroom/compaction/model adjustment | SSTables/read + pending compaction + write/disk cost |
| tombstone scans | retention/bucketing/TWCS/delete pattern | tombstones/read + p99 + repair/gc_grace safety |
| GC/heap pressure | allocation/cache/JVM/container adjustment | GC pause/heap/RSS/page-cache + p99 |
| disk queue/saturation | faster/more disks/nodes or less amplification | disk latency/throughput + compaction/repair RTO |
| network recovery bottleneck | bandwidth/topology/throttle/capacity | stream bytes/time + application SLO |
| driver timeouts/retries | fix server/load first; adjust idempotent policy if justified | retry/speculation counts + errors + server load |
| SAI/vector cost | index/query/model/capacity change | index bytes/build/query p99/recall and write impact |
Every tuning change is a controlled experiment with a rollback. If the “fix” improves average latency while p99, errors, recovery time or disk headroom worsens, it does not pass the capstone.
9. Production-defense scorecard
MODEL[ ] every production query has a bounded Cassandra access path[ ] partition p95/max rows+bytes and retention are measured and within design[ ] denormalized projections have write/reconciliation ownershipSLO / LOAD[ ] representative warm/steady workload reports p50/p95/p99/max, throughput, errors[ ] maintenance (compaction/repair) and N-1 phases stay within defined SLO[ ] driver local-DC routing/timeouts/retries/idempotency/speculation are testedCAPACITY[ ] heap/off-heap/page-cache, disk, network and temporary rewrite headroom measured[ ] repair/replacement can finish inside operational window/RTO[ ] growth trigger and lead time are documentedRELIABILITY[ ] hints/read-repair are not used as substitutes for scheduled anti-entropy repair[ ] backup restore is reproduced with measured RPO/RTO and application cutover[ ] multi-DC failover/recovery/repair game day passes business criteriaPROTECTION[ ] guardrails reject known unsafe patterns and are consistent on every node[ ] authentication/authorization/TLS/JMX/secrets/backup/audit gates pass[ ] SAI/vector indexes (if any) have capacity, failure and quality evidenceCHANGE[ ] Cassandra/JDK/image/driver versions and upgrade compatibility are documented[ ] rolling canary/schema-agreement/driver tests/rollback gates are executable[ ] migrations/topology changes have capacity and rollback plansOPERATIONS[ ] dashboards/alerts have node/DC/table scope and actionable runbooks[ ] owners/on-call/break-glass access are tested[ ] all runtime-dependent evidence is timestamped and reproducible
Go/no-go is business-specific. A recommendation is defensible only when each waived item has an owner, risk statement, compensating control and deadline; “we will monitor it” is not a capacity/repair/restore/security plan.
10. Cleanup and course completion
# Restore the course guardrail default only if you are continuing to use this lab.for n in 1 2 3; do docker exec atlasmart-cap-$n nodetool setguardrailsconfig -- allow_filtering_enabled true 2>/dev/null || true docker exec atlasmart-cap-$n nodetool clearsnapshot -t capstone-acceptance atlasmart_capstone 2>/dev/null || truedone# Final destructive cleanup is optional after evidence/backup exercises are complete.docker rm -f atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3docker volume rm atlasmart-cap-1-data atlasmart-cap-2-data atlasmart-cap-3-datadocker network rm atlasmart-cassandra-capstone 2>/dev/null || true
Check your understanding
- Why must the capstone use the same workload across normal, maintenance and N-1 phases?
- Why is a successful LOCAL_QUORUM write during one-node loss not convergence proof?
- Why is security re-benchmarked?
- What is the strongest backup acceptance evidence?
- What makes the production architecture defensible?
Review the answers
1. Comparable offered workload makes the latency/error/resource delta attributable to the changed system state rather than a different test.
2. It proves enough acknowledgments for the requested CL; the missed replica can still be stale until hints/read repair/repair converge it.
3. Authentication, TLS, authorization and audit add real CPU/connection/I/O work and can change failure behavior; production capacity must include them.
4. A separate restore that reaches known business state, repair/security readiness and application cutover within measured RPO/RTO.
5. Every guarantee, limit, capacity assumption, failure/recovery path, security boundary, change procedure and rollback is tied to reproducible evidence and an owner.
Production judgment
Do not promote one local run into a universal Cassandra tuning rule. Record workload fit and non-goals; partition cardinality/rows/bytes and retention; read/write mix and p50/p95/p99/max latency; RF/CL and coordinator/replica failure behavior; JVM heap/GC, off-heap/page-cache, disk capacity/latency/IOPS/throughput and network bandwidth/packet loss; SSTable/read amplification, compaction backlog, tombstones and repair state; SAI/vector build/query/write and recall costs where used; authentication/authorization/TLS/JMX/secret/tenant boundaries; driver local-DC routing, timeout/retry/idempotency/speculation behavior; metrics/log/tracing coverage; backup RPO/RTO/restore proof; and the operator skill/runbook needed to perform repair, topology, upgrade and recovery safely.
Managed Cassandra services may hide disks, JMX, repair, backup, upgrade sequencing or metric names, and their quotas/cost model can change capacity decisions. Translate the same evidence questions into provider-native signals; do not assume the provider removes application data-model, driver, consistency, SLO, security, migration or rollback responsibility. This is the final Cassandra lesson. The next step is not another feature chapter: keep the scorecard/runbooks alive as workload, versions, topology, security policy, driver behavior and recovery objectives change, and rerun the relevant game days before those assumptions expire.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Apache Cassandra course curriculum.
Authoritative references
Re-check these version-sensitive sources before a real upgrade, capacity commitment, security change or game day.
- Apache Cassandra 5.0 release/download baseline
- Monitoring metrics
- Troubleshooting with nodetool histograms
- nodetool tablestats
- nodetool tablehistograms
- nodetool proxyhistograms
- cassandra-stress user mode
- cassandra.yaml including guardrails/storage compatibility
- nodetool setguardrailsconfig
- Repair operations
- Backups
- Security
- Java Driver core documentation
- Production recommendations
- Topology changes
- Storage-Attached Indexing
- Vector search