Chapter 27 · Observability, Performance, Capacity, Guardrails, Upgrades, Multi-DC Resilience, and Capstone

Capstone: Model, Deploy, Load-Test, Fail, Repair, Recover, Secure, Tune, and Defend a Production Cassandra Cluster

Defend the complete AtlasMart Cassandra production design through measurable load, failure, repair, restore, security, tuning and change gates.

Advanced · Production capstone220–360 minutesFull production-defense capstoneApache Cassandra 5.0.9 · Java 17 · Java Driver 4.19.3 · cqlsh/nodetool/cassandra-stress · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart is ready for production review. The architecture diagram looks polished, but the approval board asks harder questions: Can the schema survive the real partition distribution? Can p99 stay inside the checkout SLO during compaction/repair and one-node loss? Can the team recover a tested backup within RTO, reject dangerous queries, rotate credentials/certificates, survive a DC loss, upgrade without schema/driver breakage, and explain every rollback? The capstone is accepted only when those claims are reproducible.

01

Turn product queries/retention/invariants into bounded Cassandra tables with an explicit RF/CL/driver contract.

02

Deploy and load-test representative distributions, report p50/p95/p99/max/throughput/errors and correlate them with Cassandra/JVM/host evidence.

03

Inject node/DC and bad-query failures, repair/converge, enforce guardrails and verify recovery rather than merely restarting processes.

04

Demonstrate a cataloged backup restore and security acceptance gates before application cutover.

05

Produce a production-defense scorecard/runbook with SLO/RPO/RTO, capacity headroom, upgrade/migration/rollback and ownership.

Chapter 27 capstone lab baseline

The mandatory single-datacenter labs use Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the official image, cqlsh/nodetool from that same image, and Apache Cassandra Java Driver 4.19.3 where client behavior matters. Use Docker network atlasmart-cassandra-capstone, cluster atlasmart-capstone, nodes atlasmart-cap-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy, replication factor (RF) 3, and application reads/writes at LOCAL_QUORUM unless the lesson deliberately changes consistency level (CL). Tables explicitly use UnifiedCompactionStrategy (UCS), default_time_to_live=0, and gc_grace_seconds=864000.

For repeatability, mandatory Lessons 1–3 keep authentication/client TLS/internode TLS disabled only inside this isolated Docker network; no Cassandra or JMX port is published on the host, and JMX remains local-only inside containers. Chapter 26 remains the production security baseline. Lesson 4 creates separate disposable upgrade/multi-DC clusters, and Lesson 5 makes the security state an explicit capstone acceptance gate. A full RF=3 two-DC game day requires six Cassandra containers and roughly 10–14 GiB of available RAM depending on container limits/JVM ergonomics; the lesson also provides a reduced four-node RF=2 simulation for constrained laptops and labels the semantic difference.

Exact token values, latencies, GC pauses, SSTable counts, compaction/repair bytes, disk throughput, network rates, failure-detection timing, guardrail messages and benchmark throughput are runtime evidence. The lesson never treats example numbers as results from your machine. Before every benchmark/failure drill capture host CPU/RAM/disk, Docker limits, Cassandra/Java/driver versions, topology, RF/CL, schema/compaction, dataset shape, concurrency, warmup and security state.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms used across the capstone

A coordinator is the Cassandra node handling one client request; a replica stores a copy of the requested partition according to the keyspace replication strategy. A partition is the rows sharing a partition key; the partitioner hashes that key to a token, and virtual nodes (vnodes) give each physical node multiple token ranges. A datacenter (DC) and rack model failure/locality domains. RF (replication factor) is the number of replicas per DC configured by NetworkTopologyStrategy; a CL (consistency level) controls how many appropriately scoped replica responses are required.

An SSTable (Sorted String Table) is an immutable on-disk data-file set. Compaction rewrites SSTables to merge versions/tombstones and manage read/space amplification. Repair is anti-entropy comparison/streaming between replicas. heap is Java-managed memory; off-heap covers native/direct structures outside the Java heap; the operating-system page cache is another memory consumer. GC (garbage collection) reclaims Java heap objects and can introduce pauses.

A histogram is a distribution, not an average. p50/p95/p99 are percentiles: p99 means 99% of observed values are at or below that value. tail latency is the high-percentile response time users often feel during saturation/failure. throughput is completed work per time; error rate is failed work divided by attempted work. A Service Level Objective (SLO) is a target such as p99 latency/availability; RPO (Recovery Point Objective) is tolerated data-loss time; RTO (Recovery Time Objective) is tolerated service-recovery time.

JMX (Java Management Extensions) exposes node-local Cassandra metrics/management operations; nodetool is itself a JMX client. An exported metric is forwarded to an external time-series system; Cassandra metrics are node-local until an operator aggregates them. A guardrail warns or rejects dangerous schema/query/operational patterns. A rolling upgrade changes one node at a time while the cluster remains available. schema agreement means nodes report the same current schema version. A driver is the client library implementing native protocol, topology discovery, load balancing, timeouts, retries, idempotency and speculative execution. SAI expands to Storage-Attached Indexing; a vector index supports approximate nearest-neighbor retrieval and has separate memory/disk/build/recall costs.

1. Model and defend the application contract

The capstone begins with queries, not nodes. AtlasMart's core reads are: latest orders for one customer/month, recent orders for one region/day bucket, and a point lookup/update for a known order. The first two become separate denormalized query tables. The monthly/day bucket bounds retention work; a small synthetic bucket on the region feed prevents one hot region/day partition from growing without limit. Every duplicated projection has an application/streaming reconciliation owner.

CQL · AtlasMart query-first capstone schema
CREATE KEYSPACE IF NOT EXISTS atlasmart_capstoneWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_customer_month (    customer_id text,    order_month date,    order_time timestamp,    order_id uuid,    region text,    status text,    total decimal,    payload text,    PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND default_time_to_live = 0  AND gc_grace_seconds = 864000;CREATE TABLE IF NOT EXISTS atlasmart_capstone.orders_by_region_day (    region text,    order_day date,    bucket tinyint,    order_time timestamp,    order_id uuid,    customer_id text,    status text,    total decimal,    PRIMARY KEY ((region,order_day,bucket),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND default_time_to_live = 0  AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;
Business contract Cassandra design Acceptance evidence
customer order history ((customer_id,order_month), order_time,order_id) partition p95/max bounded; slice query p99
regional recent feed ((region,order_day,bucket), order_time,order_id) known fan-out across bounded buckets; merged API limit
availability RF=3 dc1; LOCAL_QUORUM normal one replica loss still satisfies 2-of-3 local quorum
durability/convergence commit log + replicas + repair cadence repair evidence before tombstone purge policy
search/filter purpose-built table first; SAI only if measured fit index build/query/write/disk/fanout metrics
vector/RAG optional fixed-dimension VECTOR + SAI ANN only for evaluated use case recall@k + p99 + index capacity; not required for checkout

Reject an unbounded “all orders by customer forever” partition, cross-partition scans, blind ALLOW FILTERING, random sharding without a routing/fan-out plan, LWT for ordinary projections, giant multi-partition batches and a single CL chosen by habit.

2. Deploy, seed, warm and load-test the exact schema

bash · create or verify the three-node capstone cluster
docker network inspect atlasmart-cassandra-capstone >/dev/null 2>&1 || docker network create atlasmart-cassandra-capstonefor n in 1 2 3; do docker volume create atlasmart-cap-$n-data; donedocker inspect atlasmart-cap-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-1 --hostname atlasmart-cap-1 \  --network atlasmart-cassandra-capstone \  -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-cap-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-2 --hostname atlasmart-cap-2 \  --network atlasmart-cassandra-capstone \  -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cap-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cap-3 --hostname atlasmart-cap-3 \  --network atlasmart-cassandra-capstone \  -e CASSANDRA_CLUSTER_NAME=atlasmart-capstone -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-cap-1 -v atlasmart-cap-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cap-1 nodetool versiondocker exec atlasmart-cap-1 java -versiondocker exec atlasmart-cap-1 nodetool statusdocker inspect atlasmart-cap-1 --format '{{json .HostConfig.PortBindings}}'
yaml · cassandra-stress user profile for the real query table
specname: atlasmart_orderskeyspace: atlasmart_capstonekeyspace_definition: |  CREATE KEYSPACE IF NOT EXISTS atlasmart_capstone WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};table: orders_by_customer_monthtable_definition: |  CREATE TABLE IF NOT EXISTS orders_by_customer_month (    customer_id text,    order_month date,    order_time timestamp,    order_id uuid,    region text,    status text,    total decimal,    payload text,    PRIMARY KEY ((customer_id,order_month),order_time,order_id)  ) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)    AND compaction = {'class':'UnifiedCompactionStrategy'};columnspec:  - name: customer_id    size: fixed(12)    population: gaussian(1..50000,10000)  - name: order_month    cluster: fixed(1)  - name: order_time    cluster: gaussian(10..500,80)  - name: payload    size: gaussian(40..1200,180)insert:  partitions: fixed(1)  select: fixed(1)/500  batchtype: UNLOGGEDqueries:  latest_orders:    cql: SELECT order_time,order_id,status,total FROM orders_by_customer_month WHERE customer_id=? AND order_month=? LIMIT 20    fields: samerow
bash · reproducible capstone workload phases
docker cp ./atlasmart-stress.yaml atlasmart-cap-1:/tmp/atlasmart-stress.yaml# A. Warmup: do not include in reported latency SLO.docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \  duration=45s "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \  -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 -rate threads=8# B. Normal load.docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \  duration=2m "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \  -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 -rate threads=16# C. Maintenance load: run the same workload while a controlled full repair is active.docker exec atlasmart-cap-2 nodetool repair --full atlasmart_capstone orders_by_customer_month &repair_pid=$!docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \  duration=90s "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \  -node atlasmart-cap-1,atlasmart-cap-2,atlasmart-cap-3 -rate threads=16wait $repair_pid || true# D. N-1 load: one replica unavailable, same offered workload.docker pause atlasmart-cap-3docker exec atlasmart-cap-1 cassandra-stress user profile=/tmp/atlasmart-stress.yaml \  duration=60s "ops(insert=3,latest_orders=7)" no-warmup cl=LOCAL_QUORUM \  -node atlasmart-cap-1,atlasmart-cap-2 -rate threads=16docker unpause atlasmart-cap-3

Record the complete cassandra-stress output for each phase, but the production SLO comes from the real application/driver workload. Cassandra-stress is a schema/load tool, not a security-hardened production traffic generator, and its generated distributions must be compared with production telemetry. The same offered load across phases is what makes the failure/maintenance delta interpretable.

3. Capture one evidence bundle for every phase

bash · capstone evidence collector
mkdir -p capstone-evidencedate -u +%FT%TZ > capstone-evidence/captured_at_utc.txtdocker exec atlasmart-cap-1 nodetool status > capstone-evidence/status.txtdocker exec atlasmart-cap-1 nodetool describecluster > capstone-evidence/describecluster.txtdocker exec atlasmart-cap-1 nodetool proxyhistograms > capstone-evidence/proxyhistograms.txtdocker exec atlasmart-cap-1 nodetool tablehistograms atlasmart_capstone orders_by_customer_month > capstone-evidence/tablehistograms.txtdocker exec atlasmart-cap-1 nodetool tablestats -F json atlasmart_capstone.orders_by_customer_month > capstone-evidence/tablestats.jsondocker exec atlasmart-cap-1 nodetool compactionstats > capstone-evidence/compactionstats.txtdocker exec atlasmart-cap-1 nodetool netstats > capstone-evidence/netstats.txtdocker exec atlasmart-cap-1 nodetool gcstats > capstone-evidence/gcstats.txtdocker exec atlasmart-cap-1 nodetool tpstats > capstone-evidence/tpstats.txtdocker exec atlasmart-cap-1 nodetool getguardrailsconfig --expand > capstone-evidence/guardrails.txtdocker exec atlasmart-cap-1 sh -lc 'df -h /var/lib/cassandra' > capstone-evidence/disk.txtdocker stats --no-stream atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3 > capstone-evidence/docker-stats.txtdocker logs --since 15m atlasmart-cap-1 > capstone-evidence/node1.log 2>&1

Attach the workload metadata to this bundle: dataset size, partition p50/p95/max, retention, query mix, concurrency, payload sizes, CL, RF, driver version/local DC/retry/idempotency/speculation policy, security state, warmup, host/container limits and whether repair/compaction/failure was active. Otherwise the numbers cannot be compared later.

4. Fail, repair and prove convergence

bash · replica failure/recovery acceptance
# Create a known business checkpoint before failure.docker exec atlasmart-cap-1 cqlsh -e \"CONSISTENCY LOCAL_QUORUM; INSERT INTO atlasmart_capstone.orders_by_customer_month (customer_id,order_month,order_time,order_id,region,status,total,payload) VALUES ('capstone-cust','2026-09-01','2026-09-08T12:00:00Z',00000000-0000-0000-0000-000000000275,'EU','READY',27.00,'checkpoint');"docker pause atlasmart-cap-3docker exec atlasmart-cap-1 cqlsh -e \"CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_capstone.orders_by_customer_month SET status='PAID' WHERE customer_id='capstone-cust' AND order_month='2026-09-01' AND order_time='2026-09-08T12:00:00Z' AND order_id=00000000-0000-0000-0000-000000000275;"docker unpause atlasmart-cap-3docker exec atlasmart-cap-1 nodetool status# Hints/read repair may help, but anti-entropy repair is the authoritative convergence step.docker exec atlasmart-cap-1 nodetool repair --full atlasmart_capstone orders_by_customer_monthfor n in 1 2 3; do  docker exec atlasmart-cap-$n cqlsh -e \  "CONSISTENCY LOCAL_ONE; SELECT status,total FROM atlasmart_capstone.orders_by_customer_month WHERE customer_id='capstone-cust' AND order_month='2026-09-01';"done

A successful LOCAL_QUORUM mutation during one replica loss proves 2-of-3 local acknowledgments were available; it does not prove the paused replica was current. The post-repair direct checks make convergence observable. In a real multi-DC deployment, repeat the Lesson 4 whole-DC scenario and repair policy before production signoff.

5. Prevent the known dangerous path instead of accepting it

bash · enforce the capstone ALLOW FILTERING policy on every node
for n in 1 2 3; do  docker exec atlasmart-cap-$n nodetool setguardrailsconfig -- allow_filtering_enabled falsedonefor n in 1 2 3; do  docker exec atlasmart-cap-$n cqlsh -e \  "SELECT customer_id FROM atlasmart_capstone.orders_by_customer_month WHERE status='PAID' ALLOW FILTERING;" || truedone

The acceptance result is the rejection plus a documented supported alternative—not the rejection alone. If status discovery is a business query, model a bounded status query table or evaluate SAI with measured selectivity/index bytes/write amplification/query p99. If vector search is introduced later, add vector dimension/index build/recall@k/failure capacity guardrails to this scorecard.

6. Backup and restore is a release gate

The capstone does not re-teach Chapter 25, but production approval requires a fresh restore drill from the same schema/version/security configuration. A snapshot directory on the source node is not enough. Capture a named snapshot on every replica, copy/catalog/checksum it off the node failure domain, restore into a separate target topology with the schema recreated, stream/import SSTables, run repair/verification and execute the same business checkpoint query.

bash · create a named capstone recovery point
for n in 1 2 3; do  docker exec atlasmart-cap-$n nodetool flush atlasmart_capstone orders_by_customer_month  docker exec atlasmart-cap-$n nodetool snapshot -t capstone-acceptance -cf orders_by_customer_month atlasmart_capstone  docker exec atlasmart-cap-$n nodetool listsnapshotsdone# Continue with the Chapter 25 catalog/checksum/off-host copy + separate# target-cluster sstableloader/repair workflow. Record timestamps:# T_backup, T_incident, T_target_ready, T_data_loaded, T_verified, T_cutover.
Recovery gate Evidence
RPO latest mutation actually recoverable from tested snapshot/incremental/PITR chain
RTO incident declaration → verified application cutover, including schema/security/repair/index readiness
integrity file inventory/checksums before loading
data known business IDs/totals/status + partition-aware samples/invariants
convergence repair/replica checks after restore path
security roles/TLS/JMX/secrets/backup key access restored before exposure
rollback original source/target authority and write-freeze/reconciliation decision documented
Deliberately incomplete capstone: “backup succeeded” but nobody can reproduce restore.

This fails the production-defense goal. A backup is accepted only after a separate restore target reaches the expected business state, security controls and application query path within the measured RPO/RTO, with rollback/authority explicit.

7. Secure the production variant

The performance lab intentionally kept auth/TLS off on an isolated network. That state is a failing production gate. Apply Chapter 26's staged security design to a disposable production candidate: PasswordAuthenticator + CassandraAuthorizer, replicated/repaired system_auth, non-default break-glass administration, inherited least-privilege application roles, client TLS and internode TLS with rotation evidence, JMX local-only or separately authenticated/TLS/network-limited, protected secrets/backups and complete per-node audit collection.

text · security acceptance evidence, not secret values
authentication    anonymous CQL denied; intended roles authenticateauthorization     API/support denial matrix passes; no application superusersystem_auth       RF/topology/repair health recordedclient TLS        active clients show TLS + validated trust/endpoint identityinternode TLS     encrypted peer path survives rolling restart/repair/streamingJMX               local-only OR dedicated authenticated+TLS management networksecrets           owner/source/rotation; no Git/CLI/history leakagebackups           encrypted/restricted/audited; restore keys availableaudit             every intended coordinator node + durable off-node collectionpatch inventory   Cassandra/JDK/image/driver maintained and digest recorded

Re-run the load/SLO tests after security is enabled because TLS/auth/audit can change CPU, connection and I/O behavior. Security is not exempt from performance/capacity evidence, and performance is not a justification for plaintext or superuser shortcuts.

8. Tune only the measured bottleneck, then repeat the same test

Observed mechanism Candidate action Required before/after proof
large/hot partition bucket/query-table redesign partition histogram + same query p99 + fanout cost
SSTables/read + compaction debt headroom/compaction/model adjustment SSTables/read + pending compaction + write/disk cost
tombstone scans retention/bucketing/TWCS/delete pattern tombstones/read + p99 + repair/gc_grace safety
GC/heap pressure allocation/cache/JVM/container adjustment GC pause/heap/RSS/page-cache + p99
disk queue/saturation faster/more disks/nodes or less amplification disk latency/throughput + compaction/repair RTO
network recovery bottleneck bandwidth/topology/throttle/capacity stream bytes/time + application SLO
driver timeouts/retries fix server/load first; adjust idempotent policy if justified retry/speculation counts + errors + server load
SAI/vector cost index/query/model/capacity change index bytes/build/query p99/recall and write impact

Every tuning change is a controlled experiment with a rollback. If the “fix” improves average latency while p99, errors, recovery time or disk headroom worsens, it does not pass the capstone.

9. Production-defense scorecard

text · final go/no-go scorecard
MODEL[ ] every production query has a bounded Cassandra access path[ ] partition p95/max rows+bytes and retention are measured and within design[ ] denormalized projections have write/reconciliation ownershipSLO / LOAD[ ] representative warm/steady workload reports p50/p95/p99/max, throughput, errors[ ] maintenance (compaction/repair) and N-1 phases stay within defined SLO[ ] driver local-DC routing/timeouts/retries/idempotency/speculation are testedCAPACITY[ ] heap/off-heap/page-cache, disk, network and temporary rewrite headroom measured[ ] repair/replacement can finish inside operational window/RTO[ ] growth trigger and lead time are documentedRELIABILITY[ ] hints/read-repair are not used as substitutes for scheduled anti-entropy repair[ ] backup restore is reproduced with measured RPO/RTO and application cutover[ ] multi-DC failover/recovery/repair game day passes business criteriaPROTECTION[ ] guardrails reject known unsafe patterns and are consistent on every node[ ] authentication/authorization/TLS/JMX/secrets/backup/audit gates pass[ ] SAI/vector indexes (if any) have capacity, failure and quality evidenceCHANGE[ ] Cassandra/JDK/image/driver versions and upgrade compatibility are documented[ ] rolling canary/schema-agreement/driver tests/rollback gates are executable[ ] migrations/topology changes have capacity and rollback plansOPERATIONS[ ] dashboards/alerts have node/DC/table scope and actionable runbooks[ ] owners/on-call/break-glass access are tested[ ] all runtime-dependent evidence is timestamped and reproducible

Go/no-go is business-specific. A recommendation is defensible only when each waived item has an owner, risk statement, compensating control and deadline; “we will monitor it” is not a capacity/repair/restore/security plan.

10. Cleanup and course completion

bash · reset the disposable capstone lab after evidence is saved
# Restore the course guardrail default only if you are continuing to use this lab.for n in 1 2 3; do  docker exec atlasmart-cap-$n nodetool setguardrailsconfig -- allow_filtering_enabled true 2>/dev/null || true  docker exec atlasmart-cap-$n nodetool clearsnapshot -t capstone-acceptance atlasmart_capstone 2>/dev/null || truedone# Final destructive cleanup is optional after evidence/backup exercises are complete.docker rm -f atlasmart-cap-1 atlasmart-cap-2 atlasmart-cap-3docker volume rm atlasmart-cap-1-data atlasmart-cap-2-data atlasmart-cap-3-datadocker network rm atlasmart-cassandra-capstone 2>/dev/null || true

Check your understanding

  1. Why must the capstone use the same workload across normal, maintenance and N-1 phases?
  2. Why is a successful LOCAL_QUORUM write during one-node loss not convergence proof?
  3. Why is security re-benchmarked?
  4. What is the strongest backup acceptance evidence?
  5. What makes the production architecture defensible?
Review the answers

1. Comparable offered workload makes the latency/error/resource delta attributable to the changed system state rather than a different test.

2. It proves enough acknowledgments for the requested CL; the missed replica can still be stale until hints/read repair/repair converge it.

3. Authentication, TLS, authorization and audit add real CPU/connection/I/O work and can change failure behavior; production capacity must include them.

4. A separate restore that reaches known business state, repair/security readiness and application cutover within measured RPO/RTO.

5. Every guarantee, limit, capacity assumption, failure/recovery path, security boundary, change procedure and rollback is tied to reproducible evidence and an owner.

Production judgment

Do not promote one local run into a universal Cassandra tuning rule. Record workload fit and non-goals; partition cardinality/rows/bytes and retention; read/write mix and p50/p95/p99/max latency; RF/CL and coordinator/replica failure behavior; JVM heap/GC, off-heap/page-cache, disk capacity/latency/IOPS/throughput and network bandwidth/packet loss; SSTable/read amplification, compaction backlog, tombstones and repair state; SAI/vector build/query/write and recall costs where used; authentication/authorization/TLS/JMX/secret/tenant boundaries; driver local-DC routing, timeout/retry/idempotency/speculation behavior; metrics/log/tracing coverage; backup RPO/RTO/restore proof; and the operator skill/runbook needed to perform repair, topology, upgrade and recovery safely.

Managed Cassandra services may hide disks, JMX, repair, backup, upgrade sequencing or metric names, and their quotas/cost model can change capacity decisions. Translate the same evidence questions into provider-native signals; do not assume the provider removes application data-model, driver, consistency, SLO, security, migration or rollback responsibility. This is the final Cassandra lesson. The next step is not another feature chapter: keep the scorecard/runbooks alive as workload, versions, topology, security policy, driver behavior and recovery objectives change, and rerun the relevant game days before those assumptions expire.

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Apache Cassandra course curriculum.

Authoritative references

Re-check these version-sensitive sources before a real upgrade, capacity commitment, security change or game day.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.