Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone
Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform
Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.
Redis Open Source 8.10.1 using the pinned
redis:8.10.1 image. The observability/benchmark
node is atlasmart-redis-ch24 on
127.0.0.1:6441, standalone topology, logical
database 0, AOF everysec plus RDB save rules,
maxmemory 0/noeviction unless a
bounded experiment says otherwise, named
academy-admin and atlasmart-app ACL
users, and fixture prefix atlasmart:ch24:*. TLS is
off only on this loopback-local disposable node; the capstone
security acceptance criteria reuse Chapter 22 TLS/ACL guidance.
Python examples target redis==8.1.0.
Search/JSON/vector/time-series/probabilistic features are
optional and must be included in capacity accounting only when
the chosen AtlasMart architecture actually uses them.
Learning outcomes
This lesson turns Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.
Explain the mechanisms and terminology behind Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform.
Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.
Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.
Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.
Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.
1. Capstone problem: can another operator recover AtlasMart from your runbook?
A production Redis design is not complete because it passes a feature checklist. It is complete when an operator can identify the topology and security boundaries, load it realistically, detect saturation, fail it safely, rediscover service, restore from backup, explain the bounded data-loss window, and return to a known-good state from written evidence. The capstone therefore treats the runbook itself as an executable artifact.
2. Freeze the architecture and assumptions before testing
| Area | AtlasMart capstone decision | Acceptance evidence |
|---|---|---|
| Topology | Choose standalone+restore, Sentinel, or Cluster based on availability/partitioning needs | Node/role/service-discovery evidence |
| Durability | Document RDB/AOF/fsync and acknowledged-write strategy | INFO persistence + recovery drill |
| Security | Named least-privilege ACLs; TLS on non-loopback paths; network isolation | ACL DRYRUN/WHOAMI + TLS verification |
| Memory | Dataset + indexes + buffers + fork/failover headroom | INFO/MEMORY + capacity worksheet |
| Workload | Measured command mix, payload, skew, concurrency, pipeline depth | Custom benchmark contract |
| SLOs | Availability + p50/p95/p99/max + RPO/RTO | Time-series evidence and error budget |
| Backup | Independent artifact and tested restore | checksums + restored invariants |
| Change | Pinned versions and rollback point | version evidence + rollback rehearsal |
3. Acceptance gate A: identity, security, configuration, and observability
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin ACL WHOAMIdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CLIENT LIST
For a real deployment add certificate-chain verification, firewall/private-network evidence, secret-source/rotation evidence, and ACL tests for every application/monitoring/replication/Sentinel/Cluster identity. Redis ACLs do not replace tenant/business authorization in the application.
4. Acceptance gate B: representative load without deleting the safety margin
Run the Chapter 24 custom workload at normal expected load, then bounded higher steps. Record p50/p95/p99/max, error rate, schedule lag, CPU/RSS, memory fragmentation, client buffers, persistence events, replication lag, evictions, and hot-key distribution. The pass criterion is not “no crash”; it is SLO compliance with sufficient headroom for failover, backup/rewrite, and demand bursts.
5. Acceptance gate C: fail over using the topology you actually chose
A standalone server cannot fail over. If AtlasMart requires automatic failover, reuse the dedicated Chapter 20 three-Sentinel topology or Chapter 21 Cluster topology and run the same application client used in acceptance. Stop the current primary, not a hard-coded old address; measure detection, promotion, client rediscovery, request errors, and acknowledged-write outcomes. Never claim zero data loss from asynchronous replication.
| Topology | Failure drill | Evidence |
|---|---|---|
| Sentinel | Stop discovered primary; require Sentinel-aware client rediscovery | SDOWN/ODOWN, promoted role, replicas reconfigured, app recovery time |
| Cluster | Stop a primary with an eligible replica; require cluster-aware client routing | cluster_state, replica promotion, MOVED/topology refresh, slot availability |
| Standalone | No failover claim; test restart + restore only | availability gap and restore RTO |
6. Acceptance gate D: recover from backup as a separate exercise
Failover and backup answer different failures. Restore the most recent accepted backup into a clean node, validate checksums and application fixtures, record RPO/RTO, and prove that the restore procedure does not rely on the failed node. If TLS certificates, ACL files, encryption keys, or application secrets are required, validate their recovery path independently without committing them to source control.
7. Acceptance gate E: tune one mechanism at a time and rerun the same test
Tuning without a before/after contract is storytelling. Candidates can include pipeline depth, key redesign, maxmemory/eviction policy, persistence settings, replica/backlog sizing, slow-command removal, local caching, or Search/index changes. Each change must state the expected mechanism, risk, rollback, and the metrics that would falsify the hypothesis.
8. Acceptance gate F: upgrade and rollback rehearsal
Pin both old and new versions, capture release/security notes, back up first, upgrade in topology-aware order, verify application and Redis state, then rehearse the rollback path on disposable data. Current course baseline is Redis 8.10.1; the Docker Official Image also publishes 8.8.2 for older-version rehearsal, but future operators must re-check the exact supported tags and compatibility at execution time.
9. SLO and error-budget worksheet
| Signal | Target/threshold | Measured | Pass? | Action if failed |
|---|---|---|---|---|
| Availability during game day | team-defined SLO | measure | yes/no | client/topology/failover repair |
| p99 latency at expected load | team-defined SLO | measure | yes/no | diagnose layer; do not tune blindly |
| Error rate/timeouts | team-defined | measure | yes/no | inspect client/server/network saturation |
| RPO | business-defined | measure acknowledged-write loss | yes/no | durability/replication design |
| RTO | business-defined | measure restore/recovery duration | yes/no | runbook/backup/topology improvement |
| Memory/fork/failover headroom | capacity model | measure | yes/no | reduce load or add capacity before release |
10. Failure timeline template
T0 baseline evidence captured; version/topology/security/workload recordedT1 representative load stable; p50/p95/p99/max and error rate capturedT2 failure injected into named/discovered current primaryT3 control plane detects failureT4 replacement primary/slot owner availableT5 application rediscovers service and successful writes resumeT6 old node repaired/rejoined; redundancy restoredT7 independent backup restored and invariants verifiedT8 post-event metrics compared with SLO/RPO/RTO and error budgetT9 rollback/cleanup completed; runbook corrections committed
11. Deliberately wrong capstone: “all commands returned OK”
A command checklist can pass while the application has stale reads, broken retries, untested restore, unsafe ACLs, no failover-capable client, hot keys, or insufficient headroom. The repair is outcome-based acceptance: workload + observability + fault injection + restore + security + rollback with business invariants and measured timelines.
12. Production runbook: minimum sections
- Architecture: node roles, ports, service discovery, slot/replication layout, dependencies, fault domains.
- Security: ACL identities, TLS/network boundaries, secret sources/rotation, emergency access, audit evidence.
- SLOs and capacity: workload contract, latency/error targets, RPO/RTO, memory/fork/failover headroom.
- Observability: INFO/latency/slowlog/client/host metrics, alerts, dashboards, log correlation, escalation thresholds.
- Failure procedures: current-primary discovery, failover/partition handling, stale-read policy, retry/idempotency boundaries.
- Backup/restore: artifact locations, checksums, retention, restore commands, application invariant checks.
- Change management: pinned versions, compatibility checks, topology-aware order, rollback boundary, post-change validation.
- Ownership: who can execute each operation, who approves risk, and where incident evidence is stored.
Check your understanding
- Why is a standalone restart not a failover test?
- Why must backup restore be tested separately from replica/Sentinel/Cluster failover?
- What makes a benchmark useful for capacity?
- When is tuning complete?
- What is the capstone’s final artifact?
Review the answers
There is no alternate serving node or service-discovery transition; it only measures outage/restart/recovery.
Replication/high availability can propagate logical damage and does not prove independent disaster recovery.
A documented representative workload, SLO-bound tail/error measurements, correlated resource evidence, and preserved headroom.
When the hypothesized mechanism changes the expected metrics, the same acceptance test passes, and rollback remains defined—not when a single number improves.
A measured, executable production runbook plus evidence that another operator can use it to operate and recover the chosen Redis topology.
13. Course completion judgment
A production Redis operator should now be able to explain data-structure semantics, key lifecycle, atomicity, Streams, Search/JSON/vector/probabilistic/time-series features, client behavior, persistence, memory/eviction, replication, Sentinel, Cluster, security, application patterns, and observability as one system of tradeoffs. The final standard is not memorizing commands: it is preserving correctness while measuring latency, capacity, durability, availability, security, recovery, and cost under the workload that actually matters.
Summary and next step
Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform completes the Redis course path. Keep the capstone evidence, recovery notes, and operating assumptions together as the final production-readiness record.
Authoritative references
docker rm -f atlasmart-redis-ch24 2>/dev/null || truedocker volume rm atlasmart-redis-ch24-data 2>/dev/null || true# Do not run docker system prune; remove only resources named by this chapter.