Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone

Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform

Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.

Advanced210–320 minutescapstone, SLO, failover, recovery, security, runbook, operationsRedis Open Source 8.10.1redis-py 8.1.0 where Python is usedDocker + redis-cli + PythonStandalone evidence node · DB 0AOF everysec + RDB · maxmemory 0/noeviction baselineNamed ACL users · TLS off only on loopbackReuses Sentinel/Cluster labs for failover acceptanceFree/local-firstLast reviewed: September 6, 2026
Reproducible Chapter 24 baseline

Redis Open Source 8.10.1 using the pinned redis:8.10.1 image. The observability/benchmark node is atlasmart-redis-ch24 on 127.0.0.1:6441, standalone topology, logical database 0, AOF everysec plus RDB save rules, maxmemory 0/noeviction unless a bounded experiment says otherwise, named academy-admin and atlasmart-app ACL users, and fixture prefix atlasmart:ch24:*. TLS is off only on this loopback-local disposable node; the capstone security acceptance criteria reuse Chapter 22 TLS/ACL guidance. Python examples target redis==8.1.0. Search/JSON/vector/time-series/probabilistic features are optional and must be included in capacity accounting only when the chosen AtlasMart architecture actually uses them.

Learning outcomes

This lesson turns Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.

01

Explain the mechanisms and terminology behind Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform.

02

Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.

03

Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.

04

Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.

05

Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.

1. Capstone problem: can another operator recover AtlasMart from your runbook?

A production Redis design is not complete because it passes a feature checklist. It is complete when an operator can identify the topology and security boundaries, load it realistically, detect saturation, fail it safely, rediscover service, restore from backup, explain the bounded data-loss window, and return to a known-good state from written evidence. The capstone therefore treats the runbook itself as an executable artifact.

2. Freeze the architecture and assumptions before testing

Area AtlasMart capstone decision Acceptance evidence
Topology Choose standalone+restore, Sentinel, or Cluster based on availability/partitioning needs Node/role/service-discovery evidence
Durability Document RDB/AOF/fsync and acknowledged-write strategy INFO persistence + recovery drill
Security Named least-privilege ACLs; TLS on non-loopback paths; network isolation ACL DRYRUN/WHOAMI + TLS verification
Memory Dataset + indexes + buffers + fork/failover headroom INFO/MEMORY + capacity worksheet
Workload Measured command mix, payload, skew, concurrency, pipeline depth Custom benchmark contract
SLOs Availability + p50/p95/p99/max + RPO/RTO Time-series evidence and error budget
Backup Independent artifact and tested restore checksums + restored invariants
Change Pinned versions and rollback point version evidence + rollback rehearsal

3. Acceptance gate A: identity, security, configuration, and observability

Shell · pre-flight evidence bundle
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin ACL WHOAMIdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CLIENT LIST

For a real deployment add certificate-chain verification, firewall/private-network evidence, secret-source/rotation evidence, and ACL tests for every application/monitoring/replication/Sentinel/Cluster identity. Redis ACLs do not replace tenant/business authorization in the application.

4. Acceptance gate B: representative load without deleting the safety margin

Run the Chapter 24 custom workload at normal expected load, then bounded higher steps. Record p50/p95/p99/max, error rate, schedule lag, CPU/RSS, memory fragmentation, client buffers, persistence events, replication lag, evictions, and hot-key distribution. The pass criterion is not “no crash”; it is SLO compliance with sufficient headroom for failover, backup/rewrite, and demand bursts.

5. Acceptance gate C: fail over using the topology you actually chose

A standalone server cannot fail over. If AtlasMart requires automatic failover, reuse the dedicated Chapter 20 three-Sentinel topology or Chapter 21 Cluster topology and run the same application client used in acceptance. Stop the current primary, not a hard-coded old address; measure detection, promotion, client rediscovery, request errors, and acknowledged-write outcomes. Never claim zero data loss from asynchronous replication.

Topology Failure drill Evidence
Sentinel Stop discovered primary; require Sentinel-aware client rediscovery SDOWN/ODOWN, promoted role, replicas reconfigured, app recovery time
Cluster Stop a primary with an eligible replica; require cluster-aware client routing cluster_state, replica promotion, MOVED/topology refresh, slot availability
Standalone No failover claim; test restart + restore only availability gap and restore RTO

6. Acceptance gate D: recover from backup as a separate exercise

Failover and backup answer different failures. Restore the most recent accepted backup into a clean node, validate checksums and application fixtures, record RPO/RTO, and prove that the restore procedure does not rely on the failed node. If TLS certificates, ACL files, encryption keys, or application secrets are required, validate their recovery path independently without committing them to source control.

7. Acceptance gate E: tune one mechanism at a time and rerun the same test

Tuning without a before/after contract is storytelling. Candidates can include pipeline depth, key redesign, maxmemory/eviction policy, persistence settings, replica/backlog sizing, slow-command removal, local caching, or Search/index changes. Each change must state the expected mechanism, risk, rollback, and the metrics that would falsify the hypothesis.

8. Acceptance gate F: upgrade and rollback rehearsal

Pin both old and new versions, capture release/security notes, back up first, upgrade in topology-aware order, verify application and Redis state, then rehearse the rollback path on disposable data. Current course baseline is Redis 8.10.1; the Docker Official Image also publishes 8.8.2 for older-version rehearsal, but future operators must re-check the exact supported tags and compatibility at execution time.

9. SLO and error-budget worksheet

Signal Target/threshold Measured Pass? Action if failed
Availability during game day team-defined SLO measure yes/no client/topology/failover repair
p99 latency at expected load team-defined SLO measure yes/no diagnose layer; do not tune blindly
Error rate/timeouts team-defined measure yes/no inspect client/server/network saturation
RPO business-defined measure acknowledged-write loss yes/no durability/replication design
RTO business-defined measure restore/recovery duration yes/no runbook/backup/topology improvement
Memory/fork/failover headroom capacity model measure yes/no reduce load or add capacity before release

10. Failure timeline template

Text · record one capstone game day
T0  baseline evidence captured; version/topology/security/workload recordedT1  representative load stable; p50/p95/p99/max and error rate capturedT2  failure injected into named/discovered current primaryT3  control plane detects failureT4  replacement primary/slot owner availableT5  application rediscovers service and successful writes resumeT6  old node repaired/rejoined; redundancy restoredT7  independent backup restored and invariants verifiedT8  post-event metrics compared with SLO/RPO/RTO and error budgetT9  rollback/cleanup completed; runbook corrections committed

11. Deliberately wrong capstone: “all commands returned OK”

Failure pattern

A command checklist can pass while the application has stale reads, broken retries, untested restore, unsafe ACLs, no failover-capable client, hot keys, or insufficient headroom. The repair is outcome-based acceptance: workload + observability + fault injection + restore + security + rollback with business invariants and measured timelines.

12. Production runbook: minimum sections

  • Architecture: node roles, ports, service discovery, slot/replication layout, dependencies, fault domains.
  • Security: ACL identities, TLS/network boundaries, secret sources/rotation, emergency access, audit evidence.
  • SLOs and capacity: workload contract, latency/error targets, RPO/RTO, memory/fork/failover headroom.
  • Observability: INFO/latency/slowlog/client/host metrics, alerts, dashboards, log correlation, escalation thresholds.
  • Failure procedures: current-primary discovery, failover/partition handling, stale-read policy, retry/idempotency boundaries.
  • Backup/restore: artifact locations, checksums, retention, restore commands, application invariant checks.
  • Change management: pinned versions, compatibility checks, topology-aware order, rollback boundary, post-change validation.
  • Ownership: who can execute each operation, who approves risk, and where incident evidence is stored.

Check your understanding

  1. Why is a standalone restart not a failover test?
  2. Why must backup restore be tested separately from replica/Sentinel/Cluster failover?
  3. What makes a benchmark useful for capacity?
  4. When is tuning complete?
  5. What is the capstone’s final artifact?
Review the answers

There is no alternate serving node or service-discovery transition; it only measures outage/restart/recovery.

Replication/high availability can propagate logical damage and does not prove independent disaster recovery.

A documented representative workload, SLO-bound tail/error measurements, correlated resource evidence, and preserved headroom.

When the hypothesized mechanism changes the expected metrics, the same acceptance test passes, and rollback remains defined—not when a single number improves.

A measured, executable production runbook plus evidence that another operator can use it to operate and recover the chosen Redis topology.

13. Course completion judgment

A production Redis operator should now be able to explain data-structure semantics, key lifecycle, atomicity, Streams, Search/JSON/vector/probabilistic/time-series features, client behavior, persistence, memory/eviction, replication, Sentinel, Cluster, security, application patterns, and observability as one system of tradeoffs. The final standard is not memorizing commands: it is preserving correctness while measuring latency, capacity, durability, availability, security, recovery, and cost under the workload that actually matters.

Summary and next step

Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform completes the Redis course path. Keep the capstone evidence, recovery notes, and operating assumptions together as the final production-readiness record.

Authoritative references

Shell · bounded cleanup
docker rm -f atlasmart-redis-ch24 2>/dev/null || truedocker volume rm atlasmart-redis-ch24-data 2>/dev/null || true# Do not run docker system prune; remove only resources named by this chapter.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.