Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone
Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans
Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.
Redis Open Source 8.10.1 using the pinned
redis:8.10.1 image. The observability/benchmark
node is atlasmart-redis-ch24 on
127.0.0.1:6441, standalone topology, logical
database 0, AOF everysec plus RDB save rules,
maxmemory 0/noeviction unless a
bounded experiment says otherwise, named
academy-admin and atlasmart-app ACL
users, and fixture prefix atlasmart:ch24:*. TLS is
off only on this loopback-local disposable node; the capstone
security acceptance criteria reuse Chapter 22 TLS/ACL guidance.
Python examples target redis==8.1.0.
Search/JSON/vector/time-series/probabilistic features are
optional and must be included in capacity accounting only when
the chosen AtlasMart architecture actually uses them.
Learning outcomes
This lesson turns Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.
Explain the mechanisms and terminology behind Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans.
Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.
Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.
Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.
Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.
1. Problem: “we have persistence” is not the same as “we have a restorable backup”
AtlasMart can have RDB snapshots and AOF files on the same node and still lose them with the node, disk, operator, or ransomware incident. A backup is an independently retained artifact plus metadata and a demonstrated restore procedure. Recovery Point Objective (RPO) is the acceptable data-loss window; Recovery Time Objective (RTO) is the acceptable recovery duration.
2. Record the persistence and version contract before copying files
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG GET dir dbfilename appendonly appenddirname appendfsync
Record Redis/redis-cli version, image digest if your release process uses one, persistence settings, topology, ACL/TLS state, key counts, and an application-level fixture checksum. Without this metadata, a future restore may load but still violate application expectations.
3. RDB path: snapshot, copy, checksum, restore
RDB is a point-in-time snapshot. Trigger BGSAVE,
wait until rdb_bgsave_in_progress:0 and
rdb_last_bgsave_status:ok, then copy the completed
file. Do not copy a half-written temporary artifact.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-App-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user atlasmart-app SET atlasmart:ch24:l4:sentinel-value rdb-restore-proofdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin BGSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencemkdir -p ch24-backup-rdbdocker cp atlasmart-redis-ch24:/data/dump.rdb ch24-backup-rdb/dump.rdbpython -c "import hashlib,pathlib; p=pathlib.Path('ch24-backup-rdb/dump.rdb'); print(hashlib.sha256(p.read_bytes()).hexdigest(), p)"
4. Redis 8.10 MP-AOF BACKUP family: online start, seal, copy
Redis 8.10 introduces
BACKUP START/LIST/STATUS/SEAL/CLEANUP. The sealed
artifact is a self-contained multi-part AOF set: BASE + INCR +
manifest. The restored state reflects the seal boundary. This is
version-specific; older Redis releases require their documented
backup procedures.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin COMMAND INFO BACKUP# BACKUP requires backupdirname configured at startup; use the dedicated restore/backup node below for a fully reproducible 8.10 exercise.
backupdirname and preload-file are
startup-only settings. The lesson does not mutate the shared
Chapter 24 node into a different persistence layout.
5. Reproducible 8.10 backup node
docker volume create atlasmart-redis-ch24-backup-datadocker run -d --name atlasmart-redis-ch24-backup --restart no -p 127.0.0.1:6442:6379 -v atlasmart-redis-ch24-backup-data:/data redis:8.10.1 redis-server --dir /data --appendonly yes --appendfsync everysec --backupdirname backupdir --requirepass AtlasMart-Ch24-Backup-Lab-Only-2026docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli SET atlasmart:ch24:l4:backup:proof before-sealdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP STARTdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP STATUS# When status is incrementing, add one more write so the INCR portion is meaningful.docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli SET atlasmart:ch24:l4:backup:after-start included-at-sealdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP SEALdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP LISTmkdir -p ch24-backup-mpaofdocker cp atlasmart-redis-ch24-backup:/data/backupdir/. ch24-backup-mpaof/python -c "import hashlib,pathlib; [print(hashlib.sha256(p.read_bytes()).hexdigest(),p) for p in sorted(pathlib.Path('ch24-backup-mpaof').iterdir()) if p.is_file()]"
6. Restore test: proof is application state, not “Redis started”
A restore must run in a different disposable node or
environment. Validate Redis startup, key count, sentinel
fixtures, representative types/TTLs, application invariants,
ACL/TLS configuration separately, and elapsed recovery time. For
the Redis 8.10 sealed backup, use startup-only
preload-file aof:/restore/appendonly.aof.manifest.
docker run -d --name atlasmart-redis-ch24-restore --restart no -p 127.0.0.1:6443:6379 -v "$(pwd)/ch24-backup-mpaof:/restore:ro" redis:8.10.1 redis-server --preload-file aof:/restore/appendonly.aof.manifest --appendonly no --save "" --requirepass AtlasMart-Ch24-Restore-Lab-Only-2026docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Restore-Lab-Only-2026 atlasmart-redis-ch24-restore redis-cli GET atlasmart:ch24:l4:backup:proofdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Restore-Lab-Only-2026 atlasmart-redis-ch24-restore redis-cli GET atlasmart:ch24:l4:backup:after-startdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Restore-Lab-Only-2026 atlasmart-redis-ch24-restore redis-cli DBSIZE
The $(pwd) bind-mount form is POSIX-shell syntax.
In PowerShell, replace it with an absolute Windows path such
as ${PWD}\ch24-backup-mpaof:/restore:ro. The
backup bytes and restore semantics are identical; only
host-path syntax differs.
7. Upgrade is a compatibility change, not image replacement
As of this review, the Docker Official Image publishes both
redis:8.8.2 and redis:8.10.1. A useful
local rehearsal is therefore 8.8.2 → 8.10.1 on disposable data.
Record release notes/security advisories, persistence
compatibility, client compatibility, feature changes, and a
rollback point before replacing the binary. Never use the moving
redis:latest tag in an upgrade runbook.
| Topology | Safer order | Verification |
|---|---|---|
| Standalone | Backup → stop → replace binary → start | INFO server/persistence + data/client acceptance |
| Replication/Sentinel | Upgrade replicas first; preserve failover capacity; promote only through planned procedure | role/offsets, Sentinel discovery, app reconnect |
| Cluster | Upgrade replicas one at a time before primaries; keep slot redundancy/headroom | INFO server, CLUSTER INFO, redis-cli --cluster check, Search readiness if used |
8. Rollback plan must include data-format direction
“Run the old image again” is not automatically safe after a newer server has rewritten persistence files or activated newer features. The rollback plan must identify the exact previous binary, the backup artifact created before the upgrade, configuration/ACL/TLS files, and the point after which rollback requires restoring the pre-upgrade data instead of directly opening newer persistence artifacts. Re-check the exact version pair every time.
9. Deliberately wrong: backup exists but restore has never been timed
An untested file copy is evidence of storage, not recoverability. The repair is a scheduled restore drill into an isolated node, with checksum, key/invariant checks, client connection tests, and measured RTO. Keep backup credentials and encryption/key material in the disaster-recovery plan without embedding real secrets in the course.
Check your understanding
- Why is a replica not a backup?
- What state does an 8.10 sealed BACKUP restore represent?
- Why upgrade replicas before primaries in replicated/cluster topologies?
- When can rollback require restoring an older backup instead of only starting the old binary?
Review the answers
Replication can propagate deletion/corruption and shares the live failure domain; a backup is independently retained and restore-tested.
BASE plus INCR through the BACKUP SEAL boundary.
It preserves serving capacity and a safer failover path while one node changes.
When newer persistence/config/features are not safely backward-compatible with the old server.
10. Production judgment and bridge
Backup success is a restore SLO, not a green cron job. Upgrade success includes application behavior, replication/failover, security, observability, and rollback readiness. The final lesson converts all course mechanisms into one production acceptance runbook.
Summary and next step
Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform.
Authoritative references
docker rm -f atlasmart-redis-ch24-restore atlasmart-redis-ch24-backup 2>/dev/null || truedocker volume rm atlasmart-redis-ch24-backup-data 2>/dev/null || true# Keep ch24-backup-* directories only if you want to inspect the artifacts; otherwise remove them manually after the lesson.