Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone

Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans

Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.

Advanced210–320 minutesRDB, AOF, BACKUP, restore, upgrade, rollbackRedis Open Source 8.10.1redis-py 8.1.0 where Python is usedDocker + redis-cli + PythonStandalone evidence node · DB 0AOF everysec + RDB · maxmemory 0/noeviction baselineNamed ACL users · TLS off only on loopbackReuses Sentinel/Cluster labs for failover acceptanceFree/local-firstLast reviewed: September 6, 2026
Reproducible Chapter 24 baseline

Redis Open Source 8.10.1 using the pinned redis:8.10.1 image. The observability/benchmark node is atlasmart-redis-ch24 on 127.0.0.1:6441, standalone topology, logical database 0, AOF everysec plus RDB save rules, maxmemory 0/noeviction unless a bounded experiment says otherwise, named academy-admin and atlasmart-app ACL users, and fixture prefix atlasmart:ch24:*. TLS is off only on this loopback-local disposable node; the capstone security acceptance criteria reuse Chapter 22 TLS/ACL guidance. Python examples target redis==8.1.0. Search/JSON/vector/time-series/probabilistic features are optional and must be included in capacity accounting only when the chosen AtlasMart architecture actually uses them.

Learning outcomes

This lesson turns Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.

01

Explain the mechanisms and terminology behind Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans.

02

Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.

03

Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.

04

Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.

05

Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.

1. Problem: “we have persistence” is not the same as “we have a restorable backup”

AtlasMart can have RDB snapshots and AOF files on the same node and still lose them with the node, disk, operator, or ransomware incident. A backup is an independently retained artifact plus metadata and a demonstrated restore procedure. Recovery Point Objective (RPO) is the acceptable data-loss window; Recovery Time Objective (RTO) is the acceptable recovery duration.

2. Record the persistence and version contract before copying files

Shell · evidence before backup
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG GET dir dbfilename appendonly appenddirname appendfsync

Record Redis/redis-cli version, image digest if your release process uses one, persistence settings, topology, ACL/TLS state, key counts, and an application-level fixture checksum. Without this metadata, a future restore may load but still violate application expectations.

3. RDB path: snapshot, copy, checksum, restore

RDB is a point-in-time snapshot. Trigger BGSAVE, wait until rdb_bgsave_in_progress:0 and rdb_last_bgsave_status:ok, then copy the completed file. Do not copy a half-written temporary artifact.

Shell · create and copy a bounded RDB backup
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-App-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user atlasmart-app SET atlasmart:ch24:l4:sentinel-value rdb-restore-proofdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin BGSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencemkdir -p ch24-backup-rdbdocker cp atlasmart-redis-ch24:/data/dump.rdb ch24-backup-rdb/dump.rdbpython -c "import hashlib,pathlib; p=pathlib.Path('ch24-backup-rdb/dump.rdb'); print(hashlib.sha256(p.read_bytes()).hexdigest(), p)"

4. Redis 8.10 MP-AOF BACKUP family: online start, seal, copy

Redis 8.10 introduces BACKUP START/LIST/STATUS/SEAL/CLEANUP. The sealed artifact is a self-contained multi-part AOF set: BASE + INCR + manifest. The restored state reflects the seal boundary. This is version-specific; older Redis releases require their documented backup procedures.

Shell · version-gated 8.10 backup workflow
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin COMMAND INFO BACKUP# BACKUP requires backupdirname configured at startup; use the dedicated restore/backup node below for a fully reproducible 8.10 exercise.
Why a dedicated node?

backupdirname and preload-file are startup-only settings. The lesson does not mutate the shared Chapter 24 node into a different persistence layout.

5. Reproducible 8.10 backup node

Shell · disposable MP-AOF backup node
docker volume create atlasmart-redis-ch24-backup-datadocker run -d --name atlasmart-redis-ch24-backup --restart no -p 127.0.0.1:6442:6379 -v atlasmart-redis-ch24-backup-data:/data redis:8.10.1 redis-server --dir /data --appendonly yes --appendfsync everysec --backupdirname backupdir --requirepass AtlasMart-Ch24-Backup-Lab-Only-2026docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli SET atlasmart:ch24:l4:backup:proof before-sealdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP STARTdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP STATUS# When status is incrementing, add one more write so the INCR portion is meaningful.docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli SET atlasmart:ch24:l4:backup:after-start included-at-sealdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP SEALdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Backup-Lab-Only-2026 atlasmart-redis-ch24-backup redis-cli BACKUP LISTmkdir -p ch24-backup-mpaofdocker cp atlasmart-redis-ch24-backup:/data/backupdir/. ch24-backup-mpaof/python -c "import hashlib,pathlib; [print(hashlib.sha256(p.read_bytes()).hexdigest(),p) for p in sorted(pathlib.Path('ch24-backup-mpaof').iterdir()) if p.is_file()]"

6. Restore test: proof is application state, not “Redis started”

A restore must run in a different disposable node or environment. Validate Redis startup, key count, sentinel fixtures, representative types/TTLs, application invariants, ACL/TLS configuration separately, and elapsed recovery time. For the Redis 8.10 sealed backup, use startup-only preload-file aof:/restore/appendonly.aof.manifest.

Shell · restore the sealed MP-AOF artifact
docker run -d --name atlasmart-redis-ch24-restore --restart no -p 127.0.0.1:6443:6379 -v "$(pwd)/ch24-backup-mpaof:/restore:ro" redis:8.10.1 redis-server --preload-file aof:/restore/appendonly.aof.manifest --appendonly no --save "" --requirepass AtlasMart-Ch24-Restore-Lab-Only-2026docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Restore-Lab-Only-2026 atlasmart-redis-ch24-restore redis-cli GET atlasmart:ch24:l4:backup:proofdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Restore-Lab-Only-2026 atlasmart-redis-ch24-restore redis-cli GET atlasmart:ch24:l4:backup:after-startdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Restore-Lab-Only-2026 atlasmart-redis-ch24-restore redis-cli DBSIZE
Windows note

The $(pwd) bind-mount form is POSIX-shell syntax. In PowerShell, replace it with an absolute Windows path such as ${PWD}\ch24-backup-mpaof:/restore:ro. The backup bytes and restore semantics are identical; only host-path syntax differs.

7. Upgrade is a compatibility change, not image replacement

As of this review, the Docker Official Image publishes both redis:8.8.2 and redis:8.10.1. A useful local rehearsal is therefore 8.8.2 → 8.10.1 on disposable data. Record release notes/security advisories, persistence compatibility, client compatibility, feature changes, and a rollback point before replacing the binary. Never use the moving redis:latest tag in an upgrade runbook.

Topology Safer order Verification
Standalone Backup → stop → replace binary → start INFO server/persistence + data/client acceptance
Replication/Sentinel Upgrade replicas first; preserve failover capacity; promote only through planned procedure role/offsets, Sentinel discovery, app reconnect
Cluster Upgrade replicas one at a time before primaries; keep slot redundancy/headroom INFO server, CLUSTER INFO, redis-cli --cluster check, Search readiness if used

8. Rollback plan must include data-format direction

“Run the old image again” is not automatically safe after a newer server has rewritten persistence files or activated newer features. The rollback plan must identify the exact previous binary, the backup artifact created before the upgrade, configuration/ACL/TLS files, and the point after which rollback requires restoring the pre-upgrade data instead of directly opening newer persistence artifacts. Re-check the exact version pair every time.

9. Deliberately wrong: backup exists but restore has never been timed

Failure pattern

An untested file copy is evidence of storage, not recoverability. The repair is a scheduled restore drill into an isolated node, with checksum, key/invariant checks, client connection tests, and measured RTO. Keep backup credentials and encryption/key material in the disaster-recovery plan without embedding real secrets in the course.

Check your understanding

  1. Why is a replica not a backup?
  2. What state does an 8.10 sealed BACKUP restore represent?
  3. Why upgrade replicas before primaries in replicated/cluster topologies?
  4. When can rollback require restoring an older backup instead of only starting the old binary?
Review the answers

Replication can propagate deletion/corruption and shares the live failure domain; a backup is independently retained and restore-tested.

BASE plus INCR through the BACKUP SEAL boundary.

It preserves serving capacity and a safer failover path while one node changes.

When newer persistence/config/features are not safely backward-compatible with the old server.

10. Production judgment and bridge

Backup success is a restore SLO, not a green cron job. Upgrade success includes application behavior, replication/failover, security, observability, and rollback readiness. The final lesson converts all course mechanisms into one production acceptance runbook.

Summary and next step

Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Capstone: Design, Load-Test, Fail Over, Recover, Secure, Tune, and Operate a Production Redis Platform.

Authoritative references

Shell · cleanup only Chapter 24 backup/restore resources
docker rm -f atlasmart-redis-ch24-restore atlasmart-redis-ch24-backup 2>/dev/null || truedocker volume rm atlasmart-redis-ch24-backup-data 2>/dev/null || true# Keep ch24-backup-* directories only if you want to inspect the artifacts; otherwise remove them manually after the lesson.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.