Chapter 17 · Persistence: RDB Snapshots, AOF, fsync, Rewrites, and Crash Recovery

RDB Snapshots: Forking, Copy-on-Write, Save Rules, Startup Load, and Data-Loss Window

Make Redis Database snapshots observable from save rules and BGSAVE through fork/copy-on-write cost, file creation, startup loading, and the recoverable point-in-time window.

Advanced180–250 minutesRDB snapshots, fork/COW, save rules, startup recoveryRedis Open Source 8.10.1Docker local labFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart can tolerate losing a short interval of derived recommendation-cache changes after a crash, but operators need to know exactly what an RDB snapshot can recover. Redis Database (RDB) persistence is a point-in-time snapshot, not a continuous journal.

01

Explain save rules and why BGSAVE is normally preferable to foreground SAVE.

02

Trace fork and copy-on-write (COW) behavior and observe the corresponding INFO persistence fields.

03

Verify dump.rdb creation, LASTSAVE, startup loading, and the post-snapshot data-loss window.

04

Separate RDB persistence from replication, availability, and off-host backup.

05

Run a bounded restart/crash experiment on a dedicated local container and verify final state.

Exact Chapter 17 lab boundary

The normal course baseline remains Redis Open Source 8.10.1 from pinned image redis:8.10.1. Because this chapter intentionally restarts/crashes persistence processes, destructive exercises use dedicated disposable Chapter 17 containers and named volumes rather than the shared atlasmart-redis-ch01 container. All ports bind only to 127.0.0.1; the temporary lab password is public classroom data, not a production secret; TLS is omitted only on loopback. Mandatory topology is standalone, logical DB 0, no explicit maxmemory/eviction policy. This lesson uses a dedicated RDB-only node on host port 6381.

1. A snapshot is a recoverable point, not a write-by-write history

RDB captures the in-memory dataset at a point in time. A save rule such as save 60 1 means Redis may start a background snapshot after at least one change in roughly sixty seconds; it does not make each write durable. The recovery point objective (RPO) is therefore bounded by when a successful snapshot actually completed, not by when an application received an OK reply.

Do not confuse ACK with persistence

A successful SET proves Redis accepted the in-memory write. With RDB-only persistence it does not prove that write is already represented in dump.rdb.

2. Build an isolated RDB-only node

Shell · create one isolated persistence node
docker volume create atlasmart-redis-ch17-rdb_datadocker run --rm -v atlasmart-redis-ch17-rdb_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch17-rdb -p 127.0.0.1:6381:6379 -v atlasmart-redis-ch17-rdb_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --appendonly no --save 60 1 --dbfilename dump.rdbdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin ACL WHOAMI# users.acl lives in the named volume, so academy-admin remains available after the deliberate restart/crash.
redis-cli · record RDB configuration and baseline
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin CONFIG GET save appendonly dir dbfilenamedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin LASTSAVE

3. BGSAVE forks; the parent keeps serving clients

On the Linux container, BGSAVE forks a child. The child serializes the snapshot while the parent continues accepting commands. Linux virtual-memory pages are initially shared; when the parent changes a page, copy-on-write creates a private copy. That is why a write-heavy workload during a large snapshot can temporarily increase physical memory.

redis-cli · create deterministic state and snapshot
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin SET atlasmart:ch17:l1:before snapshot-v1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin HSET atlasmart:ch17:l1:order:1001 status paid total_cents 2599docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin BGSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin LASTSAVE
INFO field Interpretation
rdb_bgsave_in_progress 1 while the child is saving
current_cow_size / current_cow_peak COW memory while a child fork is active
rdb_last_cow_size COW memory consumed by the last RDB save
rdb_last_bgsave_status Whether the last background save succeeded
rdb_changes_since_last_save Writes since the last successful snapshot

4. Inspect the immutable finished artifact

Redis writes an RDB to a temporary file and atomically renames it to the configured dbfilename after completion. Once produced, the finished RDB is not modified in place, which is why copying a completed RDB is safer than copying a live AOF set without a consistency plan.

Shell · inspect only the dedicated volume
docker exec atlasmart-redis-ch17-rdb sh -lc 'ls -lh /data/dump.rdb && stat /data/dump.rdb'docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin LASTSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin INFO persistence

5. Prove the RDB data-loss window

After the successful snapshot, write a second key but do not trigger another snapshot. Then stop the process abruptly and start the same container again. Because the volume survives, Redis can load the previous dump.rdb; the later unsnapshotted key is expected to disappear. That is a controlled demonstration of RPO, not a claim about exact seconds of data loss in production.

Shell · controlled crash and restart
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin SET atlasmart:ch17:l1:after unsnapshotted-v2docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin GET atlasmart:ch17:l1:afterdocker kill atlasmart-redis-ch17-rdbdocker start atlasmart-redis-ch17-rdbsleep 1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin MGET atlasmart:ch17:l1:before atlasmart:ch17:l1:afterdocker logs --tail 40 atlasmart-redis-ch17-rdb# Expected: before survives; after may be absent because it was written after the last completed RDB.

6. Startup load evidence and integrity

After restart, inspect INFO persistence and logs. Fields such as rdb_last_load_keys_loaded and rdb_last_load_keys_expired describe the most recent RDB load. A successful load proves that Redis parsed this artifact; it still does not prove the file was copied off-host, encrypted, retained, or periodically restore-tested.

redis-cli · recovery evidence
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin DBSIZEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin GET atlasmart:ch17:l1:beforedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-rdb redis-cli --user academy-admin GET atlasmart:ch17:l1:after

7. Wrong approach: treat replicas or a local dump.rdb as backup

A replica normally reflects primary mutations and deletions, so application mistakes can propagate. A local RDB can also disappear with its host or volume. Persistence improves process-restart recovery; replication improves availability/read scaling; backup requires an independently retained artifact plus a tested restore path.

Three separate questions

Persistence: can this node reconstruct data after restart? Replication: can another node serve after a failure? Backup: can operators restore an independent historical copy after corruption, deletion, or site loss?

8. Verification checklist and cleanup

Verify the exact image/version, save rule, completed BGSAVE status, RDB file size, LASTSAVE change, expected key comparison after restart, and cleanup of only this node.

Shell · bounded cleanup
docker rm -f atlasmart-redis-ch17-rdbdocker volume rm atlasmart-redis-ch17-rdb_data# Removes only the named Chapter 17 container/volume. Never use docker system prune for this lab.

9. Production judgment

RDB is attractive when compact snapshots, fast restart, portable point-in-time files, and bounded data loss fit the workload. Size the node for fork latency and COW headroom; schedule snapshots with storage/CPU contention in mind; monitor latest_fork_usec, RDB duration/status, COW bytes, disk space, and snapshot age. For tighter RPO, pair RDB with AOF or another durable system rather than pretending a more aggressive save rule is free.

Check your understanding

  1. Does SET returning OK mean an RDB-only write is on disk?
  2. Why can BGSAVE increase RSS during heavy writes?
  3. What does LASTSAVE represent?
  4. Is a replica a backup?
  5. What should be measured before reducing the save interval?
Review the answers

No. It is durable only after a successful snapshot includes it.

The child and parent initially share pages; parent writes trigger copy-on-write page copies.

The Unix timestamp of the last successful RDB save, not the time of the newest in-memory write.

No. Replication and backup address different failure classes.

Snapshot duration, fork latency, COW memory, disk throughput/headroom, workload write rate, and acceptable RPO.

10. Summary and bridge

RDB gives compact point-in-time recovery with an explicit window between successful snapshots. Lesson 2 moves to AOF, where each write is logged but durability still depends on the fsync policy and storage behavior.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.