Chapter 18 · Memory Management, Eviction, Expiration, Fragmentation, and OOM Prevention
Build a Memory Budget Including Data, Indexes, Replication Buffers, Fork Headroom, and Failover
Assemble a production memory budget that includes dataset/index memory, allocator/process overhead, client and replication/AOF buffers, fork/COW headroom, failover behavior, and container/OS safety margin.
Learning outcomes
AtlasMart must choose a memory limit for a production-like Redis node before enabling replicas, Search indexes, and persistence. “Dataset is 20 GiB, so allocate 20 GiB” omits exactly the transients that appear during backup, synchronization, rewrite, failover, and workload spikes.
Construct a process memory budget from measured Redis and platform components.
Separate maxmemory from container/node RAM and excluded buffers.
Budget Search/index and integrated-data-type memory instead of counting only source values.
Use measured RDB/AOF COW and replication buffer signals as headroom inputs.
Explain replica maxmemory behavior and how promotion/failover changes the operating envelope.
The course baseline remains Redis Open Source
8.10.1 from pinned image
redis:8.10.1. Memory-pressure experiments use
dedicated disposable Chapter 18 containers, named volumes, and
loopback-only host ports 6391–6395 so the shared Chapter 01
lab is never pushed toward eviction or out-of-memory behavior.
The temporary academy-admin password is
classroom-only. TLS is omitted only on loopback. Mandatory
topology is standalone, logical DB 0. Each lesson states its
own persistence, maxmemory, eviction, and container-memory
settings. Lesson 5 uses a 48 MiB maxmemory node inside a 192
MiB container solely to demonstrate the measurement worksheet.
AOF everysec is enabled and BGSAVE is permitted so actual
local COW/buffer metrics can be collected; this ratio is not a
production recommendation.
1. Start from the process envelope, not from maxmemory
A practical budget begins with the smallest hard ceiling:
container memory limit, cgroup, VM RAM available to Redis, or
managed-service database limit. From that ceiling reserve
OS/platform requirements and transient Redis process headroom.
Only then choose a maxmemory that leaves room for
what eviction does not control.
| Budget layer | Evidence source | Why it can grow |
|---|---|---|
| Dataset + Redis key/object overhead | used_memory_dataset, MEMORY STATS, key samples | Cardinality, value size, encodings, expirations. |
| Indexes/integrated features | INFO modules/feature metrics, FT.INFO where applicable, measured MEMORY deltas | Search/vector/time-series/probabilistic secondary structures. |
| Redis internal overhead | used_memory_overhead, clients, scripts/functions, hash templates | Connections, dictionaries, scripts, compact-hash templates. |
| Excluded buffers | mem_not_counted_for_evict, mem_aof_buffer, replication metrics | AOF/replica traffic and backlog. |
| Allocator/RSS slack | allocator_* and used_memory_rss | Fragmentation and resident allocator pages. |
| Fork/COW peak | rdb_last_cow_size, aof_last_cow_size | Writes during BGSAVE/rewrite/full synchronization. |
| Platform safety margin | Container/OS metrics | Kernel, libraries, monitoring agents, failure transients. |
2. Build a measured budget fixture
docker volume create atlasmart-redis-ch18-budget_datadocker run --rm -v atlasmart-redis-ch18-budget_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch18-budget --memory 192m -p 127.0.0.1:6395:6379 -v atlasmart-redis-ch18-budget_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --appendonly yes --appendfsync everysec --save '' --maxmemory 48mb --maxmemory-policy noevictiondocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin ACL WHOAMI
import os, redisHOST = "127.0.0.1"PORT = 6395r = redis.Redis(host=HOST, port=PORT, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=False)assert r.ping()payload = b"catalog:" + b"x" * 4096pipe = r.pipeline(transaction=False)for i in range(2500): pipe.set(f"atlasmart:ch18:l5:catalog:{i}", payload)pipe.execute()for i in range(200): r.hset(f"atlasmart:ch18:l5:inventory:{i}", mapping={b"available": b"100", b"reserved": b"0"})print("keys", r.dbsize())
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin MEMORY STATSdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin CONFIG GET maxmemory maxmemory-policy appendonly appendfsyncdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin DBSIZE
3. Index memory belongs in the budget even when source documents look small
Redis Search, vector indexes, JSON structures, time-series metadata, probabilistic structures, and compact-hash templates can add memory beyond the source payload. Do not estimate those structures from source-file size. Measure the feature actually used, including index build/rebuild peaks.
used_memory_hash_templates and
MEMORY STATS hash.templates expose compact-hash
template memory. Redis 8.10 MEMORY USAGE also
allocates a proportional template share to each compact-hash
key.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO memory# If AtlasMart uses Redis Search, capture FT.INFO for each index and compare INFO MEMORY before/after representative indexing.# If vectors/time series/probabilistic structures are used, load production-like cardinality and measure their actual process footprint.
4. Measure fork/COW rather than reserving a folklore percentage
Chapter 17 showed that BGSAVE and AOF rewrite fork a child on the Linux-based container. Copy-on-write grows with pages modified while the child is alive. The correct headroom is workload- and duration-dependent.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin BGSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO persistence# Wait until rdb_bgsave_in_progress returns 0, then record rdb_last_cow_size.# Repeat under representative write concurrency on a disposable environment; do not invent a percentage.
| Metric | Use in budget |
|---|---|
| rdb_last_cow_size | Observed COW from last RDB background save. |
| aof_last_cow_size | Observed COW from last AOF rewrite when applicable. |
| current_cow_size/current_cow_peak | In-progress fork COW observation. |
| aof_rewrite_buffer_length / mem_aof_buffer where present | Write amplification/buffer evidence during persistence work. |
5. Replication/failover can change who must carry the memory budget
Redis replication is asynchronous and replicas have their own
buffers and full-sync transients. Current Redis documentation
also notes that replicas ignore maxmemory by
default for normal replication behavior; after promotion, the
new primary must enforce the configured policy while also
handling new client/write pressure. That is why a primary-only
steady-state budget is insufficient for Sentinel or Cluster
failover planning.
Can every eligible replica become primary while remaining within its node/container memory envelope, including replication buffers, client reconnection bursts, persistence work, and the chosen maxmemory policy?
6. A reusable worksheet
Hard process/container/node envelope = ______ bytes- platform/OS/agent reserve = ______- measured RSS/allocator slack peak = ______- mem_not_counted_for_evict peak = ______- client/output buffer peak = ______- fork/COW peak (RDB/AOF/full sync) = ______- failover/reconnect/sync contingency = ______= maximum safe Redis operating envelope = ______Choose maxmemory below that envelope, then verify dataset + index growth fits the chosen eviction semantics.
Percent rules are attractive because they are easy, but fork duration, write rate, allocator behavior, replica backlog, index build, client buffers, and container limits vary. Use measured peaks plus an explicit safety margin.
7. Export facts; calculate decisions outside Redis
The server exposes measurements, but the capacity decision belongs in an auditable worksheet, infrastructure-as-code variable set, or monitoring calculation. Keep the raw evidence with the chosen limit so future operators know why it exists.
import os, redisHOST = "127.0.0.1"PORT = 6395r = redis.Redis(host=HOST, port=PORT, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=False)assert r.ping()m = r.info("memory")persistence = r.info("persistence")replication = r.info("replication")fields = { "used_memory": m.get("used_memory"), "used_memory_dataset": m.get("used_memory_dataset"), "used_memory_rss": m.get("used_memory_rss"), "used_memory_overhead": m.get("used_memory_overhead"), "mem_not_counted_for_evict": m.get("mem_not_counted_for_evict"), "mem_clients_normal": m.get("mem_clients_normal"), "mem_replication_backlog": m.get("mem_replication_backlog"), "mem_aof_buffer": m.get("mem_aof_buffer"), "allocator_frag_bytes": m.get("allocator_frag_bytes"), "rdb_last_cow_size": persistence.get("rdb_last_cow_size"), "aof_last_cow_size": persistence.get("aof_last_cow_size"), "role": replication.get("role"),}for k,v in fields.items(): print(k, v)
8. Failure modes the budget must survive
| Failure/mode | What can spike | Required evidence |
|---|---|---|
| Traffic burst | Client buffers, hot-set growth, miss/repopulation writes | Connection/output-buffer metrics, write rate, p95/p99. |
| BGSAVE/AOF rewrite | Fork/COW and persistence buffers | COW fields, rewrite duration, disk latency. |
| Replica full sync | Replication backlog/buffers plus RDB/full-sync work | INFO replication + memory during sync. |
| Sentinel/Cluster failover | Promotion, reconnect storm, new primary writes | Per-node memory before/after promotion. |
| Search/index rebuild | Index memory plus build transient | Feature/index metrics and RSS/used_memory delta. |
| Large delete/expiry wave | Lazyfree backlog or synchronous free latency | lazyfree_pending_objects and latency monitor. |
9. Deliberately wrong budget: dataset bytes + 10%
A fixed percentage can accidentally work in one environment and fail in another. If the write rate during a fork doubles, a replica reconnects and needs a full synchronization, or a Search index rebuilds, the previously “safe” margin may disappear. The repair is to budget each transient from observed production-representative evidence and test the same node role during failover.
10. Acceptance criteria for an AtlasMart capacity decision
- The hard process/container/node ceiling and every subtraction are documented.
-
maxmemoryis lower than the hard ceiling by measured non-evictable/transient headroom. - Dataset and all index/data-type structures fit the intended eviction semantics at expected peak cardinality.
- RDB/AOF/full-sync/failover COW and buffers were exercised on a production-representative disposable environment.
- p50/p95/p99 latency, hit/miss rate, evictions, rejected writes, RSS, and allocator metrics remain within service objectives during pressure.
- Every replica/failover candidate can assume the primary role without relying on the old primary's memory envelope.
- Capacity alarms trigger before the kernel/container OOM killer becomes the control plane.
docker rm -f atlasmart-redis-ch18-budgetdocker volume rm atlasmart-redis-ch18-budget_data# Removes only this named Chapter 18 node and volume.
Check your understanding
- Why should maxmemory be lower than a container memory limit?
- Why must indexes be measured separately?
- Why does failover change the capacity question?
- What is the safest universal COW headroom percentage?
Review the answers
Because excluded replication/AOF buffers, allocator/RSS slack, client buffers, fork/COW, executable/library pages, and other process overhead can exist beyond the eviction budget.
Secondary indexes and integrated data types add real process memory beyond source values, and build/rebuild transients can increase peak usage.
A replica may become the write-serving primary, receive reconnect bursts, enforce maxmemory, and perform persistence/replication work under a different load mix.
There is none. Measure COW under representative dataset size, write rate, persistence/full-sync duration, and platform conditions.
Production judgment and next bridge
Chapter 18 ends with a memory budget that survives more than steady-state dataset growth. Chapter 19 now makes replication itself observable: initial sync, partial resynchronization, backlog sizing, replica reads, lag, and failure recovery will turn several “future buffer” lines in this worksheet into measured topology-specific values.
Summary and next step
Build a Memory Budget Including Data, Indexes, Replication Buffers, Fork Headroom, and Failover is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Primary/Replica Replication, Asynchronous Semantics, Initial Sync, and Command Propagation.
Authoritative references
- Redis key eviction
- Redis memory optimization
- INFO
- MEMORY STATS
- MEMORY USAGE
- MEMORY DOCTOR
- MEMORY PURGE
- MEMORY MALLOC-STATS
- OBJECT FREQ
- OBJECT IDLETIME
- UNLINK
- Redis CLI keyspace analysis
- Redis latency monitoring
- Redis persistence
- Redis replication
- Redis Sentinel
- Redis Cluster specification
- Redis Search and query
- Redis 8.10 release notes
- Redis 8.10 whats new
- Redis ACLs
- Redis security
- Redis licenses
- Docker Official Redis image