Chapter 18 · Memory Management, Eviction, Expiration, Fragmentation, and OOM Prevention

Build a Memory Budget Including Data, Indexes, Replication Buffers, Fork Headroom, and Failover

Assemble a production memory budget that includes dataset/index memory, allocator/process overhead, client and replication/AOF buffers, fork/COW headroom, failover behavior, and container/OS safety margin.

Advanced180–250 minutescapacity budget, indexes, buffers, fork COW, failoverRedis Open Source 8.10.1Docker + redis-py 8.1.0Free/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart must choose a memory limit for a production-like Redis node before enabling replicas, Search indexes, and persistence. “Dataset is 20 GiB, so allocate 20 GiB” omits exactly the transients that appear during backup, synchronization, rewrite, failover, and workload spikes.

01

Construct a process memory budget from measured Redis and platform components.

02

Separate maxmemory from container/node RAM and excluded buffers.

03

Budget Search/index and integrated-data-type memory instead of counting only source values.

04

Use measured RDB/AOF COW and replication buffer signals as headroom inputs.

05

Explain replica maxmemory behavior and how promotion/failover changes the operating envelope.

Exact Chapter 18 lab boundary

The course baseline remains Redis Open Source 8.10.1 from pinned image redis:8.10.1. Memory-pressure experiments use dedicated disposable Chapter 18 containers, named volumes, and loopback-only host ports 6391–6395 so the shared Chapter 01 lab is never pushed toward eviction or out-of-memory behavior. The temporary academy-admin password is classroom-only. TLS is omitted only on loopback. Mandatory topology is standalone, logical DB 0. Each lesson states its own persistence, maxmemory, eviction, and container-memory settings. Lesson 5 uses a 48 MiB maxmemory node inside a 192 MiB container solely to demonstrate the measurement worksheet. AOF everysec is enabled and BGSAVE is permitted so actual local COW/buffer metrics can be collected; this ratio is not a production recommendation.

1. Start from the process envelope, not from maxmemory

A practical budget begins with the smallest hard ceiling: container memory limit, cgroup, VM RAM available to Redis, or managed-service database limit. From that ceiling reserve OS/platform requirements and transient Redis process headroom. Only then choose a maxmemory that leaves room for what eviction does not control.

Budget layer Evidence source Why it can grow
Dataset + Redis key/object overhead used_memory_dataset, MEMORY STATS, key samples Cardinality, value size, encodings, expirations.
Indexes/integrated features INFO modules/feature metrics, FT.INFO where applicable, measured MEMORY deltas Search/vector/time-series/probabilistic secondary structures.
Redis internal overhead used_memory_overhead, clients, scripts/functions, hash templates Connections, dictionaries, scripts, compact-hash templates.
Excluded buffers mem_not_counted_for_evict, mem_aof_buffer, replication metrics AOF/replica traffic and backlog.
Allocator/RSS slack allocator_* and used_memory_rss Fragmentation and resident allocator pages.
Fork/COW peak rdb_last_cow_size, aof_last_cow_size Writes during BGSAVE/rewrite/full synchronization.
Platform safety margin Container/OS metrics Kernel, libraries, monitoring agents, failure transients.

2. Build a measured budget fixture

Shell · create one isolated memory node
docker volume create atlasmart-redis-ch18-budget_datadocker run --rm -v atlasmart-redis-ch18-budget_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch18-budget --memory 192m -p 127.0.0.1:6395:6379 -v atlasmart-redis-ch18-budget_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --appendonly yes --appendfsync everysec --save '' --maxmemory 48mb --maxmemory-policy noevictiondocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin ACL WHOAMI
Python · representative bounded AtlasMart data
import os, redisHOST = "127.0.0.1"PORT = 6395r = redis.Redis(host=HOST, port=PORT, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=False)assert r.ping()payload = b"catalog:" + b"x" * 4096pipe = r.pipeline(transaction=False)for i in range(2500): pipe.set(f"atlasmart:ch18:l5:catalog:{i}", payload)pipe.execute()for i in range(200): r.hset(f"atlasmart:ch18:l5:inventory:{i}", mapping={b"available": b"100", b"reserved": b"0"})print("keys", r.dbsize())
redis-cli · capture a budget snapshot
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin MEMORY STATSdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin CONFIG GET maxmemory maxmemory-policy appendonly appendfsyncdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin DBSIZE

3. Index memory belongs in the budget even when source documents look small

Redis Search, vector indexes, JSON structures, time-series metadata, probabilistic structures, and compact-hash templates can add memory beyond the source payload. Do not estimate those structures from source-file size. Measure the feature actually used, including index build/rebuild peaks.

Redis 8.10 example

used_memory_hash_templates and MEMORY STATS hash.templates expose compact-hash template memory. Redis 8.10 MEMORY USAGE also allocates a proportional template share to each compact-hash key.

redis-cli · optional feature evidence
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO memory# If AtlasMart uses Redis Search, capture FT.INFO for each index and compare INFO MEMORY before/after representative indexing.# If vectors/time series/probabilistic structures are used, load production-like cardinality and measure their actual process footprint.

4. Measure fork/COW rather than reserving a folklore percentage

Chapter 17 showed that BGSAVE and AOF rewrite fork a child on the Linux-based container. Copy-on-write grows with pages modified while the child is alive. The correct headroom is workload- and duration-dependent.

redis-cli · bounded snapshot/COW observation
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin BGSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-budget redis-cli --user academy-admin INFO persistence# Wait until rdb_bgsave_in_progress returns 0, then record rdb_last_cow_size.# Repeat under representative write concurrency on a disposable environment; do not invent a percentage.
Metric Use in budget
rdb_last_cow_size Observed COW from last RDB background save.
aof_last_cow_size Observed COW from last AOF rewrite when applicable.
current_cow_size/current_cow_peak In-progress fork COW observation.
aof_rewrite_buffer_length / mem_aof_buffer where present Write amplification/buffer evidence during persistence work.

5. Replication/failover can change who must carry the memory budget

Redis replication is asynchronous and replicas have their own buffers and full-sync transients. Current Redis documentation also notes that replicas ignore maxmemory by default for normal replication behavior; after promotion, the new primary must enforce the configured policy while also handling new client/write pressure. That is why a primary-only steady-state budget is insufficient for Sentinel or Cluster failover planning.

Failover design question

Can every eligible replica become primary while remaining within its node/container memory envelope, including replication buffers, client reconnection bursts, persistence work, and the chosen maxmemory policy?

6. A reusable worksheet

Text · memory budget worksheet
Hard process/container/node envelope = ______ bytes- platform/OS/agent reserve             = ______- measured RSS/allocator slack peak      = ______- mem_not_counted_for_evict peak         = ______- client/output buffer peak              = ______- fork/COW peak (RDB/AOF/full sync)       = ______- failover/reconnect/sync contingency     = ______= maximum safe Redis operating envelope   = ______Choose maxmemory below that envelope, then verify dataset + index growth fits the chosen eviction semantics.
No universal headroom percentage

Percent rules are attractive because they are easy, but fork duration, write rate, allocator behavior, replica backlog, index build, client buffers, and container limits vary. Use measured peaks plus an explicit safety margin.

7. Export facts; calculate decisions outside Redis

The server exposes measurements, but the capacity decision belongs in an auditable worksheet, infrastructure-as-code variable set, or monitoring calculation. Keep the raw evidence with the chosen limit so future operators know why it exists.

Python · emit budget inputs without inventing headroom
import os, redisHOST = "127.0.0.1"PORT = 6395r = redis.Redis(host=HOST, port=PORT, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=False)assert r.ping()m = r.info("memory")persistence = r.info("persistence")replication = r.info("replication")fields = {    "used_memory": m.get("used_memory"),    "used_memory_dataset": m.get("used_memory_dataset"),    "used_memory_rss": m.get("used_memory_rss"),    "used_memory_overhead": m.get("used_memory_overhead"),    "mem_not_counted_for_evict": m.get("mem_not_counted_for_evict"),    "mem_clients_normal": m.get("mem_clients_normal"),    "mem_replication_backlog": m.get("mem_replication_backlog"),    "mem_aof_buffer": m.get("mem_aof_buffer"),    "allocator_frag_bytes": m.get("allocator_frag_bytes"),    "rdb_last_cow_size": persistence.get("rdb_last_cow_size"),    "aof_last_cow_size": persistence.get("aof_last_cow_size"),    "role": replication.get("role"),}for k,v in fields.items(): print(k, v)

8. Failure modes the budget must survive

Failure/mode What can spike Required evidence
Traffic burst Client buffers, hot-set growth, miss/repopulation writes Connection/output-buffer metrics, write rate, p95/p99.
BGSAVE/AOF rewrite Fork/COW and persistence buffers COW fields, rewrite duration, disk latency.
Replica full sync Replication backlog/buffers plus RDB/full-sync work INFO replication + memory during sync.
Sentinel/Cluster failover Promotion, reconnect storm, new primary writes Per-node memory before/after promotion.
Search/index rebuild Index memory plus build transient Feature/index metrics and RSS/used_memory delta.
Large delete/expiry wave Lazyfree backlog or synchronous free latency lazyfree_pending_objects and latency monitor.

9. Deliberately wrong budget: dataset bytes + 10%

A fixed percentage can accidentally work in one environment and fail in another. If the write rate during a fork doubles, a replica reconnects and needs a full synchronization, or a Search index rebuilds, the previously “safe” margin may disappear. The repair is to budget each transient from observed production-representative evidence and test the same node role during failover.

10. Acceptance criteria for an AtlasMart capacity decision

  • The hard process/container/node ceiling and every subtraction are documented.
  • maxmemory is lower than the hard ceiling by measured non-evictable/transient headroom.
  • Dataset and all index/data-type structures fit the intended eviction semantics at expected peak cardinality.
  • RDB/AOF/full-sync/failover COW and buffers were exercised on a production-representative disposable environment.
  • p50/p95/p99 latency, hit/miss rate, evictions, rejected writes, RSS, and allocator metrics remain within service objectives during pressure.
  • Every replica/failover candidate can assume the primary role without relying on the old primary's memory envelope.
  • Capacity alarms trigger before the kernel/container OOM killer becomes the control plane.
Shell · bounded cleanup
docker rm -f atlasmart-redis-ch18-budgetdocker volume rm atlasmart-redis-ch18-budget_data# Removes only this named Chapter 18 node and volume.

Check your understanding

  1. Why should maxmemory be lower than a container memory limit?
  2. Why must indexes be measured separately?
  3. Why does failover change the capacity question?
  4. What is the safest universal COW headroom percentage?
Review the answers

Because excluded replication/AOF buffers, allocator/RSS slack, client buffers, fork/COW, executable/library pages, and other process overhead can exist beyond the eviction budget.

Secondary indexes and integrated data types add real process memory beyond source values, and build/rebuild transients can increase peak usage.

A replica may become the write-serving primary, receive reconnect bursts, enforce maxmemory, and perform persistence/replication work under a different load mix.

There is none. Measure COW under representative dataset size, write rate, persistence/full-sync duration, and platform conditions.

Production judgment and next bridge

Chapter 18 ends with a memory budget that survives more than steady-state dataset growth. Chapter 19 now makes replication itself observable: initial sync, partial resynchronization, backlog sizing, replica reads, lag, and failure recovery will turn several “future buffer” lines in this worksheet into measured topology-specific values.

Summary and next step

Build a Memory Budget Including Data, Indexes, Replication Buffers, Fork Headroom, and Failover is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Primary/Replica Replication, Asynchronous Semantics, Initial Sync, and Command Propagation.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.