Chapter 18 · Memory Management, Eviction, Expiration, Fragmentation, and OOM Prevention

Big Keys, Hot Keys, Large Collections, Lazy Freeing, Active Defrag, and Memory Diagnostics

Diagnose big and hot keys, large collections, asynchronous freeing, allocator fragmentation, and active defragmentation without turning diagnostics into destructive tuning.

Advanced180–250 minutesbig/hot keys, lazy free, active defrag, diagnosticsRedis Open Source 8.10.1Docker + redis-py 8.1.0Free/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart is within its total memory budget, yet one shard/client path experiences latency spikes. Aggregate memory can look healthy while a single giant collection, hot key, slow free, or fragmented allocator causes operational pain.

01

Distinguish a big key from a hot key and explain why either can concentrate cost.

02

Use MEMORY USAGE plus redis-cli keyspace analyzers without KEYS/production-wide blocking scans.

03

Explain DEL versus UNLINK and observe lazyfree_pending_objects.

04

Interpret MEMORY DOCTOR, allocator metrics, and MEMORY PURGE cautiously.

05

Treat active defragmentation as measured CPU/memory tradeoff, not a blanket fix.

Exact Chapter 18 lab boundary

The course baseline remains Redis Open Source 8.10.1 from pinned image redis:8.10.1. Memory-pressure experiments use dedicated disposable Chapter 18 containers, named volumes, and loopback-only host ports 6391–6395 so the shared Chapter 01 lab is never pushed toward eviction or out-of-memory behavior. The temporary academy-admin password is classroom-only. TLS is omitted only on loopback. Mandatory topology is standalone, logical DB 0. Each lesson states its own persistence, maxmemory, eviction, and container-memory settings. Lesson 4 uses a 64 MiB maxmemory / allkeys-lfu node so --hotkeys is available. Persistence is disabled. The container limit is 192 MiB; all data is disposable.

1. Big and hot are different failure shapes

Shape What it means Typical pressure
Big key One key contains many elements or many bytes. Long O(N) operations, large replies, slow deletion/freeing, network and memory spikes.
Hot key One key receives disproportionate request rate. CPU/event-loop concentration, one Cluster slot/shard becoming the bottleneck.
Many small keys Large cardinality with per-key/object overhead. Metadata/hashtable overhead and allocator pressure.
Fragmentation Allocator active/resident pages exceed current allocations. RSS remains high even after logical data shrinks.

A key can be both big and hot, but the remediation differs. Splitting a 10-million-member collection addresses size; local caching/read replication or redesign may address read hotness.

2. Build a bounded diagnostic dataset

Shell · create one isolated memory node
docker volume create atlasmart-redis-ch18-diagnostics_datadocker run --rm -v atlasmart-redis-ch18-diagnostics_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch18-diagnostics --memory 192m -p 127.0.0.1:6394:6379 -v atlasmart-redis-ch18-diagnostics_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --save '' --appendonly no --maxmemory 64mb --maxmemory-policy allkeys-lfudocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin ACL WHOAMI
Python · create many small keys plus one bounded large string/list
import os, redisHOST = "127.0.0.1"PORT = 6394r = redis.Redis(host=HOST, port=PORT, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=False)assert r.ping()pipe = r.pipeline(transaction=False)for i in range(1500): pipe.set(f"atlasmart:ch18:l4:small:{i}", b"s"*1024)pipe.execute()r.set("atlasmart:ch18:l4:large:string", b"L" * (2 * 1024 * 1024))for start in range(0, 20000, 1000):    r.rpush("atlasmart:ch18:l4:large:list", *[f"item-{i}" for i in range(start,start+1000)])for _ in range(500): r.get("atlasmart:ch18:l4:small:0")print("done")

3. Measure individual keys before redesigning them

redis-cli · exact key-level evidence
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY USAGE atlasmart:ch18:l4:large:stringdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin STRLEN atlasmart:ch18:l4:large:stringdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY USAGE atlasmart:ch18:l4:large:list SAMPLES 0docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin LLEN atlasmart:ch18:l4:large:listdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin OBJECT FREQ atlasmart:ch18:l4:small:0

MEMORY USAGE estimates nested aggregate memory using five samples by default; SAMPLES 0 examines all nested elements and can therefore be more expensive on a genuinely huge collection. Use it deliberately.

4. Use redis-cli SCAN-based analyzers, not KEYS

--bigkeys, --memkeys, and newer --keystats scan the keyspace incrementally. --hotkeys is available only when the configured policy is LFU.

Shell · bounded prefix analysis
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin --memkeys --pattern "atlasmart:ch18:l4:*" --count 100docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin --bigkeys --pattern "atlasmart:ch18:l4:*" --count 100docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin --hotkeys --pattern "atlasmart:ch18:l4:*" --count 100
Operational caution

These modes use SCAN rather than KEYS, but scanning a large production keyspace still consumes CPU/network and may traverse the full dataset. Rate-limit and schedule diagnostics according to your environment.

5. UNLINK moves reclamation off the main deletion path

DEL synchronously frees the value. UNLINK removes keys from the keyspace and lets background lazy-free workers reclaim memory. That can reduce main-thread deletion cost for large values, but it does not make the work disappear; CPU and memory remain consumed until reclamation catches up.

redis-cli · observe asynchronous freeing
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin UNLINK atlasmart:ch18:l4:large:listdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin EXISTS atlasmart:ch18:l4:large:listdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memory# Compare lazyfree_pending_objects and lazyfreed_objects. Do not expect RSS to drop immediately.

6. MEMORY DOCTOR is a hint engine, not a diagnosis substitute

redis-cli · memory diagnostic layers
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY DOCTORdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY STATSdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin CONFIG GET activedefrag active-defrag-ignore-bytes active-defrag-threshold-lower active-defrag-threshold-upper active-defrag-cycle-min active-defrag-cycle-max

A small lab may return that there is too little data to diagnose, which is a valid result. Production decisions should use absolute fragmentation bytes, allocator ratios, RSS history, workload phase, and latency—not a single advisory string.

7. Active defragmentation trades CPU for allocator compaction

When supported and enabled, Redis active defragmentation moves allocations so allocator pages can become reclaimable. active_defrag_running exposes current activity/target CPU percentage. Enabling or tuning it without measured fragmentation can add CPU work without meaningful benefit.

Safe lab choice

This mandatory lab inspects active-defrag configuration and metrics but does not force it on. If you test it, do so only on an isolated load-representative node, record allocator fragmentation bytes and p50/p95/p99 latency before/after, and restore the original configuration.

redis-cli · allocator-specific optional observation
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY PURGEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memory# MEMORY PURGE is allocator-dependent and can itself be slow; it is not a periodic production maintenance recipe.

8. Big/hot-key repair patterns

Problem Safer design direction Tradeoff to measure
Large list/hash/zset Partition by bounded business dimension or time bucket. More keys/coordination and changed access patterns.
Hot read key Application-local cache, replica reads where staleness is acceptable, or data fan-out. Freshness and invalidation complexity.
Hot write key Shard the logical counter/workload if semantics allow, or redesign contention point. Aggregation/ordering complexity.
Slow deletion UNLINK or lazyfree configuration after measurement. Background CPU/memory backlog.
Fragmentation Allow allocator reuse; consider defrag/purge only when supported and justified. CPU spikes and incomplete RSS reduction.

9. Verification checklist

  • Identify key size using both logical length/cardinality and MEMORY USAGE.
  • Use SCAN-based analyzers with a bounded prefix in the lab.
  • Confirm --hotkeys requires LFU policy.
  • Observe lazyfree metrics after UNLINK.
  • Do not force active defrag merely because mem_fragmentation_ratio is high.
  • Record absolute fragmentation bytes and tail latency before any tuning.
Shell · bounded cleanup
docker rm -f atlasmart-redis-ch18-diagnosticsdocker volume rm atlasmart-redis-ch18-diagnostics_data# Removes only this named Chapter 18 node and volume.

Check your understanding

  1. What is the difference between a big key and a hot key?
  2. Why can UNLINK improve latency without reducing total work?
  3. Why is --hotkeys policy-dependent?
  4. Should active defrag be enabled whenever RSS exceeds used_memory?
Review the answers

A big key is large in bytes/elements; a hot key is accessed disproportionately often. One key can be both.

It removes the key immediately but shifts value reclamation to background lazy-free workers; CPU/memory work still occurs.

It relies on LFU frequency metadata, so redis-cli supports it only with volatile-lfu or allkeys-lfu.

No. First distinguish allocator fragmentation from other RSS overhead and measure absolute bytes and latency; active defrag costs CPU.

Production judgment and next bridge

Memory pathologies are about shape and traffic as much as total bytes. Lesson 5 assembles the measurements from Chapters 17–18 into one process/container capacity budget that remains valid during forks, replication/failover, and index growth.

Summary and next step

Big Keys, Hot Keys, Large Collections, Lazy Freeing, Active Defrag, and Memory Diagnostics is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Build a Memory Budget Including Data, Indexes, Replication Buffers, Fork Headroom, and Failover.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.