Chapter 23 · Application Patterns: Caching, Rate Limiting, Locks, Idempotency, and Hot-Key Control
Hot Keys, Fan-Out, Replication Read Scaling, Sharding, Local Caches, and Workload Redesign
Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.
Redis Open Source 8.10.1 using the pinned
redis:8.10.1 image, exposed only at
127.0.0.1:6431. Standalone topology, logical
database 0, AOF everysec plus an RDB save rule,
maxmemory 0 unless a lesson explicitly changes a
setting on this disposable node, default ACL user disabled,
named academy-admin and
atlasmart-app users, and fixture prefix
atlasmart:ch23:*. TLS is intentionally off only
because this mandatory lab is loopback-local; production traffic
must follow Chapter 22 network/TLS guidance. Python examples
target redis==8.1.0.
Search/JSON/vector/time-series/probabilistic features are not
required.
Learning outcomes
This lesson turns Hot Keys, Fan-Out, Replication Read Scaling, Sharding, Local Caches, and Workload Redesign into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.
Explain the mechanisms and terminology behind Hot Keys, Fan-Out, Replication Read Scaling, Sharding, Local Caches, and Workload Redesign.
Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.
Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.
Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.
Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.
1. Problem: the database is fine, Redis is fine, one key is not
AtlasMart’s flash-sale product receives 80% of catalog reads. Aggregate CPU and memory can look acceptable while a single key, hash slot, or primary absorbs disproportionate command/network traffic. A hot key is about access concentration; a big key is about value/collection size. One key can be either, both, or neither.
2. Measure concentration before changing architecture
Instrument the application because it knows logical key/request
identity with minimal Redis overhead. For LFU-aware Redis
observation, OBJECT FREQ exposes the probabilistic
access counter only under LFU policies, and
redis-cli --hotkeys is similarly LFU-dependent.
MONITOR can be intrusive and is not the first
choice on a busy production server.
3. Generate a deterministic 80/20 workload and compute top-key share
This harness does not invent a latency threshold. It prints the actual per-key distribution and latency observed on the learner’s node.
from collections import Counterimport random, statistics, timeimport redisr=redis.Redis(host="127.0.0.1", port=6431, username="atlasmart-app", password="AtlasMart-Ch23-App-Lab-Only-2026", decode_responses=True)random.seed(2305)keys=["atlasmart:ch23:hot:sku-42"]+[f"atlasmart:ch23:hot:sku-{i}" for i in range(100,120)]for k in keys: r.set(k, "x"*256)requests=[]lat=[]for _ in range(2000): k=keys[0] if random.random()<0.80 else random.choice(keys[1:]) t0=time.perf_counter_ns(); r.get(k); lat.append((time.perf_counter_ns()-t0)/1e6); requests.append(k)c=Counter(requests)print("top5", c.most_common(5))print("top_key_share", c[keys[0]]/len(requests))print("unique_keys", len(c))print("latency_ms_p50", statistics.median(lat), "max", max(lat))
4. Optional disposable LFU evidence
Because the Chapter 23 node is isolated, Lesson 5 may
temporarily switch only this node to allkeys-lfu to
populate/access LFU counters. maxmemory remains 0,
so no eviction is triggered merely by choosing the policy.
Record the old policy and restore it afterward.
OLD=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user academy-admin --raw CONFIG GET maxmemory-policy | tail -1)docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user academy-admin CONFIG SET maxmemory-policy allkeys-lfu# Re-run the Python workload several times, then:docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user academy-admin OBJECT FREQ atlasmart:ch23:hot:sku-42docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user academy-admin --hotkeys -i 0.02docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user academy-admin CONFIG SET maxmemory-policy "$OLD"
5. Fan-out multiplies read pressure
One user request can fan out into many Redis reads—for example, a product page reading product, stock badge, recommendations, promotions, reviews, and personalization separately. Reduce unnecessary round trips with appropriate data modeling/pipelining, but do not combine unrelated lifecycle/security domains into a giant key just to reduce command count.
6. Replication can scale some reads, but it does not split hot writes
Read replicas can absorb stale-tolerant reads, but Redis replication is asynchronous. A correctness-sensitive read after write may need the primary or explicit consistency handling. All writes still converge on the key’s primary, so adding replicas does not solve a hot write key. Failover also changes where the hot workload lands; reserve headroom on every promotion candidate.
7. Sharding only helps if the business key can actually be partitioned
Redis Cluster maps each key to one slot and one primary. One indivisible hot key remains on one primary. If the data semantics allow, redesign one aggregate into independently addressable shards—such as counters by region/bucket—with an explicit aggregation path. Avoid global hash tags that force unrelated hot keys into one slot. Sharding changes atomicity and read complexity, so prove the new invariants.
8. A tiny application-local cache can remove read heat
For read-mostly data with a short accepted stale window, a per-process local cache avoids even the Redis network hop. It reduces global Redis load but introduces per-process copies and invalidation/freshness complexity. Start with a short TTL and measure stale risk before adding pub/sub invalidations.
import timeclass TinyTTLCache: def __init__(self, ttl=0.25): self.ttl=ttl; self.d={} def get(self,key,loader): now=time.monotonic(); row=self.d.get(key) if row and row[0] > now: return row[1], True value=loader(key); self.d[key]=(now+self.ttl,value); return value, False
9. Deliberately wrong: “add replicas” to fix a hot write counter
Replicas can serve selected reads but cannot distribute writes to the authoritative key. Repair may require partitioning the logical counter, changing the write path, buffering/aggregating updates, or moving the invariant to a system designed for that write pattern. Measure write/read ratio and per-key command concentration before choosing.
10. Hot-key design matrix
Different causes need different repairs.
| Observed cause | Possible response | New cost/boundary |
|---|---|---|
| Read-mostly immutable-ish key | Short application-local cache | Stale window and invalidation complexity. |
| Read-heavy, stale-tolerant | Replica reads | Replication lag and failover routing. |
| Many commands per request | Pipeline/model redesign | Larger payloads/coupling; no atomicity by itself. |
| One hot aggregate write | Logical sharding/bucketing | Aggregation and cross-shard atomicity tradeoffs. |
| Cluster hot slot from hash tags | Revisit key/hash-tag model | May lose multi-key locality. |
Check your understanding
- What is the difference between a hot key and a big key?
- Will adding read replicas solve one hot write key?
- Why measure application-side per-key counts?
- What new risk comes with a local in-process cache?
Review the answers
Hot refers to access concentration; big refers to memory/value or collection size. They are independent dimensions.
No. Writes still target the primary owner of that key.
They expose logical access distribution directly without requiring intrusive server tracing.
Independent stale copies and invalidation/freshness behavior across application instances.
11. Production judgment and chapter close
Hot-key control starts with measurement: top-key share, read/write ratio, command type, payload bytes, p95/p99 latency, CPU/network on the owning node, replica lag, and Cluster slot distribution. Prefer workload redesign over infrastructure-only scaling when writes concentrate on one logical key. Chapter 24 turns these patterns into a broader operational discipline with INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, benchmarking, backup/restore, and upgrade evidence.
Summary and next step
Hot Keys, Fan-Out, Replication Read Scaling, Sharding, Local Caches, and Workload Redesign is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics.
Authoritative references
- Redis cache-aside
- Redis cache-aside with redis-py
- Redis rate limiter
- Rate limiter with redis-py
- INCR — rate limiter pattern
- SET
- DELEX
- Distributed locks with Redis
- EVAL
- TIME
- EXPIRE
- PTTL
- ZADD
- ZREMRANGEBYSCORE
- HSET
- HGETALL
- OBJECT FREQ
- redis-cli
- Redis eviction reference
- Redis replication
- Redis Cluster specification
- Redis 8.10 release notes
- Redis licenses
- redis-py 8.1.0
12. Cleanup
Restore any temporary policy before removing the isolated node, then clean only Chapter 23 resources.
docker rm -f atlasmart-redis-ch23 2>/dev/null || truedocker volume rm atlasmart-redis-ch23-data 2>/dev/null || true# Remove only the local ch23-lab directory after keeping any evidence you need.# Never use docker system prune as Chapter 23 cleanup.