Chapter 23 · Application Patterns: Caching, Rate Limiting, Locks, Idempotency, and Hot-Key Control

Cache-Aside with TTL/Jitter, Negative Caching, Stampede Protection, and Stale-While-Revalidate

Build a cache-aside read path with explicit freshness bounds, randomized expiration, negative caching, single-flight stampede control, and stale-while-revalidate behavior.

Advanced190–280 minutescache-aside, TTL jitter, negative cache, stampede, SWRRedis Open Source 8.10.1redis-py 8.1.0 where Python is usedDocker + redis-cli + Python stdlibStandalone loopback lab · DB 0AOF everysec + RDB · maxmemory 0/noeviction baselineNamed ACL users · TLS off only on loopbackFree/local-firstLast reviewed: September 6, 2026
Reproducible Chapter 23 baseline

Redis Open Source 8.10.1 using the pinned redis:8.10.1 image, exposed only at 127.0.0.1:6431. Standalone topology, logical database 0, AOF everysec plus an RDB save rule, maxmemory 0 unless a lesson explicitly changes a setting on this disposable node, default ACL user disabled, named academy-admin and atlasmart-app users, and fixture prefix atlasmart:ch23:*. TLS is intentionally off only because this mandatory lab is loopback-local; production traffic must follow Chapter 22 network/TLS guidance. Python examples target redis==8.1.0. Search/JSON/vector/time-series/probabilistic features are not required.

Learning outcomes

This lesson turns Cache-Aside with TTL/Jitter, Negative Caching, Stampede Protection, and Stale-While-Revalidate into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.

01

Explain the mechanisms and terminology behind Cache-Aside with TTL/Jitter, Negative Caching, Stampede Protection, and Stale-While-Revalidate.

02

Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.

03

Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.

04

Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.

05

Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.

1. Problem: one popular product expires and 200 workers hit the database

AtlasMart product sku-42 is requested far more often than ordinary products. A cache-aside design reads Redis first and loads the authoritative origin only on a miss. That reduces origin load, but it introduces a freshness boundary: Redis is a derived copy and can be stale until invalidation or expiration. If many copies of the same popular key expire together, every worker can miss simultaneously and create a cache stampede.

A time to live (TTL) is the remaining expiry time stored by Redis. TTL bounds cache lifetime; it is not a promise that business state is fresh until that exact instant. Explicit invalidation on writes, version checks, or shorter soft freshness windows may be required.

2. Cache-aside is a state machine, not just GET then SET

The basic state machine is: cache hit → serve cached value; cache miss → read authoritative source → populate cache → serve; authoritative write → update the source first, then invalidate or update the derived cache. A database commit and Redis invalidation are not one atomic transaction, so applications must tolerate or repair failure between those steps.

State Action Failure boundary
Hit Serve Redis value Can be stale within the chosen freshness policy.
Miss Load origin then populate Redis Concurrent misses can stampede the origin.
Negative hit Return “not found” briefly A newly created object may remain hidden until the short negative TTL ends.
Stale-but-servable Serve stale while one worker refreshes Only appropriate where bounded stale data is acceptable.

3. TTL jitter spreads synchronized expiry

Giving every product an exact 300-second TTL can synchronize misses after a deployment or bulk warmup. Jitter means adding a bounded random component, such as base TTL + uniform 0…J seconds, so expirations spread over time. Jitter does not fix bad invalidation, and it should not make the maximum stale window exceed the product requirement. Measure the resulting TTL distribution instead of copying a folklore percentage.

Shell · create a small observable TTL distribution
for i in 1 2 3 4 5 6 7 8; do  ttl=$((20 + i % 5))  docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-App-Lab-Only-2026 atlasmart-redis-ch23     redis-cli --user atlasmart-app SET "atlasmart:ch23:cache:jitter:$i" value EX "$ttl" >/dev/nulldonefor i in 1 2 3 4 5 6 7 8; do  docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-App-Lab-Only-2026 atlasmart-redis-ch23     redis-cli --user atlasmart-app TTL "atlasmart:ch23:cache:jitter:$i"done

4. Negative caching must be shorter and narrower than success caching

A missing SKU can also be hot. Without negative caching, every “not found” request can hit the origin. Store an explicit negative marker with a short TTL, not an error forever. Distinguish authoritative absence (404 after a successful origin lookup) from transient origin failure (timeout/500), which usually should not be cached as “does not exist.” Scope negative keys by tenant/authorization where necessary so one caller cannot poison another caller’s view.

5. Stampede protection uses single-flight ownership, not an infinite lock wait

One worker can acquire a short SET key token NX PX lease refresh lock; losers briefly poll for the newly populated value or follow a bounded fallback. The lease must exceed normal origin time with margin, but no fixed number is universally safe. Always use a unique token and owner-aware release so a delayed worker cannot delete a lock acquired later by another worker.

6. Stale-while-revalidate separates soft and hard freshness

Stale-while-revalidate (SWR) keeps two boundaries: a soft TTL after which a cached value is considered stale and one caller refreshes it, and a hard TTL after which Redis removes it and stale serving is no longer possible. SWR is valuable for product descriptions or recommendations where bounded stale reads are acceptable; do not use it casually for stock decrements, authorization, payment status, or other invariants.

7. Reproducible local setup

The Chapter 23 node is isolated from every earlier chapter. No production endpoint, shared Chapter 01 node, or global Docker cleanup is touched. The first block is POSIX-shell compatible (Linux/macOS/WSL/Git Bash). The second is a PowerShell equivalent for Windows Docker Desktop.

Shell · start only the Chapter 23 node
mkdir -p ch23-lab && cd ch23-labcat > users.acl <<'EOF'user default offuser academy-admin on >AtlasMart-Ch23-Admin-Lab-Only-2026 ~* &* +@alluser atlasmart-app on >AtlasMart-Ch23-App-Lab-Only-2026 ~atlasmart:ch23:* &atlasmart:ch23:* +@read +@write +@connection +@scripting +time -@admin -@dangerousEOFdocker rm -f atlasmart-redis-ch23 2>/dev/null || truedocker volume create atlasmart-redis-ch23-data >/dev/nulldocker run -d --name atlasmart-redis-ch23 \  -p 127.0.0.1:6431:6379 \  -v "$PWD/users.acl:/usr/local/etc/redis/users.acl:ro" \  -v atlasmart-redis-ch23-data:/data \  redis:8.10.1 redis-server \  --bind 0.0.0.0 --protected-mode yes \  --aclfile /usr/local/etc/redis/users.acl \  --appendonly yes --appendfsync everysec --save 300 10 \  --maxmemory 0 --maxmemory-policy noevictiondocker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 \  redis-cli --user academy-admin PING
PowerShell · Windows Docker Desktop equivalent
$Lab = Join-Path $PWD 'ch23-lab'New-Item -ItemType Directory -Force $Lab | Out-NullSet-Location $Lab@'user default offuser academy-admin on >AtlasMart-Ch23-Admin-Lab-Only-2026 ~* &* +@alluser atlasmart-app on >AtlasMart-Ch23-App-Lab-Only-2026 ~atlasmart:ch23:* &atlasmart:ch23:* +@read +@write +@connection +@scripting +time -@admin -@dangerous'@ | Set-Content -Encoding ascii .\users.acldocker rm -f atlasmart-redis-ch23 2>$nulldocker volume create atlasmart-redis-ch23-data | Out-Null$Acl = (Resolve-Path .\users.acl).Pathdocker run -d --name atlasmart-redis-ch23 `  -p 127.0.0.1:6431:6379 `  --mount "type=bind,source=$Acl,target=/usr/local/etc/redis/users.acl,readonly" `  -v atlasmart-redis-ch23-data:/data `  redis:8.10.1 redis-server `  --bind 0.0.0.0 --protected-mode yes `  --aclfile /usr/local/etc/redis/users.acl `  --appendonly yes --appendfsync everysec --save 300 10 `  --maxmemory 0 --maxmemory-policy noevictiondocker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 `  redis-cli --user academy-admin PING

8. Run a concurrent cold-start experiment and count origin amplification

The Python harness below creates a 48-request cold burst, measures cache/origin counters, proves negative caching, and prints the actual TTL. It does not contain expected benchmark numbers because CPU scheduling and Docker latency differ by machine.

Python · cache-aside + jitter + negative cache + SWR + single-flight
from __future__ import annotationsimport json, random, threading, timefrom concurrent.futures import ThreadPoolExecutorimport redisr = redis.Redis(host="127.0.0.1", port=6431, username="atlasmart-app",                password="AtlasMart-Ch23-App-Lab-Only-2026", decode_responses=True)origin = {"sku-42": {"name": "Atlas Keyboard", "price": 129}}metrics = {"hit": 0, "miss": 0, "negative": 0, "origin": 0, "stale": 0}mu = threading.Lock()BASE_TTL = 8JITTER = 3NEGATIVE_TTL = 2LOCK_MS = 1500SOFT_TTL = 4def cache_key(sku): return f"atlasmart:ch23:cache:product:{sku}"def lock_key(sku): return f"atlasmart:ch23:cache:lock:{sku}"def load_origin(sku):    with mu: metrics["origin"] += 1    time.sleep(0.08)  # deterministic-ish simulated origin cost; measure on your machine    return origin.get(sku)def ttl_with_jitter():    return BASE_TTL + random.randint(0, JITTER)def get_product(sku):    key = cache_key(sku)    raw = r.get(key)    if raw:        doc = json.loads(raw)        if doc.get("negative"):            with mu: metrics["negative"] += 1            return None        age = time.time() - doc["cached_at"]        if age <= SOFT_TTL:            with mu: metrics["hit"] += 1            return doc["value"]        # stale-while-revalidate: serve stale only inside hard Redis TTL.        with mu: metrics["stale"] += 1        token = f"refresh-{threading.get_ident()}-{time.time_ns()}"        if r.set(lock_key(sku), token, nx=True, px=LOCK_MS):            try:                fresh = load_origin(sku)                if fresh is not None:                    r.set(key, json.dumps({"value": fresh, "cached_at": time.time()}), ex=ttl_with_jitter())            finally:                # Redis 8.4+ conditional delete; Lua compare-delete is the older portable fallback.                r.execute_command("DELEX", lock_key(sku), "IFEQ", token)        return doc["value"]    with mu: metrics["miss"] += 1    token = f"load-{threading.get_ident()}-{time.time_ns()}"    if r.set(lock_key(sku), token, nx=True, px=LOCK_MS):        try:            value = load_origin(sku)            if value is None:                r.set(key, json.dumps({"negative": True}), ex=NEGATIVE_TTL)                return None            r.set(key, json.dumps({"value": value, "cached_at": time.time()}), ex=ttl_with_jitter())            return value        finally:            r.execute_command("DELEX", lock_key(sku), "IFEQ", token)    # A loser does not hammer the origin; it waits briefly for the winner.    deadline = time.monotonic() + LOCK_MS / 1000    while time.monotonic() < deadline:        time.sleep(0.02)        raw = r.get(key)        if raw:            doc = json.loads(raw)            return None if doc.get("negative") else doc["value"]    return load_origin(sku)  # bounded fallback policy; choose deliberately in productionr.delete(cache_key("sku-42"), lock_key("sku-42"), cache_key("missing"), lock_key("missing"))with ThreadPoolExecutor(max_workers=24) as ex:    list(ex.map(lambda _: get_product("sku-42"), range(48)))print("metrics_after_cold_burst", metrics)print("ttl", r.ttl(cache_key("sku-42")))print("negative_first", get_product("missing"))print("negative_second", get_product("missing"))print("metrics_final", metrics)

9. Deliberately wrong: synchronized TTL plus “just let every miss load”

A naïve miss path can produce one origin query per concurrent request. Reproduce it only with a bounded in-memory origin simulation, not against a production database. Repair with jitter, a bounded single-flight lock, short negative caching for true absence, and SWR only where stale data is acceptable. Verify by comparing origin-call counts during the same cold burst before and after the repair.

10. Evidence to record

Capture cache hits, misses, stale serves, negative hits, origin calls, lock-acquire wins/losses, TTL distribution, origin p95/p99 if available, and cache key memory. A high hit rate alone can hide stale-data defects; a low origin count alone can hide overlong TTLs. Freshness and origin protection must be measured together.

Check your understanding

  1. Why add TTL jitter?
  2. Should a database timeout be cached as a negative product lookup?
  3. What does stale-while-revalidate require from the business domain?
  4. Why must a cache-refresh lock have a unique value?
Review the answers

To spread expirations and reduce synchronized misses; it does not replace invalidation or define a universal safe TTL.

Usually no. Negative caching is for an authoritative absence, not an unknown result caused by an origin failure.

A clearly acceptable bounded stale window. It is unsuitable where the response must reflect a current invariant.

So release can verify ownership and avoid deleting a lock that expired and was acquired by another worker.

11. Production judgment and bridge

Cache-aside is appropriate for derived, read-heavy state with explicit freshness semantics. It does not provide transactional coherence with the origin. Size cache objects and response bodies, apply tenant-aware keys, protect against hot keys, set client timeouts, and test origin failure, Redis failure, invalidation loss, and stampedes. Lesson 2 applies the same “read-decide-update must be atomic” principle to rate limiting.

Summary and next step

Cache-Aside with TTL/Jitter, Negative Caching, Stampede Protection, and Stale-While-Revalidate is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Fixed/Sliding Window, Token Bucket, and Leaky-Bucket Rate Limiting with Atomic Operations.

Authoritative references

12. Cleanup

Remove only the disposable Chapter 23 resources.

Shell · cleanup
docker rm -f atlasmart-redis-ch23 2>/dev/null || truedocker volume rm atlasmart-redis-ch23-data 2>/dev/null || true# Remove only the local ch23-lab directory after keeping any evidence you need.# Never use docker system prune as Chapter 23 cleanup.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.