Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone
redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency
Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.
Redis Open Source 8.10.1 using the pinned
redis:8.10.1 image. The observability/benchmark
node is atlasmart-redis-ch24 on
127.0.0.1:6441, standalone topology, logical
database 0, AOF everysec plus RDB save rules,
maxmemory 0/noeviction unless a
bounded experiment says otherwise, named
academy-admin and atlasmart-app ACL
users, and fixture prefix atlasmart:ch24:*. TLS is
off only on this loopback-local disposable node; the capstone
security acceptance criteria reuse Chapter 22 TLS/ACL guidance.
Python examples target redis==8.1.0.
Search/JSON/vector/time-series/probabilistic features are
optional and must be included in capacity accounting only when
the chosen AtlasMart architecture actually uses them.
Learning outcomes
This lesson turns redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.
Explain the mechanisms and terminology behind redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency.
Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.
Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.
Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.
Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.
1. Problem: a benchmark says 400k ops/s, but checkout misses its SLO
A throughput number is only meaningful for the workload that produced it. AtlasMart has mixed reads/writes, hundreds-to-thousands-of-byte payloads, hot-key skew, connection reuse, AOF everysec, and a tail-latency SLO. A synthetic benchmark that uses 3-byte values and uniform keys answers a different question.
2. Define the workload contract before running a tool
| Dimension | Record explicitly | Why it changes the result |
|---|---|---|
| Command mix | GET/SET/hash/stream/search proportions | Different CPU/memory/complexity costs |
| Payload | Request and response bytes | Network, allocator, copy cost |
| Key distribution | Uniform vs measured skew/hot set | Cache locality and hotspot behavior |
| Concurrency | Clients/threads/connections | Queueing and saturation |
| Pipeline depth | 1, 4, 8, … based on application | Round trips and buffering |
| Durability | RDB/AOF/fsync settings | Disk and fork behavior |
| Warmup | Duration/requests before measurement | Caches/JIT/connections stabilize |
| Percentiles | p50/p95/p99/max + errors | Tail behavior is visible |
3. redis-benchmark is a useful baseline when its assumptions are written down
Redis documents -c clients,
-n requests, -d payload size,
-r random keyspace, and -P pipeline
depth. Use exact parameters and never publish the tool defaults
as production capacity. Run only against the disposable Chapter
24 node.
docker exec atlasmart-redis-ch24 redis-benchmark -h 127.0.0.1 -p 6379 --user academy-admin -a AtlasMart-Ch24-Admin-Lab-Only-2026 -t get,set -n 20000 -c 20 -d 512 -r 10000 -P 1 --csvdocker exec atlasmart-redis-ch24 redis-benchmark -h 127.0.0.1 -p 6379 --user academy-admin -a AtlasMart-Ch24-Admin-Lab-Only-2026 -t get,set -n 20000 -c 20 -d 512 -r 10000 -P 8 --csv
The second run tests a different transport shape. Higher
throughput with -P 8 does not prove the application
can or should batch eight independent operations.
4. Custom AtlasMart workload: realistic payload, deterministic 80/20 skew
import random, statistics, time, redisrng = random.Random(2403)r = redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")hot = [f"atlasmart:ch24:bench:product:{i}" for i in range(20)]cold = [f"atlasmart:ch24:bench:product:{i}" for i in range(20,1000)]payload = b"p" * 512p = r.pipeline(transaction=False)for k in hot + cold: p.set(k, payload)p.execute()def choose_key(): return rng.choice(hot if rng.random() < 0.80 else cold)for _ in range(500): r.get(choose_key()) # warmuplat=[]; errors=0; start=time.perf_counter()for _ in range(5000): t0=time.perf_counter_ns() try: r.get(choose_key()) except Exception: errors += 1 lat.append((time.perf_counter_ns()-t0)/1e6)elapsed=time.perf_counter()-starts=sorted(lat)def pct(q): return s[min(len(s)-1, int((len(s)-1)*q))]print({"rps": round(len(lat)/elapsed,1), "p50_ms": pct(.50), "p95_ms": pct(.95), "p99_ms": pct(.99), "max_ms": max(s), "errors": errors})
5. Pipeline depth must preserve result mapping and memory bounds
Pipelining is transport batching, not atomicity. Compare the same logical work at depth 1, 4, and 8; verify response ordering and equality; measure client memory and server output buffers if batch size grows. A giant pipeline can improve throughput while worsening per-batch latency and memory pressure.
import time, redisr = redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")keys=[f"atlasmart:ch24:bench:product:{i}" for i in range(1000)]for depth in (1,4,8): t0=time.perf_counter(); count=0 for i in range(0, 4000, depth): p=r.pipeline(transaction=False) batch=[keys[(i+j)%len(keys)] for j in range(depth)] for k in batch: p.get(k) out=p.execute(); assert len(out)==len(batch); count += len(out) print(depth, round(count/(time.perf_counter()-t0),1), "ops/s")
6. Coordinated omission: a closed-loop client can hide queueing pain
A client that sends the next request only after the previous one completes automatically reduces offered load during a pause. That can under-sample the requests users would have attempted during the stall. At minimum, record scheduled-start lateness or use a load generator capable of open-loop arrivals. The following local harness records both response latency and schedule lag; it is educational, not a replacement for a mature performance framework.
import time, redisr=redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")interval=0.002 # 500 offered requests/s in this bounded examplen=2000; base=time.perf_counter(); response=[]; lateness=[]for i in range(n): target=base+i*interval now=time.perf_counter() if now < target: time.sleep(target-now) actual=time.perf_counter(); lateness.append(max(0.0, actual-target)*1000) t0=time.perf_counter(); r.get(f"atlasmart:ch24:bench:product:{i%1000}"); response.append((time.perf_counter()-t0)*1000)print("max_schedule_lag_ms", max(lateness), "max_response_ms", max(response))
7. Capacity is an SLO-constrained region, not the crash point
Increase offered load in bounded steps and stop when tail latency, errors/timeouts, CPU, memory headroom, replication lag, or durability behavior crosses an acceptance threshold. Keep a reserve for failover, fork copy-on-write, AOF rewrite, backups, resharding, or traffic bursts. Dataset bytes alone are not capacity.
8. Deliberately wrong: publish redis-benchmark defaults as production capacity
A default benchmark can use tiny payloads, uniform keys, no application serialization, different pipelining, and no upstream work. The repair is to publish the complete workload contract, run a custom representative test, correlate Redis and host metrics, and state the SLO-constrained capacity range rather than one universal ops/s number.
9. Benchmark evidence worksheet
| Run | Workload | Durability | p50 | p95 | p99 | max | errors | CPU/RSS | Decision |
|---|---|---|---|---|---|---|---|---|---|
| Baseline A | redis-benchmark P=1, 512 B, c=20 | AOF everysec | measure | measure | measure | measure | measure | capture | synthetic baseline only |
| Baseline B | same, P=8 | AOF everysec | measure | measure | measure | measure | measure | capture | transport sensitivity |
| AtlasMart | 80/20 skew, 512 B, app client | AOF everysec | measure | measure | measure | measure | measure | capture | capacity evidence |
Check your understanding
- Why is pipeline depth part of the benchmark contract?
- Why is p99 more useful than average alone for an SLO?
- What does key skew change?
- What is coordinated omission in this context?
Review the answers
It changes round trips, throughput, buffering, and batch latency.
Averages can hide rare but user-visible stalls; tail percentiles expose them.
It changes locality, hot-key concentration, cache behavior, and potentially shard distribution.
A closed-loop generator may stop offering requests during pauses and therefore under-report queueing users would have experienced.
10. Production judgment and bridge
Benchmark versions, payloads, durability, client libraries, topology, warmup, errors, and host conditions alongside the numbers. Never compare runs that changed multiple dimensions without labeling the difference. The next lesson proves that backup and upgrade safety are also measured properties, not configuration checkboxes.
Summary and next step
redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans.
Authoritative references
import redisr=redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")batch=[]for k in r.scan_iter(match="atlasmart:ch24:bench:*"): batch.append(k) if len(batch)>=200: r.unlink(*batch); batch=[]if batch: r.unlink(*batch)