Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone

redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency

Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.

Advanced210–320 minutesredis-benchmark, custom load, pipeline depth, skew, p99, capacityRedis Open Source 8.10.1redis-py 8.1.0 where Python is usedDocker + redis-cli + PythonStandalone evidence node · DB 0AOF everysec + RDB · maxmemory 0/noeviction baselineNamed ACL users · TLS off only on loopbackReuses Sentinel/Cluster labs for failover acceptanceFree/local-firstLast reviewed: September 6, 2026
Reproducible Chapter 24 baseline

Redis Open Source 8.10.1 using the pinned redis:8.10.1 image. The observability/benchmark node is atlasmart-redis-ch24 on 127.0.0.1:6441, standalone topology, logical database 0, AOF everysec plus RDB save rules, maxmemory 0/noeviction unless a bounded experiment says otherwise, named academy-admin and atlasmart-app ACL users, and fixture prefix atlasmart:ch24:*. TLS is off only on this loopback-local disposable node; the capstone security acceptance criteria reuse Chapter 22 TLS/ACL guidance. Python examples target redis==8.1.0. Search/JSON/vector/time-series/probabilistic features are optional and must be included in capacity accounting only when the chosen AtlasMart architecture actually uses them.

Learning outcomes

This lesson turns redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.

01

Explain the mechanisms and terminology behind redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency.

02

Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.

03

Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.

04

Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.

05

Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.

1. Problem: a benchmark says 400k ops/s, but checkout misses its SLO

A throughput number is only meaningful for the workload that produced it. AtlasMart has mixed reads/writes, hundreds-to-thousands-of-byte payloads, hot-key skew, connection reuse, AOF everysec, and a tail-latency SLO. A synthetic benchmark that uses 3-byte values and uniform keys answers a different question.

2. Define the workload contract before running a tool

Dimension Record explicitly Why it changes the result
Command mix GET/SET/hash/stream/search proportions Different CPU/memory/complexity costs
Payload Request and response bytes Network, allocator, copy cost
Key distribution Uniform vs measured skew/hot set Cache locality and hotspot behavior
Concurrency Clients/threads/connections Queueing and saturation
Pipeline depth 1, 4, 8, … based on application Round trips and buffering
Durability RDB/AOF/fsync settings Disk and fork behavior
Warmup Duration/requests before measurement Caches/JIT/connections stabilize
Percentiles p50/p95/p99/max + errors Tail behavior is visible

3. redis-benchmark is a useful baseline when its assumptions are written down

Redis documents -c clients, -n requests, -d payload size, -r random keyspace, and -P pipeline depth. Use exact parameters and never publish the tool defaults as production capacity. Run only against the disposable Chapter 24 node.

Shell · two explicit redis-benchmark baselines
docker exec atlasmart-redis-ch24 redis-benchmark -h 127.0.0.1 -p 6379 --user academy-admin -a AtlasMart-Ch24-Admin-Lab-Only-2026 -t get,set -n 20000 -c 20 -d 512 -r 10000 -P 1 --csvdocker exec atlasmart-redis-ch24 redis-benchmark -h 127.0.0.1 -p 6379 --user academy-admin -a AtlasMart-Ch24-Admin-Lab-Only-2026 -t get,set -n 20000 -c 20 -d 512 -r 10000 -P 8 --csv

The second run tests a different transport shape. Higher throughput with -P 8 does not prove the application can or should batch eight independent operations.

4. Custom AtlasMart workload: realistic payload, deterministic 80/20 skew

Python · representative workload with tail percentiles
import random, statistics, time, redisrng = random.Random(2403)r = redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")hot = [f"atlasmart:ch24:bench:product:{i}" for i in range(20)]cold = [f"atlasmart:ch24:bench:product:{i}" for i in range(20,1000)]payload = b"p" * 512p = r.pipeline(transaction=False)for k in hot + cold: p.set(k, payload)p.execute()def choose_key(): return rng.choice(hot if rng.random() < 0.80 else cold)for _ in range(500): r.get(choose_key())  # warmuplat=[]; errors=0; start=time.perf_counter()for _ in range(5000):    t0=time.perf_counter_ns()    try: r.get(choose_key())    except Exception: errors += 1    lat.append((time.perf_counter_ns()-t0)/1e6)elapsed=time.perf_counter()-starts=sorted(lat)def pct(q): return s[min(len(s)-1, int((len(s)-1)*q))]print({"rps": round(len(lat)/elapsed,1), "p50_ms": pct(.50), "p95_ms": pct(.95), "p99_ms": pct(.99), "max_ms": max(s), "errors": errors})

5. Pipeline depth must preserve result mapping and memory bounds

Pipelining is transport batching, not atomicity. Compare the same logical work at depth 1, 4, and 8; verify response ordering and equality; measure client memory and server output buffers if batch size grows. A giant pipeline can improve throughput while worsening per-batch latency and memory pressure.

Python · bounded pipeline comparison
import time, redisr = redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")keys=[f"atlasmart:ch24:bench:product:{i}" for i in range(1000)]for depth in (1,4,8):    t0=time.perf_counter(); count=0    for i in range(0, 4000, depth):        p=r.pipeline(transaction=False)        batch=[keys[(i+j)%len(keys)] for j in range(depth)]        for k in batch: p.get(k)        out=p.execute(); assert len(out)==len(batch); count += len(out)    print(depth, round(count/(time.perf_counter()-t0),1), "ops/s")

6. Coordinated omission: a closed-loop client can hide queueing pain

A client that sends the next request only after the previous one completes automatically reduces offered load during a pause. That can under-sample the requests users would have attempted during the stall. At minimum, record scheduled-start lateness or use a load generator capable of open-loop arrivals. The following local harness records both response latency and schedule lag; it is educational, not a replacement for a mature performance framework.

Python · fixed-rate schedule with lateness evidence
import time, redisr=redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")interval=0.002  # 500 offered requests/s in this bounded examplen=2000; base=time.perf_counter(); response=[]; lateness=[]for i in range(n):    target=base+i*interval    now=time.perf_counter()    if now < target: time.sleep(target-now)    actual=time.perf_counter(); lateness.append(max(0.0, actual-target)*1000)    t0=time.perf_counter(); r.get(f"atlasmart:ch24:bench:product:{i%1000}"); response.append((time.perf_counter()-t0)*1000)print("max_schedule_lag_ms", max(lateness), "max_response_ms", max(response))

7. Capacity is an SLO-constrained region, not the crash point

Increase offered load in bounded steps and stop when tail latency, errors/timeouts, CPU, memory headroom, replication lag, or durability behavior crosses an acceptance threshold. Keep a reserve for failover, fork copy-on-write, AOF rewrite, backups, resharding, or traffic bursts. Dataset bytes alone are not capacity.

8. Deliberately wrong: publish redis-benchmark defaults as production capacity

Failure pattern

A default benchmark can use tiny payloads, uniform keys, no application serialization, different pipelining, and no upstream work. The repair is to publish the complete workload contract, run a custom representative test, correlate Redis and host metrics, and state the SLO-constrained capacity range rather than one universal ops/s number.

9. Benchmark evidence worksheet

Run Workload Durability p50 p95 p99 max errors CPU/RSS Decision
Baseline A redis-benchmark P=1, 512 B, c=20 AOF everysec measure measure measure measure measure capture synthetic baseline only
Baseline B same, P=8 AOF everysec measure measure measure measure measure capture transport sensitivity
AtlasMart 80/20 skew, 512 B, app client AOF everysec measure measure measure measure measure capture capacity evidence

Check your understanding

  1. Why is pipeline depth part of the benchmark contract?
  2. Why is p99 more useful than average alone for an SLO?
  3. What does key skew change?
  4. What is coordinated omission in this context?
Review the answers

It changes round trips, throughput, buffering, and batch latency.

Averages can hide rare but user-visible stalls; tail percentiles expose them.

It changes locality, hot-key concentration, cache behavior, and potentially shard distribution.

A closed-loop generator may stop offering requests during pauses and therefore under-report queueing users would have experienced.

10. Production judgment and bridge

Benchmark versions, payloads, durability, client libraries, topology, warmup, errors, and host conditions alongside the numbers. Never compare runs that changed multiple dimensions without labeling the difference. The next lesson proves that backup and upgrade safety are also measured properties, not configuration checkboxes.

Summary and next step

redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Back Up RDB/AOF, Test Restore, Perform Rolling/Topology-Aware Upgrades, and Maintain Rollback Plans.

Authoritative references

Python · cleanup benchmark fixtures with SCAN/UNLINK
import redisr=redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")batch=[]for k in r.scan_iter(match="atlasmart:ch24:bench:*"):    batch.append(k)    if len(batch)>=200: r.unlink(*batch); batch=[]if batch: r.unlink(*batch)

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.