A benchmark result without a workload contract is not portable evidence.

Benchmark Workload Shapes, Not Marketing Numbers: Keys, Values, Zipf Skew, Concurrency, and Durability

Design repeatable benchmark contracts that disclose dataset, key/value size, access distribution, read/write mix, concurrency, consistency, durability, indexes, replication, and failure state.

Advanced125–165 minutesWorkload-shape benchmark labPython 3.13+ · standard libraryVendor-neutral · free/local mandatory pathLast reviewed: August 2026
01

Define a benchmark contract including dataset, key/value size, operation mix, request distribution, concurrency, durability/consistency, indexes, replication, and failure state.

02

Generate uniform and Zipf-like key distributions and explain how skew changes contention, cache locality, shard load, and tail behavior.

03

Distinguish measured values from modeled values and reject headline ops/sec numbers whose workload/environment is undisclosed.

04

Design reproducible before/after performance tests with warmup, fixed versions, controlled datasets, and comparable correctness guarantees.

1. “100,000 operations per second” is not a workload

Performance numbers are conditional statements. AtlasMart needs to know whether those operations are 1 KiB point reads or 2 MiB writes, whether keys are uniform or Zipf-like, whether reads hit cache, whether writes fsync, whether three replicas acknowledge, whether indexes are maintained, whether the dataset fits RAM, and whether a node is failed or rebuilding. Change any of those and the comparison may answer a different question.

A benchmark contract freezes the material variables: engine/version/edition, client/driver, host/container/cloud shape, node topology and failure domains, dataset cardinality and bytes, key/value sizes, read/write/scan mix, request distribution, concurrency/offered rate, batch size, consistency/ack level, durability/fsync behavior, indexes, replication factor, compression, warmup, runtime duration, and active maintenance/failure state.

2. Uniform traffic and Zipf-like traffic exercise different systems

Uniform selection gives keys roughly equal probability. A Zipf-like distribution assigns much more probability to the most popular ranks, resembling catalogs, social objects, or “celebrity” keys. Skew can improve cache locality for hot reads but overload one shard, increase lock/conditional-write contention, or create one partition whose replicas dominate network and disk work. Therefore a benchmark that hashes uniformly can miss the production bottleneck created by the real access distribution.

Contract field AtlasMart example Why it matters
Dataset 100 logical keys in the toy lab; production must disclose total bytes Memory/cache fit and index/storage footprint
Values 1 KiB Network, serialization, cache, and storage bandwidth
Operation mix 80% read / 20% write Read vs write path and background maintenance
Distribution uniform vs Zipf-like α=1.1 Hot-key concentration and cache/shard skew
Concurrency 64 clients Queueing, pools, CPU, lock/replica pressure
Correctness RF=3, quorum write ack, durability modeled as fsync Latency is inseparable from guarantees
Failure state healthy in this lab A healthy-only result does not prove degraded-state capacity
Measure the benchmark harness too

If the load generator reaches 100% CPU, exhausts ephemeral ports, serializes requests unintentionally, or shares a noisy host with the database, it may become the bottleneck. Record client CPU, network, pool occupancy, rejected work, and offered-vs-completed rate.

3. Deliberately wrong approach: publish the larger ops/sec number and omit request distribution

The deterministic lab holds operation count, record count, value size, nominal concurrency, replication, and ack policy constant. Only key distribution changes. Uniform traffic has a ~1.1% top-key share; the Zipf-like sample gives the hottest key ~23.9% and the top ten ~63.6%. A toy contention model then predicts far lower service capacity. The modeled ops/sec is intentionally not a product benchmark—the lesson is that shape changes the answer.

4. AtlasMart lab: same count, different workload shape

Mandatory lab environment

Python 3.13+ standard library only. Random seed is fixed. “modeled_service_ms” and “modeled_single_worker_ops_s” are teaching arithmetic, not hardware measurements.

python · AtlasMart deterministic simulation
import random
from collections import Counter

random.seed(2303)
KEYS = [f"k{i:03d}" for i in range(1,101)]
N = 20000

def sample_uniform():
    return [random.choice(KEYS) for _ in range(N)]

def sample_zipf(alpha=1.1):
    weights = [1/(rank**alpha) for rank in range(1,len(KEYS)+1)]
    return random.choices(KEYS, weights=weights, k=N)

def summarize(name, keys):
    c = Counter(keys)
    top1 = c.most_common(1)[0]
    top10 = sum(n for _,n in c.most_common(10))
    # Teaching service model: base 0.20ms + contention penalty from repeated access concentration.
    penalty = sum((n/N)**2 for n in c.values()) * 18
    service_ms = 0.20 + penalty
    est_ops = 1000/service_ms
    print(name)
    print(" top1:", top1, "share=", f"{top1[1]/N:.1%}")
    print(" top10_share=", f"{top10/N:.1%}")
    print(" concentration_index=", round(sum((n/N)**2 for n in c.values()),4))
    print(" modeled_service_ms=", round(service_ms,3), "modeled_single_worker_ops_s=", round(est_ops))

print("BENCHMARK CONTRACT")
print({"records":100, "operations":N, "value_bytes":1024, "read_write":"80/20", "concurrency":64,
       "replication_factor":3, "write_ack":"quorum", "durability":"fsync modeled", "indexes":1, "failure_state":"healthy"})
print()
summarize("UNIFORM", sample_uniform())
print()
summarize("ZIPF-LIKE alpha=1.1", sample_zipf())
print("\nSame operation count, record size, and nominal concurrency; request shape alone changes concentration and modeled cost.")
Expected evidence

Uniform requests spread the top ten keys across about 11% of operations; the Zipf-like sample concentrates roughly 64% there. The same headline operation count therefore drives very different contention and routing pressure. This proves workload shape matters; it does not predict a real database throughput value.

5. Production judgment: compare guarantees and environments, not just speed

Use the same dataset snapshot, schema/indexes, replication, consistency/durability, client version, warmup policy, test duration, offered rate, and failure state for A/B comparisons. Disclose whether compaction, repair, checkpoints, backups, or rebalance ran. Report distributions, useful throughput, errors, and freshness—not only successful ops/sec. If one candidate disables fsync or lowers acknowledgements, label the changed durability window rather than presenting it as a free speedup.

YCSB is useful precisely because workload properties such as read/write proportions and request distribution are explicit. It is not a substitute for your workload. The next lesson examines another benchmark trap: a closed-loop generator can stop sending during a stall and thereby omit latencies that real scheduled arrivals would have experienced.

Check your understanding

  1. Why is ops/sec meaningless without workload shape?
  2. What does a Zipf-like distribution change?
  3. Why must consistency/durability settings be part of a benchmark contract?
  4. What is the purpose of a fixed random seed in a teaching workload?
  5. Why report offered and completed load separately?
Review the answers

1. The same operation count can represent different bytes, key skew, read/write mix, durability, replication, concurrency, indexes, and failure states, producing incomparable resource costs.

2. It concentrates requests on high-ranked keys, affecting cache locality, shard hot spots, contention, and tail latency.

3. They change message, replica, WAL/fsync, and failure semantics; comparing speed while changing guarantees is not an apples-to-apples result.

4. It makes the generated sequence reproducible so changes are attributable to the model/config rather than a different random sample.

5. A saturated system may reject, queue, or shed requests; completed throughput can plateau while clients continue offering more work.

References

Foundational claims use primary research or current official documentation where practical. Product references are implementation anchors only; the mandatory labs are vendor-neutral.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.