A benchmark result without a workload contract is not portable evidence.
Benchmark Workload Shapes, Not Marketing Numbers: Keys, Values, Zipf Skew, Concurrency, and Durability
Design repeatable benchmark contracts that disclose dataset, key/value size, access distribution, read/write mix, concurrency, consistency, durability, indexes, replication, and failure state.
Define a benchmark contract including dataset, key/value size, operation mix, request distribution, concurrency, durability/consistency, indexes, replication, and failure state.
Generate uniform and Zipf-like key distributions and explain how skew changes contention, cache locality, shard load, and tail behavior.
Distinguish measured values from modeled values and reject headline ops/sec numbers whose workload/environment is undisclosed.
Design reproducible before/after performance tests with warmup, fixed versions, controlled datasets, and comparable correctness guarantees.
1. “100,000 operations per second” is not a workload
Performance numbers are conditional statements. AtlasMart needs to know whether those operations are 1 KiB point reads or 2 MiB writes, whether keys are uniform or Zipf-like, whether reads hit cache, whether writes fsync, whether three replicas acknowledge, whether indexes are maintained, whether the dataset fits RAM, and whether a node is failed or rebuilding. Change any of those and the comparison may answer a different question.
A benchmark contract freezes the material variables: engine/version/edition, client/driver, host/container/cloud shape, node topology and failure domains, dataset cardinality and bytes, key/value sizes, read/write/scan mix, request distribution, concurrency/offered rate, batch size, consistency/ack level, durability/fsync behavior, indexes, replication factor, compression, warmup, runtime duration, and active maintenance/failure state.
2. Uniform traffic and Zipf-like traffic exercise different systems
Uniform selection gives keys roughly equal probability. A Zipf-like distribution assigns much more probability to the most popular ranks, resembling catalogs, social objects, or “celebrity” keys. Skew can improve cache locality for hot reads but overload one shard, increase lock/conditional-write contention, or create one partition whose replicas dominate network and disk work. Therefore a benchmark that hashes uniformly can miss the production bottleneck created by the real access distribution.
| Contract field | AtlasMart example | Why it matters |
|---|---|---|
| Dataset | 100 logical keys in the toy lab; production must disclose total bytes | Memory/cache fit and index/storage footprint |
| Values | 1 KiB | Network, serialization, cache, and storage bandwidth |
| Operation mix | 80% read / 20% write | Read vs write path and background maintenance |
| Distribution | uniform vs Zipf-like α=1.1 | Hot-key concentration and cache/shard skew |
| Concurrency | 64 clients | Queueing, pools, CPU, lock/replica pressure |
| Correctness | RF=3, quorum write ack, durability modeled as fsync | Latency is inseparable from guarantees |
| Failure state | healthy in this lab | A healthy-only result does not prove degraded-state capacity |
If the load generator reaches 100% CPU, exhausts ephemeral ports, serializes requests unintentionally, or shares a noisy host with the database, it may become the bottleneck. Record client CPU, network, pool occupancy, rejected work, and offered-vs-completed rate.
3. Deliberately wrong approach: publish the larger ops/sec number and omit request distribution
The deterministic lab holds operation count, record count, value size, nominal concurrency, replication, and ack policy constant. Only key distribution changes. Uniform traffic has a ~1.1% top-key share; the Zipf-like sample gives the hottest key ~23.9% and the top ten ~63.6%. A toy contention model then predicts far lower service capacity. The modeled ops/sec is intentionally not a product benchmark—the lesson is that shape changes the answer.
4. AtlasMart lab: same count, different workload shape
Python 3.13+ standard library only. Random seed is fixed. “modeled_service_ms” and “modeled_single_worker_ops_s” are teaching arithmetic, not hardware measurements.
import random
from collections import Counter
random.seed(2303)
KEYS = [f"k{i:03d}" for i in range(1,101)]
N = 20000
def sample_uniform():
return [random.choice(KEYS) for _ in range(N)]
def sample_zipf(alpha=1.1):
weights = [1/(rank**alpha) for rank in range(1,len(KEYS)+1)]
return random.choices(KEYS, weights=weights, k=N)
def summarize(name, keys):
c = Counter(keys)
top1 = c.most_common(1)[0]
top10 = sum(n for _,n in c.most_common(10))
# Teaching service model: base 0.20ms + contention penalty from repeated access concentration.
penalty = sum((n/N)**2 for n in c.values()) * 18
service_ms = 0.20 + penalty
est_ops = 1000/service_ms
print(name)
print(" top1:", top1, "share=", f"{top1[1]/N:.1%}")
print(" top10_share=", f"{top10/N:.1%}")
print(" concentration_index=", round(sum((n/N)**2 for n in c.values()),4))
print(" modeled_service_ms=", round(service_ms,3), "modeled_single_worker_ops_s=", round(est_ops))
print("BENCHMARK CONTRACT")
print({"records":100, "operations":N, "value_bytes":1024, "read_write":"80/20", "concurrency":64,
"replication_factor":3, "write_ack":"quorum", "durability":"fsync modeled", "indexes":1, "failure_state":"healthy"})
print()
summarize("UNIFORM", sample_uniform())
print()
summarize("ZIPF-LIKE alpha=1.1", sample_zipf())
print("\nSame operation count, record size, and nominal concurrency; request shape alone changes concentration and modeled cost.")
Uniform requests spread the top ten keys across about 11% of operations; the Zipf-like sample concentrates roughly 64% there. The same headline operation count therefore drives very different contention and routing pressure. This proves workload shape matters; it does not predict a real database throughput value.
5. Production judgment: compare guarantees and environments, not just speed
Use the same dataset snapshot, schema/indexes, replication, consistency/durability, client version, warmup policy, test duration, offered rate, and failure state for A/B comparisons. Disclose whether compaction, repair, checkpoints, backups, or rebalance ran. Report distributions, useful throughput, errors, and freshness—not only successful ops/sec. If one candidate disables fsync or lowers acknowledgements, label the changed durability window rather than presenting it as a free speedup.
YCSB is useful precisely because workload properties such as read/write proportions and request distribution are explicit. It is not a substitute for your workload. The next lesson examines another benchmark trap: a closed-loop generator can stop sending during a stall and thereby omit latencies that real scheduled arrivals would have experienced.
Check your understanding
- Why is ops/sec meaningless without workload shape?
- What does a Zipf-like distribution change?
- Why must consistency/durability settings be part of a benchmark contract?
- What is the purpose of a fixed random seed in a teaching workload?
- Why report offered and completed load separately?
Review the answers
1. The same operation count can represent different bytes, key skew, read/write mix, durability, replication, concurrency, indexes, and failure states, producing incomparable resource costs.
2. It concentrates requests on high-ranked keys, affecting cache locality, shard hot spots, contention, and tail latency.
3. They change message, replica, WAL/fsync, and failure semantics; comparing speed while changing guarantees is not an apples-to-apples result.
4. It makes the generated sequence reproducible so changes are attributable to the model/config rather than a different random sample.
5. A saturated system may reject, queue, or shed requests; completed throughput can plateau while clients continue offering more work.
References
Foundational claims use primary research or current official documentation where practical. Product references are implementation anchors only; the mandatory labs are vendor-neutral.
- Yahoo! Cloud Serving Benchmark (YCSB) — Open benchmark suite; Apache-2.0 repository and workload machinery.
- YCSB Workload Template — Current workload properties include request distribution and histogram/raw measurement options.
- YCSB Core Workloads — Examples of read/write/scan mixes and application-oriented workload definitions.
- Cooper et al. — Benchmarking Cloud Serving Systems with YCSB — Foundational YCSB paper defining a framework for evaluating cloud data-serving systems under specified workloads.