A benchmark is useful only when workload shape, offered load, durability, concurrency, tails, and measurement bias are explicit.
Benchmark Reads / Writes / Aggregations with Realistic Document Shapes, Skew, Durability, and Concurrency
Run a repeatable MongoDB workload benchmark that preserves document shape, skew, durability, concurrency, tail latency, and overload visibility.
Learning objectives
Build a benchmark from observed AtlasMart document shapes and access skew rather than uniform synthetic CRUD.
Separate warm-up from measurement and report p50/p95/p99, errors, offered load, completed load, and scheduler lag.
Use an open-loop or scheduled-arrival model so overload is visible instead of hidden by coordinated omission.
Hold durability, indexes, client pool, topology, and dataset constant while changing one factor at a time.
Correlate client measurements with server metrics and explain evidence before drawing a conclusion.
This chapter pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where client behavior or load
generation matters. Lesson 3 uses a disposable one-member
replica set so majority write concern is
syntactically/semantically valid while keeping the lab small.
This does not reproduce multi-member majority durability or
network failover. Host exposure is loopback-only on
127.0.0.1:27206. Authentication and TLS are
disabled only for disposable local labs; production security
remains the Chapter 22 prerequisite. Default read/write concern
and primary read preference are used unless a step states
otherwise.
FCV is observed and never changed in mandatory labs.
Atlas, Enterprise Advanced, Search, Vector Search, and KMS are
optional unless explicitly labeled. The benchmark code uses a
fixed random seed, a skewed tenant distribution, realistic
payload sizes, a bounded connection pool, warm-up, scheduled
arrivals, and explicit percentile calculation. It is a
measurement harness—not a universal performance claim. Runtime
performance, failover, capacity, and upgrade labs were not
executed in the generation environment, so metric values,
latency distributions, queue depths, oplog windows, replication
lag, incident times, and upgrade durations must be measured
locally rather than copied as invented output.
1. AtlasMart problem: a benchmark can lie without fabricating a single number
A benchmark that issues the next request only after the previous request finishes automatically reduces offered load when the database slows. That coordinated omission can hide overload: the system appears to have tolerable latency because the client stopped asking at the intended rate. Likewise, uniform keys and tiny documents can erase the hotspot and cache behavior of the real application.
Before benchmarking, write down the workload contract: document sizes, tenant/key skew, query mix, indexes, read/write concern, read preference, pool size, concurrency, offered rate, warm/cold cache, aggregation shapes, and failure-free or failure-injected topology.
| Control | Lab choice | Why record it |
|---|---|---|
| Dataset | 50k orders with payload variation | Working set and BSON shape influence cache/I/O |
| Skew | few hot tenants plus long tail | Uniform keys hide hotspots |
| Durability | majority writes on one-member RS | Keeps the acknowledgement contract fixed; not a HA model |
| Pool/concurrency | bounded PyMongo pool and worker count | Unbounded clients can benchmark connection creation instead of database work |
| Arrival model | scheduled arrivals | Makes scheduler delay/overload visible |
| Output | p50/p95/p99/max + errors + completed/offered | Averages alone hide tails and dropped work |
2. Prepare the indexed data set
docker rm -f atlasmart-ch26-l3 2>/dev/null || truedocker volume rm atlasmart-ch26-l3-db 2>/dev/null || truedocker run -d --name atlasmart-ch26-l3 -p 127.0.0.1:27206:27017 -v atlasmart-ch26-l3-db:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet rs26l3 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27206/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27206/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"rs26l3",members:[{_id:0,host:"atlasmart-ch26-l3:27017"}]})'until mongosh "mongodb://127.0.0.1:27206/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; donepython -m pip install --disable-pip-version-check 'pymongo==4.17.0'
const d=db.getSiblingDB("atlasmart");d.orders_ch26_l3.drop();d.orders_ch26_l3.createIndex({tenantId:1,status:1,createdAt:-1},{name:"idx_tenant_status_created"});d.orders_ch26_l3.createIndex({customerId:1,createdAt:-1},{name:"idx_customer_created"});for (let b=0;b<50;b++) { const docs=[]; for (let i=0;i<1000;i++) { const n=b*1000+i; const tenant=n%10<6?`tenant-hot-${n%3}`:`tenant-tail-${n%97}`; docs.push({orderId:`O-${String(n).padStart(7,"0")}`,tenantId:tenant,customerId:`C-${n%7000}`,status:["paid","packed","shipped","cancelled"][n%4],totalCents:1000+n%60000,createdAt:new Date(Date.UTC(2026,8,1)+n*250),items:Array.from({length:1+n%5},(_,j)=>({sku:`SKU-${(n+j)%5000}`,qty:1+(j%3)})),notes:"x".repeat(100+n%900)}); } d.orders_ch26_l3.insertMany(docs,{writeConcern:{w:"majority"}});}printjson({count:d.orders_ch26_l3.countDocuments(),indexes:d.orders_ch26_l3.getIndexes().map(x=>x.name)});
3. Run a scheduled-arrival mixed workload
from pymongo import MongoClient, WriteConcernfrom concurrent.futures import ThreadPoolExecutorfrom time import perf_counter, sleepimport random, statistics, threadingURI="mongodb://127.0.0.1:27206/?replicaSet=rs26l3&directConnection=true"client=MongoClient(URI,maxPoolSize=32,waitQueueTimeoutMS=2000,serverSelectionTimeoutMS=3000)db=client.atlasmartcoll=db.get_collection("orders_ch26_l3",write_concern=WriteConcern("majority"))rng=random.Random(2603)lock=threading.Lock(); rows=[]def tenant(): return f"tenant-hot-{rng.randrange(3)}" if rng.random()<0.7 else f"tenant-tail-{rng.randrange(97)}"def one(kind,scheduled): start=perf_counter(); sched_lag_ms=max(0.0,(start-scheduled)*1000) ok=True try: t=tenant() if kind<0.65: list(coll.find({"tenantId":t,"status":"paid"},{"_id":0,"orderId":1,"totalCents":1}).sort("createdAt",-1).limit(25)) elif kind<0.85: coll.update_one({"tenantId":t,"status":"packed"},{"$set":{"benchTouchedAt":__import__('datetime').datetime.now(__import__('datetime').timezone.utc)}}) else: list(coll.aggregate([{"$match":{"tenantId":t}},{"$group":{"_id":"$status","n":{"$sum":1},"revenue":{"$sum":"$totalCents"}}}])) except Exception: ok=False end=perf_counter() with lock: rows.append(((end-start)*1000,sched_lag_ms,ok))def phase(seconds,rate_per_sec,workers): rows.clear(); base=perf_counter(); futures=[] with ThreadPoolExecutor(max_workers=workers) as pool: total=int(seconds*rate_per_sec) for i in range(total): scheduled=base+i/rate_per_sec delay=scheduled-perf_counter() if delay>0: sleep(delay) futures.append(pool.submit(one,rng.random(),scheduled)) for f in futures: f.result() return list(rows), total# Warm-up: discard these results.phase(5,80,16)measured,offered=phase(20,120,16)lat=sorted(x[0] for x in measured if x[2]); lag=sorted(x[1] for x in measured)def pct(a,p): return a[min(len(a)-1,max(0,int(round((len(a)-1)*p))))] if a else Noneprint({ "offered":offered,"completed":len(measured),"ok":sum(1 for x in measured if x[2]),"errors":sum(1 for x in measured if not x[2]), "latency_ms":{"p50":pct(lat,.50),"p95":pct(lat,.95),"p99":pct(lat,.99),"max":max(lat) if lat else None}, "scheduler_lag_ms":{"p95":pct(lag,.95),"p99":pct(lag,.99),"max":max(lag) if lag else None}})
If scheduler lag rises, the harness is no longer delivering work at the intended schedule. Report that as overload evidence instead of silently calling the reduced achieved rate “the benchmark throughput.”
4. Correlate the same window with server evidence
const a=db.getSiblingDB("admin");const d=db.getSiblingDB("atlasmart");const s=a.serverStatus();printjson({opcounters:s.opcounters,connections:s.connections,queues:s.queues,network:s.network,wtCache:{bytes:s.wiredTiger.cache["bytes currently in the cache"],pagesRead:s.wiredTiger.cache["pages read into cache"],pagesWritten:s.wiredTiger.cache["pages written from cache"]}});printjson(d.orders_ch26_l3.find({tenantId:"tenant-hot-1",status:"paid"}).sort({createdAt:-1}).limit(25).explain("executionStats"));
Snapshot counters before and after each measured phase if you need rates. An explain plan validates one query shape’s access path but does not substitute for workload-level latency measurement.
5. Change exactly one factor
Good experiments preserve the dataset, offered load, random seed, topology, durability, connection pool, and warm-up method while changing one controlled factor. Examples include removing an unnecessary index, changing one query shape, increasing one hot-tenant index prefix, or comparing two tested concurrency levels. Do not change indexes, cache, pool size, document shape, and durability at the same time and then attribute improvement to one of them.
| Bad experiment | Repair |
|---|---|
| Uniform random tenants instead of production skew | Encode a measured or explicitly synthetic skew distribution |
| Closed-loop client hides overload | Schedule arrivals and report delivery lag/dropped work |
| Only average latency | Report p50/p95/p99/max and errors |
| Change five knobs at once | Change one factor and repeat the identical phase |
| Run once on a cold laptop and call it capacity | Repeat, record environment/cache state, and use staging hardware for decisions |
6. Production judgment
Benchmark only questions you can operationalize: “Can this topology meet the SLO at the observed peak mix with one member unavailable?” is stronger than “How many ops/sec can MongoDB do?” Preserve client and server evidence, report confidence/repeatability, and keep coordinated omission visible. The next lesson uses the same discipline for upgrades: binary compatibility, FCV, driver compatibility, backup, and rollback are separate gates.
Check your understanding
- What is coordinated omission?
- Why use a fixed random seed?
- Why report scheduler lag?
- Why can a one-member majority-write lab not prove production durability?
- What must stay constant in a one-factor experiment?
Review the answers
1. A measurement error where the load generator stops scheduling requests during slow periods, hiding the latency users would have experienced at the intended arrival rate.
2. It makes the synthetic workload distribution repeatable across controlled comparisons.
3. It shows whether the harness itself fell behind the intended arrival schedule under overload.
4. The acknowledgement still depends on one physical host and has no independent majority failure domains.
5. Dataset, workload schedule/skew, durability, topology, client pool, warm-up, and measurement method unless one of those is the factor being tested.
Authoritative references
Operational fields, thresholds, upgrade paths, FCV behavior, Atlas metrics, and driver compatibility evolve. Re-check the current documentation for the exact server patch, deployment topology, driver, Atlas tier, and target upgrade/downgrade path before changing production systems.
- serverStatus command
- db.stats() / dbStats
- $collStats aggregation stage
- Database Profiler
- db.setProfilingLevel()
- $currentOp aggregation stage
- MongoDB Log Messages
- Explain Results
- Replication
- Replica Set Oplog
- Check Replica Set Replication Lag
- WiredTiger Storage Engine
- Atlas Monitoring and Alerts
- Atlas Monitoring and Alert Guidance
- Atlas Alert Basics
- Atlas Metrics
- PyMongo Release Notes
- PyMongo Upgrade Guidance
- MongoDB 8.3 Release Notes
- Upgrade 8.2 to 8.3
- Upgrade 8.2 Replica Set to 8.3
- Upgrade 8.2 Sharded Cluster to 8.3
- MongoDB 8.3 Compatibility Changes
- Downgrade 8.3 to 8.2
- MongoDB Versioning
- Backup Methods
- mongosh Release Notes