Correlate application latency with cache, journal, checkpoint, index, and host-I/O evidence, then change one cause at a time and re-measure.

Diagnose Disk, Cache, Checkpoint, and Write Amplification Problems with Server Metrics

AtlasMart has a write-heavy order workload with several secondary indexes. The team needs a repeatable diagnostic sequence that distinguishes disk stalls, cache churn, checkpoint pressure, journal sync cost, and index-driven write amplification before attempting tuning.

Advanced120–200 minutesEvidence-driven storage diagnosis labMongoDB 8.3.8 · WiredTiger · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Create a baseline that joins application latency distributions with WiredTiger cache, journal, checkpoint, queue, and collection/index statistics.

02

Demonstrate index-driven write amplification with identical logical writes to lean and heavily indexed collections.

03

Distinguish database-side evidence from host/storage-device evidence and identify what MongoDB metrics cannot prove alone.

04

Use a change-one-factor diagnostic loop rather than adjusting cache, tickets, checkpoint interval, and compression simultaneously.

05

Produce a production-ready evidence checklist that bridges into backup and restore planning.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where a driver workload is useful. Topology: disposable standalone. The host publishes only 127.0.0.1:27190. Authentication and TLS are disabled only for this isolated disposable lab; production security remains the Chapter 22 prerequisite. Default read/write concern and primary read preference are used unless a comparison says otherwise. FCV is observed and never changed. Atlas/Search/Vector Search/KMS/Enterprise capabilities are not required. WiredTiger internals are treated as version-sensitive implementation details; use supported MongoDB commands and metrics instead of editing .wt files or undocumented knobs. The lab compares two synthetic collections and does not manipulate the host filesystem, firewall, system clock, disk capacity, or undocumented WiredTiger parameters. Optional host tools are read-only. Product runtime labs were not executed in the generation environment, so cache ratios, checkpoint durations, journal sync times, disk bytes, and latency percentiles must be measured locally rather than copied as invented values.

1. A storage incident needs a timeline, not a favorite metric

Suppose AtlasMart checkout p99 climbs from its normal baseline while CPU, disk, cache, and write volume are all changing. The useful question is not “is WiredTiger slow?” but which mechanism became the bottleneck during the same interval. Capture application latency, operation mix, cache page traffic, dirty bytes, checkpoint duration, journal sync work, execution queueing, collection/index size, and host storage latency/throughput. Then alter one cause and repeat the exact workload.

2. Start the server and create lean versus heavily indexed order collections

disposable diagnostic server
docker rm -f atlasmart-ch24-l5 2>/dev/null || truedocker volume rm atlasmart-ch24-l5-db 2>/dev/null || truedocker run -d --name atlasmart-ch24-l5 \  -p 127.0.0.1:27190:27017 \  -v atlasmart-ch24-l5-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27190/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27190/admin" --quiet --eval '  printjson(db.adminCommand({buildInfo:1}));  printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}));  printjson(db.serverStatus().storageEngine);'
create identical logical schemas with different index write cost
const d=db.getSiblingDB("atlasmart");for (const n of ["orders_lean_ch24_l5","orders_heavy_ch24_l5"]) d[n].drop();d.createCollection("orders_lean_ch24_l5"); d.createCollection("orders_heavy_ch24_l5");d.orders_lean_ch24_l5.createIndex({tenantId:1,status:1,createdAt:-1},{name:"idx_workload"});const h=d.orders_heavy_ch24_l5;h.createIndex({tenantId:1,status:1,createdAt:-1},{name:"idx_workload"});h.createIndex({customerId:1},{name:"idx_customer"});h.createIndex({paymentId:1},{name:"idx_payment"});h.createIndex({warehouseId:1,status:1},{name:"idx_warehouse_status"});h.createIndex({totalCents:1},{name:"idx_total"});h.createIndex({updatedAt:-1},{name:"idx_updated"});printjson({lean:d.orders_lean_ch24_l5.getIndexes().map(x=>x.name),heavy:h.getIndexes().map(x=>x.name)});

3. Use one snapshot function for every diagnostic phase

MongoDB-side evidence snapshot
const admin=db.getSiblingDB("admin"); const d=db.getSiblingDB("atlasmart");function snap(label) {  const s=admin.serverStatus(), c=s.wiredTiger.cache, l=s.wiredTiger.log, t=s.wiredTiger.transaction;  function cs(name) { const x=d.getCollection(name).stats(); return {size:x.size,storageSize:x.storageSize,totalIndexSize:x.totalIndexSize,indexSizes:x.indexSizes}; }  printjson({label,    cache:{used:c["bytes currently in the cache"],dirty:c["tracked dirty bytes in the cache"],pagesRead:c["pages read into cache"],pagesWritten:c["pages written from cache"],evicted:c["unmodified pages evicted"],appEvictionMicros:c["application thread time evicting (usecs)"]},    journal:{bytes:l["log bytes written"],syncOps:l["log sync operations"],syncMicros:l["log sync time duration (usecs)"]},    checkpointMs:t["transaction checkpoint most recent time (msecs)"],    executionQueue:s.queues?.execution,    lean:cs("orders_lean_ch24_l5"), heavy:cs("orders_heavy_ch24_l5")  });}snap("before");

4. Same logical writes, different index amplification

PyMongo workload with insert and update distributions
from statistics import medianfrom time import perf_counterfrom pymongo import MongoClient, WriteConcernclient=MongoClient("mongodb://127.0.0.1:27190")db=client.atlasmartdef run(name):    c=db.get_collection(name,write_concern=WriteConcern(w=1,j=True))    insert_ms=[]; update_ms=[]    for batch in range(40):        docs=[]        for j in range(250):            i=batch*250+j            docs.append({"_id":i,"tenantId":f"tenant-{i%8}","status":"paid","createdAt":i,                         "updatedAt":i,"customerId":f"C-{i%3000}","paymentId":f"P-{i}",                         "warehouseId":f"W-{i%20}","totalCents":1000+i%20000,"payload":"x"*300})        t=perf_counter(); c.insert_many(docs); insert_ms.append((perf_counter()-t)*1000)    for r in range(30):        t=perf_counter(); c.update_many({"tenantId":f"tenant-{r%8}"},{"$inc":{"totalCents":1},"$set":{"updatedAt":100000+r}}); update_ms.append((perf_counter()-t)*1000)    def q(v,p):        s=sorted(v); return s[int(p*(len(s)-1))]    return {"insert_p50_ms":median(insert_ms),"insert_p95_ms":q(insert_ms,.95),"update_p50_ms":median(update_ms),"update_p95_ms":q(update_ms,.95)}for name in ["orders_lean_ch24_l5","orders_heavy_ch24_l5"]:    print(name,run(name))

The heavily indexed collection must maintain more B-tree structures per write, but the magnitude of latency or bytes written is environment-dependent. Use the post-snapshot deltas rather than inserting a canned “6× slower” claim.

post-workload storage snapshot
const admin=db.getSiblingDB("admin"), d=db.getSiblingDB("atlasmart"), s=admin.serverStatus();const c=s.wiredTiger.cache,l=s.wiredTiger.log,t=s.wiredTiger.transaction;for (const name of ["orders_lean_ch24_l5","orders_heavy_ch24_l5"]) {  const st=d[name].stats(); printjson({name,count:st.count,storageSize:st.storageSize,totalIndexSize:st.totalIndexSize,indexSizes:st.indexSizes});}printjson({cache:{dirty:c["tracked dirty bytes in the cache"],pagesWritten:c["pages written from cache"],pagesRead:c["pages read into cache"]},journal:{bytes:l["log bytes written"],syncOps:l["log sync operations"],syncMicros:l["log sync time duration (usecs)"]},checkpointMs:t["transaction checkpoint most recent time (msecs)"],executionQueue:s.queues?.execution});

5. What each symptom suggests—and what it cannot prove alone

Observed pattern Plausible mechanism Next evidence / safer action
Rising pages-read + eviction + p95/p99 Working set/churn Confirm access pattern, index selectivity, host memory and filesystem-cache pressure; improve query/index/data locality or capacity
Dirty bytes + pages-written + checkpoint time rise Write/checkpoint/reconciliation pressure Check write volume, indexes, storage latency, checkpoint trend; do not change checkpoint interval first
Journal sync time rises with j:true workload Durability path may be storage-limited Check device latency/queue depth and journal/data contention; keep the required write concern
Persistent queues.execution + latency Storage-engine admission/backpressure under resource pressure Correlate CPU, disk, cache, locks, and workload concurrency; avoid manual ticket changes without proof
Heavy collection has much larger index bytes/write latency Index write amplification Remove only truly redundant indexes after query evidence/hide testing; do not sacrifice required invariants
CPU increases after stronger compressor Compression cost may matter Compare storage/I/O reduction and end-to-end latency with representative data before reverting

6. Add host/device evidence without changing the host

optional read-only host observations
docker stats --no-stream atlasmart-ch24-l5# If sysstat is already installed on the host, collect device latency/queue evidence:command -v iostat >/dev/null && iostat -xz 1 5 || echo "iostat not installed; use your platform/cloud disk telemetry instead"# Do not install tools or alter I/O schedulers merely to complete this lesson.

MongoDB’s serverStatus can reveal time spent synchronizing the journal and checkpoint duration, but it cannot by itself prove the physical device is the root cause. Cloud disk throttling, filesystem behavior, RAID, virtualization, noisy neighbors, and host memory pressure require platform evidence.

7. The change-one-factor loop

  1. Freeze the workload definition: data shape, indexes, read/write concern, concurrency, client pool, cache warmup, and query mix.
  2. Capture a baseline interval with application p50/p95/p99 and MongoDB/host metric deltas.
  3. State one causal hypothesis, such as “two unused indexes are increasing write amplification.”
  4. Apply one reversible change in the disposable/staging environment.
  5. Rerun the identical workload and compare both latency and mechanism metrics.
  6. Keep or roll back the change based on the invariant and evidence—not because one metric moved in the desired direction.
Deliberately wrong incident response

Changing cache size, ticket parameters, compression, journal interval, and indexes at the same time can improve or worsen latency without revealing which mechanism mattered. It also creates rollback ambiguity. One controlled change per hypothesis is the safer operational pattern.

8. Production judgment and bridge to backup

WiredTiger tuning should be the end of a diagnosis, not the beginning. Prefer correct query/index design, capacity, healthy storage, bounded concurrency, and appropriate durability settings. Treat internal statistics as version-sensitive and build monitoring around documented MongoDB interfaces. Crucially, storage-engine durability is still not backup: a healthy checkpoint and journal can faithfully persist an application mistake. Chapter 25 therefore moves from local durability mechanics to independent logical/physical backups, consistent snapshots, point-in-time recovery, and restore drills.

Check your understanding

  1. What proves index write amplification in this lab?
  2. Can a high checkpoint duration alone prove the disk is slow?
  3. Why should write concern not be weakened just because journal sync is expensive?
  4. What is the purpose of the change-one-factor loop?
  5. Why does Chapter 25 follow this chapter?
Review the answers

1. The heavy collection maintains more indexes; compare index bytes, write latency distributions, and storage/journal/cache deltas under the same logical writes.

2. No. Correlate its trend with dirty/page-write pressure and host/device latency or throttling evidence.

3. Write concern is chosen from durability/invariant requirements; first fix capacity/storage/workload causes or explicitly accept the changed failure guarantee.

4. It preserves causal interpretability and makes rollback decisions evidence-based.

5. Journals/checkpoints provide local durability and recovery, but independent backup/restore is required for logical corruption, deletion, disasters, and recovery objectives.

Authoritative references

WiredTiger metrics and internal field names are implementation- and version-sensitive. The lesson uses documented MongoDB interfaces for evidence and requires re-checking the current server manual before relying on exact metric names or defaults in a later release.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.