Follow one AtlasMart document from the MongoDB collection abstraction into WiredTiger tables, pages, cache, and on-disk files without treating implementation internals as an application API.

WiredTiger Tables, B-Trees, Pages, Cache, File Layout, and the Storage-Engine Boundary

AtlasMart sees query latency rise as its catalog and orders grow. Before changing cache sizes or storage settings, the team needs a precise model of which bytes live in MongoDB documents, WiredTiger pages, the internal cache, the operating-system filesystem cache, and data files.

Advanced120–200 minutesStorage-path observability labMongoDB 8.3.8 · WiredTiger · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Define the storage-engine boundary between MongoDB logical collections/indexes and WiredTiger tables, B-trees, pages, cache, filesystem cache, and files.

02

Read documented serverStatus and collection statistics without treating internal WiredTiger field names as stable application contracts.

03

Distinguish compressed on-disk representations from collection data in the WiredTiger internal cache.

04

Connect working-set size and page traffic to observable cache/I/O evidence before considering any cache setting.

05

Avoid direct manipulation of WiredTiger files and explain why replication is not a substitute for storage-engine understanding or backup.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where a driver workload is useful. Topology: disposable standalone. The host publishes only 127.0.0.1:27186. Authentication and TLS are disabled only for this isolated disposable lab; production security remains the Chapter 22 prerequisite. Default read/write concern and primary read preference are used unless a comparison says otherwise. FCV is observed and never changed. Atlas/Search/Vector Search/KMS/Enterprise capabilities are not required. WiredTiger internals are treated as version-sensitive implementation details; use supported MongoDB commands and metrics instead of editing .wt files or undocumented knobs. The default cache is left unchanged in Lesson 1 so the learner observes the server-selected maximum before the bounded-cache experiment in Lesson 4. Product runtime labs were not executed in the generation environment, so cache ratios, checkpoint durations, journal sync times, disk bytes, and latency percentiles must be measured locally rather than copied as invented values.

1. AtlasMart problem: “the database is on disk” is not a useful performance model

An AtlasMart order document is a logical BSON object to the application. WiredTiger, MongoDB’s default storage engine, persists collection and index data in storage-engine tables implemented with B-tree structures. Those tables contain internal and leaf pages; pages move between disk and the WiredTiger internal cache as the server reads and modifies data. Separately, the operating system can keep compressed data-file blocks in its filesystem cache. The same logical document therefore crosses several representations, and a symptom such as “high memory” or “high disk I/O” cannot be diagnosed until the layer is identified.

Layer What it contains Compression / lifetime Supported evidence
MongoDB logical layer BSON documents, collection/index definitions, query plans Logical view; not a file format contract listCollections, getIndexes(), explain()
WiredTiger internal cache Pages/updates used by the running mongod Collection data is represented uncompressed for manipulation; indexes use an internal representation serverStatus().wiredTiger.cache
Filesystem cache Operating-system cached file blocks Same compressed representation as data files Host memory/I/O telemetry; not a MongoDB collection quota
WiredTiger files Collection/index tables plus metadata and journal under dbPath Version-sensitive internal format stats().wiredTiger.uri, creation string, safe directory listing
Journal Write-ahead log records between checkpoints Compressed separately; recovery input serverStatus().wiredTiger.log, journal directory listing

2. Start the lab and capture the storage-engine identity before load

start a loopback-only Community standalone
docker rm -f atlasmart-ch24-l1 2>/dev/null || truedocker volume rm atlasmart-ch24-l1-db 2>/dev/null || truedocker run -d --name atlasmart-ch24-l1 \  -p 127.0.0.1:27186:27017 \  -v atlasmart-ch24-l1-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27186/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27186/admin" --quiet --eval '  printjson(db.adminCommand({buildInfo:1}));  printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}));  printjson(db.serverStatus().storageEngine);'
inspect supported storage evidence
const admin = db.getSiblingDB("admin");const atlas = db.getSiblingDB("atlasmart");printjson(admin.serverStatus().storageEngine);const ss = admin.serverStatus();printjson({  wtVersion: ss.wiredTiger?.version,  maximumCacheBytes: ss.wiredTiger?.cache?.["maximum bytes configured"],  bytesInCache: ss.wiredTiger?.cache?.["bytes currently in the cache"],  dirtyBytes: ss.wiredTiger?.cache?.["tracked dirty bytes in the cache"]});printjson(admin.runCommand({getCmdLineOpts:1}).parsed?.storage || {});

The exact wiredTiger subfields vary by server version and platform. The stable lesson is the relationship: cache maximum, current cache bytes, dirty bytes, page reads/writes, eviction, log, and checkpoint metrics form evidence about the storage path. A monitoring system should tolerate absent/new fields instead of failing when a minor release adds statistics.

3. Create deterministic AtlasMart data and connect logical objects to WiredTiger statistics

seed orders and indexes
const d = db.getSiblingDB("atlasmart");d.orders_ch24_l1.drop();const bulk=[];for (let i=0;i<30000;i++) {  bulk.push({    orderId:`O-${String(i).padStart(6,"0")}`,    tenantId:`tenant-${i%8}`,    customerId:`C-${i%5000}`,    status:["paid","packed","shipped","returned"][i%4],    createdAt:new Date(Date.UTC(2026,7,1,0,0,i%60)),    totalCents:1000+(i%25000),    note:"atlasmart-order-" + "x".repeat(180)  });  if (bulk.length===1000) { d.orders_ch24_l1.insertMany(bulk); bulk.length=0; }}if (bulk.length) d.orders_ch24_l1.insertMany(bulk);d.orders_ch24_l1.createIndex({tenantId:1,status:1,createdAt:-1},{name:"idx_tenant_status_created"});d.orders_ch24_l1.createIndex({customerId:1},{name:"idx_customer"});const st=d.orders_ch24_l1.stats({indexDetails:true});printjson({  count:st.count,  logicalBytes:st.size,  storageBytes:st.storageSize,  totalIndexBytes:st.totalIndexSize,  wiredTigerUri:st.wiredTiger?.uri,  collectionCreationString:st.wiredTiger?.creationString,  indexNames:Object.keys(st.indexDetails||{})});

The collection creationString can expose values such as format=btree and block_compressor=snappy. That is valuable evidence for this pinned server, but it remains a storage-engine implementation string rather than an application contract. storageSize also should not be interpreted as “all bytes currently occupying RAM.”

4. Observe page traffic across a cold-to-warm workload

take before/after cache snapshots around repeated reads
const admin=db.getSiblingDB("admin");const c=db.getSiblingDB("atlasmart").orders_ch24_l1;function wtSnapshot(label) {  const s=admin.serverStatus(); const x=s.wiredTiger.cache;  printjson({label,    max:x["maximum bytes configured"],    used:x["bytes currently in the cache"],    dirty:x["tracked dirty bytes in the cache"],    pagesRead:x["pages read into cache"],    pagesWritten:x["pages written from cache"],    cleanEvicted:x["unmodified pages evicted"],    appEvictionMicros:x["application thread time evicting (usecs)"]  });}wtSnapshot("before");for (let r=0;r<8;r++) {  c.find({tenantId:`tenant-${r}`}).sort({createdAt:-1}).limit(1000).toArray();}wtSnapshot("after-first-pass");for (let r=0;r<8;r++) {  c.find({tenantId:`tenant-${r}`}).sort({createdAt:-1}).limit(1000).toArray();}wtSnapshot("after-second-pass");

Do not predeclare that the second pass must be faster: the dataset may already be warm, the operating-system cache may dominate, or the query may touch pages outside the WiredTiger cache. The evidence is the measured latency and deltas in page reads/evictions, not a textbook warm-cache story.

5. File layout is observable, but files are not an application API

list the disposable dbPath without editing it
docker exec atlasmart-ch24-l1 sh -lc '  echo "--- dbPath files ---"  find /data/db -maxdepth 2 -type f -printf "%P %s bytes\n" | sort | head -80  echo "--- journal ---"  find /data/db/journal -maxdepth 1 -type f -printf "%f %s bytes\n" 2>/dev/null | sort' 
Deliberately wrong approach

Do not rename, delete, copy individual collection-*.wt files into another running deployment, or derive business identity from their filenames. MongoDB’s catalog and WiredTiger metadata coordinate those files. Use supported backup/restore or replication procedures; Chapter 25 covers those boundaries.

6. Production judgment

A storage diagnosis begins by naming the layer and collecting deltas over the same workload window: logical size/indexes, cache occupancy and dirty bytes, page traffic, checkpoint duration, journal activity, and host I/O. A large database can perform well when the active working set and indexes fit the memory hierarchy; a smaller database can perform poorly when access is random and churns pages. Avoid changing cache size merely because “cache used” is high—cache is supposed to be used. In containers, verify the memory limit MongoDB sees because the default cache calculation may need explicit sizing when the process cannot use host RAM. The next lesson adds the durability timeline: journal and checkpoint are related, but not interchangeable.

Check your understanding

  1. Why can compressed storage size not be used as the amount of WiredTiger cache required?
  2. What does format=btree in a creation string prove?
  3. Does high cache usage by itself prove cache pressure?
  4. Why should application code not map collection names to .wt filenames?
  5. What evidence should be captured before changing cache settings?
Review the answers

1. Collection data uses an uncompressed internal representation in the WiredTiger cache, while on-disk/filesystem-cache blocks retain compression benefits.

2. It is evidence about the pinned WiredTiger table implementation, not a stable schema or application API.

3. No. Correlate working-set latency with page reads, eviction, dirty data, queueing, and host resources.

4. Those names/catalog mappings are storage-engine internals and can change; supported MongoDB commands are the contract.

5. Cache maximum/used/dirty bytes, page reads/writes, eviction, workload latency, host memory/I/O, and the working-set shape.

Authoritative references

WiredTiger metrics and internal field names are implementation- and version-sensitive. The lesson uses documented MongoDB interfaces for evidence and requires re-checking the current server manual before relying on exact metric names or defaults in a later release.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.