Follow one AtlasMart document from the MongoDB collection abstraction into WiredTiger tables, pages, cache, and on-disk files without treating implementation internals as an application API.
WiredTiger Tables, B-Trees, Pages, Cache, File Layout, and the Storage-Engine Boundary
AtlasMart sees query latency rise as its catalog and orders grow. Before changing cache sizes or storage settings, the team needs a precise model of which bytes live in MongoDB documents, WiredTiger pages, the internal cache, the operating-system filesystem cache, and data files.
Learning objectives
Define the storage-engine boundary between MongoDB logical collections/indexes and WiredTiger tables, B-trees, pages, cache, filesystem cache, and files.
Read documented serverStatus and collection
statistics without treating internal WiredTiger field names
as stable application contracts.
Distinguish compressed on-disk representations from collection data in the WiredTiger internal cache.
Connect working-set size and page traffic to observable cache/I/O evidence before considering any cache setting.
Avoid direct manipulation of WiredTiger files and explain why replication is not a substitute for storage-engine understanding or backup.
This lesson pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where a driver workload is useful.
Topology: disposable standalone. The host publishes only
127.0.0.1:27186. Authentication and TLS are
disabled only for this isolated disposable lab; production
security remains the Chapter 22 prerequisite. Default read/write
concern and primary read preference are used unless a comparison
says otherwise.
FCV is observed and never changed.
Atlas/Search/Vector Search/KMS/Enterprise capabilities are not
required. WiredTiger internals are treated as version-sensitive
implementation details; use supported MongoDB commands and
metrics instead of editing .wt files or
undocumented knobs. The default cache is left unchanged in
Lesson 1 so the learner observes the server-selected maximum
before the bounded-cache experiment in Lesson 4. Product runtime
labs were not executed in the generation environment, so cache
ratios, checkpoint durations, journal sync times, disk bytes,
and latency percentiles must be measured locally rather than
copied as invented values.
1. AtlasMart problem: “the database is on disk” is not a useful performance model
An AtlasMart order document is a logical BSON object to the application. WiredTiger, MongoDB’s default storage engine, persists collection and index data in storage-engine tables implemented with B-tree structures. Those tables contain internal and leaf pages; pages move between disk and the WiredTiger internal cache as the server reads and modifies data. Separately, the operating system can keep compressed data-file blocks in its filesystem cache. The same logical document therefore crosses several representations, and a symptom such as “high memory” or “high disk I/O” cannot be diagnosed until the layer is identified.
| Layer | What it contains | Compression / lifetime | Supported evidence |
|---|---|---|---|
| MongoDB logical layer | BSON documents, collection/index definitions, query plans | Logical view; not a file format contract |
listCollections, getIndexes(),
explain()
|
| WiredTiger internal cache |
Pages/updates used by the running mongod
|
Collection data is represented uncompressed for manipulation; indexes use an internal representation | serverStatus().wiredTiger.cache |
| Filesystem cache | Operating-system cached file blocks | Same compressed representation as data files | Host memory/I/O telemetry; not a MongoDB collection quota |
| WiredTiger files |
Collection/index tables plus metadata and journal under
dbPath
|
Version-sensitive internal format |
stats().wiredTiger.uri, creation string,
safe directory listing
|
| Journal | Write-ahead log records between checkpoints | Compressed separately; recovery input |
serverStatus().wiredTiger.log, journal
directory listing
|
2. Start the lab and capture the storage-engine identity before load
docker rm -f atlasmart-ch24-l1 2>/dev/null || truedocker volume rm atlasmart-ch24-l1-db 2>/dev/null || truedocker run -d --name atlasmart-ch24-l1 \ -p 127.0.0.1:27186:27017 \ -v atlasmart-ch24-l1-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27186/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27186/admin" --quiet --eval ' printjson(db.adminCommand({buildInfo:1})); printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1})); printjson(db.serverStatus().storageEngine);'
const admin = db.getSiblingDB("admin");const atlas = db.getSiblingDB("atlasmart");printjson(admin.serverStatus().storageEngine);const ss = admin.serverStatus();printjson({ wtVersion: ss.wiredTiger?.version, maximumCacheBytes: ss.wiredTiger?.cache?.["maximum bytes configured"], bytesInCache: ss.wiredTiger?.cache?.["bytes currently in the cache"], dirtyBytes: ss.wiredTiger?.cache?.["tracked dirty bytes in the cache"]});printjson(admin.runCommand({getCmdLineOpts:1}).parsed?.storage || {});
The exact wiredTiger subfields vary by server
version and platform. The stable lesson is the relationship:
cache maximum, current cache bytes, dirty bytes, page
reads/writes, eviction, log, and checkpoint metrics form
evidence about the storage path. A monitoring system should
tolerate absent/new fields instead of failing when a minor
release adds statistics.
3. Create deterministic AtlasMart data and connect logical objects to WiredTiger statistics
const d = db.getSiblingDB("atlasmart");d.orders_ch24_l1.drop();const bulk=[];for (let i=0;i<30000;i++) { bulk.push({ orderId:`O-${String(i).padStart(6,"0")}`, tenantId:`tenant-${i%8}`, customerId:`C-${i%5000}`, status:["paid","packed","shipped","returned"][i%4], createdAt:new Date(Date.UTC(2026,7,1,0,0,i%60)), totalCents:1000+(i%25000), note:"atlasmart-order-" + "x".repeat(180) }); if (bulk.length===1000) { d.orders_ch24_l1.insertMany(bulk); bulk.length=0; }}if (bulk.length) d.orders_ch24_l1.insertMany(bulk);d.orders_ch24_l1.createIndex({tenantId:1,status:1,createdAt:-1},{name:"idx_tenant_status_created"});d.orders_ch24_l1.createIndex({customerId:1},{name:"idx_customer"});const st=d.orders_ch24_l1.stats({indexDetails:true});printjson({ count:st.count, logicalBytes:st.size, storageBytes:st.storageSize, totalIndexBytes:st.totalIndexSize, wiredTigerUri:st.wiredTiger?.uri, collectionCreationString:st.wiredTiger?.creationString, indexNames:Object.keys(st.indexDetails||{})});
The collection creationString can expose values
such as format=btree and
block_compressor=snappy. That is valuable evidence
for this pinned server, but it remains a storage-engine
implementation string rather than an application contract.
storageSize also should not be interpreted as “all
bytes currently occupying RAM.”
4. Observe page traffic across a cold-to-warm workload
const admin=db.getSiblingDB("admin");const c=db.getSiblingDB("atlasmart").orders_ch24_l1;function wtSnapshot(label) { const s=admin.serverStatus(); const x=s.wiredTiger.cache; printjson({label, max:x["maximum bytes configured"], used:x["bytes currently in the cache"], dirty:x["tracked dirty bytes in the cache"], pagesRead:x["pages read into cache"], pagesWritten:x["pages written from cache"], cleanEvicted:x["unmodified pages evicted"], appEvictionMicros:x["application thread time evicting (usecs)"] });}wtSnapshot("before");for (let r=0;r<8;r++) { c.find({tenantId:`tenant-${r}`}).sort({createdAt:-1}).limit(1000).toArray();}wtSnapshot("after-first-pass");for (let r=0;r<8;r++) { c.find({tenantId:`tenant-${r}`}).sort({createdAt:-1}).limit(1000).toArray();}wtSnapshot("after-second-pass");
Do not predeclare that the second pass must be faster: the dataset may already be warm, the operating-system cache may dominate, or the query may touch pages outside the WiredTiger cache. The evidence is the measured latency and deltas in page reads/evictions, not a textbook warm-cache story.
5. File layout is observable, but files are not an application API
docker exec atlasmart-ch24-l1 sh -lc ' echo "--- dbPath files ---" find /data/db -maxdepth 2 -type f -printf "%P %s bytes\n" | sort | head -80 echo "--- journal ---" find /data/db/journal -maxdepth 1 -type f -printf "%f %s bytes\n" 2>/dev/null | sort'
Do not rename, delete, copy individual
collection-*.wt files into another running
deployment, or derive business identity from their filenames.
MongoDB’s catalog and WiredTiger metadata coordinate those
files. Use supported backup/restore or replication procedures;
Chapter 25 covers those boundaries.
6. Production judgment
A storage diagnosis begins by naming the layer and collecting deltas over the same workload window: logical size/indexes, cache occupancy and dirty bytes, page traffic, checkpoint duration, journal activity, and host I/O. A large database can perform well when the active working set and indexes fit the memory hierarchy; a smaller database can perform poorly when access is random and churns pages. Avoid changing cache size merely because “cache used” is high—cache is supposed to be used. In containers, verify the memory limit MongoDB sees because the default cache calculation may need explicit sizing when the process cannot use host RAM. The next lesson adds the durability timeline: journal and checkpoint are related, but not interchangeable.
Check your understanding
- Why can compressed storage size not be used as the amount of WiredTiger cache required?
-
What does
format=btreein a creation string prove? - Does high cache usage by itself prove cache pressure?
-
Why should application code not map collection names to
.wtfilenames? - What evidence should be captured before changing cache settings?
Review the answers
1. Collection data uses an uncompressed internal representation in the WiredTiger cache, while on-disk/filesystem-cache blocks retain compression benefits.
2. It is evidence about the pinned WiredTiger table implementation, not a stable schema or application API.
3. No. Correlate working-set latency with page reads, eviction, dirty data, queueing, and host resources.
4. Those names/catalog mappings are storage-engine internals and can change; supported MongoDB commands are the contract.
5. Cache maximum/used/dirty bytes, page reads/writes, eviction, workload latency, host memory/I/O, and the working-set shape.
Authoritative references
WiredTiger metrics and internal field names are implementation- and version-sensitive. The lesson uses documented MongoDB interfaces for evidence and requires re-checking the current server manual before relying on exact metric names or defaults in a later release.
- WiredTiger Storage Engine
- Storage FAQ
- serverStatus
- Self-Managed Diagnostics FAQ
- Journaling
- Configure Journaling
- Write Concern
- Write Operation Performance
- Configuration File Options
- Server Parameters
- mongod Options
- db.collection.stats()
- $collStats
- dbStats
- db.createCollection() Storage Engine Options
- Create Indexes and Storage Engine Options
- Performance Tuning
- Production Notes
- Log Messages
- MongoDB 8.3 Release Notes
- mongosh Release Notes
- PyMongo Release Notes