Diagnose cache pressure from dirty bytes, page traffic, eviction, queueing, and workload latency instead of changing cache or ticket parameters by superstition.
WiredTiger Cache Pressure, Eviction, Dirty Data, Tickets/Concurrency, and Diagnostic Metrics
AtlasMart’s working set no longer fits comfortably in a deliberately small lab cache. The resulting eviction and page traffic must be distinguished from normal cache turnover, ticket dynamics, CPU saturation, and slow storage.
Learning objectives
Identify cache occupancy, dirty bytes, page reads/writes,
eviction, and application-thread eviction metrics in
serverStatus.
Use a deliberately bounded 256 MiB cache to create a working set that can demonstrate churn without modifying production settings.
Interpret MongoDB 7.0+ dynamic read/write ticket behavior
through queues.execution rather than static
ticket folklore.
Distinguish normal cache use from sustained pressure using metric deltas plus workload latency.
Recognize rare cache-pressure errors such as
TemporarilyUnavailable without tuning
undocumented/internal parameters.
This lesson pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where a driver workload is useful.
Topology: disposable standalone. The host publishes only
127.0.0.1:27189. Authentication and TLS are
disabled only for this isolated disposable lab; production
security remains the Chapter 22 prerequisite. Default read/write
concern and primary read preference are used unless a comparison
says otherwise.
FCV is observed and never changed.
Atlas/Search/Vector Search/KMS/Enterprise capabilities are not
required. WiredTiger internals are treated as version-sensitive
implementation details; use supported MongoDB commands and
metrics instead of editing .wt files or
undocumented knobs. This lesson intentionally starts its
disposable mongod with the documented minimum
--wiredTigerCacheSizeGB 0.256 so the synthetic
working set can exceed the cache. That setting is a laboratory
constraint, not a recommended production cache size. Product
runtime labs were not executed in the generation environment, so
cache ratios, checkpoint durations, journal sync times, disk
bytes, and latency percentiles must be measured locally rather
than copied as invented values.
1. Cache “fullness” is not the diagnosis; stalled work is
WiredTiger is expected to use its cache. The meaningful question is whether the working set causes sustained page reads, evictions, dirty-data pressure, application-thread eviction work, queueing, and tail-latency growth. MongoDB 7.0+ dynamically adjusts the maximum number of storage-engine read/write transactions. Consequently, a moment with no available tickets is not by itself an overload signal; persistent execution queues and latency/resource saturation are stronger evidence.
2. Start a bounded-cache server and create an intentionally oversized working set
docker rm -f atlasmart-ch24-l4 2>/dev/null || truedocker volume rm atlasmart-ch24-l4-db 2>/dev/null || truedocker run -d --name atlasmart-ch24-l4 \ -p 127.0.0.1:27189:27017 \ -v atlasmart-ch24-l4-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_all --port 27017 --wiredTigerCacheSizeGB 0.256until mongosh "mongodb://127.0.0.1:27189/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27189/admin" --quiet --eval ' printjson(db.adminCommand({buildInfo:1})); printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1})); printjson(db.serverStatus().storageEngine);'
import randomfrom pymongo import MongoClientrng=random.Random(2404)client=MongoClient("mongodb://127.0.0.1:27189")c=client.atlasmart.cache_pressure_ch24_l4c.drop()for batch in range(120): docs=[] for j in range(250): i=batch*250+j payload=rng.randbytes(8192) docs.append({"_id":i,"tenantId":f"tenant-{i%16}","state":i%7,"payload":payload}) c.insert_many(docs,ordered=True)c.create_index([("tenantId",1),("state",1)],name="idx_tenant_state")print(c.database.command("collStats",c.name)["size"])
The exact logical size depends on BSON overhead but should exceed the 256 MiB cache. If the learner machine is resource-constrained, reduce batches and retain the same method; do not fill the host disk or allocate memory until the machine becomes unstable.
3. Build a version-tolerant metric snapshot
const admin=db.getSiblingDB("admin");function snapshot(label) { const s=admin.serverStatus(); const c=s.wiredTiger.cache; printjson({label, cache:{ max:c["maximum bytes configured"], used:c["bytes currently in the cache"], dirty:c["tracked dirty bytes in the cache"], pagesRead:c["pages read into cache"], pagesWritten:c["pages written from cache"], cleanEvicted:c["unmodified pages evicted"], appEvictionMicros:c["application thread time evicting (usecs)"], checkpointBlockedEviction:c["checkpoint blocked page eviction"] }, executionQueue:s.queues?.execution, checkpointMs:s.wiredTiger.transaction["transaction checkpoint most recent time (msecs)"], logSyncMicros:s.wiredTiger.log["log sync time duration (usecs)"] });}snapshot("baseline");
4. Drive random reads and small updates, then compare metric deltas with latency
from concurrent.futures import ThreadPoolExecutorfrom statistics import medianfrom time import perf_counterfrom random import Randomfrom pymongo import MongoClientclient=MongoClient("mongodb://127.0.0.1:27189",maxPoolSize=32)c=client.atlasmart.cache_pressure_ch24_l4def worker(worker_id): rng=Random(240400+worker_id); samples=[] for n in range(400): _id=rng.randrange(0,30000) t=perf_counter() if n%10: c.find_one({"_id":_id},{"payload":1}) else: c.update_one({"_id":_id},{"$inc":{"touches":1}}) samples.append((perf_counter()-t)*1000) return samplesall_samples=[]with ThreadPoolExecutor(max_workers=8) as ex: for group in ex.map(worker,range(8)): all_samples.extend(group)q=sorted(all_samples)print({"ops":len(q),"p50_ms":median(q),"p95_ms":q[int(.95*(len(q)-1))],"p99_ms":q[int(.99*(len(q)-1))]})
const s=db.getSiblingDB("admin").serverStatus();const c=s.wiredTiger.cache;printjson({ used:c["bytes currently in the cache"], dirty:c["tracked dirty bytes in the cache"], pagesRead:c["pages read into cache"], pagesWritten:c["pages written from cache"], cleanEvicted:c["unmodified pages evicted"], appEvictionMicros:c["application thread time evicting (usecs)"], executionQueue:s.queues?.execution, checkpointMs:s.wiredTiger.transaction["transaction checkpoint most recent time (msecs)"]});
Subtract cumulative counters from the baseline snapshot. A strong cache-pressure story requires a coherent pattern: working set exceeds the configured cache, page reads/evictions keep increasing, application eviction work may rise, and latency/tail latency degrades under the same access pattern. If only cache-used bytes are high while page churn and latency are stable, the cache is simply doing its job.
5. Ticket/concurrency signals: observe dynamic admission instead of hard-coding 128
MongoDB dynamically manages concurrent storage-engine
transactions beginning in 7.0, with an upper bound of 128 read
and 128 write tickets per node. In MongoDB 8.x,
serverStatus().queues.execution exposes execution
queueing. A persistent queue together with CPU/disk/cache
pressure and rising latency is actionable; “available tickets
reached zero once” is not. Manually changing
storageEngineConcurrentReadTransactions or
storageEngineConcurrentWriteTransactions without a
proven bottleneck can make throughput worse.
MongoDB can return TemporarilyUnavailable in rare
cache-pressure cases and records related diagnostics. This lab
is not designed to force that error. Do not keep increasing load
until the host becomes unhealthy merely to make an error appear.
6. Production judgment
Capacity-plan the working set, indexes, process overhead, filesystem cache, and other host consumers together. In a container, confirm the memory limit MongoDB sees before choosing an explicit cache size. Favor workload/query/index changes or additional resources over undocumented eviction tuning. Monitor trends and distributions, not one point. Lesson 5 turns these signals into a root-cause workflow that also includes journal, checkpoint, and index write amplification.
Check your understanding
- Why is high WiredTiger cache usage not automatically bad?
- What makes the 256 MiB cache acceptable in this lesson?
- Why is “tickets available = 0” insufficient evidence of overload in MongoDB 7.0+?
- Which cache counters help show churn?
-
Should the lab force
TemporarilyUnavailable?
Review the answers
1. A cache exists to hold useful data; pressure is demonstrated by churn, eviction work, queueing, and latency/resource symptoms.
2. It is an explicit disposable lab constraint used to make the synthetic working set exceed cache; it is not a production recommendation.
3. The server dynamically adjusts admission tickets; persistent execution queues plus latency and resource saturation are the meaningful context.
4. Pages read into cache, pages written from cache, unmodified pages evicted, dirty bytes, and application-thread eviction time, interpreted as deltas.
5. No. It is enough to explain and observe safely if it occurs; deliberately destabilizing the host is not required.
Authoritative references
WiredTiger metrics and internal field names are implementation- and version-sensitive. The lesson uses documented MongoDB interfaces for evidence and requires re-checking the current server manual before relying on exact metric names or defaults in a later release.
- WiredTiger Storage Engine
- Storage FAQ
- serverStatus
- Self-Managed Diagnostics FAQ
- Journaling
- Configure Journaling
- Write Concern
- Write Operation Performance
- Configuration File Options
- Server Parameters
- mongod Options
- db.collection.stats()
- $collStats
- dbStats
- db.createCollection() Storage Engine Options
- Create Indexes and Storage Engine Options
- Performance Tuning
- Production Notes
- Log Messages
- MongoDB 8.3 Release Notes
- mongosh Release Notes
- PyMongo Release Notes