Measure the application and the database in the same time window before deciding what is slow, saturated, or risky.
serverStatus, dbStats/collStats, Logs, Profiler, currentOp, Atlas Metrics, and Alert Design
Build a layered MongoDB observability workflow from application SLOs through serverStatus, collection statistics, profiler/log evidence, current operations, Atlas metrics, and actionable alerts.
Learning objectives
Define a service-level objective (SLO) and baseline workload before interpreting MongoDB metrics or changing configuration.
Correlate serverStatus,
db.stats(), $collStats, logs,
profiler data, and $currentOp with
client-visible latency.
Use bounded profiler filters and restore the profiler state after the lab.
Distinguish point-in-time operation evidence from cumulative counters and histograms.
Design alerts around sustained risk, tail latency, queueing, and business impact instead of one noisy average.
This chapter pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where client behavior or load
generation matters. Lesson 1 uses one disposable single-member
replica set so replication and storage metrics are present while
keeping the local topology small. Host exposure is loopback-only
on 127.0.0.1:27202. Authentication and TLS are
disabled only for disposable local labs; production security
remains the Chapter 22 prerequisite. Default read/write concern
and primary read preference are used unless a step states
otherwise.
FCV is observed and never changed in mandatory labs.
Atlas, Enterprise Advanced, Search, Vector Search, and KMS are
optional unless explicitly labeled. MongoDB 8.3 adds slow
in-progress query logging via slowinprogms and
exposes tracked-memory fields such as
inUseTrackedMemBytes/peakTrackedMemBytes
in $currentOp. These fields are version-sensitive
and should not become hard-coded cross-version assumptions.
Runtime performance, failover, capacity, and upgrade labs were
not executed in the generation environment, so metric values,
latency distributions, queue depths, oplog windows, replication
lag, incident times, and upgrade durations must be measured
locally rather than copied as invented output.
1. AtlasMart problem: “MongoDB is slow” is not a diagnosis
AtlasMart receives a page saying “checkout latency doubled.” That is a symptom, not a mechanism. A useful investigation starts with a service-level objective (SLO): a measurable reliability or latency target such as “99% of checkout reads complete within the application’s agreed budget over the evaluation window.” The lesson uses illustrative lab budgets only; production values must come from the product contract and real workload.
A database metric has meaning only when aligned to the same time window as client observations. A cumulative counter that rose over six hours cannot explain a 30-second latency spike unless you calculate its change over that 30-second window. A mean latency can also hide a severe tail. Record p50, p95, p99, error rate, offered load, completed load, and queueing together.
| Signal | Question it answers | Common misuse |
|---|---|---|
| Client p95/p99 | Are users seeing a tail-latency problem? | Looking only at server averages |
serverStatus() deltas |
Did global work, connections, queueing, cache, or replication change? | Treating cumulative totals as rates |
$collStats |
Which namespace shows latency/scan/storage evidence? | Assuming collection latency histograms equal end-to-end request latency |
| Profiler / logs | Which concrete operation shapes were slow or expensive? | Leaving invasive profiling enabled indefinitely |
$currentOp |
What is active now? | Using one snapshot to infer a historical incident |
2. Establish the baseline and seed an observable workload
docker rm -f atlasmart-ch26-l1 2>/dev/null || truedocker volume rm atlasmart-ch26-l1-db 2>/dev/null || truedocker run -d --name atlasmart-ch26-l1 \ -p 127.0.0.1:27202:27017 -v atlasmart-ch26-l1-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --replSet rs26l1 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27202/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27202/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"rs26l1",members:[{_id:0,host:"atlasmart-ch26-l1:27017"}]})'until mongosh "mongodb://127.0.0.1:27202/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; done
const d=db.getSiblingDB("atlasmart");d.orders_ch26_l1.drop();const batch=[];for (let i=0;i<20000;i++) batch.push({ orderId:`O-${String(i).padStart(6,"0")}`, tenantId:`tenant-${i%20}`, customerId:`C-${i%4000}`, status:["paid","packed","shipped","cancelled"][i%4], totalCents:1000+(i%25000), createdAt:new Date(Date.UTC(2026,8,1)+i*1000), notes:`atlasmart order ${i} `+"x".repeat(i%80)});d.orders_ch26_l1.insertMany(batch);d.orders_ch26_l1.createIndex({tenantId:1,status:1,createdAt:-1},{name:"idx_tenant_status_created"});printjson({count:d.orders_ch26_l1.countDocuments(),indexes:d.orders_ch26_l1.getIndexes().map(x=>x.name)});// Indexed shape.for (let i=0;i<50;i++) d.orders_ch26_l1.find({tenantId:"tenant-3",status:"paid"}).sort({createdAt:-1}).limit(25).toArray();// Deliberately weak shape for evidence: unanchored regex on notes.for (let i=0;i<8;i++) d.orders_ch26_l1.find({notes:/order 19/}).limit(20).toArray();
The weak regex query is safe only because the collection is disposable and small. The point is not to manufacture a universal latency number; it is to create two different query shapes and then correlate their evidence.
3. Read global, database, and collection evidence
const admin=db.getSiblingDB("admin");const d=db.getSiblingDB("atlasmart");const s=admin.serverStatus();printjson({ version:s.version, uptime:s.uptime, connections:s.connections, opcounters:s.opcounters, mem:s.mem, queues:s.queues, wtCache:s.wiredTiger?.cache && { bytesInCache:s.wiredTiger.cache["bytes currently in the cache"], dirtyBytes:s.wiredTiger.cache["tracked dirty bytes in the cache"], pagesRead:s.wiredTiger.cache["pages read into cache"], pagesWritten:s.wiredTiger.cache["pages written from cache"] }});printjson(d.stats({scale:1024*1024}));printjson(d.orders_ch26_l1.aggregate([{$collStats:{ latencyStats:{histograms:true},storageStats:{},queryExecStats:{},count:{}}}]).next());printjson(admin.runCommand({getParameter:1,featureCompatibilityVersion:1}));
$collStats.latencyStats accumulates operations
since process start and reflects collection lock acquisition,
not full application request time.
queryExecStats.collectionScans tells you that
collection scans occurred, but not which business request caused
each one. Use these fields to narrow the search, then join them
to query/log/client evidence.
4. Use profiler and logs as bounded evidence—not permanent debug mode
The database profiler writes operation details to
system.profile and can add overhead or expose
sensitive query data. MongoDB 8.3 also separates the ordinary
slow-operation threshold from slowinprogms, the
threshold for logging a query while it is still running. In
production, prefer the least invasive evidence that answers the
question and control retention/access to diagnostics.
const d=db.getSiblingDB("atlasmart");const before=d.getProfilingStatus();printjson({before});// Filtered profiling: only operations on this disposable lesson collection over 1 ms.d.setProfilingLevel(1,{filter:{ns:"atlasmart.orders_ch26_l1",millis:{$gt:1}}});for (let i=0;i<5;i++) d.orders_ch26_l1.find({notes:/order 19/}).limit(20).toArray();printjson(d.system.profile.find({ns:"atlasmart.orders_ch26_l1"},{op:1,ns:1,millis:1,docsExamined:1,keysExamined:1,planSummary:1,queryHash:1,planCacheKey:1}).sort({ts:-1}).limit(10).toArray());d.setProfilingLevel(0);printjson({after:d.getProfilingStatus()});
docker logs --since 5m atlasmart-ch26-l1 2>&1 | tail -n 120
When a profiler filter is present,
slowms and sampleRate do not control
which operations match that filter. Also remember that profiling
is unavailable on mongos; on a router, the related
settings influence diagnostic logging instead.
5. Ask “what is running now?” with $currentOp
const admin=db.getSiblingDB("admin");printjson(admin.aggregate([ {$currentOp:{allUsers:true,idleConnections:false,idleSessions:false}}, {$match:{active:true}}, {$project:{_id:0,opid:1,op:1,ns:1,secs_running:1,microsecs_running:1,queryShapeHash:1,numYields:1,writeConflicts:1,inUseTrackedMemBytes:1,peakTrackedMemBytes:1,command:1}}, {$sort:{microsecs_running:-1}}, {$limit:20}]).toArray());
$currentOp must be the first aggregation stage and
is the preferred interface over the deprecated
currentOp command. MongoDB 8.3 adds tracked-memory
fields that can help identify memory-heavy active operations. A
zero-row result simply means no matching long-lived operation
existed at that instant.
6. Alert design: sustained risk beats a single average
Atlas exposes managed metrics and more than two hundred alertable event types, but an alert is useful only if it maps to an operator action. For MongoDB 7.0+ do not use the dynamic “tickets available” value as the primary overload alert; sustained queued readers/writers is more meaningful. Likewise, monitor oplog window together with replication headroom rather than alerting on replication lag alone.
| Risk | Better evidence combination | Runbook question |
|---|---|---|
| Tail latency | application p95/p99 + queueing + slow shapes | Which operation shape or resource wait changed? |
| Connection storm | connections/current + app pool metrics + admission queue | Did clients multiply connections or did requests get stuck? |
| Replication risk | lag + oplog window + headroom + secondary disk/cache | Can the secondary catch up before history rolls over? |
| Disk pressure | free bytes + growth rate + checkpoint/journal latency | When will the disk cross the safe operating floor? |
| Query regression | targeting/scan ratios + profiler/log + deploy marker | Which release/query shape changed the access path? |
from statistics import mediansamples=[ {"minute":1,"p99_ms":38,"queued":0}, {"minute":2,"p99_ms":42,"queued":1}, {"minute":3,"p99_ms":145,"queued":22}, {"minute":4,"p99_ms":152,"queued":25}, {"minute":5,"p99_ms":161,"queued":31},]window=samples[-3:]alert=all(x["p99_ms"]>120 and x["queued"]>10 for x in window)print({"window":window,"median_p99":median(x["p99_ms"] for x in window),"alert":alert})
Check your understanding
-
Why take deltas of cumulative
serverStatuscounters? - Why can p99 be more useful than an average for customer-facing latency?
- Why restore the profiler to level 0 after the lab?
-
What does an empty
$currentOpresult prove? - Why avoid tickets-available alerts on MongoDB 7.0+?
Review the answers
1. Because totals since startup do not directly describe the rate during the incident window.
2. Averages can hide a small but operationally important tail of very slow requests.
3. Profiling can add overhead and capture sensitive operation data; it should be bounded to the diagnostic need.
4. Only that no matching operation was active at the sampling instant; it does not disprove a historical incident.
5. Admission tickets are dynamically adjusted; sustained queueing is the more useful overload signal.
7. Production judgment
Build one evidence timeline: client SLOs, release/configuration markers, server metrics, namespace statistics, active operations, slow-operation evidence, logs, and infrastructure metrics. Prefer trends and distributions over one samples, and keep diagnostic collection proportionate to its privacy and overhead cost. The next lesson converts these observations into capacity headroom rather than waiting for a metric to hit a hard limit.
Authoritative references
Operational fields, thresholds, upgrade paths, FCV behavior, Atlas metrics, and driver compatibility evolve. Re-check the current documentation for the exact server patch, deployment topology, driver, Atlas tier, and target upgrade/downgrade path before changing production systems.
- serverStatus command
- db.stats() / dbStats
- $collStats aggregation stage
- Database Profiler
- db.setProfilingLevel()
- $currentOp aggregation stage
- MongoDB Log Messages
- Explain Results
- Replication
- Replica Set Oplog
- Check Replica Set Replication Lag
- WiredTiger Storage Engine
- Atlas Monitoring and Alerts
- Atlas Monitoring and Alert Guidance
- Atlas Alert Basics
- Atlas Metrics
- PyMongo Release Notes
- PyMongo Upgrade Guidance
- MongoDB 8.3 Release Notes
- Upgrade 8.2 to 8.3
- Upgrade 8.2 Replica Set to 8.3
- Upgrade 8.2 Sharded Cluster to 8.3
- MongoDB 8.3 Compatibility Changes
- Downgrade 8.3 to 8.2
- MongoDB Versioning
- Backup Methods
- mongosh Release Notes