Measure the application and the database in the same time window before deciding what is slow, saturated, or risky.

serverStatus, dbStats/collStats, Logs, Profiler, currentOp, Atlas Metrics, and Alert Design

Build a layered MongoDB observability workflow from application SLOs through serverStatus, collection statistics, profiler/log evidence, current operations, Atlas metrics, and actionable alerts.

Advanced120–220 minutesObservability and alert-design labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Define a service-level objective (SLO) and baseline workload before interpreting MongoDB metrics or changing configuration.

02

Correlate serverStatus, db.stats(), $collStats, logs, profiler data, and $currentOp with client-visible latency.

03

Use bounded profiler filters and restore the profiler state after the lab.

04

Distinguish point-in-time operation evidence from cumulative counters and histograms.

05

Design alerts around sustained risk, tail latency, queueing, and business impact instead of one noisy average.

Reproducible lab baseline

This chapter pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where client behavior or load generation matters. Lesson 1 uses one disposable single-member replica set so replication and storage metrics are present while keeping the local topology small. Host exposure is loopback-only on 127.0.0.1:27202. Authentication and TLS are disabled only for disposable local labs; production security remains the Chapter 22 prerequisite. Default read/write concern and primary read preference are used unless a step states otherwise. FCV is observed and never changed in mandatory labs. Atlas, Enterprise Advanced, Search, Vector Search, and KMS are optional unless explicitly labeled. MongoDB 8.3 adds slow in-progress query logging via slowinprogms and exposes tracked-memory fields such as inUseTrackedMemBytes/peakTrackedMemBytes in $currentOp. These fields are version-sensitive and should not become hard-coded cross-version assumptions. Runtime performance, failover, capacity, and upgrade labs were not executed in the generation environment, so metric values, latency distributions, queue depths, oplog windows, replication lag, incident times, and upgrade durations must be measured locally rather than copied as invented output.

1. AtlasMart problem: “MongoDB is slow” is not a diagnosis

AtlasMart receives a page saying “checkout latency doubled.” That is a symptom, not a mechanism. A useful investigation starts with a service-level objective (SLO): a measurable reliability or latency target such as “99% of checkout reads complete within the application’s agreed budget over the evaluation window.” The lesson uses illustrative lab budgets only; production values must come from the product contract and real workload.

A database metric has meaning only when aligned to the same time window as client observations. A cumulative counter that rose over six hours cannot explain a 30-second latency spike unless you calculate its change over that 30-second window. A mean latency can also hide a severe tail. Record p50, p95, p99, error rate, offered load, completed load, and queueing together.

Signal Question it answers Common misuse
Client p95/p99 Are users seeing a tail-latency problem? Looking only at server averages
serverStatus() deltas Did global work, connections, queueing, cache, or replication change? Treating cumulative totals as rates
$collStats Which namespace shows latency/scan/storage evidence? Assuming collection latency histograms equal end-to-end request latency
Profiler / logs Which concrete operation shapes were slow or expensive? Leaving invasive profiling enabled indefinitely
$currentOp What is active now? Using one snapshot to infer a historical incident

2. Establish the baseline and seed an observable workload

start the monitoring lab
docker rm -f atlasmart-ch26-l1 2>/dev/null || truedocker volume rm atlasmart-ch26-l1-db 2>/dev/null || truedocker run -d --name atlasmart-ch26-l1 \  -p 127.0.0.1:27202:27017 -v atlasmart-ch26-l1-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --replSet rs26l1 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27202/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27202/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"rs26l1",members:[{_id:0,host:"atlasmart-ch26-l1:27017"}]})'until mongosh "mongodb://127.0.0.1:27202/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; done
seed indexed and intentionally inefficient query shapes
const d=db.getSiblingDB("atlasmart");d.orders_ch26_l1.drop();const batch=[];for (let i=0;i<20000;i++) batch.push({  orderId:`O-${String(i).padStart(6,"0")}`,  tenantId:`tenant-${i%20}`,  customerId:`C-${i%4000}`,  status:["paid","packed","shipped","cancelled"][i%4],  totalCents:1000+(i%25000),  createdAt:new Date(Date.UTC(2026,8,1)+i*1000),  notes:`atlasmart order ${i} `+"x".repeat(i%80)});d.orders_ch26_l1.insertMany(batch);d.orders_ch26_l1.createIndex({tenantId:1,status:1,createdAt:-1},{name:"idx_tenant_status_created"});printjson({count:d.orders_ch26_l1.countDocuments(),indexes:d.orders_ch26_l1.getIndexes().map(x=>x.name)});// Indexed shape.for (let i=0;i<50;i++) d.orders_ch26_l1.find({tenantId:"tenant-3",status:"paid"}).sort({createdAt:-1}).limit(25).toArray();// Deliberately weak shape for evidence: unanchored regex on notes.for (let i=0;i<8;i++) d.orders_ch26_l1.find({notes:/order 19/}).limit(20).toArray();

The weak regex query is safe only because the collection is disposable and small. The point is not to manufacture a universal latency number; it is to create two different query shapes and then correlate their evidence.

3. Read global, database, and collection evidence

capture a compact server/database/collection snapshot
const admin=db.getSiblingDB("admin");const d=db.getSiblingDB("atlasmart");const s=admin.serverStatus();printjson({  version:s.version,  uptime:s.uptime,  connections:s.connections,  opcounters:s.opcounters,  mem:s.mem,  queues:s.queues,  wtCache:s.wiredTiger?.cache && {    bytesInCache:s.wiredTiger.cache["bytes currently in the cache"],    dirtyBytes:s.wiredTiger.cache["tracked dirty bytes in the cache"],    pagesRead:s.wiredTiger.cache["pages read into cache"],    pagesWritten:s.wiredTiger.cache["pages written from cache"]  }});printjson(d.stats({scale:1024*1024}));printjson(d.orders_ch26_l1.aggregate([{$collStats:{  latencyStats:{histograms:true},storageStats:{},queryExecStats:{},count:{}}}]).next());printjson(admin.runCommand({getParameter:1,featureCompatibilityVersion:1}));

$collStats.latencyStats accumulates operations since process start and reflects collection lock acquisition, not full application request time. queryExecStats.collectionScans tells you that collection scans occurred, but not which business request caused each one. Use these fields to narrow the search, then join them to query/log/client evidence.

4. Use profiler and logs as bounded evidence—not permanent debug mode

The database profiler writes operation details to system.profile and can add overhead or expose sensitive query data. MongoDB 8.3 also separates the ordinary slow-operation threshold from slowinprogms, the threshold for logging a query while it is still running. In production, prefer the least invasive evidence that answers the question and control retention/access to diagnostics.

enable a narrow profiler filter and restore it afterward
const d=db.getSiblingDB("atlasmart");const before=d.getProfilingStatus();printjson({before});// Filtered profiling: only operations on this disposable lesson collection over 1 ms.d.setProfilingLevel(1,{filter:{ns:"atlasmart.orders_ch26_l1",millis:{$gt:1}}});for (let i=0;i<5;i++) d.orders_ch26_l1.find({notes:/order 19/}).limit(20).toArray();printjson(d.system.profile.find({ns:"atlasmart.orders_ch26_l1"},{op:1,ns:1,millis:1,docsExamined:1,keysExamined:1,planSummary:1,queryHash:1,planCacheKey:1}).sort({ts:-1}).limit(10).toArray());d.setProfilingLevel(0);printjson({after:d.getProfilingStatus()});
inspect diagnostic logs without rewriting the server configuration
docker logs --since 5m atlasmart-ch26-l1 2>&1 | tail -n 120

When a profiler filter is present, slowms and sampleRate do not control which operations match that filter. Also remember that profiling is unavailable on mongos; on a router, the related settings influence diagnostic logging instead.

5. Ask “what is running now?” with $currentOp

project only fields needed for an operational snapshot
const admin=db.getSiblingDB("admin");printjson(admin.aggregate([  {$currentOp:{allUsers:true,idleConnections:false,idleSessions:false}},  {$match:{active:true}},  {$project:{_id:0,opid:1,op:1,ns:1,secs_running:1,microsecs_running:1,queryShapeHash:1,numYields:1,writeConflicts:1,inUseTrackedMemBytes:1,peakTrackedMemBytes:1,command:1}},  {$sort:{microsecs_running:-1}},  {$limit:20}]).toArray());

$currentOp must be the first aggregation stage and is the preferred interface over the deprecated currentOp command. MongoDB 8.3 adds tracked-memory fields that can help identify memory-heavy active operations. A zero-row result simply means no matching long-lived operation existed at that instant.

6. Alert design: sustained risk beats a single average

Atlas exposes managed metrics and more than two hundred alertable event types, but an alert is useful only if it maps to an operator action. For MongoDB 7.0+ do not use the dynamic “tickets available” value as the primary overload alert; sustained queued readers/writers is more meaningful. Likewise, monitor oplog window together with replication headroom rather than alerting on replication lag alone.

Risk Better evidence combination Runbook question
Tail latency application p95/p99 + queueing + slow shapes Which operation shape or resource wait changed?
Connection storm connections/current + app pool metrics + admission queue Did clients multiply connections or did requests get stuck?
Replication risk lag + oplog window + headroom + secondary disk/cache Can the secondary catch up before history rolls over?
Disk pressure free bytes + growth rate + checkpoint/journal latency When will the disk cross the safe operating floor?
Query regression targeting/scan ratios + profiler/log + deploy marker Which release/query shape changed the access path?
simulate a windowed alert instead of alerting on one sample
from statistics import mediansamples=[  {"minute":1,"p99_ms":38,"queued":0},  {"minute":2,"p99_ms":42,"queued":1},  {"minute":3,"p99_ms":145,"queued":22},  {"minute":4,"p99_ms":152,"queued":25},  {"minute":5,"p99_ms":161,"queued":31},]window=samples[-3:]alert=all(x["p99_ms"]>120 and x["queued"]>10 for x in window)print({"window":window,"median_p99":median(x["p99_ms"] for x in window),"alert":alert})

Check your understanding

  1. Why take deltas of cumulative serverStatus counters?
  2. Why can p99 be more useful than an average for customer-facing latency?
  3. Why restore the profiler to level 0 after the lab?
  4. What does an empty $currentOp result prove?
  5. Why avoid tickets-available alerts on MongoDB 7.0+?
Review the answers

1. Because totals since startup do not directly describe the rate during the incident window.

2. Averages can hide a small but operationally important tail of very slow requests.

3. Profiling can add overhead and capture sensitive operation data; it should be bounded to the diagnostic need.

4. Only that no matching operation was active at the sampling instant; it does not disprove a historical incident.

5. Admission tickets are dynamically adjusted; sustained queueing is the more useful overload signal.

7. Production judgment

Build one evidence timeline: client SLOs, release/configuration markers, server metrics, namespace statistics, active operations, slow-operation evidence, logs, and infrastructure metrics. Prefer trends and distributions over one samples, and keep diagnostic collection proportionate to its privacy and overhead cost. The next lesson converts these observations into capacity headroom rather than waiting for a metric to hit a hard limit.

Authoritative references

Operational fields, thresholds, upgrade paths, FCV behavior, Atlas metrics, and driver compatibility evolve. Re-check the current documentation for the exact server patch, deployment topology, driver, Atlas tier, and target upgrade/downgrade path before changing production systems.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.