Trace AtlasMart documents through ordered aggregation stages, making stage semantics, cardinality changes, index-eligible early work, blocking-stage memory, and explain evidence concrete.
Pipeline Mental Model: Documents Flow Through Ordered Stages
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
AtlasMart wants a dashboard answering a simple business question: among paid orders for one tenant, which product categories generated the most line revenue? The answer is not stored in one document. MongoDB therefore evaluates an aggregation pipeline: an ordered sequence of stages where each stage consumes documents, transforms or filters them, and passes documents to the next stage. The order is part of the program.
Trace concrete AtlasMart documents after every pipeline stage instead of treating the pipeline as one opaque query.
Distinguish streaming stages from stages that can change cardinality or must buffer substantial state.
See why stage order can change both semantics and resource cost.
Recognize when an early $match and compatible index can reduce work before expensive stages.
Read explain evidence conservatively without assuming a fixed explain JSON shape.
This lesson pins MongoDB Community Server
8.3.8 using
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, a disposable standalone mongod published only
on loopback 127.0.0.1:27052, and mongosh
2.10.0. The database is atlasmart;
the collection is orders_ch08_l1. Authentication
and TLS are disabled only for this isolated learning
container. The lab reads Feature Compatibility Version (FCV)
and allowDiskUseByDefault but does not change
them. Default read/write concern and primary read preference
are used on the standalone. Atlas, Search, KMS, Enterprise
Advanced, and paid services are not required. Commands were
reviewed against current official documentation; runtime
output was not generated here because
Docker/mongod/mongosh/PyMongo are unavailable in this
environment.
Expected result fragments are documentation-derived shapes and invariants, not copied from a generation-time MongoDB process. Exact explain trees, optimizer rewrites, field ordering, cursor IDs, execution counters, memory/disk-spill detail, and error text can vary with server patch, FCV, index state, dataset, and topology. Verify the logical result and the named execution evidence rather than comparing output byte-for-byte.
1. A pipeline is an ordered document-flow program
A stage is a document whose single top-level
key names an aggregation operation such as $match,
$unwind, $set, $group,
$sort, or $limit. A
field-path expression such as
"$lines.qty" reads a field from the current
pipeline document. An accumulator such as
$sum combines values across documents inside a
group. These are MongoDB Query Language concepts, not
client-side JavaScript loops.
| Stage | Input → output effect | Typical resource question |
|---|---|---|
$match |
Keeps only documents satisfying a query predicate. | Can it run early and use an index? |
$unwind |
One array-bearing document can become zero, one, or many documents. | How much does cardinality multiply? |
$set |
Adds or replaces fields using expressions. | Are computed values type-safe and needed downstream? |
$group |
Many documents become one document per grouping key. | How many groups must be held; can memory spill? |
$sort |
Reorders the stream. | Can an index provide order, or must MongoDB buffer/sort? |
$limit |
Stops after a bounded number of documents. | Can it be moved/coalesced without changing semantics? |
2. Seed a fixture whose cardinality changes are visible
docker rm -f atlasmart-mongo-ch08-l1 2>/dev/null || truedocker volume rm atlasmart-mongo-ch08-l1-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch08-l1 \ -p 127.0.0.1:27052:27017 \ -v atlasmart-mongo-ch08-l1-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27052/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1,allowDiskUseByDefault:1}))'
const c = db.getCollection("orders_ch08_l1");c.drop();c.insertMany([ { _id:"o-801", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T09:00:00Z"), lines:[{sku:"BOOK-1",category:"books",qty:2,unitPriceCents:1500},{sku:"USB-1",category:"electronics",qty:1,unitPriceCents:2500}] }, { _id:"o-802", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T10:00:00Z"), lines:[{sku:"BOOK-2",category:"books",qty:1,unitPriceCents:4200}] }, { _id:"o-803", tenantId:"tenant-a", status:"cancelled", region:"west", createdAt:ISODate("2026-08-31T11:00:00Z"), lines:[{sku:"BOOK-1",category:"books",qty:5,unitPriceCents:1500}] }, { _id:"o-804", tenantId:"tenant-a", status:"paid", region:"east", createdAt:ISODate("2026-08-31T12:00:00Z"), lines:[{sku:"USB-1",category:"electronics",qty:3,unitPriceCents:2500},{sku:"BOOK-3",category:"books",qty:1,unitPriceCents:3000}] }, { _id:"o-805", tenantId:"tenant-b", status:"paid", region:"west", createdAt:ISODate("2026-08-31T13:00:00Z"), lines:[{sku:"BOOK-1",category:"books",qty:10,unitPriceCents:1500}] }, { _id:"o-806", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T14:00:00Z"), lines:[] }, { _id:"o-807", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T15:00:00Z") }]);c.createIndex({tenantId:1,status:1,createdAt:1});print("seeded", c.countDocuments({}));
The fixture intentionally contains an empty
lines array and one missing
lines field. With the simple field-path form of
$unwind, both disappear at that stage. Lesson 4
revisits how to preserve them deliberately.
3. Build the pipeline from business meaning, not syntax memorization
The dashboard first selects the tenant and paid state, then
turns each order line into its own pipeline document, computes
revenue per line, groups by category, sorts groups, and caps the
result set. Each stage consumes the shape produced by the
previous stage. Moving $group before
$unwind, for example, would not mean “the same
thing more efficiently”; the grouping expression would see
arrays rather than one line at a time.
const p = [ { $match: { tenantId:"tenant-a", status:"paid" } }, { $unwind: "$lines" }, { $set: { lineRevenueCents: { $multiply:["$lines.qty","$lines.unitPriceCents"] } } }, { $group: { _id:"$lines.category", revenueCents:{ $sum:"$lineRevenueCents" }, lineCount:{ $sum:1 } } }, { $sort: { revenueCents:-1, _id:1 } }, { $limit: 10 }];printjson(c.aggregate(p).toArray());
[ { _id: "books", revenueCents: 10200, lineCount: 3 }, { _id: "electronics", revenueCents: 10000, lineCount: 2 }]
For this deterministic fixture, the two resulting groups are
books and electronics. The exact BSON
display format is shell-dependent; the invariant to verify is
the arithmetic from the surviving tenant-a paid line items.
4. Trace state after every stage
During development, a long aggregation is easier to reason about when you execute prefixes of the pipeline and inspect the intermediate documents. This proves shape and cardinality independently from final totals. It also exposes surprises such as a field disappearing after projection or an array multiplying rows more than expected.
function trace(prefix, label) { const docs = c.aggregate(prefix).toArray(); print("\n--", label, "count=", docs.length); printjson(docs.slice(0, 8));}const stages = [ { $match: { tenantId:"tenant-a", status:"paid" } }, { $unwind: "$lines" }, { $set: { lineRevenueCents: { $multiply:["$lines.qty","$lines.unitPriceCents"] } } }, { $group: { _id:"$lines.category", revenueCents:{ $sum:"$lineRevenueCents" }, lineCount:{ $sum:1 } } }, { $sort: { revenueCents:-1, _id:1 } }, { $limit: 10 }];for (let i=0; i<stages.length; i++) trace(stages.slice(0,i+1), `after stage ${i+1}`);
after stage 1 count= 5after stage 2 count= 5after stage 3 count= 5after stage 4 count= 2after stage 5 count= 2after stage 6 count= 2
After $match, five tenant-a paid orders remain.
After $unwind, the two-line orders expand while
the empty/missing arrays disappear, leaving five line
documents. $group then collapses those five line
documents into two category documents. A correct final result
does not by itself prove the intermediate cardinality was
cheap.
5. Stage order can change the answer
MongoDB may perform semantics-preserving optimizer rewrites, but application authors still specify a logically ordered pipeline. A classic correctness mistake is limiting before sorting when the requirement is “latest two.” The first pipeline below chooses two documents from the pre-sort stream and then sorts only those two; the second chooses the latest two from the full matched set.
// Same operators, different order, different meaning.const wrong = c.aggregate([ { $match:{tenantId:"tenant-a",status:"paid"} }, { $limit:2 }, { $sort:{createdAt:-1,_id:1} }, { $project:{_id:1,createdAt:1} }]).toArray();const correct = c.aggregate([ { $match:{tenantId:"tenant-a",status:"paid"} }, { $sort:{createdAt:-1,_id:1} }, { $limit:2 }, { $project:{_id:1,createdAt:1} }]).toArray();print("wrong limit-then-sort:"); printjson(wrong);print("correct sort-then-limit:"); printjson(correct);
wrong limit-then-sort: [o-802, o-801]correct sort-then-limit: [o-807, o-806]
Use deterministic tie-breakers such as _id when
sort keys are not unique. Otherwise repeated executions can
return different members around a tie even though every returned
document satisfies the sort value.
6. Early reduction and blocking stages
$match is usually valuable as early as semantics
allow because it reduces the number of documents entering later
work; when it is the first stage it can use a suitable index
like a normal find query. $sort and
$group are examples of stages that can need
substantial memory because they may have to retain state before
producing final output. MongoDB documents a 100 MB per-stage
memory threshold governed by
allowDiskUseByDefault and the per-command
allowDiskUse option; spill behavior is therefore a
server/configuration fact, not something to guess from pipeline
syntax.
The seven-document fixture is intentionally too small to
demonstrate disk spilling. The setup prints
allowDiskUseByDefault; production investigations
should also inspect profiler/diagnostic evidence such as
usedDisk when a stage actually spills. A small
successful lab proves semantics, not high-volume memory
safety.
7. Explain the early index-eligible prefix
explain("executionStats") runs the operation for
measurement and returns planning plus execution evidence. The
exact explain JSON format is not an API-stable schema, so the
lesson focuses on durable questions: was an index used, how many
keys/documents were examined, and how many documents survived
the indexed prefix?
const e = c.explain("executionStats").aggregate([ { $match:{tenantId:"tenant-a",status:"paid"} }, { $sort:{createdAt:1} }, { $limit:3 }]);const q = e.stages?.find(s => s.$cursor)?.$cursor ?? e;printjson({ explainVersion:e.explainVersion, winningPlan:q.queryPlanner?.winningPlan, totalKeysExamined:q.executionStats?.totalKeysExamined, totalDocsExamined:q.executionStats?.totalDocsExamined, nReturnedFromQueryPrefix:q.executionStats?.nReturned});
With the fixture index, MongoDB has an index that begins with
the equality fields and then createdAt, so it has a
plausible path to support both filtering and order. Verify the
actual winning plan rather than assuming the index name
guarantees use. Chapter 10 will treat index design
systematically.
8. Verification, cleanup, and production judgment
Verification checklist
- The seed count is seven and includes paid, cancelled, other-tenant, empty-array, and missing-array cases.
-
Pipeline-prefix tracing shows the expected cardinality
expansion at
$unwindand collapse at$group. - The category totals match arithmetic from the surviving tenant-a paid lines.
- The limit-before-sort example differs from sort-before-limit for the fixture.
- Explain shows whether the early match/sort prefix used the intended index and reports key/document examination counts.
- The lesson does not claim a disk spill from the tiny fixture.
Aggregation is appropriate when the server can transform data close to storage and return a smaller, purpose-built result. It does not create stronger durability, consistency, or authorization guarantees than the reads it performs. Expensive pipelines can increase tail latency, cache pressure, temporary disk I/O, and contention with foreground traffic; on sharded deployments, stage placement and merge behavior can also move work across shards and routers. Tenant predicates must remain part of the server-side filter rather than relying on client post-filtering. Observe latency percentiles, examined-to-returned ratios, spill indicators, group cardinality, and result size. Test realistic distributions, not only average documents.
Chapter 08 remains read-only, so rollback is usually removal/reversion of application pipeline code rather than data restoration. Chapter 09 later introduces write-back stages where rollback becomes materially different. The next lesson zooms into the everyday shaping stages and the difference between logical early reduction and cargo-cult stage reordering.
docker rm -f atlasmart-mongo-ch08-l1docker volume rm atlasmart-mongo-ch08-l1-data
Check your understanding
- Why is an aggregation pipeline more than a list of independent operators?
- What happens to an order with an empty lines array under the simple $unwind form?
- Why can limit-before-sort return the wrong “latest N” result?
- What does a small successful pipeline prove about memory spilling?
- Which explain evidence is more useful than merely seeing an index name in collection metadata?
Review the answers
Each stage consumes the documents and shape produced by previous stages, so ordering defines both meaning and resource flow.
By default it produces no output document for a missing, null, or empty-array path; Lesson 4 shows preservation options.
The limit first selects an arbitrary/preexisting subset of the matched stream, and the sort can only order that subset rather than the full candidate set.
Only logical correctness for that scale. It does not demonstrate whether a blocking stage spills, fits in production memory, or has acceptable tail latency.
The actual winning plan plus execution counters such as keys examined, documents examined, and returned documents for the measured execution.
Authoritative references
- MongoDB 8.3 release notes — Current stable/minor series and patch status; re-check before reproducing the lab.
- MongoDB versioning — Release-series and compatibility context for the pinned server.
- mongosh release notes — Current mongosh release used by the lesson commands.
- PyMongo release notes — Current official Python driver line used where driver cursor behavior is demonstrated.
- Aggregation pipeline — Core ordered-stage document-flow model and expression concepts.
- Aggregation pipeline limits — Stage count, 16 MiB output-document, memory, allowDiskUse, and disk-spill rules.
- $match stage — Early filtering semantics and index eligibility.
- $unwind stage — Array deconstruction and default missing/empty behavior.
- $group stage — Grouping semantics, accumulators, and memory restrictions.
- $limit stage — Limit semantics and sort+limit coalescence.
- Explain command — Verbosity modes, execution statistics, and explain-output stability warning.