Trace AtlasMart documents through ordered aggregation stages, making stage semantics, cardinality changes, index-eligible early work, blocking-stage memory, and explain evidence concrete.

Pipeline Mental Model: Documents Flow Through Ordered Stages

Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.

Intermediate90–120 minutesStage-by-stage document-flow labMongoDB 8.3.8 · mongosh 2.10.0Last reviewed: September 2026

Learning outcomes

AtlasMart wants a dashboard answering a simple business question: among paid orders for one tenant, which product categories generated the most line revenue? The answer is not stored in one document. MongoDB therefore evaluates an aggregation pipeline: an ordered sequence of stages where each stage consumes documents, transforms or filters them, and passes documents to the next stage. The order is part of the program.

01

Trace concrete AtlasMart documents after every pipeline stage instead of treating the pipeline as one opaque query.

02

Distinguish streaming stages from stages that can change cardinality or must buffer substantial state.

03

See why stage order can change both semantics and resource cost.

04

Recognize when an early $match and compatible index can reduce work before expensive stages.

05

Read explain evidence conservatively without assuming a fixed explain JSON shape.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 using mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, a disposable standalone mongod published only on loopback 127.0.0.1:27052, and mongosh 2.10.0. The database is atlasmart; the collection is orders_ch08_l1. Authentication and TLS are disabled only for this isolated learning container. The lab reads Feature Compatibility Version (FCV) and allowDiskUseByDefault but does not change them. Default read/write concern and primary read preference are used on the standalone. Atlas, Search, KMS, Enterprise Advanced, and paid services are not required. Commands were reviewed against current official documentation; runtime output was not generated here because Docker/mongod/mongosh/PyMongo are unavailable in this environment.

How to read expected output

Expected result fragments are documentation-derived shapes and invariants, not copied from a generation-time MongoDB process. Exact explain trees, optimizer rewrites, field ordering, cursor IDs, execution counters, memory/disk-spill detail, and error text can vary with server patch, FCV, index state, dataset, and topology. Verify the logical result and the named execution evidence rather than comparing output byte-for-byte.

1. A pipeline is an ordered document-flow program

A stage is a document whose single top-level key names an aggregation operation such as $match, $unwind, $set, $group, $sort, or $limit. A field-path expression such as "$lines.qty" reads a field from the current pipeline document. An accumulator such as $sum combines values across documents inside a group. These are MongoDB Query Language concepts, not client-side JavaScript loops.

Stage Input → output effect Typical resource question
$match Keeps only documents satisfying a query predicate. Can it run early and use an index?
$unwind One array-bearing document can become zero, one, or many documents. How much does cardinality multiply?
$set Adds or replaces fields using expressions. Are computed values type-safe and needed downstream?
$group Many documents become one document per grouping key. How many groups must be held; can memory spill?
$sort Reorders the stream. Can an index provide order, or must MongoDB buffer/sort?
$limit Stops after a bounded number of documents. Can it be moved/coalesced without changing semantics?

2. Seed a fixture whose cardinality changes are visible

bash · isolated Chapter 08 Lesson 1 lab setup
docker rm -f atlasmart-mongo-ch08-l1 2>/dev/null || truedocker volume rm atlasmart-mongo-ch08-l1-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch08-l1 \  -p 127.0.0.1:27052:27017 \  -v atlasmart-mongo-ch08-l1-data:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27052/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1,allowDiskUseByDefault:1}))' 
javascript · seed paid/cancelled orders with arrays, empty arrays, and a missing lines field
const c = db.getCollection("orders_ch08_l1");c.drop();c.insertMany([  { _id:"o-801", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T09:00:00Z"), lines:[{sku:"BOOK-1",category:"books",qty:2,unitPriceCents:1500},{sku:"USB-1",category:"electronics",qty:1,unitPriceCents:2500}] },  { _id:"o-802", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T10:00:00Z"), lines:[{sku:"BOOK-2",category:"books",qty:1,unitPriceCents:4200}] },  { _id:"o-803", tenantId:"tenant-a", status:"cancelled", region:"west", createdAt:ISODate("2026-08-31T11:00:00Z"), lines:[{sku:"BOOK-1",category:"books",qty:5,unitPriceCents:1500}] },  { _id:"o-804", tenantId:"tenant-a", status:"paid", region:"east", createdAt:ISODate("2026-08-31T12:00:00Z"), lines:[{sku:"USB-1",category:"electronics",qty:3,unitPriceCents:2500},{sku:"BOOK-3",category:"books",qty:1,unitPriceCents:3000}] },  { _id:"o-805", tenantId:"tenant-b", status:"paid", region:"west", createdAt:ISODate("2026-08-31T13:00:00Z"), lines:[{sku:"BOOK-1",category:"books",qty:10,unitPriceCents:1500}] },  { _id:"o-806", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T14:00:00Z"), lines:[] },  { _id:"o-807", tenantId:"tenant-a", status:"paid", region:"west", createdAt:ISODate("2026-08-31T15:00:00Z") }]);c.createIndex({tenantId:1,status:1,createdAt:1});print("seeded", c.countDocuments({}));

The fixture intentionally contains an empty lines array and one missing lines field. With the simple field-path form of $unwind, both disappear at that stage. Lesson 4 revisits how to preserve them deliberately.

3. Build the pipeline from business meaning, not syntax memorization

The dashboard first selects the tenant and paid state, then turns each order line into its own pipeline document, computes revenue per line, groups by category, sorts groups, and caps the result set. Each stage consumes the shape produced by the previous stage. Moving $group before $unwind, for example, would not mean “the same thing more efficiently”; the grouping expression would see arrays rather than one line at a time.

javascript · complete category-revenue pipeline
const p = [  { $match: { tenantId:"tenant-a", status:"paid" } },  { $unwind: "$lines" },  { $set: { lineRevenueCents: { $multiply:["$lines.qty","$lines.unitPriceCents"] } } },  { $group: { _id:"$lines.category", revenueCents:{ $sum:"$lineRevenueCents" }, lineCount:{ $sum:1 } } },  { $sort: { revenueCents:-1, _id:1 } },  { $limit: 10 }];printjson(c.aggregate(p).toArray());
text · expected logical result
[  { _id: "books", revenueCents: 10200, lineCount: 3 },  { _id: "electronics", revenueCents: 10000, lineCount: 2 }]

For this deterministic fixture, the two resulting groups are books and electronics. The exact BSON display format is shell-dependent; the invariant to verify is the arithmetic from the surviving tenant-a paid line items.

4. Trace state after every stage

During development, a long aggregation is easier to reason about when you execute prefixes of the pipeline and inspect the intermediate documents. This proves shape and cardinality independently from final totals. It also exposes surprises such as a field disappearing after projection or an array multiplying rows more than expected.

javascript · run every prefix and print document counts plus samples
function trace(prefix, label) {  const docs = c.aggregate(prefix).toArray();  print("\n--", label, "count=", docs.length);  printjson(docs.slice(0, 8));}const stages = [  { $match: { tenantId:"tenant-a", status:"paid" } },  { $unwind: "$lines" },  { $set: { lineRevenueCents: { $multiply:["$lines.qty","$lines.unitPriceCents"] } } },  { $group: { _id:"$lines.category", revenueCents:{ $sum:"$lineRevenueCents" }, lineCount:{ $sum:1 } } },  { $sort: { revenueCents:-1, _id:1 } },  { $limit: 10 }];for (let i=0; i<stages.length; i++) trace(stages.slice(0,i+1), `after stage ${i+1}`);
text · expected stage cardinalities
after stage 1 count= 5after stage 2 count= 5after stage 3 count= 5after stage 4 count= 2after stage 5 count= 2after stage 6 count= 2
Expected cardinality story

After $match, five tenant-a paid orders remain. After $unwind, the two-line orders expand while the empty/missing arrays disappear, leaving five line documents. $group then collapses those five line documents into two category documents. A correct final result does not by itself prove the intermediate cardinality was cheap.

5. Stage order can change the answer

MongoDB may perform semantics-preserving optimizer rewrites, but application authors still specify a logically ordered pipeline. A classic correctness mistake is limiting before sorting when the requirement is “latest two.” The first pipeline below chooses two documents from the pre-sort stream and then sorts only those two; the second chooses the latest two from the full matched set.

javascript · controlled wrong order: limit before sort
// Same operators, different order, different meaning.const wrong = c.aggregate([  { $match:{tenantId:"tenant-a",status:"paid"} },  { $limit:2 },  { $sort:{createdAt:-1,_id:1} },  { $project:{_id:1,createdAt:1} }]).toArray();const correct = c.aggregate([  { $match:{tenantId:"tenant-a",status:"paid"} },  { $sort:{createdAt:-1,_id:1} },  { $limit:2 },  { $project:{_id:1,createdAt:1} }]).toArray();print("wrong limit-then-sort:"); printjson(wrong);print("correct sort-then-limit:"); printjson(correct);
text · expected order difference
wrong limit-then-sort:  [o-802, o-801]correct sort-then-limit: [o-807, o-806]

Use deterministic tie-breakers such as _id when sort keys are not unique. Otherwise repeated executions can return different members around a tie even though every returned document satisfies the sort value.

6. Early reduction and blocking stages

$match is usually valuable as early as semantics allow because it reduces the number of documents entering later work; when it is the first stage it can use a suitable index like a normal find query. $sort and $group are examples of stages that can need substantial memory because they may have to retain state before producing final output. MongoDB documents a 100 MB per-stage memory threshold governed by allowDiskUseByDefault and the per-command allowDiskUse option; spill behavior is therefore a server/configuration fact, not something to guess from pipeline syntax.

Do not fake a spill in a tiny lesson dataset

The seven-document fixture is intentionally too small to demonstrate disk spilling. The setup prints allowDiskUseByDefault; production investigations should also inspect profiler/diagnostic evidence such as usedDisk when a stage actually spills. A small successful lab proves semantics, not high-volume memory safety.

7. Explain the early index-eligible prefix

explain("executionStats") runs the operation for measurement and returns planning plus execution evidence. The exact explain JSON format is not an API-stable schema, so the lesson focuses on durable questions: was an index used, how many keys/documents were examined, and how many documents survived the indexed prefix?

javascript · explain an indexed match/sort/limit prefix
const e = c.explain("executionStats").aggregate([  { $match:{tenantId:"tenant-a",status:"paid"} },  { $sort:{createdAt:1} },  { $limit:3 }]);const q = e.stages?.find(s => s.$cursor)?.$cursor ?? e;printjson({  explainVersion:e.explainVersion,  winningPlan:q.queryPlanner?.winningPlan,  totalKeysExamined:q.executionStats?.totalKeysExamined,  totalDocsExamined:q.executionStats?.totalDocsExamined,  nReturnedFromQueryPrefix:q.executionStats?.nReturned});

With the fixture index, MongoDB has an index that begins with the equality fields and then createdAt, so it has a plausible path to support both filtering and order. Verify the actual winning plan rather than assuming the index name guarantees use. Chapter 10 will treat index design systematically.

8. Verification, cleanup, and production judgment

Verification checklist

  • The seed count is seven and includes paid, cancelled, other-tenant, empty-array, and missing-array cases.
  • Pipeline-prefix tracing shows the expected cardinality expansion at $unwind and collapse at $group.
  • The category totals match arithmetic from the surviving tenant-a paid lines.
  • The limit-before-sort example differs from sort-before-limit for the fixture.
  • Explain shows whether the early match/sort prefix used the intended index and reports key/document examination counts.
  • The lesson does not claim a disk spill from the tiny fixture.

Aggregation is appropriate when the server can transform data close to storage and return a smaller, purpose-built result. It does not create stronger durability, consistency, or authorization guarantees than the reads it performs. Expensive pipelines can increase tail latency, cache pressure, temporary disk I/O, and contention with foreground traffic; on sharded deployments, stage placement and merge behavior can also move work across shards and routers. Tenant predicates must remain part of the server-side filter rather than relying on client post-filtering. Observe latency percentiles, examined-to-returned ratios, spill indicators, group cardinality, and result size. Test realistic distributions, not only average documents.

Chapter 08 remains read-only, so rollback is usually removal/reversion of application pipeline code rather than data restoration. Chapter 09 later introduces write-back stages where rollback becomes materially different. The next lesson zooms into the everyday shaping stages and the difference between logical early reduction and cargo-cult stage reordering.

bash · cleanup / full reset
docker rm -f atlasmart-mongo-ch08-l1docker volume rm atlasmart-mongo-ch08-l1-data

Check your understanding

  1. Why is an aggregation pipeline more than a list of independent operators?
  2. What happens to an order with an empty lines array under the simple $unwind form?
  3. Why can limit-before-sort return the wrong “latest N” result?
  4. What does a small successful pipeline prove about memory spilling?
  5. Which explain evidence is more useful than merely seeing an index name in collection metadata?
Review the answers

Each stage consumes the documents and shape produced by previous stages, so ordering defines both meaning and resource flow.

By default it produces no output document for a missing, null, or empty-array path; Lesson 4 shows preservation options.

The limit first selects an arbitrary/preexisting subset of the matched stream, and the sort can only order that subset rather than the full candidate set.

Only logical correctness for that scale. It does not demonstrate whether a blocking stage spills, fits in production memory, or has acceptable tail latency.

The actual winning plan plus execution counters such as keys examined, documents examined, and returned documents for the measured execution.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.