Use MongoDB filtering, shaping, sorting, limiting, skipping, and computed fields with correctness-first early reduction, deterministic ordering, and measured index evidence.
$match, $project/$set, $unset, $sort, $limit, $skip, and Early-Reduction Strategy
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
AtlasMart needs a compact “latest paid orders” API. The naïve implementation can return correct data while still reading far more documents than necessary, sorting in memory, transporting private debug fields, or paginating nondeterministically. This lesson separates logical shaping from physical efficiency and shows what each common stage actually guarantees.
Use $match, $project, $set, $unset, $sort, $limit, and $skip with exact input/output semantics.
Apply selective filters and bounded results early when doing so preserves the required meaning.
Avoid manual projection folklore: MongoDB can optimize field propagation, so early $project is not automatically faster.
Use deterministic compound sorting and understand the cost model of skip-based pagination.
Prove index-supported filtering/sorting with explain rather than by intuition.
This lesson pins MongoDB Community Server
8.3.8 using
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, a disposable standalone mongod published only
on loopback 127.0.0.1:27053, and mongosh
2.10.0. The database is atlasmart;
the collection is orders_ch08_l2. Authentication
and TLS are disabled only for this isolated learning
container. The lab reads Feature Compatibility Version (FCV)
and allowDiskUseByDefault but does not change
them. Default read/write concern and primary read preference
are used on the standalone. Atlas, Search, KMS, Enterprise
Advanced, and paid services are not required. Commands were
reviewed against current official documentation; runtime
output was not generated here because
Docker/mongod/mongosh/PyMongo are unavailable in this
environment.
Expected result fragments are documentation-derived shapes and invariants, not copied from a generation-time MongoDB process. Exact explain trees, optimizer rewrites, field ordering, cursor IDs, execution counters, memory/disk-spill detail, and error text can vary with server patch, FCV, index state, dataset, and topology. Verify the logical result and the named execution evidence rather than comparing output byte-for-byte.
1. Each shaping stage changes a different dimension
| Stage | What it changes | What it does not promise |
|---|---|---|
$match |
Which documents continue. | It does not project fields or guarantee order. |
$project |
Returned fields and/or computed values. | Early placement is not automatically a performance win; MongoDB can optimize field propagation. |
$set |
Adds/replaces fields while retaining the rest. | It does not persist the computed field to the collection in a read aggregation. |
$unset |
Removes named fields; it is a projection alias. | It does not delete those fields from stored documents. |
$sort |
Order of downstream documents. | A stable business order requires a unique tie-breaker when values can tie. |
$skip |
Discards N documents from the current ordered stream. | It does not seek directly to a logical key; large offsets can still require significant work. |
$limit |
Maximum documents passed downstream. | Placed before sort/filter it can change the answer. |
2. Seed enough documents to make selectivity visible
docker rm -f atlasmart-mongo-ch08-l2 2>/dev/null || truedocker volume rm atlasmart-mongo-ch08-l2-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch08-l2 \ -p 127.0.0.1:27053:27017 \ -v atlasmart-mongo-ch08-l2-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27053/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1,allowDiskUseByDefault:1}))'
const c = db.getCollection("orders_ch08_l2");c.drop();const docs=[];for (let i=1;i<=60;i++) { docs.push({ _id:`o-${String(i).padStart(3,"0")}`, tenantId:i<=48?"tenant-a":"tenant-b", status:i%5===0?"cancelled":"paid", createdAt:new Date(Date.UTC(2026,7,1,0,i,0)), subtotalCents:1000+i*37, taxCents:100+(i%7)*11, internal:{riskScore:i%13,debugNote:`trace-${i}`}, customer:{id:`c-${i%9}`,segment:i%3===0?"business":"consumer"} });}c.insertMany(docs);c.createIndex({tenantId:1,status:1,createdAt:-1,_id:1});print("seeded",c.countDocuments({}));
The compound index begins with the tenant and status equality
predicates, followed by the requested descending date and
deterministic _id order. Chapter 10 will formalize
index ordering; here the index exists only to make the
early-reduction evidence observable.
3. Filter, order, bound, then shape the response
The pipeline below first decides which documents qualify, then establishes the business order, then bounds work to eight documents. Only after those decisions does it compute the API field and remove internal/source fields. This ordering is easy to reason about and gives the optimizer an index-eligible prefix.
const p=[ { $match:{tenantId:"tenant-a",status:"paid"} }, { $sort:{createdAt:-1,_id:1} }, { $limit:8 }, { $set:{grossCents:{$add:["$subtotalCents","$taxCents"]}} }, { $unset:["internal","subtotalCents","taxCents"] }, { $project:{_id:1,createdAt:1,"customer.segment":1,grossCents:1} }];printjson(c.aggregate(p).toArray());
o-048 grossCents=2942o-047 grossCents=2894o-046 grossCents=2846o-044 grossCents=2750o-043 grossCents=2702o-042 grossCents=2654o-041 grossCents=2683o-039 grossCents=2587
$set computes grossCents only in the
pipeline output. $unset removes fields from the
flowing documents, not from storage. The final
$project is the API contract: it deliberately
returns only the identifiers, timestamp, segment, and derived
amount.
4. Projection is not a magic “make everything earlier faster” stage
A common rule of thumb says to project away unused fields at the start of every pipeline. MongoDB documentation explicitly warns that this is unlikely to improve performance because the database can automatically determine which fields later stages need. Projection is still vital for result contracts, privacy, and semantics—but move it early only for a concrete reason and verify the plan.
If a stage truly removes or renames a field that a later stage references, the later expression sees the new shape. Optimizer rewrites preserve semantics; they do not restore a field your logical pipeline no longer contains.
const broken=c.aggregate([ { $project:{_id:1,createdAt:1,status:1} }, { $match:{tenantId:"tenant-a",status:"paid"} }]).toArray();print("broken count",broken.length);// Repair: keep tenantId until after filtering, or—preferably—filter first.print("repaired count",c.aggregate([ { $match:{tenantId:"tenant-a",status:"paid"} }, { $project:{_id:1,createdAt:1,status:1} }]).toArray().length);
broken count 0repaired count 39
5. Sorting, limiting, and deterministic pagination
Sorting by createdAt alone is ambiguous when two
orders share the same timestamp. Adding _id as a
tie-breaker creates a total order for this fixture. When a
$sort directly precedes $limit without
intervening cardinality-changing stages, MongoDB can coalesce
the limit into the sort so the sort retains only the top N
results. That is an optimizer behavior to verify, not a license
to put $limit before $sort.
const page2=c.aggregate([ { $match:{tenantId:"tenant-a",status:"paid"} }, { $sort:{createdAt:-1,_id:1} }, { $skip:8 }, { $limit:8 }, { $project:{_id:1,createdAt:1} }]).toArray();printjson(page2);
o-038, o-037, o-036, o-034, o-033, o-032, o-031, o-029
$skip is simple and useful for shallow pages, but a
large offset means the pipeline still walks/discards earlier
results. For deep pagination, a seek/range predicate based on
the last sort key is often a better design; that pattern is
revisited when indexes are taught in depth.
6. Explain proves the early prefix, not your intentions
const e=c.explain("executionStats").aggregate([ { $match:{tenantId:"tenant-a",status:"paid"} }, { $sort:{createdAt:-1,_id:1} }, { $limit:8 }, { $project:{_id:1,createdAt:1,status:1} }]);const q=e.stages?.find(s => s.$cursor)?.$cursor ?? e;printjson({ winningPlan:q.queryPlanner?.winningPlan, keys:q.executionStats?.totalKeysExamined, docs:q.executionStats?.totalDocsExamined, returnedFromQueryPrefix:q.executionStats?.nReturned});
Inspect the actual winning plan and execution counters. An index
existing in getIndexes() does not prove it was
chosen. Likewise, a low docsExamined count on sixty
documents does not prove production latency; it proves how this
fixture was accessed under this plan.
7. Missing fields and computed expressions
Aggregation expressions operate on the current document shape.
Missing values can propagate as null/missing depending on the
operator, so business calculations should define behavior
explicitly rather than assume every historical document has
every field. Chapter 07 established schema/version discipline;
aggregation code should still tolerate the versions it claims to
support. Use $type, $ifNull, or
explicit predicates where the business contract requires a
defined fallback.
Using $unset or $project to remove
internal fields is good response hygiene, but authorization
must prevent unauthorized reads at the database/application
boundary. Projection is not a substitute for tenant isolation,
field-level policy, or least-privilege credentials.
8. Verification, cleanup, and production judgment
Verification checklist
- The collection contains sixty deterministic orders and the compound index exists.
-
The result contains only tenant-a paid orders, sorted
newest-first with
_idtie-breaking. -
The result exposes
grossCentsbut notinternal,subtotalCents, ortaxCents. - The broken projection-before-match example demonstrates a shape dependency and the repaired pipeline returns the expected qualifying count.
- Page 2 follows the same deterministic order and contains at most eight documents.
- Explain shows the actual chosen access path and measured key/document counts.
Early reduction is a correctness-preserving strategy, not a universal stage-order recipe. Push tenant/security predicates and selective business filters toward indexable early stages when their semantics are independent of later computed fields. Keep result shaping explicit even if the optimizer can trim fields internally. Watch examined-to-returned ratios, in-memory sort/spill evidence, response size, deep-skip latency, and plan changes after data distribution shifts.
On replica sets and sharded deployments, read preference/read
concern and shard targeting affect where work occurs and what
data can be observed; this standalone lab makes none of those
distributed guarantees. No stage here changes stored data, so
rollback is application-code rollback. The next lesson
introduces $group, where choosing the wrong
aggregation key can produce a perfectly deterministic but
business-invalid answer.
docker rm -f atlasmart-mongo-ch08-l2docker volume rm atlasmart-mongo-ch08-l2-data
Check your understanding
- Why is an early $project not automatically a performance optimization?
- Why add _id to a sort that already uses createdAt?
- What does $unset change in a read aggregation?
- Why can deep $skip pagination become expensive?
- What should you inspect before claiming the compound index made the pipeline efficient?
Review the answers
MongoDB can automatically determine and propagate needed fields, so manual early projection often adds no benefit. Use it for semantics/contracts or after measured evidence.
It creates a deterministic total order when multiple documents have the same createdAt value.
Only the flowing/result document shape. It does not remove the stored fields from the collection.
The server generally must traverse and discard the earlier ordered results rather than jumping to an application-defined logical offset.
The actual explain winning plan and execution statistics such as keys/documents examined and returned documents, plus realistic workload latency.
Authoritative references
- MongoDB 8.3 release notes — Current stable/minor series and patch status; re-check before reproducing the lab.
- MongoDB versioning — Release-series and compatibility context for the pinned server.
- mongosh release notes — Current mongosh release used by the lesson commands.
- PyMongo release notes — Current official Python driver line used where driver cursor behavior is demonstrated.
- Aggregation pipeline — Core ordered-stage document-flow model and expression concepts.
- Aggregation pipeline limits — Stage count, 16 MiB output-document, memory, allowDiskUse, and disk-spill rules.
- $project stage — Projection semantics and result shaping.
- $set stage — Computed/additional fields in aggregation results.
- $unset stage — Field removal and its equivalence to exclusion projection.
- Query optimization — Projection guidance, ESR-oriented example, and limiting results.
- Use indexes to sort — Index-supported sort versus in-memory sort behavior.
- $limit stage — Top-N coalescence when sort directly precedes limit.
- Aggregation optimization — Optimizer rewrites and index-eligible early pipeline stages.