Build compact multi-dimensional catalog summaries with $facet, explicit and automatic buckets, index-aware stage order, and hard memory/output limits.
$facet, $bucket/$bucketAuto, Multi-Dimensional Analysis, and Memory Considerations
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning objectives
Use $facet to run independent sub-pipelines over the same input set without confusing it with a sequence of dependent branches.
Define exact $bucket boundaries and handle out-of-range/missing values deliberately with a default bucket.
Use $bucketAuto for distribution-oriented buckets while understanding that boundaries are data-dependent and approximate the requested bucket count.
Explain why $facet has a non-spillable 100 MB result limit and a final 16 MiB BSON document limit.
Preserve index eligibility with selective stages before $facet instead of placing $facet first.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim
and mongosh 2.10.0. The server is a disposable
standalone published only on loopback
127.0.0.1:27059. Authentication and TLS are
disabled only for this isolated lab. Feature Compatibility
Version (FCV) and allowDiskUseByDefault are
observed but never changed. Default read/write concern and
primary read preference apply. Atlas, Search, KMS, Enterprise
Advanced, and paid services are not required. Runtime output
shown as “expected” is documentation-derived because this
generation environment has no Docker/mongod/mongosh runtime.
The exact optimizer tree, execution counters, spill fields, and stage-specific explain shape can vary with patch version, FCV, indexes, data distribution, and topology. The lesson therefore names the invariant to verify—matched documents, traversal set/depth, facet counts, window values, indexes used, disk-use evidence, and target collection state—instead of requiring byte-for-byte explain output.
1. AtlasMart problem: one filtered catalog, several dimensions
A category page often needs several summaries—price bands, brand
counts, and rating distribution—over the same tenant/category
result set. $facet sends the same input documents
into independent sub-pipelines and returns one document whose
fields are arrays containing each facet’s results. The output of
one facet branch is not the input of another.
| Construct | Key semantic |
|---|---|
| $facet | Runs multiple independent sub-pipelines on the same incoming documents. |
| $bucket | Uses explicit ordered boundaries; lower bound is inclusive and upper bound is exclusive. |
| $bucket default | Captures values outside boundaries or missing values; without it, an out-of-range input causes an error. |
| $bucketAuto | Chooses boundaries from data in an attempt to distribute documents across the requested number of buckets. |
| facet memory |
Each facet result document is limited to 100 MB and cannot
spill to disk; allowDiskUse does not change
this.
|
| final BSON | The document emitted from the pipeline must still satisfy the 16 MiB BSON document limit. |
2. Seed values that sit exactly on bucket boundaries
docker rm -f atlasmart-mongo-ch09-l3 2>/dev/null || truedocker volume rm atlasmart-mongo-ch09-l3-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch09-l3 \ -p 127.0.0.1:27059:27017 \ -v atlasmart-mongo-ch09-l3-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27059/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(),hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1,allowDiskUseByDefault:1}))'
const c=db.products_ch09_l3;c.drop();c.insertMany([ {_id:"p1",tenantId:"tenant-a",category:"electronics",brand:"Aster",priceCents:1999,rating:4.8}, {_id:"p2",tenantId:"tenant-a",category:"electronics",brand:"Aster",priceCents:2500,rating:4.2}, {_id:"p3",tenantId:"tenant-a",category:"electronics",brand:"Bora",priceCents:4999,rating:3.9}, {_id:"p4",tenantId:"tenant-a",category:"electronics",brand:"Bora",priceCents:5000,rating:4.5}, {_id:"p5",tenantId:"tenant-a",category:"electronics",brand:"Coda",priceCents:8999,rating:4.7}, {_id:"p6",tenantId:"tenant-a",category:"electronics",brand:"Coda",priceCents:12000,rating:3.6}, {_id:"p7",tenantId:"tenant-a",category:"electronics",brand:"Aster",priceCents:25000,rating:4.1}, {_id:"p8",tenantId:"tenant-a",category:"books",brand:"Leaf",priceCents:1800,rating:4.9}, {_id:"p9",tenantId:"tenant-b",category:"electronics",brand:"Other",priceCents:2000,rating:5.0}]);c.createIndex({tenantId:1,category:1,priceCents:1});printjson({count:c.countDocuments({}),indexes:c.getIndexes()});
p2 is exactly 2500 and p4 exactly
5000. Because each bucket is [lower, upper), 2500
belongs to the second bucket and 5000 to the third.
p7 at 25000 requires the default bucket in the
chosen boundaries.
3. One facet, three independent analyses
const p=[ {$match:{tenantId:"tenant-a",category:"electronics"}}, {$facet:{ priceBands:[{$bucket:{groupBy:"$priceCents",boundaries:[0,2500,5000,10000,20000],default:"20000+",output:{count:{$sum:1},products:{$push:"$_id"}}}}], brands:[{$sortByCount:"$brand"}], ratingAuto:[{$bucketAuto:{groupBy:"$rating",buckets:3,output:{count:{$sum:1},avgPrice:{$avg:"$priceCents"}}}}] }}];printjson(c.aggregate(p).toArray());
priceBands counts: [0,2500) -> 1 (p1) [2500,5000) -> 2 (p2,p3) [5000,10000) -> 2 (p4,p5) [10000,20000) -> 1 (p6) 20000+ -> 1 (p7)brands: Aster=3, Bora=2, Coda=2ratingAuto boundaries are data-dependent; verify counts and min/max output rather than hard-coding exact cut points.
$bucketAuto is useful when the business request is
“roughly N distribution buckets.” It is the wrong tool when
user-facing labels or pricing rules require stable contractual
boundaries.
4. Controlled failure: omit the default bucket
Because one product is above the highest boundary, omitting
default causes the bucket operation to fail rather
than silently discarding the product. That is a valuable
failure: it prevents a histogram from lying by omission.
try { printjson(c.aggregate([ {$match:{tenantId:"tenant-a",category:"electronics"}}, {$bucket:{groupBy:"$priceCents",boundaries:[0,2500,5000,10000,20000]}} ]).toArray());} catch(e) { print("expected bucket failure:",e.codeName||e.name,e.message); }print("repair: add default:'20000+' for values >= 20000");
5. Index use depends on pipeline order
If $facet is the first stage, MongoDB documents
that it performs a collection scan and cannot use indexes. If an
earlier $match or $sort has already
used an index, $facet does not itself trigger a
collection scan. The two pipelines below produce equivalent
business summaries, but the second gives the selective predicate
a chance to use the compound index once before the facet.
const facetFirst=[{$facet:{ price:[{$match:{tenantId:"tenant-a",category:"electronics"}},{$bucket:{groupBy:"$priceCents",boundaries:[0,5000,10000,20000,30000]}}], brands:[{$match:{tenantId:"tenant-a",category:"electronics"}},{$sortByCount:"$brand"}]}}];const matchFirst=[{$match:{tenantId:"tenant-a",category:"electronics"}},{$facet:{ price:[{$bucket:{groupBy:"$priceCents",boundaries:[0,5000,10000,20000,30000]}}], brands:[{$sortByCount:"$brand"}]}}];print("facet-first explain"); printjson(c.explain("executionStats").aggregate(facetFirst));print("match-first explain"); printjson(c.explain("executionStats").aggregate(matchFirst));
With nine documents, wall-clock time is noise. Compare winning plans and examined counts. In production, also compare result size, branch cardinality, spill indicators in stages that can spill, and p95/p99 latency.
6. Memory rules: facet is intentionally different
Many aggregation stages can spill temporary state when disk use
is allowed. $facet is different: its intermediate
result document is capped at 100 MB and cannot spill, and the
pipeline’s final BSON document is capped at 16 MiB.
$bucket and $bucketAuto outside a
facet are memory-sensitive stages that can use disk under the
normal aggregation memory rules, but putting them inside a facet
does not remove the facet’s own limit.
Do not put unbounded arrays of product IDs or giant documents into facet outputs. Return compact counts/top-N summaries, or run separate queries/materialize a dedicated analytic result when the dimensional output itself is large.
7. Verification, cleanup, and production judgment
Verification checklist
- Nine products exist, with seven tenant-a electronics products.
- Explicit boundary counts match the listed expected result.
- 2500 and 5000 demonstrate inclusive-lower/exclusive-upper semantics.
- The no-default bucket fails because 25000 is outside the explicit ranges.
-
The match-first pipeline is checked for index use before
$facet. -
No claim is made that
allowDiskUsecan bypass the facet 100 MB limit.
Production judgment. $facet is
powerful when several compact summaries share the same selective
input set. It is risky when branches accumulate large arrays,
when $facet is placed first and forces a scan, or
when the one-document output approaches memory/BSON limits.
Track input cardinality, per-branch output bytes, facet
failures, examined counts, and latency tails. Keep tenant
filtering before faceting and treat bucket labels as API
contracts when users depend on them. If the workload is
repeatedly expensive, precompute a materialized summary rather
than forcing every request to rebuild it.
Lesson 4 keeps the documents instead of collapsing them into
facets and adds calculations over ordered partitions with
$setWindowFields.
docker rm -f atlasmart-mongo-ch09-l3docker volume rm atlasmart-mongo-ch09-l3-data
Check your understanding
- What happens to a value exactly equal to an internal bucket boundary?
- Why can omitting a default bucket be safer than silently ignoring out-of-range data?
- Can allowDiskUse make a $facet result exceed 100 MB?
- Why should a selective $match usually precede $facet?
- When is $bucketAuto inappropriate?
Review the answers
It belongs to the bucket for which that value is the inclusive lower bound; the preceding bucket’s upper bound is exclusive.
The operation fails and exposes the mismatch between data and declared boundaries rather than producing an incomplete histogram.
No. The current documentation says $facet cannot spill and its result document is limited to 100 MB.
A first-stage facet performs a collection scan; an earlier match can use an index and reduce the documents sent to every facet branch.
When boundaries themselves are contractual/stable, such as pricing bands or compliance categories, because bucketAuto boundaries depend on the data distribution.
Authoritative references
- MongoDB 8.3 release notes — Current 8.3 behavior and version-sensitive aggregation changes; re-check before reproducing.
- MongoDB aggregation pipeline — Ordered-stage execution model used throughout the chapter.
- Aggregation pipeline limits — Memory, disk-spill, stage-count, and 16 MiB output-document constraints.
- mongosh release notes — mongosh version used for the chapter commands.
- $facet stage — Independent sub-pipelines, 100 MB non-spillable limit, 16 MiB final output, and index-use behavior.
- $bucket stage — Explicit boundary semantics, default handling, and memory behavior.
- $bucketAuto stage — Automatic boundary selection, distribution caveats, and memory behavior.