Build correct MongoDB group keys and accumulators for AtlasMart analytics while preventing tenant-boundary mistakes, undefined first/last semantics, and high-cardinality memory surprises.
$group, Accumulators, Distinct-Like Analysis, and Correct Aggregation Keys
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
AtlasMart finance asks for revenue by tenant and category. A
pipeline can execute flawlessly yet produce a dangerous result
if its aggregation key omits a business
dimension. $group is therefore as much a
data-model/invariant decision as a syntax feature: every input
document is assigned to the group identified by the evaluated
_id expression, and accumulators summarize the
documents in that group.
Choose grouping keys that preserve tenant and business boundaries.
Use core accumulators such as $sum, $avg, $min, $max, $addToSet, $first, and $last with defined semantics.
Build distinct-like analysis with grouping while understanding when the distinct command may be simpler.
Make ordering explicit before order-sensitive accumulators such as $first/$last.
Connect group cardinality to memory pressure and verify selective upstream index use.
This lesson pins MongoDB Community Server
8.3.8 using
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, a disposable standalone mongod published only
on loopback 127.0.0.1:27054, and mongosh
2.10.0. The database is atlasmart;
the collection is order_lines_ch08_l3.
Authentication and TLS are disabled only for this isolated
learning container. The lab reads Feature Compatibility
Version (FCV) and allowDiskUseByDefault but does
not change them. Default read/write concern and primary read
preference are used on the standalone. Atlas, Search, KMS,
Enterprise Advanced, and paid services are not required.
Commands were reviewed against current official documentation;
runtime output was not generated here because
Docker/mongod/mongosh/PyMongo are unavailable in this
environment.
Expected result fragments are documentation-derived shapes and invariants, not copied from a generation-time MongoDB process. Exact explain trees, optimizer rewrites, field ordering, cursor IDs, execution counters, memory/disk-spill detail, and error text can vary with server patch, FCV, index state, dataset, and topology. Verify the logical result and the named execution evidence rather than comparing output byte-for-byte.
1. Grouping collapses documents around an explicit key
The $group stage outputs one document per distinct
value of its _id expression. That value can be a
scalar or a compound document. Accumulators then combine values
from every input document assigned to that group. This changes
cardinality: six input line documents might become two, four, or
six output groups depending entirely on the chosen key.
| Accumulator | Meaning in this lesson | Important boundary |
|---|---|---|
$sum |
Adds revenue or counts documents with
{$sum:1}.
|
Numeric type/overflow semantics still matter; do not silently mix units. |
$avg |
Mean of numeric inputs in a group. | Average of line revenue is not average order value unless the grouping/input grain matches orders. |
$min/$max |
Minimum/maximum value observed. | These describe input values, not chronology unless the value itself is temporal. |
$addToSet |
Unique values per group. | Output array order is unspecified; it can grow large. |
$first/$last |
First/last document/value in the current defined order. | Meaningful chronological use requires an explicit preceding sort or another documented ordering source. |
2. Seed data where a missing tenant key creates a visible bug
docker rm -f atlasmart-mongo-ch08-l3 2>/dev/null || truedocker volume rm atlasmart-mongo-ch08-l3-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch08-l3 \ -p 127.0.0.1:27054:27017 \ -v atlasmart-mongo-ch08-l3-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27054/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1,allowDiskUseByDefault:1}))'
const c=db.getCollection("order_lines_ch08_l3");c.drop();c.insertMany([ {_id:"l1",tenantId:"tenant-a",orderId:"o1",category:"books",region:"west",qty:2,revenueCents:3000,discountCents:0,createdAt:ISODate("2026-08-01T10:00:00Z")}, {_id:"l2",tenantId:"tenant-a",orderId:"o1",category:"electronics",region:"west",qty:1,revenueCents:2500,discountCents:200,createdAt:ISODate("2026-08-01T10:00:00Z")}, {_id:"l3",tenantId:"tenant-a",orderId:"o2",category:"books",region:"west",qty:1,revenueCents:4200,discountCents:0,createdAt:ISODate("2026-08-02T10:00:00Z")}, {_id:"l4",tenantId:"tenant-a",orderId:"o3",category:"books",region:"east",qty:3,revenueCents:9000,discountCents:500,createdAt:ISODate("2026-08-03T10:00:00Z")}, {_id:"l5",tenantId:"tenant-b",orderId:"o9",category:"books",region:"west",qty:10,revenueCents:15000,discountCents:0,createdAt:ISODate("2026-08-04T10:00:00Z")}, {_id:"l6",tenantId:"tenant-b",orderId:"o10",category:"electronics",region:"west",qty:2,revenueCents:5000,discountCents:0,createdAt:ISODate("2026-08-05T10:00:00Z")}]);c.createIndex({tenantId:1,region:1,createdAt:1});print("seeded",c.countDocuments({}));
3. Deliberately wrong: group only by category
The query below is syntactically valid and deterministic, but it
merges revenue from two tenants. In a multi-tenant system that
is a data-isolation/reporting bug, not a minor reporting
discrepancy. The database cannot infer that
tenantId belongs in the business key.
printjson(c.aggregate([ { $group:{ _id:"$category", revenueCents:{ $sum:"$revenueCents" }, lines:{ $sum:1 } }}, { $sort:{_id:1} }]).toArray());// This merges tenant-a and tenant-b into the same business group.
[ { _id: "books", revenueCents: 31200, lines: 4 }, { _id: "electronics", revenueCents: 7500, lines: 2 }]
A faster aggregation of the wrong groups is still wrong. Define the input grain and every business partition in the group key before tuning memory or indexes.
4. Repair with a compound aggregation key and explicit accumulators
printjson(c.aggregate([ { $group:{ _id:{tenantId:"$tenantId",category:"$category"}, lineCount:{ $sum:1 }, units:{ $sum:"$qty" }, revenueCents:{ $sum:"$revenueCents" }, avgRevenueCents:{ $avg:"$revenueCents" }, minRevenueCents:{ $min:"$revenueCents" }, maxRevenueCents:{ $max:"$revenueCents" }, orderIds:{ $addToSet:"$orderId" } }}, { $sort:{"_id.tenantId":1,"_id.category":1} }]).toArray());
tenant-a / books: lineCount=3 units=6 revenueCents=16200 avg=5400 min=3000 max=9000 orderIds=[o1,o2,o3]tenant-a / electronics: lineCount=1 units=1 revenueCents=2500 avg=2500 min=2500 max=2500 orderIds=[o1]tenant-b / books: lineCount=1 units=10 revenueCents=15000tenant-b / electronics: lineCount=1 units=2 revenueCents=5000
The output key is now a document containing both tenant and
category. orderIds is a distinct-like set only
within that business group. If the number of distinct order IDs
can become very large, materializing the entire set is itself a
growth/memory decision; a count of distinct items may require a
different pipeline shape.
5. Distinct-like analysis and the grain of the question
printjson(c.aggregate([ { $match:{tenantId:"tenant-a"} }, { $group:{_id:"$category"} }, { $sort:{_id:1} }]).toArray());
This pipeline illustrates the mechanism but does not imply that
$group should replace every
distinct command. Use the simplest operation that
expresses the requirement and benchmark the real query shape.
The important idea is that grouping creates one output document
per key and therefore can be used to deduplicate values when no
additional accumulator is needed.
6. $first and $last require a defined order
Without a meaningful order, “first” and “last” are properties of
the pipeline’s current stream, not business chronology.
AtlasMart explicitly sorts tenant-a lines by
createdAt and _id before grouping by
region, so the order-sensitive accumulators have a defined
interpretation.
printjson(c.aggregate([ { $match:{tenantId:"tenant-a"} }, { $sort:{createdAt:1,_id:1} }, { $group:{ _id:"$region", firstLine:{ $first:"$_id" }, lastLine:{ $last:"$_id" }, firstAt:{ $first:"$createdAt" }, lastAt:{ $last:"$createdAt" } }}, { $sort:{_id:1} }]).toArray());
7. Memory pressure is driven by groups and accumulator state
$group is a blocking stage: it must maintain state
for the distinct groups until results can be emitted. MongoDB
documents a 100 MB memory threshold for such stages, with
temporary-file spill behavior controlled by
allowDiskUseByDefault and per-command
allowDiskUse. High-cardinality grouping keys and
accumulators that retain arrays/sets can therefore be much more
expensive than a low-cardinality sum.
This six-document fixture cannot meaningfully exercise spill
behavior. The lesson intentionally measures grouping
correctness and the indexed prefix; production testing should
use representative group cardinality/value distributions and
observe usedDisk, temporary I/O, latency
percentiles, and resource contention.
8. Explain the selective prefix before the group
const e=c.explain("executionStats").aggregate([ { $match:{tenantId:"tenant-a",region:"west"} }, { $group:{_id:"$category",revenueCents:{$sum:"$revenueCents"}} }, { $sort:{revenueCents:-1,_id:1} }]);const q=e.stages?.find(s => s.$cursor)?.$cursor ?? e;printjson({winningPlan:q.queryPlanner?.winningPlan,keys:q.executionStats?.totalKeysExamined,docs:q.executionStats?.totalDocsExamined});
The $match can use the compound index to reduce
documents before grouping. The group itself is not “made free”
by the index; it still builds aggregate state for the surviving
categories. Explain is evidence about the chosen execution, not
a promise that future data distribution will produce the same
cost.
9. Verification, cleanup, and production judgment
Verification checklist
- The wrong category-only group visibly combines tenant-a and tenant-b revenue.
- The repaired compound key emits separate tenant/category groups.
-
lineCount, units, revenue, min/max/avg, and order ID sets match the six seeded facts. - The distinct-like pipeline returns tenant-a categories only.
- The first/last pipeline has an explicit deterministic sort before grouping.
- Explain demonstrates whether the tenant/region prefix uses the fixture index.
Use grouping when the business question truly requires
consolidation across documents. Treat the group key as an
invariant boundary, particularly for tenant, currency, region,
time bucket, or data-version dimensions. Large group cardinality
increases memory and spill risk; large
$addToSet/$push accumulator state can
also create large output documents subject to the 16 MiB
result-document limit. Watch spill indicators, temporary disk,
examined documents, group counts, output-document size, and
latency tails.
Grouping is a read transformation and does not atomically “lock”
the source dataset into a snapshot beyond the read
concern/topology semantics of the operation. If exact financial
reporting requires a specific consistency point, design that
requirement explicitly. The next lesson returns to arrays and
shows how $unwind can silently discard empty, null,
or missing cases unless preservation is requested.
docker rm -f atlasmart-mongo-ch08-l3docker volume rm atlasmart-mongo-ch08-l3-data
Check your understanding
- Why is grouping only by category unsafe in the fixture?
- What determines how many documents $group outputs?
- Why does $first need an explicit sort for chronological meaning?
- Why can $addToSet be more memory-sensitive than a simple $sum?
- What does an indexed $match before $group improve—and what does it not eliminate?
Review the answers
Because category is not the full business partition: it merges facts from tenant-a and tenant-b.
The number of distinct evaluated values of the $group _id expression among the input documents.
$first means first in the current input order; without a defined order that is not necessarily earliest by business time.
It retains distinct values for each group, so accumulator state can grow with cardinality instead of remaining a fixed-size numeric total.
It can reduce the number of documents reaching the group. The group still has to maintain aggregate state for the surviving grouping keys.
Authoritative references
- MongoDB 8.3 release notes — Current stable/minor series and patch status; re-check before reproducing the lab.
- MongoDB versioning — Release-series and compatibility context for the pinned server.
- mongosh release notes — Current mongosh release used by the lesson commands.
- PyMongo release notes — Current official Python driver line used where driver cursor behavior is demonstrated.
- Aggregation pipeline — Core ordered-stage document-flow model and expression concepts.
- Aggregation pipeline limits — Stage count, 16 MiB output-document, memory, allowDiskUse, and disk-spill rules.
- $group stage — Grouping key semantics, accumulator usage, memory restrictions, and optimizations.
- Aggregation accumulators — Reference for accumulator operators and stage availability.
- $first accumulator — Order-sensitive first-value semantics.
- $addToSet accumulator — Unique-set accumulator behavior and unspecified result ordering.
- $sort stage — Ordering semantics used before first/last.
- Explain command — Execution evidence for the selective prefix and full aggregation.