Build correct MongoDB group keys and accumulators for AtlasMart analytics while preventing tenant-boundary mistakes, undefined first/last semantics, and high-cardinality memory surprises.

$group, Accumulators, Distinct-Like Analysis, and Correct Aggregation Keys

Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.

Intermediate95–125 minutes$group + accumulator correctness labMongoDB 8.3.8 · mongosh 2.10.0Last reviewed: September 2026

Learning outcomes

AtlasMart finance asks for revenue by tenant and category. A pipeline can execute flawlessly yet produce a dangerous result if its aggregation key omits a business dimension. $group is therefore as much a data-model/invariant decision as a syntax feature: every input document is assigned to the group identified by the evaluated _id expression, and accumulators summarize the documents in that group.

01

Choose grouping keys that preserve tenant and business boundaries.

02

Use core accumulators such as $sum, $avg, $min, $max, $addToSet, $first, and $last with defined semantics.

03

Build distinct-like analysis with grouping while understanding when the distinct command may be simpler.

04

Make ordering explicit before order-sensitive accumulators such as $first/$last.

05

Connect group cardinality to memory pressure and verify selective upstream index use.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 using mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, a disposable standalone mongod published only on loopback 127.0.0.1:27054, and mongosh 2.10.0. The database is atlasmart; the collection is order_lines_ch08_l3. Authentication and TLS are disabled only for this isolated learning container. The lab reads Feature Compatibility Version (FCV) and allowDiskUseByDefault but does not change them. Default read/write concern and primary read preference are used on the standalone. Atlas, Search, KMS, Enterprise Advanced, and paid services are not required. Commands were reviewed against current official documentation; runtime output was not generated here because Docker/mongod/mongosh/PyMongo are unavailable in this environment.

How to read expected output

Expected result fragments are documentation-derived shapes and invariants, not copied from a generation-time MongoDB process. Exact explain trees, optimizer rewrites, field ordering, cursor IDs, execution counters, memory/disk-spill detail, and error text can vary with server patch, FCV, index state, dataset, and topology. Verify the logical result and the named execution evidence rather than comparing output byte-for-byte.

1. Grouping collapses documents around an explicit key

The $group stage outputs one document per distinct value of its _id expression. That value can be a scalar or a compound document. Accumulators then combine values from every input document assigned to that group. This changes cardinality: six input line documents might become two, four, or six output groups depending entirely on the chosen key.

Accumulator Meaning in this lesson Important boundary
$sum Adds revenue or counts documents with {$sum:1}. Numeric type/overflow semantics still matter; do not silently mix units.
$avg Mean of numeric inputs in a group. Average of line revenue is not average order value unless the grouping/input grain matches orders.
$min/$max Minimum/maximum value observed. These describe input values, not chronology unless the value itself is temporal.
$addToSet Unique values per group. Output array order is unspecified; it can grow large.
$first/$last First/last document/value in the current defined order. Meaningful chronological use requires an explicit preceding sort or another documented ordering source.

2. Seed data where a missing tenant key creates a visible bug

bash · isolated Chapter 08 Lesson 3 lab setup
docker rm -f atlasmart-mongo-ch08-l3 2>/dev/null || truedocker volume rm atlasmart-mongo-ch08-l3-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch08-l3 \  -p 127.0.0.1:27054:27017 \  -v atlasmart-mongo-ch08-l3-data:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27054/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1,allowDiskUseByDefault:1}))' 
javascript · seed line-level facts for two tenants
const c=db.getCollection("order_lines_ch08_l3");c.drop();c.insertMany([  {_id:"l1",tenantId:"tenant-a",orderId:"o1",category:"books",region:"west",qty:2,revenueCents:3000,discountCents:0,createdAt:ISODate("2026-08-01T10:00:00Z")},  {_id:"l2",tenantId:"tenant-a",orderId:"o1",category:"electronics",region:"west",qty:1,revenueCents:2500,discountCents:200,createdAt:ISODate("2026-08-01T10:00:00Z")},  {_id:"l3",tenantId:"tenant-a",orderId:"o2",category:"books",region:"west",qty:1,revenueCents:4200,discountCents:0,createdAt:ISODate("2026-08-02T10:00:00Z")},  {_id:"l4",tenantId:"tenant-a",orderId:"o3",category:"books",region:"east",qty:3,revenueCents:9000,discountCents:500,createdAt:ISODate("2026-08-03T10:00:00Z")},  {_id:"l5",tenantId:"tenant-b",orderId:"o9",category:"books",region:"west",qty:10,revenueCents:15000,discountCents:0,createdAt:ISODate("2026-08-04T10:00:00Z")},  {_id:"l6",tenantId:"tenant-b",orderId:"o10",category:"electronics",region:"west",qty:2,revenueCents:5000,discountCents:0,createdAt:ISODate("2026-08-05T10:00:00Z")}]);c.createIndex({tenantId:1,region:1,createdAt:1});print("seeded",c.countDocuments({}));

3. Deliberately wrong: group only by category

The query below is syntactically valid and deterministic, but it merges revenue from two tenants. In a multi-tenant system that is a data-isolation/reporting bug, not a minor reporting discrepancy. The database cannot infer that tenantId belongs in the business key.

javascript · wrong grouping key merges tenants
printjson(c.aggregate([  { $group:{      _id:"$category",      revenueCents:{ $sum:"$revenueCents" },      lines:{ $sum:1 }  }},  { $sort:{_id:1} }]).toArray());// This merges tenant-a and tenant-b into the same business group.
text · expected wrong-key result
[  { _id: "books", revenueCents: 31200, lines: 4 },  { _id: "electronics", revenueCents: 7500, lines: 2 }]
Correctness before performance

A faster aggregation of the wrong groups is still wrong. Define the input grain and every business partition in the group key before tuning memory or indexes.

4. Repair with a compound aggregation key and explicit accumulators

javascript · correct per-tenant per-category rollup
printjson(c.aggregate([  { $group:{      _id:{tenantId:"$tenantId",category:"$category"},      lineCount:{ $sum:1 },      units:{ $sum:"$qty" },      revenueCents:{ $sum:"$revenueCents" },      avgRevenueCents:{ $avg:"$revenueCents" },      minRevenueCents:{ $min:"$revenueCents" },      maxRevenueCents:{ $max:"$revenueCents" },      orderIds:{ $addToSet:"$orderId" }  }},  { $sort:{"_id.tenantId":1,"_id.category":1} }]).toArray());
text · expected tenant-preserving groups
tenant-a / books:       lineCount=3 units=6 revenueCents=16200 avg=5400 min=3000 max=9000 orderIds=[o1,o2,o3]tenant-a / electronics: lineCount=1 units=1 revenueCents=2500 avg=2500 min=2500 max=2500 orderIds=[o1]tenant-b / books:       lineCount=1 units=10 revenueCents=15000tenant-b / electronics: lineCount=1 units=2 revenueCents=5000

The output key is now a document containing both tenant and category. orderIds is a distinct-like set only within that business group. If the number of distinct order IDs can become very large, materializing the entire set is itself a growth/memory decision; a count of distinct items may require a different pipeline shape.

5. Distinct-like analysis and the grain of the question

javascript · distinct categories for one tenant via group
printjson(c.aggregate([  { $match:{tenantId:"tenant-a"} },  { $group:{_id:"$category"} },  { $sort:{_id:1} }]).toArray());

This pipeline illustrates the mechanism but does not imply that $group should replace every distinct command. Use the simplest operation that expresses the requirement and benchmark the real query shape. The important idea is that grouping creates one output document per key and therefore can be used to deduplicate values when no additional accumulator is needed.

6. $first and $last require a defined order

Without a meaningful order, “first” and “last” are properties of the pipeline’s current stream, not business chronology. AtlasMart explicitly sorts tenant-a lines by createdAt and _id before grouping by region, so the order-sensitive accumulators have a defined interpretation.

javascript · sort before first/last accumulators
printjson(c.aggregate([  { $match:{tenantId:"tenant-a"} },  { $sort:{createdAt:1,_id:1} },  { $group:{      _id:"$region",      firstLine:{ $first:"$_id" },      lastLine:{ $last:"$_id" },      firstAt:{ $first:"$createdAt" },      lastAt:{ $last:"$createdAt" }  }},  { $sort:{_id:1} }]).toArray());

7. Memory pressure is driven by groups and accumulator state

$group is a blocking stage: it must maintain state for the distinct groups until results can be emitted. MongoDB documents a 100 MB memory threshold for such stages, with temporary-file spill behavior controlled by allowDiskUseByDefault and per-command allowDiskUse. High-cardinality grouping keys and accumulators that retain arrays/sets can therefore be much more expensive than a low-cardinality sum.

Safe lab boundary

This six-document fixture cannot meaningfully exercise spill behavior. The lesson intentionally measures grouping correctness and the indexed prefix; production testing should use representative group cardinality/value distributions and observe usedDisk, temporary I/O, latency percentiles, and resource contention.

8. Explain the selective prefix before the group

javascript · explain match then group then sort
const e=c.explain("executionStats").aggregate([  { $match:{tenantId:"tenant-a",region:"west"} },  { $group:{_id:"$category",revenueCents:{$sum:"$revenueCents"}} },  { $sort:{revenueCents:-1,_id:1} }]);const q=e.stages?.find(s => s.$cursor)?.$cursor ?? e;printjson({winningPlan:q.queryPlanner?.winningPlan,keys:q.executionStats?.totalKeysExamined,docs:q.executionStats?.totalDocsExamined});

The $match can use the compound index to reduce documents before grouping. The group itself is not “made free” by the index; it still builds aggregate state for the surviving categories. Explain is evidence about the chosen execution, not a promise that future data distribution will produce the same cost.

9. Verification, cleanup, and production judgment

Verification checklist

  • The wrong category-only group visibly combines tenant-a and tenant-b revenue.
  • The repaired compound key emits separate tenant/category groups.
  • lineCount, units, revenue, min/max/avg, and order ID sets match the six seeded facts.
  • The distinct-like pipeline returns tenant-a categories only.
  • The first/last pipeline has an explicit deterministic sort before grouping.
  • Explain demonstrates whether the tenant/region prefix uses the fixture index.

Use grouping when the business question truly requires consolidation across documents. Treat the group key as an invariant boundary, particularly for tenant, currency, region, time bucket, or data-version dimensions. Large group cardinality increases memory and spill risk; large $addToSet/$push accumulator state can also create large output documents subject to the 16 MiB result-document limit. Watch spill indicators, temporary disk, examined documents, group counts, output-document size, and latency tails.

Grouping is a read transformation and does not atomically “lock” the source dataset into a snapshot beyond the read concern/topology semantics of the operation. If exact financial reporting requires a specific consistency point, design that requirement explicitly. The next lesson returns to arrays and shows how $unwind can silently discard empty, null, or missing cases unless preservation is requested.

bash · cleanup / full reset
docker rm -f atlasmart-mongo-ch08-l3docker volume rm atlasmart-mongo-ch08-l3-data

Check your understanding

  1. Why is grouping only by category unsafe in the fixture?
  2. What determines how many documents $group outputs?
  3. Why does $first need an explicit sort for chronological meaning?
  4. Why can $addToSet be more memory-sensitive than a simple $sum?
  5. What does an indexed $match before $group improve—and what does it not eliminate?
Review the answers

Because category is not the full business partition: it merges facts from tenant-a and tenant-b.

The number of distinct evaluated values of the $group _id expression among the input documents.

$first means first in the current input order; without a defined order that is not necessarily earliest by business time.

It retains distinct values for each group, so accumulator state can grow with cardinality instead of remaining a fixed-size numeric total.

It can reduce the number of documents reaching the group. The group still has to maintain aggregate state for the surviving grouping keys.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.