Deconstruct AtlasMart arrays with $unwind while preserving empty, null, and missing parent cases when required, measuring cardinality, and detecting legacy non-array shapes.
$unwind Arrays Without Losing Empty/Missing Cases
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
AtlasMart shipment documents embed an items array.
Analysts want item-level rows, but the collection also contains
empty shipments, explicit nulls, missing fields from an old
writer, and one legacy scalar object. $unwind is
the stage that deconstructs an array into one pipeline document
per element, and its edge-case options determine whether those
nonstandard inputs survive.
Predict $unwind cardinality for arrays, empty arrays, null, missing fields, and non-array scalar/object values.
Use preserveNullAndEmptyArrays when business semantics require retaining zero-item or unknown-item parent documents.
Use includeArrayIndex as explicit positional evidence rather than assuming original array position survives elsewhere.
Avoid double-counting parent entities after cardinality expansion.
Normalize or reject legacy non-array values deliberately instead of silently treating them as equivalent arrays.
This lesson pins MongoDB Community Server
8.3.8 using
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, a disposable standalone mongod published only
on loopback 127.0.0.1:27055, and mongosh
2.10.0. The database is atlasmart;
the collection is shipments_ch08_l4.
Authentication and TLS are disabled only for this isolated
learning container. The lab reads Feature Compatibility
Version (FCV) and allowDiskUseByDefault but does
not change them. Default read/write concern and primary read
preference are used on the standalone. Atlas, Search, KMS,
Enterprise Advanced, and paid services are not required.
Commands were reviewed against current official documentation;
runtime output was not generated here because
Docker/mongod/mongosh/PyMongo are unavailable in this
environment.
Expected result fragments are documentation-derived shapes and invariants, not copied from a generation-time MongoDB process. Exact explain trees, optimizer rewrites, field ordering, cursor IDs, execution counters, memory/disk-spill detail, and error text can vary with server patch, FCV, index state, dataset, and topology. Verify the logical result and the named execution evidence rather than comparing output byte-for-byte.
1. $unwind changes the grain of the stream
Before unwinding, the grain is one pipeline document per
shipment. After unwinding items, the grain is one
document per item element for normal arrays. That means every
downstream count, sum, group, and join must be interpreted at
the new grain. Many aggregation bugs come from forgetting that
cardinality and entity identity changed.
| Input value at items | Default $unwind | With preserveNullAndEmptyArrays:true |
|---|---|---|
| Two-element array | Two output documents. | Two output documents. |
| One-element array | One output document. | One output document. |
| Empty array | No output document. | Parent is preserved; output field may be absent. |
| Explicit null | No output document. | Parent is preserved with null. |
| Missing field | No output document. | Parent is preserved; path remains missing. |
| Non-array, non-null object/scalar | One output document using the value as-is. | Same single-document treatment; array index is null when requested. |
2. Seed every important boundary case
docker rm -f atlasmart-mongo-ch08-l4 2>/dev/null || truedocker volume rm atlasmart-mongo-ch08-l4-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch08-l4 \ -p 127.0.0.1:27055:27017 \ -v atlasmart-mongo-ch08-l4-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27055/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1,allowDiskUseByDefault:1}))'
const c=db.getCollection("shipments_ch08_l4");c.drop();c.insertMany([ {_id:"s1",tenantId:"tenant-a",items:[{sku:"A",qty:2},{sku:"B",qty:1}]}, {_id:"s2",tenantId:"tenant-a",items:[{sku:"C",qty:4}]}, {_id:"s3",tenantId:"tenant-a",items:[]}, {_id:"s4",tenantId:"tenant-a",items:null}, {_id:"s5",tenantId:"tenant-a"}, {_id:"s6",tenantId:"tenant-a",items:{sku:"LEGACY",qty:1}}]);print("seeded",c.countDocuments({}));
3. Default unwind silently removes zero-element/null/missing parents
printjson(c.aggregate([ { $unwind:"$items" }, { $project:{_id:1,items:1} }, { $sort:{_id:1} }]).toArray());
The output contains two rows for s1, one for
s2, and one for the legacy non-array
s6. The empty-array, null, and missing cases do not
produce output documents. That behavior may be exactly right for
an “item facts” report, but it is wrong for a report whose
denominator must include shipments with zero or unknown items.
4. Preserve parent cases and expose original array position
printjson(c.aggregate([ { $unwind:{path:"$items",preserveNullAndEmptyArrays:true,includeArrayIndex:"itemIndex"} }, { $project:{_id:1,items:1,itemIndex:1,itemType:{$type:"$items"}} }, { $sort:{_id:1,itemIndex:1} }]).toArray());
For real arrays, itemIndex records the zero-based
position. For non-array inputs it is null. The
index is evidence about this unwind operation; it is not a
durable identifier for the line item unless the application
chooses to model it as such.
5. Quantify cardinality instead of guessing
const before=c.countDocuments({tenantId:"tenant-a"});const simpleCount=c.aggregate([{ $match:{tenantId:"tenant-a"} },{ $unwind:"$items" },{ $count:"n" }]).toArray()[0]?.n ?? 0;const preservedCount=c.aggregate([{ $match:{tenantId:"tenant-a"} },{ $unwind:{path:"$items",preserveNullAndEmptyArrays:true} },{ $count:"n" }]).toArray()[0]?.n ?? 0;printjson({inputShipments:before,simpleUnwindRows:simpleCount,preservedUnwindRows:preservedCount});
{ inputShipments: 6, simpleUnwindRows: 4, preservedUnwindRows: 7 }
Cardinality growth depends on the array-length distribution, not just the number of parent documents. A million shipments with one item each behave differently from a million shipments whose heavy tail includes arrays with thousands of elements. Production tests should record p50/p95/p99 array lengths and the maximum, not only the average.
6. Deliberately wrong: count parent entities after changing the grain
print("WRONG: shipments per tenant after unwind");printjson(c.aggregate([ { $unwind:"$items" }, { $group:{_id:"$tenantId",shipmentCount:{$sum:1}} }]).toArray());print("RIGHT if the question is shipment count: count before unwind");printjson(c.aggregate([ { $group:{_id:"$tenantId",shipmentCount:{$sum:1}} }]).toArray());
WRONG after simple unwind: shipmentCount = 4RIGHT before unwind: shipmentCount = 6
After unwind, {$sum:1} counts item-level output
rows, not shipments. If the requirement is item count, that is
correct; if the requirement is shipment count, it is a grain
error. Another valid approach after unwind is to group by
shipment ID first and then count distinct parents, but doing
extra work to recover the original grain is usually less clear
than counting before expansion when the question allows it.
7. Legacy non-array values need an explicit policy
MongoDB treats a non-array, non-null, non-empty value as a
single unwind value. That compatibility behavior is convenient
but can hide schema drift: s6.items is an object
where the target model expects an array. The pipeline below
records the original type and chooses to normalize only true
arrays; the legacy object yields a preserved parent with no
normalized item. A different business policy could
quarantine/reject the document instead.
printjson(c.aggregate([ { $set:{ itemsType:{ $type:"$items" }, normalizedItems:{ $cond:[ { $eq:[{$type:"$items"},"array"] }, "$items", [] ] } }}, { $unwind:{path:"$normalizedItems",preserveNullAndEmptyArrays:true,includeArrayIndex:"itemIndex"} }, { $project:{_id:1,itemsType:1,normalizedItems:1,itemIndex:1} }, { $sort:{_id:1,itemIndex:1} }]).toArray());
Chapter 07’s validator and migration techniques are the stronger long-term fix. Aggregation-time normalization is not permission to leave malformed storage indefinitely.
8. Verification, cleanup, and production judgment
Verification checklist
- Six parent shipment documents exist.
- Default unwind produces item-level rows for normal arrays plus one row for the legacy object, while dropping empty/null/missing cases.
- Preserved unwind retains every parent case and exposes array indexes for actual array elements.
- The count comparison demonstrates the change in stream grain.
- The wrong shipment-count example overcounts shipments that contain multiple items and omits dropped parents.
- The type-aware variant identifies/contains the non-array legacy shape instead of silently treating it as a clean array.
$unwind is appropriate when downstream logic truly
needs element-level grain. Its principal cost is cardinality
multiplication: more documents flow into group/sort/lookup
stages, potentially increasing CPU, memory, network, and spill
pressure. In sharded environments, expansion can occur before or
after merge depending on pipeline planning, so inspect explain
rather than assume where the work runs. Preserve
null/empty/missing parents only when the business denominator
requires them; preserving everything “just in case” also
increases downstream work.
Tenant predicates should be applied before expansion whenever semantics permit. Treat malformed non-array values as a schema-quality signal. The next lesson combines these fundamentals with explain analysis and demonstrates the most important production habit: a pipeline can be logically correct and still operationally expensive.
docker rm -f atlasmart-mongo-ch08-l4docker volume rm atlasmart-mongo-ch08-l4-data
Check your understanding
- What is the default $unwind behavior for an empty array?
- What does preserveNullAndEmptyArrays change?
- How does $unwind treat a non-array, non-null object?
- Why can {$sum:1} after unwind count the wrong business entity?
- What production distribution matters more than average array length alone?
Review the answers
It produces no output document for that parent.
It retains parents whose path is null, missing, or an empty array, using the documented output shape for each case.
It treats the value as a single unwind value and outputs one document; includeArrayIndex is null for such non-array input.
Because the stream grain changed from parent documents to array elements, so the count measures element rows unless you explicitly reconstruct distinct parents.
The full/high-tail array-cardinality distribution—especially p95/p99/max—because a small number of huge arrays can dominate expansion cost.
Authoritative references
- MongoDB 8.3 release notes — Current stable/minor series and patch status; re-check before reproducing the lab.
- MongoDB versioning — Release-series and compatibility context for the pinned server.
- mongosh release notes — Current mongosh release used by the lesson commands.
- PyMongo release notes — Current official Python driver line used where driver cursor behavior is demonstrated.
- Aggregation pipeline — Core ordered-stage document-flow model and expression concepts.
- Aggregation pipeline limits — Stage count, 16 MiB output-document, memory, allowDiskUse, and disk-spill rules.
- $unwind stage — Default behavior, preserveNullAndEmptyArrays, includeArrayIndex, and non-array handling.
- $type expression — Observable null/missing/array/object type distinctions inside aggregation expressions.
- $count stage — Counting the current pipeline grain.
- Schema validation — Long-term enforcement for expected array types rather than permanent aggregation-time repair.
- Aggregation optimization — Context for selective stages and execution planning around cardinality-changing work.