Understand how arrays expand multikey indexes, constrain compound keys, affect sorting/coverage, and interact with same-element query semantics.
Multikey Indexes for Arrays: Index-Key Expansion, Constraints, and Query Shapes
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning objectives
Explain automatic multikey conversion when an indexed path contains arrays.
Derive conceptual index-key expansion from distinct array elements and relate expansion to storage/write cost.
Apply the compound-multikey restriction that each indexed document can have at most one indexed array field in the compound key.
Understand when a multikey index can cover a query and why returning the array field or using $elemMatch changes coverage.
Connect $elemMatch query semantics to multikey index bounds without confusing logical correctness with index coverage.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim
and mongosh 2.10.0. The server is a disposable
standalone published only on loopback
127.0.0.1:27064. Authentication and TLS are
disabled only for this isolated lab. Feature Compatibility
Version (FCV) is observed but never changed. Default read/write
concern and primary read preference apply. Atlas, Search, KMS,
Enterprise Advanced, and paid services are not required. Runtime
output shown as “expected” is documentation-derived because this
generation environment has no Docker/mongod/mongosh runtime.
Index behavior depends on query shape, projection, sort, data
distribution, planner choice, cache state, topology, and patch
version. The labs therefore inspect winningPlan,
totalKeysExamined, totalDocsExamined,
index definitions/sizes, and target query results. Small
fixtures prove semantics, not production latency. Timing
snippets are comparative demonstrations only, not benchmarks.
1. One document can produce multiple index entries
When an indexed field contains an array, MongoDB automatically marks the index as multikey. Each distinct array element contributes a key entry pointing back to the same document. You do not request “multikey mode”; it follows from the indexed data shape. Duplicate values inside the same array do not create duplicate entries for that same index key.
| Document shape | Conceptual indexed tuples for {tenantId,tags,priceCents} |
|---|---|
| P1 tags=[usb, travel, usb] | (tenant-a, usb, 3900) and (tenant-a, travel, 3900) |
| P2 tags=[power, travel] | (tenant-a, power, 2900) and (tenant-a, travel, 2900) |
| P5 tags=[] | No tag-derived entry for this array path; query behavior must account for missing/empty cases separately. |
2. Seed products and create one compound multikey index
docker rm -f atlasmart-mongo-ch10-l3 2>/dev/null || truedocker volume rm atlasmart-mongo-ch10-l3-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch10-l3 \ -p 127.0.0.1:27064:27017 \ -v atlasmart-mongo-ch10-l3-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27064/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(),hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1}))'
const c=db.products_ch10_l3;c.drop();c.insertMany([ {_id:1,tenantId:"tenant-a",sku:"P1",name:"USB-C Hub",tags:["usb","travel","usb"],priceCents:3900,warehouses:["w1","w2"]}, {_id:2,tenantId:"tenant-a",sku:"P2",name:"Travel Charger",tags:["power","travel"],priceCents:2900,warehouses:["w2"]}, {_id:3,tenantId:"tenant-a",sku:"P3",name:"USB Cable",tags:["usb","cable"],priceCents:1200,warehouses:["w1","w3"]}, {_id:4,tenantId:"tenant-b",sku:"P4",name:"USB Stand",tags:["usb","desk"],priceCents:2200,warehouses:["w9"]}, {_id:5,tenantId:"tenant-a",sku:"P5",name:"Notebook",tags:[],priceCents:700,warehouses:[]}]);printjson(c.find({},{_id:0,tenantId:1,sku:1,tags:1,warehouses:1}).toArray());
const c=db.products_ch10_l3;print(c.createIndex({tenantId:1,tags:1,priceCents:1,sku:1},{name:"idx_tenant_tags_price_sku"}));printjson(c.getIndexes());printjson(c.find({tenantId:"tenant-a",tags:"usb"},{_id:0,sku:1,priceCents:1}).hint("idx_tenant_tags_price_sku").explain("executionStats"));
const c=db.products_ch10_l3;printjson(c.aggregate([ {$project:{_id:0,sku:1,distinctTags:{$setUnion:["$tags",[]]},conceptualKeyCount:{$size:{$setUnion:["$tags",[]]}}}}, {$sort:{sku:1}}]).toArray());
The projected conceptualKeyCount counts distinct
tag values for teaching. Physical index byte size includes key
encoding, record references, tree overhead, compression, and
allocation behavior; use indexSizes/totalIndexSize()
for actual storage evidence.
3. Multikey coverage has extra restrictions
A compound multikey index can cover some queries if the
projection does not return the array field, the query does not
use $elemMatch, and normal covered-query
requirements are met. Returning tags requires
reconstructing the original array from the document; the
multiple index keys do not preserve the array as one original
value.
const c=db.products_ch10_l3;print("projection excludes array field");printjson(c.find({tenantId:"tenant-a",tags:"usb"},{_id:0,sku:1,priceCents:1}).hint("idx_tenant_tags_price_sku").explain("executionStats"));print("projection returns array field");printjson(c.find({tenantId:"tenant-a",tags:"usb"},{_id:0,sku:1,tags:1}).hint("idx_tenant_tags_price_sku").explain("executionStats"));
4. Hard constraint: one indexed array field per document in a compound multikey index
The fixture has both tags and
warehouses arrays. A compound index that tries to
index both paths for those same documents would require a
Cartesian product of parallel arrays. MongoDB rejects that
design.
const c=db.products_ch10_l3;try{ c.createIndex({tags:1,warehouses:1},{name:"idx_parallel_arrays"}); print("UNEXPECTED: parallel-array compound index created");}catch(e){ printjson({code:e.code,codeName:e.codeName,message:e.message});}
Choose one array path for the compound index, remodel the relationship, or use separate indexes/query shapes. Do not flatten two independent arrays into one compound key just to satisfy an index idea.
5. $elemMatch preserves same-element logic but changes index/coverage behavior
Arrays of embedded offer documents need same-element semantics:
warehouse and quantity must belong to the same offer.
$elemMatch provides that logical contract. MongoDB
can use multikey bounds for appropriate embedded-array indexes,
but a query containing $elemMatch is not eligible
for the covered-query case described above.
const e=db.offers_ch10_l3;e.drop();e.insertMany([ {_id:1,tenantId:"tenant-a",sku:"P1",offers:[{warehouse:"w1",qty:2,priceCents:3900},{warehouse:"w2",qty:20,priceCents:4100}]}, {_id:2,tenantId:"tenant-a",sku:"P2",offers:[{warehouse:"w1",qty:15,priceCents:5000},{warehouse:"w2",qty:1,priceCents:2900}]}]);e.createIndex({tenantId:1,"offers.warehouse":1,"offers.qty":1},{name:"idx_offer_elem"});printjson(e.find({tenantId:"tenant-a",offers:{$elemMatch:{warehouse:"w1",qty:{$gte:10}}}},{_id:0,sku:1}).explain("executionStats"));
An incorrect query that matches warehouse in one array element
and quantity in another is not “faster”; it is wrong. Start
with the correct $elemMatch predicate, then
design and measure an index for that predicate.
6. Sorting and multikey indexes need explicit evidence
Sorting on an indexed array field commonly introduces an
in-memory sort unless MongoDB can use unbounded index bounds for
all sort fields and no multikey-indexed boundary shares the sort
path prefix. Treat array sorts as a specific explain problem,
not as “the field is indexed, so sorting is free.” Similarly,
$expr does not support multikey indexes.
A product with 3 distinct tags contributes more index entries than a product with 1 tag. High-cardinality arrays amplify writes and index size even when the parent document count is stable. Monitor array-length distributions, not just collection document count.
7. Verification, cleanup, and production judgment
Verification checklist
- The tag index reports multikey behavior for the array path.
-
Repeated
usbwithin P1 does not count as a second distinct conceptual tag key. - The query returning only sku/price can potentially avoid fetching the array field.
-
Returning
tagsdestroys multikey covered-query eligibility. - The parallel-array compound index attempt fails safely on the disposable collection.
-
The offer query uses
$elemMatchto preserve same-element semantics.
Production judgment. Multikey indexes are essential for array-heavy document models but their cost scales with indexed array cardinality. They can increase storage, write work, cache churn, build time, and query-plan complexity. Compound multikey restrictions can become a modeling constraint, not merely an indexing detail. On sharded deployments, a multikey index cannot itself be the shard-key index, though a trailing non-shard-key field can make a compound index multikey after the shard-key prefix. Lesson 4 moves from array multiplicity to uniqueness, where null/missing and collation semantics can enforce—or accidentally block—business invariants.
docker rm -f atlasmart-mongo-ch10-l3docker volume rm atlasmart-mongo-ch10-l3-data
Check your understanding
- When does MongoDB make an index multikey?
- Does ["usb","usb"] create two identical multikey entries for the same document?
- Can one compound index index two independent array fields in the same document?
- Why can returning the array field prevent coverage?
- What correctness problem does $elemMatch solve?
Review the answers
1. Automatically when an indexed path contains an array value.
2. No. Repeated values within the same array do not create duplicate index entries for that same key.
3. No. Each indexed document can have at most one indexed array field in a compound multikey index.
4. Separate index entries do not reconstruct the original array value; MongoDB must fetch the document.
5. It requires multiple predicates to match the same array element rather than different elements independently.
Authoritative references
- MongoDB 8.3 release notes — Current 8.3 baseline and patch-sensitive behavior; re-check before reproduction.
- Indexes overview — Index concepts, names, build considerations, and index-management overview.
- Explain results — IXSCAN/FETCH/COLLSCAN evidence, covered-query plans, keys/documents examined, and execution statistics.
- Measure index use — Using $indexStats and explain evidence instead of intuition to manage indexes.
- mongosh release notes — mongosh version used for chapter commands.
- Multikey indexes — Automatic multikey conversion, distinct array entries, compound restrictions, sort behavior, covered-query rules, and shard-key limitations.
- Multikey index bounds — How array predicates and $elemMatch can combine index bounds.