Chapter 02 · BSON, Documents, Collections, Flexible Schema, and Data Types
_id, ObjectId, Field Names, Embedded Documents, Arrays, and Document Size Limits
Make document identity and structural limits concrete: _id/ObjectId, embedded data, arrays, field-name constraints, 16 MiB size, and nesting boundaries.
Learning outcomes
AtlasMart can now preserve BSON types, but a document also has
structural rules. The team wants product identifiers, nested
seller/dimensions data, tags and variants as arrays, and
occasionally category-specific field names. A schema that works
at 2 KB can fail completely at 17 MiB; an _id that
looks like a timestamp is not a globally ordered clock; and
field names containing dots or dollar signs are technically
storable yet operationally constrained. This lesson turns those
boundaries into observable tests.
Explain the required unique _id field, its
immutability, and valid/invalid value classes.
Describe the 12-byte ObjectId structure and why its timestamp makes it roughly time-ordered but not a strict global sequence.
Model embedded documents and arrays while recognizing unbounded-growth and multikey/index consequences.
Apply current field-name rules, including null-character prohibition, duplicate-field hazards, and dot/dollar restrictions.
Prove the 16 MiB BSON document limit and 100-level nesting limit without risking unrelated data.
Mandatory server examples use a disposable loopback-only
standalone
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim. Driver examples pin pymongo==4.17.0. The lab
is intentionally unauthenticated only because it is
short-lived and published to 127.0.0.1; Chapter
01 already demonstrated the authenticated reusable lab. Do not
expose this container on an untrusted interface. Atlas is
optional and not required.
The build environment does not provide Docker, mongod, or mongosh, and cannot install PyMongo from the network. Product/driver commands were reviewed against current official documentation but were not executed here. Expected-output blocks describe stable evidence shape, not fabricated captured output. Learners should record their actual versions and outputs.
1. Every standard collection document has one primary identity: _id
MongoDB requires a unique _id field for every
document in a standard collection. If an insert omits it, a
driver normally generates one—commonly an ObjectId.
MongoDB creates a unique index on _id for normal
collections. The _id value is immutable after
insertion. It may be many BSON types, but it cannot be an array
or regular-expression value; if _id is itself an
embedded document, its subfield names have additional
restrictions such as not beginning with $.
A meaningful application key such as sku can be the
_id when it is truly stable, unique, bounded, and
operationally suitable. Otherwise keep an ObjectId identity and
enforce the business key with a unique index later. Do not
overload _id with a mutable email address or a
value whose case/collation semantics are unclear.
use atlasmartdb.document_lab.drop();const r = db.document_lab.insertOne({sku:"sku-lamp-01", name:"Atlas Lamp"});printjson(r.insertedId);printjson(db.document_lab.findOne({sku:"sku-lamp-01"}));print("timestamp:", r.insertedId.getTimestamp());// _id cannot be changed after insert.try { db.document_lab.updateOne({_id:r.insertedId}, {$set:{_id:ObjectId()}})} catch (e) { print(e.codeName || e.code, e.message); }
The timestamp extracted from an ObjectId is creation-time metadata with one-second resolution. ObjectIds generated in the same second are not guaranteed to be ordered by creation time, and clients can have different clocks. Use an explicit business timestamp or sequence when strict ordering is required.
2. Embedded documents and arrays create locality—and growth obligations
An embedded document nests a document as a field value. An array preserves an ordered list of values and can contain scalars, documents, or mixed BSON types. Embedding seller snapshot data or a bounded list of variant options can make the common product read local to one document and preserves single-document atomicity for updates to that aggregate.
The edge case is growth. MongoDB's maximum BSON document size is 16 mebibytes (MiB), and BSON documents support no more than 100 levels of nesting; every object or array adds a level. Long before 16 MiB, an unbounded review/event/history array can create write amplification, cache pressure, contention, and index growth. The limit is a hard correctness boundary, not a target size.
| Shape | Good fit | Warning signal |
|---|---|---|
dimensions embedded object |
Small, owned by product, read together | Different lifecycle/owner needs independent updates |
tags array |
Bounded set of labels | Thousands/millions of append-only values |
variants array of documents |
Bounded options managed with product | Variant cardinality grows without a defined ceiling |
| reviews embedded in product | Rarely: only tiny bounded snapshot | User-generated, independently growing history—usually reference/separate collection |
3. Field names are flexible, not consequence-free
Field names cannot contain the null character. Each field name
should be unique within a document; duplicate field names are
unsupported and can produce inconsistent driver/query/update
behavior. Modern MongoDB servers can store field names
containing periods (.) or dollar signs
($), but official guidance says to avoid them
because important capabilities are restricted: they cannot be
indexed, used as shard keys, validated with
$jsonSchema, or used with Field Level Encryption,
among other constraints. Dot characters also conflict with
normal path syntax and require helpers such as
$getField/$setField.
There is also an Extended JSON ambiguity: a user field such as
$date can collide with an Extended JSON type
wrapper. The ability to persist a field is therefore not the
same as having a robust application/tooling contract.
use atlasmartdb.document_lab.insertOne({ _id:"field-edge", "price.usd": Decimal128("18.00"), "$source": "legacy-import"});printjson(db.document_lab.aggregate([ {$match:{_id:"field-edge"}}, {$project:{ _id:0, dotted:{$getField:{field:"price.usd", input:"$$CURRENT"}}, source:{$getField:{field:{$literal:"$source"}, input:"$$CURRENT"}} }}]).toArray());
This demonstrates that storage is possible; it does not make those names good schema design. For new AtlasMart data, prefer ordinary field names that work with indexes, validators, encryption, imports/exports, and dot-path queries.
4. Measure size before the hard limit becomes an incident
Use the server's $bsonSize or a driver BSON encoder
to measure real encoded size. The following lab intentionally
creates a document just below the limit and then one above it in
a disposable collection. Do not paste it into a production
shell: it allocates multi-megabyte strings and exists only to
expose the boundary.
docker rm -f atlasmart-mongo-ch02-l2 2>/dev/null || truedocker run --name atlasmart-mongo-ch02-l2 -p 127.0.0.1:27023:27017 -d mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27023/atlasmart?directConnection=true"
db.size_lab.drop();const mib = 1024 * 1024;const safe = {_id:"under-limit", payload:"x".repeat(15 * mib)};db.size_lab.insertOne(safe);print("stored under-limit BSON bytes:");printjson(db.size_lab.aggregate([ {$match:{_id:"under-limit"}}, {$project:{_id:0, bytes:{$bsonSize:"$$ROOT"}}}]).toArray());const tooLarge = {_id:"over-limit", payload:"x".repeat(17 * mib)};try { db.size_lab.insertOne(tooLarge); }catch (e) { print(e.codeName || e.code, e.message); }print("over-limit persisted:", db.size_lab.countDocuments({_id:"over-limit"}));db.size_lab.deleteMany({_id:{$in:["under-limit","over-limit"]}});
The client or server should reject the oversized document; exact error text is version/client dependent. The important evidence is that BSON size crosses the documented 16 MiB ceiling and the document is not persisted. If your client rejects it before sending, that still proves the application cannot store it as one BSON document. Use separate documents or GridFS for files/large payloads rather than trying to tune away the limit.
5. Production judgment: choose boundaries that stay bounded
Use ObjectId when an opaque distributed client-generated identity is convenient, but do not infer strict event order from it. Use embedded documents and arrays when they represent one bounded aggregate and the dominant access pattern benefits from locality. Separate independently growing histories, high-cardinality relationships, and large binary content. Keep field naming conservative even when the server accepts unusual characters.
Monitor document-size distributions, not only averages. A collection can look healthy until one “celebrity” product accumulates thousands of elements and becomes the first document to hit growth/index/cache limits. Schema tests should include maximum expected cardinality and a rollback/migration path.
The next lesson turns the type catalog into application semantics: numeric widths, decimal money, UTC dates versus BSON Timestamp, binary subtypes, regex values, ObjectId, and the crucial distinction between explicit null and a missing field.
Cleanup
docker rm -f atlasmart-mongo-ch02-l2
Verification checklist
- You know why
_idis unique and immutable. - You do not treat ObjectId ordering as a strict global sequence.
- You can explain the 16 MiB and 100-level boundaries.
- You avoid unbounded embedded arrays.
- You understand why dots/dollar signs are legal in some contexts yet poor default field names.
Check your understanding
- What happens when an insert omits _id?
- Why is ObjectId only approximately ordered by creation time?
- What are the current BSON document size and nesting limits?
- Why avoid a field named price.usd even though current MongoDB can store it?
- What is the design response to an array whose size is not naturally bounded?
Review the answers
A MongoDB driver normally adds an _id value, commonly an ObjectId, before the document is inserted. Standard collection documents require a unique _id.
Its timestamp component has one-second resolution and ObjectIds can be generated by different clients with different clocks; values created within the same second do not have a guaranteed creation order.
A document can be at most 16 MiB, and BSON documents support no more than 100 nesting levels, with each object or array adding a level.
Dot-containing fields have query/tooling restrictions and cannot participate in capabilities such as indexes, shard keys, JSON Schema validation, or Field Level Encryption. Normal dot syntax also interprets the name as a path.
Model the growing items as separate documents/collection (or another bounded pattern) and keep only bounded summary/snapshot data in the parent document.
Authoritative references
- MongoDB release notes — Official current stable server series and patch notes.
- MongoDB 8.3 release notes — Official 8.3 changes; 8.3.8 is the latest released patch at review time and 8.3.9 is upcoming.
- MongoDB Extended JSON v2 — Canonical and Relaxed Extended JSON representations and type-preservation rules.
- BSON types — Official BSON type definitions and ObjectId notes.
- PyMongo BSON data formats — Official Python driver mapping between Python dictionaries/types and BSON.
- MongoDB limits and thresholds — Official 16 MiB document, 100-level nesting, naming, and _id restrictions.
- MongoDB documents — Official document field order, duplicate-field, and size behavior.
- Field names with periods and dollar signs — Official restrictions and tool/interchange caveats for dot/$ field names.