Prompt 19 · Lesson 01 · Time-series storage model
Time-Series Collection Model: timeField, metaField, Granularity, and Internal Bucketing
Understand the public measurement model and the internal bucket mechanics before tuning it.
Learning objectives
Explain timeField, metaField, measurement, series, bucket, bucket catalog, and why the public collection is not stored as one physical BSON document per measurement.
Choose granularity or custom bucket intervals from ingestion/query shape instead of treating the default as universally optimal.
Observe collection options, automatic indexes, bucket diagnostics, and conceptual bucket boundaries without making system.buckets an application dependency.
Diagnose metadata-cardinality and granularity mistakes that create too many sparse buckets or force excessive bucket filtering.
Build and reset a deterministic AtlasMart telemetry collection while distinguishing documented invariants from runtime measurements.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where a driver is used. The mandatory lab is
a disposable standalone on loopback port
27160 named atlasmart-ch19-l1;
replication or sharding is not required to learn the time-series
mechanisms in this chapter. Authentication and TLS are disabled
only for the isolated local lab. Feature Compatibility Version
(FCV) is inspected.
FCV is observed and never changed. Default
read/write concern and primary read preference apply. Atlas,
Search, Vector Search, KMS, and Enterprise Advanced are not
mandatory. Internal system.buckets.* data and
time-series diagnostic counters are used only for observation;
they are not application APIs. Product commands were not
executed in this generation environment because Docker, mongod,
mongosh, and PyMongo are unavailable here, so
environment-dependent bucket counts, storage ratios, explain
plans, and throughput/latency values must be measured on the
learner machine rather than copied as invented output. MongoDB
8.3 additionally rejects a timeField name beginning
with $.
1. AtlasMart telemetry is not just “documents with dates”
AtlasMart stores checkout latency, freezer temperature, power draw, and queue depth as measurements over time. A normal collection can hold those documents, but the workload has a special shape: many append-like measurements share stable source metadata, queries usually constrain a time interval, recent values arrive together, and retention is often bounded. A time-series collection presents normal measurement documents to the application while MongoDB stores them through an internal bucket collection optimized for this shape.
The timeField names the BSON Date that defines
the measurement's time coordinate. The optional
metaField labels the series—for example tenant,
store, and sensor identity—and should rarely change. Metric
fields such as temperatureC and
powerW are measurements. MongoDB groups
measurements with identical metadata and nearby times into
internal buckets. The in-memory
bucket catalog tracks open/reopenable buckets
so concurrent inserts can be routed efficiently.
| Term | AtlasMart example | Design consequence |
|---|---|---|
timeField |
eventAt |
Must be a BSON Date; it drives bucket time and collection-level TTL if configured. |
metaField |
{tenantId, storeId, sensorId} |
Stable identity used to group series; high uniqueness/change rate creates more buckets. |
| measurement | temperature/power at one instant | Public document seen by queries and drivers. |
| bucket | internal group of nearby measurements | Implementation/storage unit; do not expose its schema to applications. |
| granularity | seconds / minutes / hours | Controls default bucket time span: up to 1h / 24h / 30d. |
2. Create the smallest observable time-series collection
docker rm -f atlasmart-ch19-l1 2>/dev/null || truedocker volume rm atlasmart-ch19-l1-data 2>/dev/null || truedocker run -d --name atlasmart-ch19-l1 \ -p 127.0.0.1:27160:27017 \ -v atlasmart-ch19-l1-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_alldocker exec atlasmart-ch19-l1 mongosh --quiet --eval 'printjson(db.adminCommand({buildInfo:1}).version);printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion);'
const d = db.getSiblingDB("atlasmart");d.telemetry_ch19_l1.drop();d.createCollection("telemetry_ch19_l1", { timeseries: { timeField: "eventAt", metaField: "meta", granularity: "seconds" }});const t0 = ISODate("2026-09-03T04:00:00Z");const docs = [];for (const sensorId of ["freezer-01", "freezer-02"]) { for (let m = 0; m < 6; m++) { docs.push({ eventAt: new Date(t0.getTime() + m*10*60*1000), meta: {tenantId:"tenant-a", storeId:"baku-01", sensorId}, temperatureC: sensorId === "freezer-01" ? -18 + m*0.2 : -17.5 + m*0.15, ingestedAt: new Date(t0.getTime() + m*10*60*1000 + 1500) }); }}printjson(d.telemetry_ch19_l1.insertMany(docs));print("measurements", d.telemetry_ch19_l1.countDocuments({}));printjson(d.runCommand({listCollections:1, filter:{name:"telemetry_ch19_l1"}}).cursor.firstBatch[0]);printjson(d.telemetry_ch19_l1.getIndexes());
The deterministic fixture inserts exactly
12 measurements. The collection metadata should
report type:"timeseries",
timeField:"eventAt", metaField:"meta",
and granularity:"seconds". On modern MongoDB
releases, a newly created time-series collection with a
metaField also receives a compound meta/time index. The exact
generated index name is evidence to inspect, not a string to
hard-code into application behavior.
3. Bucket boundaries are a storage mechanism, not an application schema
With granularity:"seconds", a bucket can cover up
to one hour of time for one identical metadata value. MongoDB
can also close a bucket earlier because of measurement count,
size, schema, cache pressure, or other bucket-catalog
conditions. Across all granularity configurations, a bucket is
capped at 1000 measurements or 125 KB,
whichever limit is reached first; high-cardinality workloads may
be capped lower to protect the WiredTiger cache.
Custom bucketing, available since MongoDB 6.3, replaces the
named granularity with equal
bucketMaxSpanSeconds and
bucketRoundingSeconds values. The rounding interval
determines the start boundary of a new bucket. Existing bucket
layout is not rewritten when you later increase granularity, and
MongoDB does not let you decrease an existing collection's
bucket span.
const d = db.getSiblingDB("atlasmart");const tsStats = d.telemetry_ch19_l1.aggregate([ {$collStats:{storageStats:{}}}]).toArray()[0].storageStats;printjson({ count: tsStats.count, storageSize: tsStats.storageSize, totalIndexSize: tsStats.totalIndexSize, timeseries: tsStats.timeseries});const buckets = d.getCollection("system.buckets.telemetry_ch19_l1");print("diagnostic bucket documents", buckets.countDocuments({}));printjson(buckets.find({}, {meta:1,"control.min.eventAt":1,"control.max.eventAt":1}).limit(5).toArray());print("WARNING: system.buckets is internal diagnostic evidence, not an application API");
$collStats exposes a
storageStats.timeseries subdocument containing
bucket statistics such as bucket count, average bucket size,
measurement commits, closures, and reopen activity. Those values
are explicitly diagnostic and may evolve. They prove how this
run was bucketed; they do not promise an identical bucket count
on every release, machine, insert ordering, or workload.
4. Deliberately break metadata design and diagnose bucket explosion
A common mistake is putting per-request or per-measurement
values into the metaField because they look like “metadata.” If
requestId changes every reading, MongoDB cannot
group those readings into the same metadata series. That creates
many short-lived, sparsely populated buckets and damages
compression/query efficiency.
const d = db.getSiblingDB("atlasmart");for (const n of ["stable_meta_ch19_l1","exploding_meta_ch19_l1"]) d[n].drop();for (const n of ["stable_meta_ch19_l1","exploding_meta_ch19_l1"]) { d.createCollection(n,{timeseries:{timeField:"eventAt",metaField:"meta",granularity:"seconds"}});}const base = ISODate("2026-09-03T00:00:00Z");for (let i=0;i<240;i++) { const eventAt = new Date(base.getTime()+i*60*1000); d.stable_meta_ch19_l1.insertOne({eventAt,meta:{tenantId:"tenant-a",sensorId:"s-1"},v:i}); d.exploding_meta_ch19_l1.insertOne({eventAt,meta:{tenantId:"tenant-a",sensorId:"s-1",requestId:`r-${i}`},v:i});}for (const n of ["stable_meta_ch19_l1","exploding_meta_ch19_l1"]) { const st=d[n].aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats; printjson({name:n,measurements:d[n].countDocuments({}),bucketCount:st.timeseries.bucketCount,storageSize:st.storageSize,totalIndexSize:st.totalIndexSize});}
Both collections contain the same 240 public measurements. The important evidence is the relative bucket count/storage/index state measured on your run. Do not copy a fixed expected compression ratio: WiredTiger cache state, allocator behavior, schema, ordering, and server release affect physical sizes.
Keep stable identity/filter dimensions in meta;
keep request IDs, firmware sample values, trace IDs, and other
measurement-specific attributes outside the metaField unless
they truly define a stable series.
5. Production judgment
Choose a time-series collection when the workload is genuinely time-ordered measurement data and most writes append new observations. The feature improves storage/index efficiency by organizing measurements into buckets, but it does not eliminate capacity planning, indexing, retention, or data-model decisions. A too-fine granularity can create too many buckets; a too-coarse one can force queries to read broad bucket spans and filter heavily. Stable metadata is often more important than raw document count.
Operationally, monitor bucket catalog diagnostics, storage/index
growth, insert latency, query plans, cache pressure, metadata
cardinality, late-arrival rates, and retention lag. Do not query
or mutate system.buckets.* from application code.
For sharded time-series collections, remember that zone sharding
is unsupported and that MongoDB 8.0 deprecated shard keys
containing the timeField.
Bridge. Lesson 2 turns these mechanics into an explicit sensor/metadata schema and compares stable-series compression against high-cardinality metadata.
docker rm -f atlasmart-ch19-l1 2>/dev/null || truedocker volume rm atlasmart-ch19-l1-data 2>/dev/null || true
Check your understanding
- What makes two measurements candidates for the same bucket?
- Why is requestId usually a bad metaField member?
- Does granularity guarantee every bucket spans exactly its documented maximum?
- Should applications query system.buckets directly?
- Can you later reduce a collection from a coarser bucket span to a finer one?
Review the answers
1. They need identical metaField values and sufficiently close timeField values, subject to bucket count/size/cache and other closure rules.
2. It changes per measurement, multiplying series/buckets and reducing bucket density and compression.
3. No. The span is an upper time window; buckets can close earlier for count, size, cache, schema, and other reasons.
4. No. It is internal implementation/diagnostic state, not the supported application schema.
5. No. MongoDB supports increasing the bucket time span, not decreasing it in place.
Authoritative references
- Time Series Collections — time-series model, writable non-materialized view, automatic index, and sharding boundaries.
- Create and Query a Time Series Collection — timeField/metaField, granularity, custom bucketing, and collection options.
- About Time Series Data and Bucketing — system.buckets organization, bucket catalog, bucket creation/closure, and out-of-order timestamps.
- Time Series Collection Considerations — metadata cardinality, bucket density, granularity, compression, and zone-sharding boundary.
- Set Granularity for Time Series Data — granularity and custom bucket span/rounding behavior.
- Add Secondary Indexes to Time Series Collections — automatic and additional indexes plus sort/query support.
- Time Series Collection Limitations — bucket limits and index/query/update restrictions.
- Automatic Removal for Time Series Collections — expireAfterSeconds, bucket-level TTL timing, and collMod.
- Time Series Compression — zstd and column-compression mechanisms and tradeoffs.
- $collStats — time-series storageStats and bucket diagnostic fields.
- serverStatus — bucket catalog counters and server-level time-series diagnostics.
- $setWindowFields — time-range windows, ordering requirements, and analytics behavior.
- Explain Results — queryPlanner/executionStats interpretation and plan-format caveats.
- MongoDB 8.3 Release Notes — current stable 8.3 series and time-series compatibility changes.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17.x driver baseline where client measurement is used.