Prompt 19 · Lesson 02 · Series and metadata design
Model Sensors, Metrics, Events, and Device Metadata for Efficient Compression and Queries
Design stable series metadata and separate event-time meaning from arrival-time evidence.
Learning objectives
Design a stable metaField for sensors/metrics that supports series grouping and common filters without embedding high-churn attributes.
Separate event time from ingest time and explain how choosing one as the timeField changes bucketing, retention, and late-arrival semantics.
Use bucket/storage diagnostics to compare stable versus high-cardinality metadata without fabricating compression ratios.
Explain time-series column/zstd compression at a practical level and why similar ordered values tend to compress well.
Handle late/out-of-order measurements and migration/backfill ordering as an operational concern rather than assuming perfect arrival order.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where a driver is used. The mandatory lab is
a disposable standalone on loopback port
27161 named atlasmart-ch19-l2;
replication or sharding is not required to learn the time-series
mechanisms in this chapter. Authentication and TLS are disabled
only for the isolated local lab. Feature Compatibility Version
(FCV) is inspected.
FCV is observed and never changed. Default
read/write concern and primary read preference apply. Atlas,
Search, Vector Search, KMS, and Enterprise Advanced are not
mandatory. Internal system.buckets.* data and
time-series diagnostic counters are used only for observation;
they are not application APIs. Product commands were not
executed in this generation environment because Docker, mongod,
mongosh, and PyMongo are unavailable here, so
environment-dependent bucket counts, storage ratios, explain
plans, and throughput/latency values must be measured on the
learner machine rather than copied as invented output. The
collection uses eventAt as the timeField and keeps
ingestedAt as an ordinary measurement field so
event-time and arrival-time semantics stay distinct.
1. Model a series around stable identity, not every label available
AtlasMart sensors emit a measurement every minute. The series
identity needed by most queries is tenant + store + sensor.
Firmware version, request ID, network route, and ingestion
timestamp may change often; placing those values in
meta would split one physical sensor into many
storage series.
A good metaField is both stable and useful for filtering. MongoDB requires exact metadata equality for bucket grouping, so even semantically equivalent object/array values that differ in representation can create separate groups. Query scalar metadata subfields rather than relying on whole-object equality when application serialization order can vary.
| Field | Role | Why |
|---|---|---|
eventAt |
timeField | When the physical/business event happened; drives bucket time and collection TTL if enabled. |
meta.tenantId |
stable metadata | Tenant partition/filter dimension. |
meta.storeId |
stable metadata | Operational location identity. |
meta.sensorId |
stable metadata | Uniquely identifies one device series. |
temperatureC, powerW |
measurements | Values that change at every observation. |
firmware, ingestedAt |
ordinary fields | Useful evidence but not stable series identity. |
2. Build stable and high-cardinality variants
docker rm -f atlasmart-ch19-l2 2>/dev/null || truedocker volume rm atlasmart-ch19-l2-data 2>/dev/null || truedocker run -d --name atlasmart-ch19-l2 \ -p 127.0.0.1:27161:27017 \ -v atlasmart-ch19-l2-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_alldocker exec atlasmart-ch19-l2 mongosh --quiet --eval 'printjson(db.adminCommand({buildInfo:1}).version);printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion);'
const d=db.getSiblingDB("atlasmart");for (const n of ["telemetry_ch19_l2","telemetry_badmeta_ch19_l2"]) d[n].drop();for (const n of ["telemetry_ch19_l2","telemetry_badmeta_ch19_l2"]) { d.createCollection(n,{timeseries:{timeField:"eventAt",metaField:"meta",granularity:"minutes"}});}const base=ISODate("2026-09-01T00:00:00Z");const good=[], bad=[];for (let sensor=0;sensor<20;sensor++) { for (let i=0;i<72;i++) { const eventAt=new Date(base.getTime()+i*20*60*1000); const common={eventAt, ingestedAt:new Date(eventAt.getTime()+2000+(i%5)*1000), temperatureC:18+(sensor%4)+(i%10)/10, powerW:80+(i%30), firmware:`1.${Math.floor(i/24)}`}; good.push({...common,meta:{tenantId:"tenant-a",storeId:"baku-01",sensorId:`sensor-${sensor.toString().padStart(2,"0")}`}}); bad.push({...common,meta:{tenantId:"tenant-a",storeId:"baku-01",sensorId:`sensor-${sensor.toString().padStart(2,"0")}`,requestId:`req-${sensor}-${i}`}}); }}d.telemetry_ch19_l2.insertMany(good);d.telemetry_badmeta_ch19_l2.insertMany(bad);printjson({good:d.telemetry_ch19_l2.countDocuments({}),bad:d.telemetry_badmeta_ch19_l2.countDocuments({})});
The fixture deterministically creates 1,440 measurements in each collection. That equal logical workload makes bucket/storage differences attributable primarily to metadata design. Do not interpret local timings as a universal benchmark; the lesson measures structure first.
const d=db.getSiblingDB("atlasmart");function evidence(n){ const s=d[n].aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats; return {name:n,count:s.count,size:s.size,storageSize:s.storageSize,totalIndexSize:s.totalIndexSize,timeseries:s.timeseries};}printjson(evidence("telemetry_ch19_l2"));printjson(evidence("telemetry_badmeta_ch19_l2"));print("Interpret relative bucket density and storage; do not hard-code a compression ratio.");
3. Compression rewards coherent series, but measure it
Time-series buckets use a compressed storage format. MongoDB time-series data uses zstd and column-oriented compression techniques such as delta encoding, run-length encoding, and object/array compression. Similar adjacent values—timestamps increasing predictably, temperatures changing gradually, repeated metadata—are good candidates for compression. High-cardinality metadata and sparse buckets weaken those patterns.
collStats.size is the logical uncompressed record
size, while storageSize reflects allocated
compressed storage. The ratio between them is environment- and
workload-specific. Index bytes are separate and must be counted
when planning total footprint.
const d=db.getSiblingDB("atlasmart");for (const n of ["telemetry_ch19_l2","telemetry_badmeta_ch19_l2"]) { const s=d[n].aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats; printjson({ name:n, logicalBytes:s.size, allocatedStorageBytes:s.storageSize, totalIndexBytes:s.totalIndexSize, logicalToAllocated:s.storageSize ? s.size/s.storageSize : null, bucketCount:s.timeseries.bucketCount, avgBucketSize:s.timeseries.avgBucketSize });}
4. Event time and ingest time answer different questions
Event time says when the measured phenomenon
occurred. Ingest time says when MongoDB
received it. AtlasMart may receive a freezer reading ten minutes
late because a gateway was offline. If eventAt is
the timeField, the measurement is placed according to its actual
observation time and may require MongoDB to reopen an older
eligible bucket. If ingestedAt were the timeField,
storage and TTL would follow arrival time instead, which can
simplify ingestion ordering but changes analytical/retention
semantics.
const d=db.getSiblingDB("atlasmart");const before=d.telemetry_ch19_l2.aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats.timeseries;const late={ eventAt:ISODate("2026-09-01T03:10:00Z"), ingestedAt:ISODate("2026-09-03T04:30:00Z"), meta:{tenantId:"tenant-a",storeId:"baku-01",sensorId:"sensor-00"}, temperatureC:17.2,powerW:83,firmware:"1.2",lateArrival:true};d.telemetry_ch19_l2.insertOne(late);const after=d.telemetry_ch19_l2.aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats.timeseries;printjson({before:{bucketCount:before.bucketCount,numBucketsReopened:before.numBucketsReopened},after:{bucketCount:after.bucketCount,numBucketsReopened:after.numBucketsReopened}});printjson(d.telemetry_ch19_l2.find({lateArrival:true}).toArray());
The document itself must be queryable at its
eventAt after insertion. Whether the diagnostic
numBucketsReopened increments depends on whether a
suitable archived bucket can actually be reopened on this run;
if not, MongoDB can create another suitable bucket. That counter
is evidence, not a guaranteed outcome. For large historical
backfills, insert oldest-to-newest where practical to improve
time-series migration behavior.
5. Wrong model: every metric name or firmware version defines metadata
Another failure mode is encoding high-churn dimensions such as firmware version, alert state, or arbitrary metric names into the metaField. This can make every state transition a new series. The repaired model keeps stable sensor identity in metadata and lets measurement fields evolve under an explicit application schema/version strategy.
Time-series optimization does not mean “no schema.” Readers must tolerate old/new measurement fields, numeric type changes, missing values, and late data. Keep a schema/version signal if application evolution needs one, but do not automatically put that version into meta if doing so fragments the series.
6. Production judgment
Model metadata around stable ownership/identity and high-value filters. Measure unique metadata values, bucket count, bucket reopen/closure counters, storage/index footprint, and query selectivity. If cardinality grows with request volume rather than with the number of real series, the model is probably wrong. For out-of-order sources, track event-to-ingest delay distributions and test backfill behavior instead of assuming chronological arrival.
Bridge. Lesson 3 builds indexes and time-window analytics on this stable event-time model, then uses explain evidence to separate a correct result from an efficient query shape.
docker rm -f atlasmart-ch19-l2 2>/dev/null || truedocker volume rm atlasmart-ch19-l2-data 2>/dev/null || true
Check your understanding
- Why should firmware usually remain outside the metaField?
- What does eventAt represent in this chapter?
- What does ingestedAt represent?
- Does a late insert always reopen an old bucket?
- Why compare size, storageSize, and totalIndexSize separately?
Review the answers
1. It can change over time; including it splits one physical source into multiple metadata series and can reduce bucket density.
2. The time the phenomenon occurred; it is the collection timeField and therefore drives bucketing and any collection-level TTL.
3. When the database/application received the measurement; it is useful for lateness analysis but is not the timeField in the main model.
4. No. MongoDB may reopen a suitable bucket or create another one; observe bucket diagnostics rather than assuming a fixed path.
5. Logical data size, compressed allocated storage, and index storage are different capacity components.
Authoritative references
- Time Series Collections — time-series model, writable non-materialized view, automatic index, and sharding boundaries.
- Create and Query a Time Series Collection — timeField/metaField, granularity, custom bucketing, and collection options.
- About Time Series Data and Bucketing — system.buckets organization, bucket catalog, bucket creation/closure, and out-of-order timestamps.
- Time Series Collection Considerations — metadata cardinality, bucket density, granularity, compression, and zone-sharding boundary.
- Set Granularity for Time Series Data — granularity and custom bucket span/rounding behavior.
- Add Secondary Indexes to Time Series Collections — automatic and additional indexes plus sort/query support.
- Time Series Collection Limitations — bucket limits and index/query/update restrictions.
- Automatic Removal for Time Series Collections — expireAfterSeconds, bucket-level TTL timing, and collMod.
- Time Series Compression — zstd and column-compression mechanisms and tradeoffs.
- $collStats — time-series storageStats and bucket diagnostic fields.
- serverStatus — bucket catalog counters and server-level time-series diagnostics.
- $setWindowFields — time-range windows, ordering requirements, and analytics behavior.
- Explain Results — queryPlanner/executionStats interpretation and plan-format caveats.
- MongoDB 8.3 Release Notes — current stable 8.3 series and time-series compatibility changes.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17.x driver baseline where client measurement is used.