Prompt 19 · Lesson 02 · Series and metadata design

Model Sensors, Metrics, Events, and Device Metadata for Efficient Compression and Queries

Design stable series metadata and separate event-time meaning from arrival-time evidence.

Intermediate–Advanced120–180 minutesMetadata/compression labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Design a stable metaField for sensors/metrics that supports series grouping and common filters without embedding high-churn attributes.

02

Separate event time from ingest time and explain how choosing one as the timeField changes bucketing, retention, and late-arrival semantics.

03

Use bucket/storage diagnostics to compare stable versus high-cardinality metadata without fabricating compression ratios.

04

Explain time-series column/zstd compression at a practical level and why similar ordered values tend to compress well.

05

Handle late/out-of-order measurements and migration/backfill ordering as an operational concern rather than assuming perfect arrival order.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where a driver is used. The mandatory lab is a disposable standalone on loopback port 27161 named atlasmart-ch19-l2; replication or sharding is not required to learn the time-series mechanisms in this chapter. Authentication and TLS are disabled only for the isolated local lab. Feature Compatibility Version (FCV) is inspected. FCV is observed and never changed. Default read/write concern and primary read preference apply. Atlas, Search, Vector Search, KMS, and Enterprise Advanced are not mandatory. Internal system.buckets.* data and time-series diagnostic counters are used only for observation; they are not application APIs. Product commands were not executed in this generation environment because Docker, mongod, mongosh, and PyMongo are unavailable here, so environment-dependent bucket counts, storage ratios, explain plans, and throughput/latency values must be measured on the learner machine rather than copied as invented output. The collection uses eventAt as the timeField and keeps ingestedAt as an ordinary measurement field so event-time and arrival-time semantics stay distinct.

1. Model a series around stable identity, not every label available

AtlasMart sensors emit a measurement every minute. The series identity needed by most queries is tenant + store + sensor. Firmware version, request ID, network route, and ingestion timestamp may change often; placing those values in meta would split one physical sensor into many storage series.

A good metaField is both stable and useful for filtering. MongoDB requires exact metadata equality for bucket grouping, so even semantically equivalent object/array values that differ in representation can create separate groups. Query scalar metadata subfields rather than relying on whole-object equality when application serialization order can vary.

Field Role Why
eventAt timeField When the physical/business event happened; drives bucket time and collection TTL if enabled.
meta.tenantId stable metadata Tenant partition/filter dimension.
meta.storeId stable metadata Operational location identity.
meta.sensorId stable metadata Uniquely identifies one device series.
temperatureC, powerW measurements Values that change at every observation.
firmware, ingestedAt ordinary fields Useful evidence but not stable series identity.

2. Build stable and high-cardinality variants

start the disposable MongoDB 8.3.8 lab
docker rm -f atlasmart-ch19-l2 2>/dev/null || truedocker volume rm atlasmart-ch19-l2-data 2>/dev/null || truedocker run -d --name atlasmart-ch19-l2 \  -p 127.0.0.1:27161:27017 \  -v atlasmart-ch19-l2-data:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_alldocker exec atlasmart-ch19-l2 mongosh --quiet --eval 'printjson(db.adminCommand({buildInfo:1}).version);printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion);' 
create two time-series schemas and load the same logical workload
const d=db.getSiblingDB("atlasmart");for (const n of ["telemetry_ch19_l2","telemetry_badmeta_ch19_l2"]) d[n].drop();for (const n of ["telemetry_ch19_l2","telemetry_badmeta_ch19_l2"]) {  d.createCollection(n,{timeseries:{timeField:"eventAt",metaField:"meta",granularity:"minutes"}});}const base=ISODate("2026-09-01T00:00:00Z");const good=[], bad=[];for (let sensor=0;sensor<20;sensor++) {  for (let i=0;i<72;i++) {    const eventAt=new Date(base.getTime()+i*20*60*1000);    const common={eventAt, ingestedAt:new Date(eventAt.getTime()+2000+(i%5)*1000), temperatureC:18+(sensor%4)+(i%10)/10, powerW:80+(i%30), firmware:`1.${Math.floor(i/24)}`};    good.push({...common,meta:{tenantId:"tenant-a",storeId:"baku-01",sensorId:`sensor-${sensor.toString().padStart(2,"0")}`}});    bad.push({...common,meta:{tenantId:"tenant-a",storeId:"baku-01",sensorId:`sensor-${sensor.toString().padStart(2,"0")}`,requestId:`req-${sensor}-${i}`}});  }}d.telemetry_ch19_l2.insertMany(good);d.telemetry_badmeta_ch19_l2.insertMany(bad);printjson({good:d.telemetry_ch19_l2.countDocuments({}),bad:d.telemetry_badmeta_ch19_l2.countDocuments({})});

The fixture deterministically creates 1,440 measurements in each collection. That equal logical workload makes bucket/storage differences attributable primarily to metadata design. Do not interpret local timings as a universal benchmark; the lesson measures structure first.

compare bucket and storage diagnostics
const d=db.getSiblingDB("atlasmart");function evidence(n){  const s=d[n].aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats;  return {name:n,count:s.count,size:s.size,storageSize:s.storageSize,totalIndexSize:s.totalIndexSize,timeseries:s.timeseries};}printjson(evidence("telemetry_ch19_l2"));printjson(evidence("telemetry_badmeta_ch19_l2"));print("Interpret relative bucket density and storage; do not hard-code a compression ratio.");

3. Compression rewards coherent series, but measure it

Time-series buckets use a compressed storage format. MongoDB time-series data uses zstd and column-oriented compression techniques such as delta encoding, run-length encoding, and object/array compression. Similar adjacent values—timestamps increasing predictably, temperatures changing gradually, repeated metadata—are good candidates for compression. High-cardinality metadata and sparse buckets weaken those patterns.

collStats.size is the logical uncompressed record size, while storageSize reflects allocated compressed storage. The ratio between them is environment- and workload-specific. Index bytes are separate and must be counted when planning total footprint.

calculate only measured, local ratios
const d=db.getSiblingDB("atlasmart");for (const n of ["telemetry_ch19_l2","telemetry_badmeta_ch19_l2"]) {  const s=d[n].aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats;  printjson({    name:n,    logicalBytes:s.size,    allocatedStorageBytes:s.storageSize,    totalIndexBytes:s.totalIndexSize,    logicalToAllocated:s.storageSize ? s.size/s.storageSize : null,    bucketCount:s.timeseries.bucketCount,    avgBucketSize:s.timeseries.avgBucketSize  });}

4. Event time and ingest time answer different questions

Event time says when the measured phenomenon occurred. Ingest time says when MongoDB received it. AtlasMart may receive a freezer reading ten minutes late because a gateway was offline. If eventAt is the timeField, the measurement is placed according to its actual observation time and may require MongoDB to reopen an older eligible bucket. If ingestedAt were the timeField, storage and TTL would follow arrival time instead, which can simplify ingestion ordering but changes analytical/retention semantics.

insert a deliberately late measurement and inspect reopen evidence
const d=db.getSiblingDB("atlasmart");const before=d.telemetry_ch19_l2.aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats.timeseries;const late={  eventAt:ISODate("2026-09-01T03:10:00Z"),  ingestedAt:ISODate("2026-09-03T04:30:00Z"),  meta:{tenantId:"tenant-a",storeId:"baku-01",sensorId:"sensor-00"},  temperatureC:17.2,powerW:83,firmware:"1.2",lateArrival:true};d.telemetry_ch19_l2.insertOne(late);const after=d.telemetry_ch19_l2.aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats.timeseries;printjson({before:{bucketCount:before.bucketCount,numBucketsReopened:before.numBucketsReopened},after:{bucketCount:after.bucketCount,numBucketsReopened:after.numBucketsReopened}});printjson(d.telemetry_ch19_l2.find({lateArrival:true}).toArray());

The document itself must be queryable at its eventAt after insertion. Whether the diagnostic numBucketsReopened increments depends on whether a suitable archived bucket can actually be reopened on this run; if not, MongoDB can create another suitable bucket. That counter is evidence, not a guaranteed outcome. For large historical backfills, insert oldest-to-newest where practical to improve time-series migration behavior.

5. Wrong model: every metric name or firmware version defines metadata

Another failure mode is encoding high-churn dimensions such as firmware version, alert state, or arbitrary metric names into the metaField. This can make every state transition a new series. The repaired model keeps stable sensor identity in metadata and lets measurement fields evolve under an explicit application schema/version strategy.

Schema evolution

Time-series optimization does not mean “no schema.” Readers must tolerate old/new measurement fields, numeric type changes, missing values, and late data. Keep a schema/version signal if application evolution needs one, but do not automatically put that version into meta if doing so fragments the series.

6. Production judgment

Model metadata around stable ownership/identity and high-value filters. Measure unique metadata values, bucket count, bucket reopen/closure counters, storage/index footprint, and query selectivity. If cardinality grows with request volume rather than with the number of real series, the model is probably wrong. For out-of-order sources, track event-to-ingest delay distributions and test backfill behavior instead of assuming chronological arrival.

Bridge. Lesson 3 builds indexes and time-window analytics on this stable event-time model, then uses explain evidence to separate a correct result from an efficient query shape.

cleanup only the chapter-specific lab
docker rm -f atlasmart-ch19-l2 2>/dev/null || truedocker volume rm atlasmart-ch19-l2-data 2>/dev/null || true

Check your understanding

  1. Why should firmware usually remain outside the metaField?
  2. What does eventAt represent in this chapter?
  3. What does ingestedAt represent?
  4. Does a late insert always reopen an old bucket?
  5. Why compare size, storageSize, and totalIndexSize separately?
Review the answers

1. It can change over time; including it splits one physical source into multiple metadata series and can reduce bucket density.

2. The time the phenomenon occurred; it is the collection timeField and therefore drives bucketing and any collection-level TTL.

3. When the database/application received the measurement; it is useful for lateness analysis but is not the timeField in the main model.

4. No. MongoDB may reopen a suitable bucket or create another one; observe bucket diagnostics rather than assuming a fixed path.

5. Logical data size, compressed allocated storage, and index storage are different capacity components.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.