Prompt 19 · Lesson 04 · Retention and time semantics

Expiration/Retention Policies and the Difference Between Event Time and Ingest Time

Make retention semantics explicit and observe why TTL is eventual, bucket-level behavior.

Advanced120–190 minutesTTL/event-time labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Explain collection-level time-series retention as timeField + expireAfterSeconds plus asynchronous bucket-level deletion.

02

Distinguish event-time retention from ingest-time retention for late-arriving telemetry and explain why they encode different business policies.

03

Observe collection options, bucket time ranges, and TTL delay without promising immediate deletion.

04

Use collMod safely in a disposable lab to change or disable retention while understanding that old data may become eligible in bulk.

05

Diagnose the mistake of using TTL as an exact scheduler, legal-hold mechanism, or backup strategy.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where a driver is used. The mandatory lab is a disposable standalone on loopback port 27163 named atlasmart-ch19-l4; replication or sharding is not required to learn the time-series mechanisms in this chapter. Authentication and TLS are disabled only for the isolated local lab. Feature Compatibility Version (FCV) is inspected. FCV is observed and never changed. Default read/write concern and primary read preference apply. Atlas, Search, Vector Search, KMS, and Enterprise Advanced are not mandatory. Internal system.buckets.* data and time-series diagnostic counters are used only for observation; they are not application APIs. Product commands were not executed in this generation environment because Docker, mongod, mongosh, and PyMongo are unavailable here, so environment-dependent bucket counts, storage ratios, explain plans, and throughput/latency values must be measured on the learner machine rather than copied as invented output. The lab uses short custom one-minute buckets so TTL behavior becomes observable without waiting for an hour-wide default seconds bucket to age out.

1. Retention is a data policy expressed through the timeField

AtlasMart wants raw freezer telemetry for 30 days, while long-term rollups live elsewhere. On a time-series collection, expireAfterSeconds defines expiration relative to the collection's timeField. That means the choice between event time and ingest time is not cosmetic: it determines which clock the retention policy follows.

MongoDB does not delete each measurement the instant it reaches the threshold. Time-series TTL removes expired buckets. A bucket is eligible only after all measurements in it are expired, then a background process removes it on a later pass. The TTL monitor runs periodically (documented as every 60 seconds), and actual deletion can take longer under workload.

2. Create event-time retention with intentionally short bucket spans

start the disposable MongoDB 8.3.8 lab
docker rm -f atlasmart-ch19-l4 2>/dev/null || truedocker volume rm atlasmart-ch19-l4-data 2>/dev/null || truedocker run -d --name atlasmart-ch19-l4 \  -p 127.0.0.1:27163:27017 \  -v atlasmart-ch19-l4-data:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_alldocker exec atlasmart-ch19-l4 mongosh --quiet --eval 'printjson(db.adminCommand({buildInfo:1}).version);printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion);' 
create a short-retention event-time collection
const d=db.getSiblingDB("atlasmart");d.telemetry_ch19_l4.drop();d.createCollection("telemetry_ch19_l4",{  timeseries:{    timeField:"eventAt",    metaField:"meta",    bucketMaxSpanSeconds:60,    bucketRoundingSeconds:60  },  expireAfterSeconds:120});const now=new Date();d.telemetry_ch19_l4.insertMany([  {eventAt:new Date(now.getTime()-10*60*1000),ingestedAt:now,meta:{tenantId:"tenant-a",sensorId:"late-old"},v:10},  {eventAt:new Date(now.getTime()-9*60*1000),ingestedAt:now,meta:{tenantId:"tenant-a",sensorId:"late-old"},v:11},  {eventAt:new Date(now.getTime()-20*1000),ingestedAt:now,meta:{tenantId:"tenant-a",sensorId:"fresh"},v:20}]);printjson(d.runCommand({listCollections:1,filter:{name:"telemetry_ch19_l4"}}).cursor.firstBatch[0]);printjson(d.getCollection("system.buckets.telemetry_ch19_l4").find({}, {meta:1,"control.min.eventAt":1,"control.max.eventAt":1}).toArray());print("public count immediately after insert",d.telemetry_ch19_l4.countDocuments({}));

The two old event-time measurements are already beyond the 120-second threshold even though ingestedAt is current. They can still be visible immediately because deletion is asynchronous and bucket-based. The fresh measurement is in a different metadata series/bucket and should not be considered expired.

3. Poll for evidence; never assert exact deletion time

observe eventual TTL removal with a bounded poll
const d=db.getSiblingDB("atlasmart");for (let i=0;i<15;i++) {  const c=d.telemetry_ch19_l4.countDocuments({});  print(new Date().toISOString(),"count",c);  if (c===1) break;  sleep(10000);}printjson(d.telemetry_ch19_l4.find({}, {_id:0,eventAt:1,ingestedAt:1,"meta.sensorId":1,v:1}).toArray());

The expected invariant is eventual removal of fully expired buckets, not “delete at exactly T+120 seconds.” If the bounded poll still shows expired data, record that as real retention lag rather than declaring the lab failed. Workload and bucket timing affect when removal becomes visible.

4. Event-time versus ingest-time retention changes the answer for late data

create an ingest-time comparison collection
const d=db.getSiblingDB("atlasmart");d.telemetry_ingest_ch19_l4.drop();d.createCollection("telemetry_ingest_ch19_l4",{  timeseries:{timeField:"ingestedAt",metaField:"meta",bucketMaxSpanSeconds:60,bucketRoundingSeconds:60},  expireAfterSeconds:120});const now=new Date();d.telemetry_ingest_ch19_l4.insertOne({eventAt:new Date(now.getTime()-24*60*60*1000),ingestedAt:now,meta:{tenantId:"tenant-a",sensorId:"gateway-replay"},v:42});printjson(d.telemetry_ingest_ch19_l4.findOne({"meta.sensorId":"gateway-replay"},{_id:0,eventAt:1,ingestedAt:1,v:1}));

The same 24-hour-old event is “new” under ingest-time retention because its timeField is the current arrival time. That may be the right policy for a replay/audit queue, but it is different from “keep measurements for 30 days after they occurred.” Choose the clock from the business rule, not from whichever timestamp is easiest to populate.

5. Change retention safely and understand bulk eligibility

collMod can enable, change, or disable expireAfterSeconds on a time-series collection. Lowering retention can suddenly make a large historical population eligible for deletion, which can create storage/I/O pressure. In production, estimate eligible volume and observe deletion throughput before tightening retention aggressively.

inspect, change, then disable retention in the disposable lab
const d=db.getSiblingDB("atlasmart");printjson(d.runCommand({listCollections:1,filter:{name:"telemetry_ch19_l4"}}).cursor.firstBatch[0].options);printjson(d.runCommand({collMod:"telemetry_ch19_l4",expireAfterSeconds:300}));printjson(d.runCommand({collMod:"telemetry_ch19_l4",expireAfterSeconds:"off"}));printjson(d.runCommand({listCollections:1,filter:{name:"telemetry_ch19_l4"}}).cursor.firstBatch[0].options);

6. Wrong assumptions: TTL is a scheduler, backup, or legal hold

TTL is inappropriate when a workflow must fire at an exact second; use an application scheduler/queue. TTL is also not backup: expiration is a delete that becomes part of database state, and independent recovery copies are still required. For legal/regulated retention, document whether “event time” or “ingest time” is the governing clock, how delayed deletion is handled, and how holds override normal expiration. A separate archive or policy-controlled collection is often clearer than dynamically defeating TTL on individual measurements.

7. Production judgment

Use time-series TTL when bounded raw-history retention aligns with bucket-level asynchronous removal. Monitor oldest retained event time, bucket ranges, TTL deletion activity, storage trends, and delayed-arrival distributions. Test retention changes against representative volumes. If stale measurements must disappear before authorization or compliance decisions, do not rely on eventual TTL deletion as the enforcement boundary—filter by time in the read path and enforce policy explicitly.

Bridge. Lesson 5 turns bucket density, index size, retention horizon, and ingestion rate into a capacity worksheet and a repeatable local load-measurement harness.

cleanup only the chapter-specific lab
docker rm -f atlasmart-ch19-l4 2>/dev/null || truedocker volume rm atlasmart-ch19-l4-data 2>/dev/null || true

Check your understanding

  1. What timestamp does collection-level time-series TTL use?
  2. Why can an expired measurement remain visible?
  3. How does event-time retention differ from ingest-time retention for a late replay?
  4. Can collMod change expireAfterSeconds?
  5. Why is TTL not a precise scheduler?
Review the answers

1. The value of the configured timeField plus expireAfterSeconds.

2. Deletion is bucket-level and asynchronous; all measurements in the bucket must be expired and the background remover must run.

3. Event-time retention may make it immediately eligible based on when it occurred; ingest-time retention starts aging from when it arrived.

4. Yes, including disabling it with off, but tightening retention can make a large population eligible at once.

5. MongoDB does not guarantee immediate deletion at the threshold; bucket eligibility and background task timing add delay.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.