Chapter 21 · TTL, Retention, Expiration, Backfills, and Lifecycle Automation
Design Retention for Sessions, Events, User Data, Audit Records, and Legal Holds Without Accidental Loss
Build a retention control plane for AtlasMart sessions, events, user data, audit records, and legal holds that prevents accidental loss and leaves verifiable lifecycle evidence.
1. AtlasMart problem: retention is now a system, not a field
By this point AtlasMart has sessions, high-volume events,
customer profile/order data, audit records, soft-deleted user
accounts, legal holds, nested resources, backups, exports, and
event-driven side effects. A single
expireAt convention cannot safely govern all of
them. The final lesson builds a retention control plane: policy
registry, document classification, eligibility calculation, hold
overrides, deletion mechanism, evidence, recovery dependency,
and periodic verification.
- Design a retention matrix for sessions, events, user data, audit data, and legal holds.
- Enforce retention transitions with trusted policy identifiers and hold overrides.
- Prove that user-visible expiration, physical deletion, archive, and recovery are separate controls.
- Create verification jobs that find stale data, missing TTL values, unsafe holds, and orphaned nested data.
- Produce audit evidence that explains why each destructive action was allowed.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps the course-wide project identity
demo-atlasmart-firestore. Mandatory work is local
and no-cost: Node.js 22+, Firebase CLI 15.30.0,
Firestore emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, and the same Standard-edition
Native-mode mental model used by earlier chapters. The default
database is (default). TTL sweeping itself is
not treated as an emulator guarantee: the lab
models eligibility and lifecycle rules deterministically,
while an optional real-project check verifies actual managed
TTL deletion only in an isolated billed project. TTL deletes
are excluded from Firestore's free usage and require billing.
Expiration is a business fact; TTL deletion is an
asynchronous storage-cleanup mechanism.
A document whose expireAt is in the past may
still exist and be returned by queries until the managed
sweeper deletes it. Therefore product behavior must filter or
reject expired state explicitly when timeliness matters. TTL
is not a scheduler, not a transaction boundary, not a
legal-hold engine, and not a substitute for backups/PITR.
2. Retention matrix
| Data class | Visibility rule | Physical lifecycle | Hold behavior | Recovery dependency |
|---|---|---|---|---|
| Sessions | invalid after business expiry | TTL cleanup | usually no hold | normally none beyond security need |
| Telemetry/events | normal until policy expiry | TTL or batch + TTL | hold only where governance says so | optional archive for analytics/compliance |
| User profile/order data | soft delete or restricted immediately after approved request | workflow after dependency/legal checks | hold can block purge | backup/PITR/export as policy requires |
| Audit records | restricted access; not user-deletable | archive/explicit retention workflow | hold extends retention | governed archive + backup |
| Nested subcollections | inherit explicit owner lifecycle, not automatic cascade | recursive workflow when required | same owner/hold rules | verification must scan nested copies |
3. Policy registry and immutable evidence
{ "version":"2026-09-17.1", "classes":{ "session-30d":{"physical":"ttl","businessExpiry":true,"hold":"deny"}, "event-90d":{"physical":"ttl","businessExpiry":false,"hold":"allow"}, "user-delete":{"physical":"workflow","businessExpiry":false,"hold":"allow"}, "audit-7y":{"physical":"archive-workflow","businessExpiry":false,"hold":"allow"} }}
Store the registry in version control and record
policyVersion on lifecycle/audit evidence. A client
may request deletion, but only trusted server logic should
assign/transition regulated retention classes or release holds.
4. Deletion authorization predicate
export function canDestroy({doc, policy, nowMs, archiveVerified}) { if (doc.legalHold) return {ok:false, reason:"LEGAL_HOLD"}; if (!policy) return {ok:false, reason:"UNKNOWN_POLICY"}; if (policy.physical === "archive-workflow" && !archiveVerified) return {ok:false, reason:"ARCHIVE_NOT_VERIFIED"}; if (doc.expireAtMs != null && doc.expireAtMs > nowMs) return {ok:false, reason:"NOT_EXPIRED"}; return {ok:true, reason:"AUTHORIZED_BY_POLICY"};}
This predicate is intentionally stricter than “timestamp is old.” It forces hold and archive requirements into the same decision surface. In production, authorization also checks tenant ownership, request identity, approvals, and retention-policy version.
5. Verification jobs: detect policy drift before data loss
Run non-destructive audits that compare document state to policy
expectations. Examples: eligible docs missing
expireAt, held docs that accidentally have
destructive TTL values, unknown retention classes,
expired-but-product-visible docs, orphaned subcollections after
parent deletion, and stale archives lacking source
hash/evidence.
const findings=[];for (const d of docs) { if (!registry.classes[d.retentionClass]) findings.push([d.id,"UNKNOWN_CLASS"]); if (d.legalHold && d.expireAtMs != null) findings.push([d.id,"HOLD_WITH_TTL"]); if (d.retentionClass === "session-30d" && d.expireAtMs == null) findings.push([d.id,"MISSING_TTL"]); if (d.deletedAtMs && d.visibleToUser) findings.push([d.id,"SOFT_DELETE_QUERY_LEAK"]);}if (findings.length) process.exitCode=1;console.log(JSON.stringify({findings}, null, 2));
6. Security boundary
Client Security Rules can help prevent untrusted users from rewriting owner/retention fields, but Admin/server SDKs bypass Rules and must enforce application authorization plus IAM. App Check can reduce abusive client traffic but is not authorization and does not govern backend TTL configuration. TTL policy administration requires IAM permissions such as Firestore index policy update/list/get; separate that capability from normal application writers.
7. Failure injection matrix
| Fault | Expected detection | Safe repair |
|---|---|---|
Client changes retentionClass from audit to
session
|
Rules/application validation fails | Immutable/allow-listed server transition |
Held audit row gets past expireAt |
Retention audit emits HOLD_WITH_TTL |
Clear TTL field, record incident, confirm no managed delete occurred |
| User-delete workflow removes parent only | Orphan scan finds subcollection data | Scoped recursive workflow + verification |
| TTL delete trigger runs twice | Idempotency marker suppresses duplicate effect | Keep event/document idempotency key |
| Query returns expired session before physical delete | Business visibility test fails | Filter/reject by business expiry independent of storage existence |
| Archive copy missing |
Destroy predicate returns
ARCHIVE_NOT_VERIFIED
|
Retry archive/verification; do not purge |
8. Operational dashboard
Track at least: candidate count by retention class; missing/invalid TTL count; hold count; soft-delete age distribution; backfill checkpoint/errors; TTL deletion count; expiration-to-deletion delay; delete-trigger failures/retries; orphan count; archive verification failures; and spend. Keep business-expiry SLOs separate from TTL cleanup metrics.
9. End-to-end local acceptance test
- Seed one active session, one expired session, one event eligible for cleanup, one user soft-delete workflow, one audit record on legal hold, and one parent with nested data.
- Run business visibility assertions.
- Run the retention audit and expect zero unsafe findings.
- Inject a hold-with-TTL bug and prove the audit fails.
- Repair it; rerun to green.
- Simulate a duplicate delete event and prove downstream effect remains single.
- Run orphan scan and verify nested lifecycle inventory.
- Reset emulator fixtures.
This local suite is the chapter's mandatory proof. The optional managed TTL test is additional evidence, not a prerequisite for learning.
Production judgment: define deletion by policy, not implementation
A durable retention design starts from obligations: who owns the data, when it becomes invalid, when it may be physically destroyed, what holds override deletion, whether export/archive is required, how recovery works, and what evidence must survive. TTL is one implementation mechanism inside that policy. If AtlasMart later changes edition, mode, region, SDK, or archive system, the policy contract should remain understandable and testable.
10. Chapter verification checklist
- All five data classes have explicit retention policies.
- Product-visible expiry is independent of physical TTL deletion.
- TTL fields are exempted from Standard indexing when safe and justified.
- Backfill is checkpointed, idempotent, and budgeted.
- Legal holds override destructive eligibility.
- Delete-trigger side effects are idempotent.
- Subcollections are included in deletion plans.
- Monitoring distinguishes business SLOs from TTL cleanup lag.
- Recovery dependency is documented for every destructive path.
11. Bridge to Chapter 22
Retention intentionally removes data. Chapter 22 asks the inverse operational question: when deletion or corruption is accidental, what backup, PITR, restore, and recovery workflow can recover the database—and how do you prove RPO/RTO rather than assuming multi-region placement is a backup?
Knowledge check
- What is the central architectural role of TTL?
- Who should be allowed to change regulated retention classes?
- What should a retention audit detect on legal-hold data?
- Why track orphaned subcollections?
- What does Chapter 22 add?
Review the answers
1. Asynchronous physical cleanup after policy eligibility, not the policy itself.
2. Trusted server/governance paths with explicit authorization; not arbitrary clients.
3. Any destructive TTL value/path or other state that could make held data eligible for removal.
4. TTL/parent deletion does not cascade, so nested data can survive and violate deletion obligations.
5. Recovery mechanisms—backups, PITR, restore drills, RPO/RTO—that TTL does not provide.
Summary and next step
This lesson established the working contract for Design Retention for Sessions, Events, User Data, Audit Records, and Legal Holds Without Accidental Loss. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Scheduled Backups, Backup Retention, Database Restore/Clone Concepts, and Access Control.
Authoritative references
- Firebase · Manage data retention with TTL policies
- Firebase · Firestore index overview and index exemptions
- Firebase · Cloud Firestore pricing and TTL billing
- Firebase · Firestore quotas and limits
- Firebase · Cloud Firestore triggers
- Firebase · Enterprise TTL indexes
- Google Cloud · MongoDB compatibility TTL indexes
- Google Cloud · MongoDB compatibility release notes
- Google Cloud · Firestore Monitoring metrics