Chapter 21 · TTL, Retention, Expiration, Backfills, and Lifecycle Automation
Backfilling TTL Fields Safely, Monitoring Deletion, Event Side Effects, and Cost
Backfill expiration fields with explicit eligibility gates, bounded write rates, progress checkpoints, idempotent delete-side effects, Monitoring metrics, and a rollback/stop plan.
1. AtlasMart problem: enabling TTL on old data creates an immediate backlog
AtlasMart has 40 million event documents and wants 90-day
retention. If the team enables a TTL policy and then writes
expireAt to every historical record, many documents
become immediately eligible. Managed deletion is smoothed and
low-priority, but the backfill itself still writes every changed
document and can affect indexes, billing, replication,
listeners, and application load. A safe rollout therefore treats
TTL enablement and field backfill as a migration.
- Plan TTL backfill in bounded, resumable batches with explicit stop conditions.
- Keep legal-hold and retention-class exclusions out of destructive eligibility.
- Monitor TTL deletion count and expiration-to-deletion delay without treating them as business timers.
- Make delete-trigger side effects idempotent and distinguish TTL deletes from business workflow intent.
- Estimate backfill writes, index effects, TTL-delete billing, and rollback semantics before rollout.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps the course-wide project identity
demo-atlasmart-firestore. Mandatory work is local
and no-cost: Node.js 22+, Firebase CLI 15.30.0,
Firestore emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, and the same Standard-edition
Native-mode mental model used by earlier chapters. The default
database is (default). TTL sweeping itself is
not treated as an emulator guarantee: the lab
models eligibility and lifecycle rules deterministically,
while an optional real-project check verifies actual managed
TTL deletion only in an isolated billed project. TTL deletes
are excluded from Firestore's free usage and require billing.
Expiration is a business fact; TTL deletion is an
asynchronous storage-cleanup mechanism.
A document whose expireAt is in the past may
still exist and be returned by queries until the managed
sweeper deletes it. Therefore product behavior must filter or
reject expired state explicitly when timeliness matters. TTL
is not a scheduler, not a transaction boundary, not a
legal-hold engine, and not a substitute for backups/PITR.
2. Two separate operations: policy creation and data backfill
Creating a TTL policy can take at least several minutes to become active. Adding past expiration values to existing documents makes them eligible once the policy is active. Keep those operations separate in your runbook so you can stop after policy creation, inspect status, and then start a controlled backfill rather than combining everything into one irreversible release.
gcloud firestore fields ttls update expireAt \ --collection-group=events --enable-ttl --async \ --project=YOUR_ISOLATED_BILLED_PROJECTgcloud firestore fields ttls list \ --collection-group=events \ --project=YOUR_ISOLATED_BILLED_PROJECTgcloud firestore operations list \ --project=YOUR_ISOLATED_BILLED_PROJECT
3. Backfill state machine
| State | Action | Evidence before next state |
|---|---|---|
| DISCOVER | Count candidates by retention class/age/hold | candidate count + sample hashes |
| DRY_RUN | Compute expireAt without writes |
distribution by future/past/hold + estimated writes |
| CANARY | Write a tiny bounded subset | write success + query regression + no protected docs changed |
| BATCHED | Run resumable batches | checkpoint cursor + writes + errors + latency |
| OBSERVE | Watch TTL metrics and app health | deletion count/delay + errors + billing trend |
| COMPLETE | Verify no unintended candidates remain | post-backfill counts + hold audit + sampled documents |
4. Mandatory local backfill harness
const policyDays = {"session-30d":30,"event-90d":90};const DAY = 86400000;export function computeExpireAt(doc) { if (doc.legalHold) return {skip:true, reason:"LEGAL_HOLD"}; const days = policyDays[doc.retentionClass]; if (!days) return {skip:true, reason:"NO_TTL_POLICY"}; return {skip:false, expireAtMs: doc.createdAtMs + days*DAY};}export async function backfillPage(rows, checkpoint) { const out=[]; for (const row of rows) { const d=computeExpireAt(row); if (!d.skip && row.expireAtMs == null) out.push({id:row.id, expireAtMs:d.expireAtMs}); } return {writes:out, nextCheckpoint:rows.at(-1)?.id ?? checkpoint};}
import assert from "node:assert/strict";const rows=[ {id:"e1",retentionClass:"event-90d",createdAtMs:0,expireAtMs:null,legalHold:false}, {id:"a1",retentionClass:"audit-7y",createdAtMs:0,expireAtMs:null,legalHold:true}];const r=await backfillPage(rows,null);assert.deepEqual(r.writes.map(x=>x.id), ["e1"]);console.log("legal-hold-exclusion: PASS");
Persist the checkpoint outside the batch itself or in a dedicated migration-control document. Reruns must be idempotent: documents already carrying the correct expiration should not be rewritten unnecessarily.
5. Monitoring: measure the sweeper without turning it into an SLA
Cloud Monitoring exposes TTL deletion count and expiration-to-deletion delay. Use these to detect backlog trends, not to promise that a particular document will disappear by a deadline. Record application error rate, write latency, trigger invocation/error counts, and billing alongside TTL metrics because a healthy sweeper can still coincide with an unhealthy application backfill.
| Signal | What it tells you | What it does not tell you |
|---|---|---|
ttl_deletion_count |
Documents physically removed by TTL | Whether every business-expired document is hidden from users |
ttl_expiration_to_deletion_delays |
Observed managed cleanup lag | A guaranteed deadline for future documents |
| Backfill checkpoint | Migration progress | TTL sweep completion |
| Delete-trigger success | Handler processed events | Exactly-once side effects unless handler is idempotent |
| Cost trend | Spend changing during rollout | Correctness |
6. Delete-trigger side effects: assume retries
TTL deletion invokes Firestore delete-trigger functions. Event-driven handlers must tolerate duplicate/retried delivery and must not interpret every physical delete as a user-initiated “delete account” command. AtlasMart records a deletion-event marker keyed by stable event/document identity before emitting any non-idempotent downstream action.
async function handleDelete(event) { const marker = db.doc(`ttlEffects/${event.id}`); await db.runTransaction(async tx => { if ((await tx.get(marker)).exists) return; tx.create(marker, {path:event.params.documentPath, processedAt:new Date()}); // Write only idempotent Firestore state here. }); // External effect should also have its own idempotency key = event.id.}
7. Cost model: backfill + eventual managed deletes
A backfill creates ordinary document writes plus any index/storage effects. Later TTL physical removal is billed as TTL/managed deletes and is not covered by the free usage tier. If functions run on deletion, they add their own invocation/compute/network costs. Estimate the candidate count before rollout and set a budget/operation cap for the optional managed exercise.
8. Controlled failure: backfill everything in one pass
Remove batch limits and checkpoints, then simulate a crash after 60%. A naive script restarts from the beginning, rewriting already-updated documents and losing evidence about what remains. Repair with a stable traversal order, checkpoint, idempotent “write only if missing/different” behavior, canary stage, and a stop switch.
Production judgment
For a large existing collection, backfill is a data migration. Separate policy activation, canary, field backfill, and TTL observation. Rate-limit based on production evidence, not a universal number. Keep protected classes out of the backfill, preserve rollback data, and do not disable the source of recovery until Chapter 22's backup/PITR plan is proven.
Verification checklist and cleanup
- Candidate count and cost estimate exist before writes.
- Dry run proves legal-hold exclusions.
- Canary is bounded and reversible.
- Checkpoint/restart is idempotent.
- TTL Monitoring metrics are observed separately from product expiry.
- Delete-trigger effects use idempotency keys.
- Cleanup removes only lab fixtures, not production evidence.
Bridge to Lesson 5
Lesson 5 turns these mechanisms into a complete retention control plane that covers multiple data classes, ownership, holds, verification, and handoff to disaster-recovery design.
Knowledge check
- Why separate TTL policy creation from backfill?
- What makes a backfill restart-safe?
- What do TTL delay metrics prove?
- Why must TTL delete-trigger handlers be idempotent?
- Are TTL deletes covered by Firestore free usage?
Review the answers
1. It creates a pause point to verify policy status before making historical documents eligible at scale.
2. Stable ordering/checkpoints plus idempotent writes that skip already-correct documents.
3. Observed managed deletion lag, not an exact deletion SLA for a particular document.
4. Event delivery can retry; duplicate external side effects are otherwise possible.
5. No. TTL/managed deletes require billing and are excluded from free usage.
Summary and next step
This lesson established the working contract for Backfilling TTL Fields Safely, Monitoring Deletion, Event Side Effects, and Cost. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Design Retention for Sessions, Events, User Data, Audit Records, and Legal Holds Without Accidental Loss.
Authoritative references
- Firebase · Manage data retention with TTL policies
- Firebase · Firestore index overview and index exemptions
- Firebase · Cloud Firestore pricing and TTL billing
- Firebase · Firestore quotas and limits
- Firebase · Cloud Firestore triggers
- Firebase · Enterprise TTL indexes
- Google Cloud · MongoDB compatibility TTL indexes
- Google Cloud · MongoDB compatibility release notes
- Google Cloud · Firestore Monitoring metrics