Chapter 11 · Subcollections, Collection Groups, Hierarchies, Deletion, and Data Lifecycle
Implement a Tenant / Data-Subject Deletion Plan that Finds and Verifies All Hierarchical Data
Implement a manifest-driven tenant/data-subject deletion workflow that finds hierarchical and duplicated data, applies exceptions, and verifies completion.
Learning outcomes
Build a deletion manifest that enumerates hierarchical data, root-level duplicates and subject-indexed copies before execution.
Separate tenant deletion from data-subject deletion and document retention exceptions without silently widening the destructive target.
Implement prepare → freeze → delete → verify → audit phases with idempotent reruns and correlation IDs.
Demonstrate post-delete verification across direct paths and collection-group/root queries rather than checking only one parent document.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart continues the same mandatory environment used in
Chapters 01–10: project ID
demo-atlasmart-firestore, Standard edition /
Native mode / (default) database, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0 with
@google-cloud/firestore 9.1.0, and Node.js 22+.
Mandatory deletion/lifecycle work remains local and isolated.
Managed bulk delete, managed export/import, PITR, production
IAM, Cloud Storage and legal/compliance procedures are
discussed but are not falsely claimed to have run in the
emulator.
The Emulator Suite is appropriate for proving path structure, document-delete versus surviving-subcollection behavior, collection-group query results, Security Rules behavior, recursive-cleanup logic and deterministic post-delete verification. It does not establish production delete throughput, billed read/delete cost, managed bulk-delete progress, export consistency, IAM behavior, backup/PITR recovery, legal-retention compliance or p95/p99 latency. Recursive deletion is multi-operation work rather than one atomic cascade; every production run needs a scoped target, project/database guard, durable audit evidence and post-delete verification.
1. The AtlasMart problem: one person exists in many shapes
User u-1001 appears as a tenant user document, a
root directory lookup, an order customer ID, a message author ID
and an activity subject. A tenant purge and a data-subject purge
overlap but are not identical. Tenant deletion may remove the
entire tenants/t-acme tree; subject deletion must
find copies by
uid/customerId/authorId/subjectUid
across multiple collections. Some financial records may be
retained under a documented exception and de-identified or
access-restricted according to organizational policy.
| Copy | Discovery key | Default action in lab | Why it cannot be forgotten |
|---|---|---|---|
| tenants/t-acme/users/u-1001 | exact path | delete | Primary tenant profile |
| userDirectory/u-1001 | exact path | delete | Root-level lookup copy outside tenant tree |
| tenants/t-acme/orders/* | customerId == u-1001 | policy decision | Orders may have retention obligations |
| **/messages/* | authorId == u-1001 | delete or redact by policy | Collection-group copy under many parents |
| activity/* | subjectUid == u-1001 | delete or pseudonymize by policy | Root audit/activity copy |
2. Phase 1 — prepare a manifest; do not delete while discovering
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, FieldValue, Timestamp } from "firebase-admin/firestore";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();
async function seedHierarchy() { const batch = db.batch(); batch.set(db.doc("tenants/t-acme"), { name: "Acme Seller", lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/users/u-1001"), { tenantId: "t-acme", uid: "u-1001", displayName: "Ava", lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1001"), { tenantId: "t-acme", customerId: "u-1001", status: "PAID", totalMinor: 12990, lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1002"), { tenantId: "t-acme", customerId: "u-1002", status: "PAID", totalMinor: 4950, lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1001/messages/m-001"), { tenantId: "t-acme", orderId: "o-1001", authorId: "u-1001", text: "Please leave at reception", sentAt: Timestamp.fromMillis(1789570800000), schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1002/messages/m-002"), { tenantId: "t-acme", orderId: "o-1002", authorId: "u-1002", text: "Thanks", sentAt: Timestamp.fromMillis(1789570860000), schemaVersion: 4 }); batch.set(db.doc("userDirectory/u-1001"), { uid: "u-1001", tenantId: "t-acme", lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("activity/a-001"), { subjectUid: "u-1001", tenantId: "t-acme", type: "ORDER_CREATED", schemaVersion: 4 }); await batch.commit();}
async function discoverSubject(uid, tenantId) { const exact = [ db.doc(`tenants/${tenantId}/users/${uid}`), db.doc(`userDirectory/${uid}`) ]; const orders = await db.collection(`tenants/${tenantId}/orders`).where("customerId","==",uid).get(); const messages = await db.collectionGroup("messages").where("authorId","==",uid).get(); const activity = await db.collection("activity").where("subjectUid","==",uid).get(); return { uid, tenantId, exact: exact.map(r => r.path), orders: orders.docs.map(d => d.ref.path).sort(), messages: messages.docs.map(d => d.ref.path).sort(), activity: activity.docs.map(d => d.ref.path).sort() };}const manifest = await discoverSubject("u-1001", "t-acme");console.log(JSON.stringify(manifest, null, 2));
Discovery and deletion are deliberately separate phases. That allows policy/hold checks, review and expected-count recording before anything destructive happens. The manifest should use stable path IDs and policy references, not copy sensitive field contents into the audit system unnecessarily.
3. Phase 2 — freeze new writes
A deletion job can never converge if the application keeps
recreating the subject’s data. AtlasMart first marks the
account/tenant lifecycle as PURGE_PENDING in
trusted state and rejects new writes that create governed data.
The exact enforcement may be Security Rules for client paths,
server authorization gates, or both.
async function assertSubjectWritable(uid) { const s = await db.doc(`subjectLifecycle/${uid}`).get(); if (s.exists && ["PURGE_PENDING","PURGED"].includes(s.get("state"))) { throw new Error("SUBJECT_NOT_WRITABLE"); }}
Do not rely only on UI disablement. Trusted server jobs and background processors must honor the same lifecycle state.
4. Phase 3 — apply retention decisions, then delete
{ "deletionId": "del-u-1001-20260916", "subjectUid": "u-1001", "tenantId": "t-acme", "state": "APPROVED", "deletePaths": [ "tenants/t-acme/users/u-1001", "userDirectory/u-1001", "tenants/t-acme/orders/o-1001/messages/m-001", "activity/a-001" ], "retain": [ { "path": "tenants/t-acme/orders/o-1001", "reasonCode": "FINANCE_RETENTION", "action": "RESTRICT_AND_MINIMIZE_IDENTITY" } ], "policyVersion": "atlasmart-retention-v3", "schemaVersion": 4}
The retained order is an intentional exception, not a failed deletion. What “minimize identity” means is organization- and jurisdiction-specific; the lab records the concept without prescribing legal compliance.
async function deleteExactPaths(paths) { const evidence=[]; for (const path of paths) { const ref=db.doc(path); const before=await ref.get(); if (before.exists) await ref.delete(); const after=await ref.get(); evidence.push({path, existedBefore:before.exists, existsAfter:after.exists}); } return evidence;}
5. Phase 4 — verify by the same discovery surfaces
async function verifySubject(uid, tenantId, retainedPaths = new Set()) { const found = await discoverSubject(uid, tenantId); const unexpected = [ ...found.exact, ...found.orders, ...found.messages, ...found.activity ].filter(path => !retainedPaths.has(path)); return { found, unexpected };}const retained = new Set(["tenants/t-acme/orders/o-1001"]);const result = await verifySubject("u-1001", "t-acme", retained);console.log(result);if (result.unexpected.length) throw new Error(`DELETE_INCOMPLETE: ${result.unexpected.join(",")}`);
Verification must search, not merely check the delete responses. If a message copy was added beneath another order or a new root collection was introduced without updating the deletion registry, the scan should expose that gap.
6. Phase 5 — durable audit without re-storing the deleted subject
{ "deletionId": "del-u-1001-20260916", "subjectKeyHash": "sha256:<training-placeholder>", "tenantId": "t-acme", "state": "VERIFIED_WITH_EXCEPTIONS", "deletedCount": 4, "retainedCount": 1, "retentionReasonCodes": ["FINANCE_RETENTION"], "policyVersion": "atlasmart-retention-v3", "verificationVersion": "subject-discovery-v2", "correlationId": "corr-delete-011", "completedAt": "server timestamp"}
An audit record should prove the process without copying the personal data back into a permanent log. Access to deletion audits is itself sensitive and must be governed separately.
7. Wrong approach: delete the user document and declare success
Deleting only tenants/t-acme/users/u-1001 leaves
the directory copy, message authorship, activity records and
order references. Deleting the whole tenant tree may also
violate retention rules or remove unrelated users. Running
auto-discovery without reviewing new collection types can be
equally risky. The repair is a versioned data-location registry
plus manifest preparation, policy review, scoped execution and
repeatable verification.
Every new collection or duplicated identity field must answer: who owns it, how is it queried, how is it retained, how is a tenant deleted, how is a data subject found, and how is deletion verified? Treat deletion-registry updates as part of schema review.
8. Tenant deletion vs subject deletion
A tenant deletion can use a recursive tree operation for tenant-owned paths plus a manifest for external copies. A subject deletion is more query-oriented: it finds records by subject identifiers across root and collection-group surfaces, then applies policy-specific delete/redact/retain actions. Neither should be delegated directly to an untrusted client. In production, the trusted service must have narrowly scoped IAM and strong operator/job authentication.
9. Reproducible end-to-end lab
{ "name": "atlasmart-firestore-ch11", "private": true, "type": "module", "engines": { "node": ">=22" }, "dependencies": { "firebase": "12.19.0", "firebase-admin": "14.4.0" }, "devDependencies": { "firebase-tools": "15.30.0", "@firebase/rules-unit-testing": "5.0.2" }}
{ "firestore": { "rules": "firestore.rules", "indexes": "firestore.indexes.json" }, "emulators": { "firestore": { "port": 8080 }, "auth": { "port": 9099 }, "ui": { "enabled": true, "port": 4000 } }}
rules_version = '2';service cloud.firestore { match /databases/{database}/documents { function signedIn() { return request.auth != null; } function sameTenant(tenantId) { return signedIn() && request.auth.token.tenantId == tenantId; } match /tenants/{tenantId} { allow read: if sameTenant(tenantId); allow write: if false; match /users/{uid} { allow read: if sameTenant(tenantId); allow write: if sameTenant(tenantId) && request.auth.uid == uid; } match /orders/{orderId} { allow read: if sameTenant(tenantId); allow write: if false; match /messages/{messageId} { allow read: if sameTenant(tenantId); allow write: if false; } } } // Required for collection-group queries over every messages collection. match /{path=**}/messages/{messageId} { allow read: if signedIn() && resource.data.tenantId == request.auth.token.tenantId; allow write: if false; } match /deletionAudits/{id} { allow read, write: if false; } match /retentionHolds/{id} { allow read, write: if false; } match /{document=**} { allow read, write: if false; } }}
{ "indexes": [ { "collectionGroup": "messages", "queryScope": "COLLECTION_GROUP", "fields": [ { "fieldPath": "tenantId", "order": "ASCENDING" }, { "fieldPath": "sentAt", "order": "DESCENDING" } ] } ], "fieldOverrides": []}
mkdir atlasmart-firestore-ch11 && cd atlasmart-firestore-ch11npm init -ynpm install firebase@12.19.0 firebase-admin@14.4.0npm install --save-dev firebase-tools@15.30.0 @firebase/rules-unit-testing@5.0.2# Save firebase.json, firestore.rules and firestore.indexes.json from this lesson.npx firebase-tools@15.30.0 emulators:start --project demo-atlasmart-firestore --only firestore,auth
Seed the fixture. Generate the subject manifest and save the
expected counts. Mark o-1001 as a retention
exception. Delete the user profile, directory copy, message and
activity copy. Run verification and assert the only
subject-linked document left is the documented retained order.
Then reset the emulator so the next chapter starts from a clean
deterministic state.
[PASS] lifecycle write gate set to PURGE_PENDING[PASS] discovery manifest captured before delete[PASS] exact delete paths removed[PASS] collection-group message scan returns zero unexpected u-1001 documents[PASS] activity scan returns zero unexpected u-1001 documents[PASS] retained order exists with documented reason code[PASS] deletion audit contains counts/policy/correlation ID, not copied subject fields[PASS] emulator reset completed
Production judgment
A deletion plan is successful only when it can prove both absence and intentional exceptions. Build deletion discoverability into the schema, freeze new governed writes, separate policy from execution, use trusted scoped tooling, verify through every discovery surface and retain minimal audit evidence. Managed exports/backups/PITR can complicate retention and erasure policies and must be governed separately; they are not silently “covered” by deleting live documents.
Chapter 12 moves from lifecycle ownership to client authorization: Security Rules foundations, where path structure and lifecycle fields become part of the proof that a request is allowed.
Knowledge check
- Why separate discovery from deletion?
- What prevents a purge from racing with recreated data?
- How do you verify deletion?
- Why is a retained document not automatically a deletion failure?
- What must schema review include for new duplicated identity data?
Review the answers
1. It lets you review scope, holds and expected counts before destructive work begins.
2. A lifecycle write freeze enforced by trusted server gates and/or client Security Rules.
3. Rerun exact-path, root-query and collection-group discovery surfaces and compare results with documented retention exceptions.
4. A policy-reviewed retention exception is an intentional outcome that should be explicit and audited.
5. Ownership, query access, retention, tenant/subject discovery, deletion action and verification logic.
Summary
Chapter 11 connects hierarchy to lifecycle. AtlasMart can now query nested data across parents, prove shallow-delete orphan behavior, recursively clean scoped trees, represent archive/retention states and execute a manifest-driven tenant/data-subject deletion with post-delete verification. Chapter 12 uses those same paths and fields to build least-privilege Security Rules.
Authoritative references
- Cloud Firestore data model — subcollections, hierarchy depth and the non-cascading document-delete warning.
- Collection group queries — querying same-ID subcollections across parents.
- Securely query data — rules are not filters and rules_version 2 collection-group patterns.
- Manage indexes — collection-group/composite index configuration.
- Delete data — collection deletion guidance and server-side recursive deletion.
- Delete collections and subcollections — recursive deletion, orphan discovery and non-atomic behavior.
- Managed bulk delete — production collection-group deletion, billing and operation behavior.
- Export and import data — billing, Cloud Storage, collection-group exports and recovery implications.
- Delete User Data extension — configured paths, recursive mode and auto-discovery considerations.
- Firebase CLI release notes — pinned CLI baseline.
- Admin Node.js release notes — pinned trusted-server baseline.