Chapter 11 · Subcollections, Collection Groups, Hierarchies, Deletion, and Data Lifecycle
Deleting Documents Does Not Automatically Delete Subcollections: Recursive Delete and Orphan Risk
Prove non-cascading document deletes, detect orphaned descendants, and implement guarded recursive cleanup with verification.
Learning outcomes
Prove with a deterministic fixture that deleting a document does not delete documents in its subcollections.
Implement a scoped recursive inventory/delete routine that is intentionally non-atomic and therefore resumable/verifiable.
Distinguish shallow document delete, recursive server-side cleanup and managed production bulk-delete behavior.
Use project/database/path safety guards and post-delete scans to reduce catastrophic deletion risk.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart continues the same mandatory environment used in
Chapters 01–10: project ID
demo-atlasmart-firestore, Standard edition /
Native mode / (default) database, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0 with
@google-cloud/firestore 9.1.0, and Node.js 22+.
Mandatory deletion/lifecycle work remains local and isolated.
Managed bulk delete, managed export/import, PITR, production
IAM, Cloud Storage and legal/compliance procedures are
discussed but are not falsely claimed to have run in the
emulator.
The Emulator Suite is appropriate for proving path structure, document-delete versus surviving-subcollection behavior, collection-group query results, Security Rules behavior, recursive-cleanup logic and deterministic post-delete verification. It does not establish production delete throughput, billed read/delete cost, managed bulk-delete progress, export consistency, IAM behavior, backup/PITR recovery, legal-retention compliance or p95/p99 latency. Recursive deletion is multi-operation work rather than one atomic cascade; every production run needs a scoped target, project/database guard, durable audit evidence and post-delete verification.
1. The AtlasMart incident: the order disappeared, the messages did not
An operator deletes
tenants/t-acme/orders/o-1001 and assumes its chat
thread is gone. Firestore document semantics do not cascade:
tenants/t-acme/orders/o-1001/messages/m-001 can
still exist even when the parent order document no longer
exists. That surviving child is commonly called an
orphan in application operations, although
Firestore still considers the path valid.
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, FieldValue, Timestamp } from "firebase-admin/firestore";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();
async function seedHierarchy() { const batch = db.batch(); batch.set(db.doc("tenants/t-acme"), { name: "Acme Seller", lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/users/u-1001"), { tenantId: "t-acme", uid: "u-1001", displayName: "Ava", lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1001"), { tenantId: "t-acme", customerId: "u-1001", status: "PAID", totalMinor: 12990, lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1002"), { tenantId: "t-acme", customerId: "u-1002", status: "PAID", totalMinor: 4950, lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1001/messages/m-001"), { tenantId: "t-acme", orderId: "o-1001", authorId: "u-1001", text: "Please leave at reception", sentAt: Timestamp.fromMillis(1789570800000), schemaVersion: 4 }); batch.set(db.doc("tenants/t-acme/orders/o-1002/messages/m-002"), { tenantId: "t-acme", orderId: "o-1002", authorId: "u-1002", text: "Thanks", sentAt: Timestamp.fromMillis(1789570860000), schemaVersion: 4 }); batch.set(db.doc("userDirectory/u-1001"), { uid: "u-1001", tenantId: "t-acme", lifecycleState: "ACTIVE", schemaVersion: 4 }); batch.set(db.doc("activity/a-001"), { subjectUid: "u-1001", tenantId: "t-acme", type: "ORDER_CREATED", schemaVersion: 4 }); await batch.commit();}
await seedHierarchy();const orderRef = db.doc("tenants/t-acme/orders/o-1001");const msgRef = db.doc("tenants/t-acme/orders/o-1001/messages/m-001");console.log("before", (await orderRef.get()).exists, (await msgRef.get()).exists);await orderRef.delete();console.log("after ", (await orderRef.get()).exists, (await msgRef.get()).exists);// Expected: before true true; after false trueif ((await orderRef.get()).exists) throw new Error("parent should be deleted");if (!(await msgRef.get()).exists) throw new Error("child should survive shallow delete");
2. Orphans remain queryable
const messages = await db.collectionGroup("messages").where("tenantId", "==", "t-acme").get();for (const m of messages.docs) { const orderRef = m.ref.parent.parent; // .../orders/{orderId} const parent = await orderRef.get(); console.log({ message: m.ref.path, parent: orderRef.path, parentExists: parent.exists });}
The message under o-1001 still appears in a
collection-group query. This matters for privacy, billing,
exports, support tools and deletion obligations: “the parent is
gone from the UI” is not evidence that descendant data is gone.
3. Build a dry-run inventory before deletion
async function inventoryDocument(ref, out = []) { const snap = await ref.get(); out.push({ path: ref.path, exists: snap.exists }); for (const coll of await ref.listCollections()) { const children = await coll.get(); for (const child of children.docs) await inventoryDocument(child.ref, out); } return out;}async function inventoryCollection(ref, out = []) { const snap = await ref.get(); for (const doc of snap.docs) await inventoryDocument(doc.ref, out); return out;}
const target = db.doc("tenants/t-acme");const inventory = await inventoryDocument(target);console.table(inventory);console.log({ target: target.path, count: inventory.length, mode: "DRY_RUN" });
The dry run is your deletion plan. Store its count and target in a durable deletion-audit record before executing destructive work. For dynamic subcollections, use trusted server APIs that can enumerate child collections. Mobile/web SDKs are not the place for unbounded recursive cleanup.
4. A deterministic recursive cleanup for the emulator
async function deleteDocumentTree(ref, evidence = []) { for (const coll of await ref.listCollections()) { const snap = await coll.get(); for (const child of snap.docs) await deleteDocumentTree(child.ref, evidence); } const exists = (await ref.get()).exists; if (exists) { await ref.delete(); evidence.push(ref.path); } return evidence;}async function guardedDeleteTenant(tenantId) { const project = process.env.GCLOUD_PROJECT; if (project !== "demo-atlasmart-firestore") throw new Error(`REFUSE project=${project}`); if (!/^t-[a-z0-9-]+$/.test(tenantId)) throw new Error("REFUSE malformed tenant id"); const target = db.doc(`tenants/${tenantId}`); const before = await inventoryDocument(target); if (before.length === 0) return { before: 0, deleted: 0 }; const deleted = await deleteDocumentTree(target); const after = await inventoryDocument(target); if (after.some(x => x.exists)) throw new Error("POST_DELETE_VERIFICATION_FAILED"); return { before: before.length, deleted: deleted.length, paths: deleted };}
This lab implementation is deliberately explicit so learners can inspect the mechanism. Current server libraries also provide recursive-deletion helpers. The Firebase CLI/managed services can delete trees or collection groups in production, but recursive deletion remains a multi-operation process that can fail partway through. Therefore safe tooling must be resumable and followed by verification.
Running recursive delete against the wrong project, database or root path can be catastrophic. Repair it with environment allowlists, exact target display, dry-run inventory, change ticket/correlation ID, bounded permissions, operator confirmation outside automation where appropriate, and a post-delete scan.
5. Partial failure is a state, not a reason to guess
Recursive deletion is not atomic. A network/process failure can leave some descendants removed and others present. The repair strategy is not “restore whatever looks missing” by intuition; it is to rerun an idempotent scoped deletion job from durable evidence and verify the target is empty. If deletion must coexist with active writers, first transition the tenant/resource into a deletion state that rejects or redirects new writes, otherwise cleanup can race with recreation.
| Failure | Observable evidence | Safe response |
|---|---|---|
| Parent removed, children remain | Orphan scan finds messages with missing parent | Run scoped recursive cleanup; verify group/path queries |
| Delete job stops halfway | Audit says RUNNING/FAILED with deleted count | Resume same scoped job; do not widen target |
| New writes race with purge | Post-delete scan finds newer updateTime | Freeze writes via lifecycle/rules/server gate, then rerun |
| Wrong project/database selected | Guard mismatch before delete | Abort before destructive call |
| Retention hold discovered | Hold manifest references target data | Exclude held records and document exception; do not hard-delete |
6. Production services are different from the lab
Managed bulk delete can remove one or more collection groups and is billing-enabled. It finds matching documents and deletes in batches; it does not automatically delete child collection groups unless they are also specified. Managed recursive/CLI deletion incurs reads/deletes and is not a single atomic operation. For a tenant-specific subtree, a trusted scoped recursive path is usually easier to reason about than a global collection-group bulk delete, but the right operational tool depends on volume, schema and retention constraints.
7. Reproducible AtlasMart lab
{ "name": "atlasmart-firestore-ch11", "private": true, "type": "module", "engines": { "node": ">=22" }, "dependencies": { "firebase": "12.19.0", "firebase-admin": "14.4.0" }, "devDependencies": { "firebase-tools": "15.30.0", "@firebase/rules-unit-testing": "5.0.2" }}
{ "firestore": { "rules": "firestore.rules", "indexes": "firestore.indexes.json" }, "emulators": { "firestore": { "port": 8080 }, "auth": { "port": 9099 }, "ui": { "enabled": true, "port": 4000 } }}
rules_version = '2';service cloud.firestore { match /databases/{database}/documents { function signedIn() { return request.auth != null; } function sameTenant(tenantId) { return signedIn() && request.auth.token.tenantId == tenantId; } match /tenants/{tenantId} { allow read: if sameTenant(tenantId); allow write: if false; match /users/{uid} { allow read: if sameTenant(tenantId); allow write: if sameTenant(tenantId) && request.auth.uid == uid; } match /orders/{orderId} { allow read: if sameTenant(tenantId); allow write: if false; match /messages/{messageId} { allow read: if sameTenant(tenantId); allow write: if false; } } } // Required for collection-group queries over every messages collection. match /{path=**}/messages/{messageId} { allow read: if signedIn() && resource.data.tenantId == request.auth.token.tenantId; allow write: if false; } match /deletionAudits/{id} { allow read, write: if false; } match /retentionHolds/{id} { allow read, write: if false; } match /{document=**} { allow read, write: if false; } }}
{ "indexes": [ { "collectionGroup": "messages", "queryScope": "COLLECTION_GROUP", "fields": [ { "fieldPath": "tenantId", "order": "ASCENDING" }, { "fieldPath": "sentAt", "order": "DESCENDING" } ] } ], "fieldOverrides": []}
mkdir atlasmart-firestore-ch11 && cd atlasmart-firestore-ch11npm init -ynpm install firebase@12.19.0 firebase-admin@14.4.0npm install --save-dev firebase-tools@15.30.0 @firebase/rules-unit-testing@5.0.2# Save firebase.json, firestore.rules and firestore.indexes.json from this lesson.npx firebase-tools@15.30.0 emulators:start --project demo-atlasmart-firestore --only firestore,auth
Run four assertions in order: seed; shallow-delete
o-1001; prove m-001 survives and
remains visible in a messages collection-group
scan; restore the fixture and recursively delete
tenants/t-acme. Then verify the tenant tree is gone
and prove userDirectory/u-1001 and
activity/a-001 still exist, because they are
outside the tree. Those leftovers motivate the manifest-based
lifecycle work in Lessons 4–5.
const tenant = await db.doc("tenants/t-acme").get();const tenantMsgs = await db.collectionGroup("messages").where("tenantId","==","t-acme").get();const directory = await db.doc("userDirectory/u-1001").get();const activity = await db.doc("activity/a-001").get();console.log({ tenant:tenant.exists, tenantMessageCount:tenantMsgs.size, directory:directory.exists, activity:activity.exists });if (tenant.exists || tenantMsgs.size !== 0) throw new Error("tenant tree not fully removed");if (!directory.exists || !activity.exists) throw new Error("external-copy fixture unexpectedly removed");
Production judgment
Never equate document deletion with object lifecycle completion. The production unit is “all paths and copies governed by the lifecycle policy, plus verified exceptions.” Recursive deletion needs operational safeguards, non-atomic retry semantics, observability and a freeze/new-write strategy. Lesson 4 turns those mechanics into an explicit lifecycle state machine with archive, soft-delete and retention decisions.
Knowledge check
- What remains after deleting an order document with messages beneath it?
- Why perform a dry run?
- Is recursive delete atomic?
- Why can external copies survive a correct subtree purge?
- What should happen if new writes continue during purge?
Review the answers
1. The message subcollection documents can remain and stay queryable.
2. To prove the exact target/path inventory and create expected-count evidence before destructive work.
3. No. It can partially complete, so the job must be resumable and verified.
4. They live outside the target path and require their own deletion-manifest entries.
5. Freeze or gate writes using lifecycle/rules/server controls, then rerun and verify cleanup.
Summary and next step
Parent deletion is shallow; reliable lifecycle cleanup is recursive, guarded, non-atomic and verified. Next we decide when data should not be immediately hard-deleted at all.
Authoritative references
- Cloud Firestore data model — subcollections, hierarchy depth and the non-cascading document-delete warning.
- Collection group queries — querying same-ID subcollections across parents.
- Securely query data — rules are not filters and rules_version 2 collection-group patterns.
- Manage indexes — collection-group/composite index configuration.
- Delete data — collection deletion guidance and server-side recursive deletion.
- Delete collections and subcollections — recursive deletion, orphan discovery and non-atomic behavior.
- Managed bulk delete — production collection-group deletion, billing and operation behavior.
- Export and import data — billing, Cloud Storage, collection-group exports and recovery implications.
- Delete User Data extension — configured paths, recursive mode and auto-discovery considerations.
- Firebase CLI release notes — pinned CLI baseline.
- Admin Node.js release notes — pinned trusted-server baseline.