Chapter 20 · Migrating MongoDB Workloads to Firestore MongoDB Compatibility
Inventory MongoDB Features, Drivers, Data Types, Indexes, Aggregations, Transactions, and Operational Dependencies
Inventory every MongoDB workload dependency before migration: drivers, BSON, IDs, indexes, queries, transactions, ODM behavior, security, change capture, and operations.
1. AtlasMart problem: moving bytes is easier than migrating behavior
AtlasMart already has a MongoDB service with product, order,
audit, and operational collections. A team can copy documents
successfully and still fail the migration because its
application depends on unsupported BSON values, a driver major
outside the supported matrix, an ODM hook, a transaction shape,
a change-stream assumption, an index option, or a runbook that
presumes access to mongod/mongos. The
first deliverable is therefore not a copy job. It is an
evidence-backed dependency inventory.
- Build a feature/driver/ODM/data/index/query/transaction/operations inventory before moving data.
- Classify every dependency as supported, refactor-required, test-required, or blocking.
- Detect BSON, _id, document-size and nesting risks before a bulk migration sidelines records.
- Separate MongoDB application semantics from cluster administration assumptions.
- Produce a migration evidence bundle that can drive later dual-run and rollback decisions.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project identity
demo-atlasmart-firestore. The mandatory lab is
local/no-cost and uses Node.js 22.23.2, the
MongoDB Node driver 6.21.0 from Chapter 19, and
deterministic Extended-JSON-like fixtures plus an append-only
change log. No real Firestore Enterprise MongoDB-compatible
database, Datastream stream, Dataflow job, Cloud Storage
bucket, service-account key, or billable migration resource is
required. The optional managed path uses an isolated
Enterprise MongoDB-compatible database, an explicitly selected
region, Google Cloud CLI 585.0.0, IAM/ADC or
SCRAM as appropriate, and a hard operation/budget/cleanup
plan. The local harness is a semantic rehearsal: it does not
prove managed service throughput, billing, network latency,
Datastream/Dataflow behavior, IAM, or production cutover
timing.
2. Edition, mode, client, security, and billing boundary
This chapter targets Firestore Enterprise edition with MongoDB compatibility. That is not Firestore Standard Native mode and it is not Enterprise Native mode. Enterprise Native exposes Core and Pipeline operations; MongoDB compatibility exposes a MongoDB-compatible endpoint and MQL/BSON surface. Do not move a Standard/Core/Pipeline assumption into this migration unless the current target documentation explicitly supports it.
| Boundary | Chapter 20 assumption |
|---|---|
| Application client | MongoDB-compatible server driver/tool path; Chapter 19 pins Node driver 6.21.0. |
| End-user auth | Firebase Authentication/App Check are not treated as the MongoDB driver authorization layer; backend application authorization remains explicit. |
| Database identity | Use current MongoDB-compatible authentication/IAM/SCRAM/OIDC guidance; never ship service-account credentials to browsers/mobile clients. |
| Security Rules | Do not assume mobile/web Firestore Security Rules authorize MongoDB-compatible driver operations. |
| Local execution | Deterministic migration harness only; it simulates contracts and CDC evidence, not a managed MongoDB-compatible service. |
| Managed execution | Enterprise target plus Datastream/Dataflow/Cloud Storage are optional and potentially billable; record region, IAM principal, budget and cleanup. |
| Performance/cost | Local operation counts and latency are not production Read/Write Unit, network, Datastream or Dataflow evidence. |
| Rollback | Source remains authoritative until the cutover state machine says otherwise; source retirement is a separate approved step. |
3. Start with five inventories, not one collection count
Inventorying only collections and document counts answers “how much data?” but not “what must still work?” AtlasMart captures five linked inventories. A single row can appear in more than one because a failure can cross layers.
| Inventory | Examples | Evidence | Migration decision |
|---|---|---|---|
| Application surface | Node driver, Mongoose, mongosh, background workers | package lock, runtime probe, command log | pin / replace / refactor |
| Data surface |
BSON types, _id, document size, nesting
|
type histogram, max-size scan, sample hashes | supported / transform / block |
| Query surface | find/update/aggregation/operators/sorts | golden query suite + expected IDs | pass / rewrite / reject |
| Consistency surface | transactions, retryable writes, change streams, idempotency | transaction tests + event semantics | preserve / redesign |
| Operations surface | backup, restore, roles, monitoring, topology/runbooks | runbook dependency map | translate to Google Cloud / delete assumption |
4. Current target constraints that belong in the inventory
As of this review, Firestore with MongoDB compatibility is an
Enterprise edition mode. The current Node driver support matrix
lists 5.x and 6.x; Chapter 19 pins mongodb@6.21.0.
The target supports many BSON types, but not DBPointer,
JavaScript/JavaScript-with-scope, Symbol, Undefined, and some
other legacy values. The top-level _id accepts only
a documented subset of BSON types and is limited to 1,500 bytes.
The target document limit is 16 MiB with a nesting-depth limit
of 20. These are migration gates, not trivia.
If a source record uses an unsupported _id type,
changing it can break references, idempotency keys, URLs,
caches, and external systems. Record an explicit mapping such
as legacyId → targetId, prove all referential
uses, and keep rollback translation.
5. Reproducible local inventory fixture
Create chapter20/source-fixture.json as a canonical
source snapshot. This is intentionally small enough to inspect,
yet includes tenant and schema fields needed by later query
regression.
[
{"_id":"p-1001","tenantId":"tenant-a","name":"Trail Camera","price":99,"stock":8,"tags":["camera","outdoor"],"schemaVersion":3},
{"_id":"p-1002","tenantId":"tenant-a","name":"USB-C Hub","price":49,"stock":3,"tags":["usb","desk"],"schemaVersion":3},
{"_id":"o-9001","tenantId":"tenant-a","kind":"order","customerId":"u-alice","total":148,"status":"PAID","schemaVersion":4},
{"_id":"o-9002","tenantId":"tenant-b","kind":"order","customerId":"u-bob","total":39,"status":"OPEN","schemaVersion":4}
]
Add two incompatible fixtures in the scanner rather than in JSON
itself: one document whose legacy field is a BSON JavaScript
Code value and one whose top-level
_id is a BSON Date. The scanner must classify both
as blockers until the mapping policy is approved.
6. Inventory the source with executable probes
The production version should read a bounded sample and aggregate schema/type statistics server-side where safe. The local version uses the same decision contract over fixtures.
import { readFile } from "node:fs/promises";
import { BSON, Code } from "mongodb";
const docs = JSON.parse(await readFile("source-fixture.json","utf8"));
const supportedIdTypes = new Set(["string","number","boolean","object"]);
function byteSize(doc) {
return BSON.serialize(doc).byteLength;
}
function inspect(doc) {
const errors = [];
if (byteSize(doc) > 16 * 1024 * 1024) errors.push("DOC_GT_16_MIB");
if (!supportedIdTypes.has(typeof doc._id)) errors.push("UNSUPPORTED_ID_TYPE");
for (const [k,v] of Object.entries(doc)) {
if (v instanceof Code) errors.push(`UNSUPPORTED_BSON_CODE:${k}`);
}
return {id:String(doc._id), bytes:byteSize(doc), errors};
}
const report = docs.map(inspect);
console.table(report);
if (report.some(r => r.errors.length)) process.exitCode = 2;
A real source scan should additionally report collection counts, largest documents, nesting depth, type frequencies by path, duplicate business keys, null/missing distinctions, array cardinality, and any field names prohibited by the target. Never dump secrets or raw sensitive values into inventory logs; report paths, types, counts, sizes, and sampled hashes.
7. Index inventory: capture intent, not just syntax
Index definitions must be translated against the target’s current index semantics. Record every source index with the query contract it serves, uniqueness/sparsity/partial behavior, sort direction, multikey use, text/geospatial/vector role, and operational dependency. An index with no known query owner is migration debt; a query with no index owner is performance risk.
| Source index | Business query | Target action |
|---|---|---|
{tenantId:1,status:1,createdAt:-1} |
recent orders per tenant/status | recreate only after target explain/query test |
{externalOrderId:1}, unique |
idempotent ingest | verify target uniqueness semantics; do not assume option parity |
| text index | catalog search | revalidate current Preview/GA status and query semantics |
| TTL index | session cleanup | translate separately; Chapter 21 covers lifecycle semantics |
8. Transactions, retries, change capture, and ODM behavior
Record every multi-document transaction, read/write concern
expectation, retry policy, idempotency key, change-stream
consumer, Mongoose plugin/hook, schema default, cast,
discriminator, middleware side effect, and raw command. A
successful CRUD smoke test says nothing about these
dependencies. In particular, Firestore MongoDB compatibility
requires retryWrites=false; application/workflow
idempotency must carry any semantics that previously relied on
retryable-write behavior.
node --version
npm ls mongodb mongoose --depth=0
# Record output in migration-evidence/runtime.txt.
# Also record the production driver options actually used.
9. Operational dependency map
Serverless removes some cluster duties but does not erase operational work. Replace source topology, replica-set, shard-balancing, node sizing, profiler, backup, restore, security-role, network, and monitoring assumptions with explicit target services and owners. “Managed” is not the same as “nothing to operate.”
| Source assumption | Target question |
|---|---|
We can inspect/configure mongod |
Is there an equivalent managed control, or must the runbook disappear? |
| Replica-set lag drives failover runbook | Which Firestore availability/region indicators and SLA apply? |
| MongoDB users/roles authorize DB access | Which IAM/SCRAM/OIDC principal and application authorization replace it? |
| Snapshot tool is our rollback | Which source-retention + target backup/export strategy supports rollback? |
10. Controlled failure: a migration plan that only counts documents
Deliberately mark every collection “ready” after matching source
counts. Then run the inventory scanner with the unsupported BSON
Code and Date _id cases enabled. The plan must
fail. Matching counts on the clean subset do not prove the
rejected documents are migratable, and they do not prove
queries, transactions, indexes, auth, or tools.
Change the readiness gate from countEqual to a
conjunction: compatibility blockers = 0; golden-query suite
passes; required indexes mapped; auth path tested;
transaction/idempotency semantics accepted; cutover/rollback
owner assigned. Keep the failure evidence rather than deleting
it.
11. Lab output: migration-inventory.json
{
"source":{"engine":"MongoDB-compatible","api":"8.0-shaped workload","driver":"mongodb 6.21.0"},
"target":{"edition":"Enterprise","mode":"MongoDB compatibility","managedExecution":"optional"},
"collections":["products","orders"],
"blockers":[
{"code":"UNSUPPORTED_BSON_CODE","policy":"convert to audited string only after owner approval"},
{"code":"UNSUPPORTED_ID_TYPE","policy":"explicit legacyId->targetId mapping + reference rewrite"}
],
"requiredEvidence":[
"counts","canonical-hash-samples","golden-query-regression","index-map",
"cdc-checkpoint","dead-letter-count","latency-baseline","cutover-checklist","rollback-checklist"
]
}
Expected state: a migration is not “ready” while any blocker lacks an owner and tested remediation. The file is configuration/evidence, not a claim that the real target has been exercised.
Production judgment
Proceed to compatibility testing only when the inventory can answer: which behavior must be identical, which may change, which requires application refactoring, which production-only properties remain untested, what downtime/RPO is acceptable, and how long the source remains available for rollback. Preserve the inventory as a versioned artifact because every later target release, driver upgrade, or application feature can invalidate it.
Verification checklist and cleanup
- Driver/ODM/tool versions recorded.
- BSON/
_id/size/nesting scan recorded. - Indexes linked to query contracts.
- Transactions, retries, change consumers and side effects inventoried.
- IAM/auth/application authorization dependencies listed.
- No production secrets copied into the lab.
- Local fixtures can be deleted and recreated deterministically.
Bridge to Lesson 2
Lesson 2 converts this inventory into an executable compatibility matrix. One incompatible behavior is deliberately refactored and then locked down with regression tests before any bulk-load step begins.
Knowledge check
- Why is a document count insufficient migration evidence?
-
Why must incompatible
_idvalues be handled explicitly? - What should happen to an unsupported record during planning?
- Why inventory ODM hooks and middleware?
- What is the role of the local harness?
Review the answers
1. It says nothing about unsupported values, IDs, query semantics, transactions, indexes, authorization, tools, or operational assumptions.
2. Changing identity can break references and external contracts; the mapping must be auditable and reversible.
3. It becomes a blocker/dead-letter candidate with an owner and remediation policy, not a silently coerced success.
4. They can encode defaults, validation and side effects that driver-level CRUD tests do not reveal.
5. To rehearse semantics and evidence generation without billing; it does not prove managed migration performance or IAM.
Summary and next step
This lesson established the working contract for Inventory MongoDB Features, Drivers, Data Types, Indexes, Aggregations, Transactions, and Operational Dependencies. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Compatibility Matrix Testing, Unsupported Features, Semantic Differences, and Application Refactoring.
Authoritative references
- Firebase · Migrate to Firestore with MongoDB compatibility
- Google Cloud · MongoDB-compatible migration process
- Google Cloud · Configure migration resources and IAM
- Google Cloud · Datastream import from MongoDB source
- Google Cloud · Dataflow write to Firestore with MongoDB compatibility
- Google Cloud · Traffic migration, freeze and cutover
- Google Cloud · Migration troubleshooting
- Google Cloud · Supported BSON types, drivers and tools
- Google Cloud · Behavior differences from MongoDB
- Google Cloud · Managed export/import for Firestore with MongoDB compatibility
- Google Cloud · Firestore with MongoDB compatibility release notes