Chapter 20 · Migrating MongoDB Workloads to Firestore MongoDB Compatibility
Bulk Migration with Export / Import or Datastream→Cloud Storage→Dataflow Patterns and Change Capture
Rehearse AtlasMart bulk loading plus change capture and understand the documented Datastream→Cloud Storage→Dataflow migration path, checkpoints, dead letters, and billing boundaries.
1. AtlasMart problem: snapshot copy plus live writes creates a moving target
After compatibility testing, AtlasMart must load existing data while production continues to change. A one-time dump can be correct at its own instant and already stale before it finishes. The documented migration architecture handles this by combining a bulk/backfill path with change capture, then freezing source writes long enough to drain the tail.
- Distinguish Firestore managed export/import from MongoDB-source migration tooling.
- Explain the documented Datastream → Cloud Storage → Dataflow → Firestore flow.
- Rehearse snapshot plus ordered/idempotent CDC locally with durable checkpoints.
- Preserve sidelined/dead-letter records for repair and replay.
- Bound optional cloud migration by IAM, billing, location, monitoring and cleanup controls.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project identity
demo-atlasmart-firestore. The mandatory lab is
local/no-cost and uses Node.js 22.23.2, the
MongoDB Node driver 6.21.0 from Chapter 19, and
deterministic Extended-JSON-like fixtures plus an append-only
change log. No real Firestore Enterprise MongoDB-compatible
database, Datastream stream, Dataflow job, Cloud Storage
bucket, service-account key, or billable migration resource is
required. The optional managed path uses an isolated
Enterprise MongoDB-compatible database, an explicitly selected
region, Google Cloud CLI 585.0.0, IAM/ADC or
SCRAM as appropriate, and a hard operation/budget/cleanup
plan. The local harness is a semantic rehearsal: it does not
prove managed service throughput, billing, network latency,
Datastream/Dataflow behavior, IAM, or production cutover
timing.
2. Edition, mode, client, security, and billing boundary
This chapter targets Firestore Enterprise edition with MongoDB compatibility. That is not Firestore Standard Native mode and it is not Enterprise Native mode. Enterprise Native exposes Core and Pipeline operations; MongoDB compatibility exposes a MongoDB-compatible endpoint and MQL/BSON surface. Do not move a Standard/Core/Pipeline assumption into this migration unless the current target documentation explicitly supports it.
| Boundary | Chapter 20 assumption |
|---|---|
| Application client | MongoDB-compatible server driver/tool path; Chapter 19 pins Node driver 6.21.0. |
| End-user auth | Firebase Authentication/App Check are not treated as the MongoDB driver authorization layer; backend application authorization remains explicit. |
| Database identity | Use current MongoDB-compatible authentication/IAM/SCRAM/OIDC guidance; never ship service-account credentials to browsers/mobile clients. |
| Security Rules | Do not assume mobile/web Firestore Security Rules authorize MongoDB-compatible driver operations. |
| Local execution | Deterministic migration harness only; it simulates contracts and CDC evidence, not a managed MongoDB-compatible service. |
| Managed execution | Enterprise target plus Datastream/Dataflow/Cloud Storage are optional and potentially billable; record region, IAM principal, budget and cleanup. |
| Performance/cost | Local operation counts and latency are not production Read/Write Unit, network, Datastream or Dataflow evidence. |
| Rollback | Source remains authoritative until the cutover state machine says otherwise; source retirement is a separate approved step. |
3. Two different meanings of “export/import”
Firestore with MongoDB compatibility has a managed export/import service for Firestore database data and recovery/offline processing. That service is not a generic MongoDB-source ingestion format. For migration from a MongoDB-compatible source, current Google guidance uses Datastream to Cloud Storage and a Dataflow pipeline to the Firestore MongoDB-compatible destination. Do not substitute one for the other because both happen to use the words “export” and “import.”
| Mechanism | Source | Destination | Use |
|---|---|---|---|
| Managed Firestore export/import | Firestore MongoDB-compatible database | Firestore MongoDB-compatible database / offline processing | backup-like copy, recovery, offline use; billing required |
mongoexport/mongodump |
MongoDB-compatible source | files | bounded tool-based loads/tests; verify compatibility |
| Datastream + Cloud Storage + Dataflow | MongoDB-compatible source | Firestore MongoDB-compatible database | documented minimal-downtime migration |
4. Documented managed pipeline
- Create the MongoDB source Datastream connection profile.
- Create the Cloud Storage destination/bucket and connection profile.
- Start a Datastream stream that captures data at rest plus change events.
- Start the Dataflow migration pipeline reading the bucket and writing the Firestore MongoDB-compatible database.
- Monitor errors, sidelined records, throughput/backlog, and data freshness.
- Freeze source writes, drain remaining changes, move reads, then enable target writes.
The current migration guide requires a source that supports MongoDB Change Streams and therefore a replica set or sharded cluster. It documents MongoDB 4.0+ with minimum patch requirements for older major/minor lines. Validate the current source-version matrix before execution; do not infer it from the application driver version.
5. Local no-cost rehearsal: snapshot + append-only change log
The mandatory lab models exactly the mechanism that can be
proven without cloud resources.
source-fixture.json is the initial snapshot;
changes.json is an ordered source change log;
target.json is the destination adapter;
checkpoint.json records the last committed event.
const changes = [
{seq:1, op:"insert", ns:"orders", id:"o-9003", doc:{_id:"o-9003",tenantId:"tenant-a",total:89,status:"OPEN"}},
{seq:2, op:"update", ns:"products", id:"p-1001", patch:{$set:{stock:7}}},
{seq:3, op:"update", ns:"orders", id:"o-9003", patch:{$set:{status:"PAID"}}},
{seq:4, op:"delete", ns:"orders", id:"o-9002"}
];
// The checkpoint is durable migration state, not an in-memory loop index.
for (const event of changes) {
if (event.seq <= checkpoint.lastAppliedSeq) continue;
await applyIdempotently(event);
checkpoint.lastAppliedSeq = event.seq;
await saveCheckpoint(checkpoint);
}
6. Why the checkpoint must advance after the durable target write
If AtlasMart saves the checkpoint first and crashes before the target mutation, the event is skipped forever. If it writes the target first and crashes before checkpoint persistence, the event can be replayed. Therefore every change application must be idempotent or otherwise detect duplicates.
async function applyAndCheckpoint(event) {
if (await alreadyApplied(event.seq)) return;
await target.transaction(async tx => {
await applyEvent(tx, event);
await tx.upsert("_migrationEvents", String(event.seq), {
seq: event.seq,
appliedAt: new Date().toISOString()
});
});
await saveCheckpoint({lastAppliedSeq:event.seq});
}
The local JSON adapter can emulate this with an atomic temp-file rename. A real database implementation must use the supported transaction/idempotency semantics of the target instead of pretending filesystem atomicity is equivalent.
7. Sidelined records are evidence, not garbage
The official migration guide warns that unsupported data types,
invalid _id values, or oversized documents can fail
writes and be sidelined. AtlasMart’s local migrator emits a
dead-letter record containing the source namespace, canonical
source ID, reason code, transform version, source hash, event
sequence, and retry count—without copying unnecessary sensitive
fields.
{
"namespace":"products",
"sourceId":"legacy-77",
"reason":"UNSUPPORTED_BSON_CODE",
"sourceHash":"sha256:...",
"transformVersion":null,
"lastEventSeq":218,
"attempts":1,
"status":"NEEDS_OWNER_DECISION"
}
8. Initial load reconciliation before live replay
After the snapshot load, compare collection counts and canonical sampled hashes before declaring the backfill complete. Then replay change events from the defined capture boundary. A count mismatch is diagnostic; a count match is necessary but not sufficient.
const report = {
sourceCount: source.length,
targetCount: target.length,
missingIds: sourceIds.filter(id => !targetIds.has(id)),
extraIds: [...targetIds].filter(id => !sourceIds.has(id)),
sampleHashMismatches: compareSamples(source, target, 50)
};
console.log(JSON.stringify(report, null, 2));
if (report.missingIds.length || report.sampleHashMismatches.length) process.exitCode = 1;
9. Optional managed path: configuration, IAM, billing, and location
The cloud migration is not a Firebase Emulator Suite exercise. It uses billed Google Cloud services and production IAM. Use an isolated project/database, explicit source and destination regions, a dedicated migration principal, budgets/alerts, bounded worker settings, and a documented deletion/stop plan. The current guide calls out Datastream, Dataflow, Cloud Storage, and Datastore/Firestore IAM roles as part of the migration setup.
# Use the current official migration guide to fill exact resource arguments.
gcloud datastream connection-profiles create mongodb SOURCE_PROFILE ...
gcloud datastream connection-profiles create cloud-storage DEST_PROFILE ...
gcloud datastream streams create STREAM ...
# Start the documented Firestore MongoDB compatibility Dataflow template.
# Monitor Datastream data freshness + Dataflow backlog/sidelined outputs.
Prefer ADC/workload identity/service accounts with least privilege. Connection secrets belong in approved secret-management paths, not shell history or repository files.
10. Controlled failure: duplicate the third change event
Append event sequence 3 twice and crash the local process after applying it but before writing the checkpoint. Restart. The repaired runner must converge to one business effect because event application is idempotent and the migration-event marker is durable.
The local design tolerates one replay pattern. It does not prove Datastream ordering, Dataflow retries, managed target transactions, or production throughput. Those need bounded cloud evidence and service metrics.
11. Performance and cost evidence
Record source read rate, destination write rate, backlog, data freshness/lag, error/sidelined count, bytes transferred, p50/p95/p99 apply latency, and cloud cost dimensions. The local harness reports operation counts and measured process latency only; it must not fabricate managed billable units.
function percentile(sorted, p) {
if (!sorted.length) return null;
return sorted[Math.min(sorted.length-1, Math.ceil(p*sorted.length)-1)];
}
latencies.sort((a,b)=>a-b);
console.log({
count:latencies.length,
p50_ms:percentile(latencies,0.50),
p95_ms:percentile(latencies,0.95),
p99_ms:percentile(latencies,0.99)
});
Production judgment
Choose a data-movement mechanism only after matching it to downtime/RPO, source topology, data volume, transform needs, supported types, cost, and rollback. For minimal downtime from a MongoDB-compatible source, the current documented Datastream/Dataflow pattern is the primary reference. For small controlled imports or recovery, supported database tools or Firestore managed export/import may serve different purposes; validate semantics and limits rather than mixing the workflows.
Verification checklist and cleanup
- Snapshot boundary and CDC start point are recorded.
- Checkpoint advances only after durable/idempotent application.
- Duplicate replay test converges.
- Dead-letter records remain inspectable/replayable.
- Counts and sampled hashes reconcile after backfill and replay.
- Optional cloud project has budget, IAM and stop/delete plan.
- No managed service is represented by fake emulator metrics.
Bridge to Lesson 4
Lesson 4 turns a continuously converging target into a traffic migration. It defines dual-run evidence, source-write freeze, lag gates, endpoint switching, target-write enablement, and a rollback window with explicit authority.
Knowledge check
- Why is Firestore managed export/import not the documented MongoDB-source migration pipeline?
- Why save the CDC checkpoint after the target mutation?
- What should happen to an unsupported target write?
- What does data freshness represent during managed migration?
- Why no production throughput claim from the local harness?
Review the answers
1. It operates on Firestore exports; the MongoDB-source path uses Datastream to Cloud Storage and Dataflow to the target.
2. Saving it first can lose an event on crash; replay after target-first must be handled idempotently.
3. Preserve a dead-letter/sidelined record with reason and repair metadata, then reprocess after an approved fix.
4. How far captured/applied change data trails the source, a key cutover gate.
5. It does not reproduce Datastream, Dataflow, network, Firestore infrastructure, IAM or billing behavior.
Summary and next step
This lesson established the working contract for Bulk Migration with Export/Import or Datastream→Cloud Storage→Dataflow Patterns and Change Capture. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Dual-Run/Cutover Planning, Validation, Lag Monitoring, Read/Write Freeze, and Rollback.
Authoritative references
- Firebase · Migrate to Firestore with MongoDB compatibility
- Google Cloud · MongoDB-compatible migration process
- Google Cloud · Configure migration resources and IAM
- Google Cloud · Datastream import from MongoDB source
- Google Cloud · Dataflow write to Firestore with MongoDB compatibility
- Google Cloud · Traffic migration, freeze and cutover
- Google Cloud · Migration troubleshooting
- Google Cloud · Supported BSON types, drivers and tools
- Google Cloud · Behavior differences from MongoDB
- Google Cloud · Managed export/import for Firestore with MongoDB compatibility
- Google Cloud · Firestore with MongoDB compatibility release notes