Chapter 20 · Migrating MongoDB Workloads to Firestore MongoDB Compatibility

Bulk Migration with Export / Import or Datastream→Cloud Storage→Dataflow Patterns and Change Capture

Rehearse AtlasMart bulk loading plus change capture and understand the documented Datastream→Cloud Storage→Dataflow migration path, checkpoints, dead letters, and billing boundaries.

Advanced · 180–240 minutesMongoDB migration · compatibility · CDC · cutover · rollbackNode 22.23.2 · mongodb 6.21.0 · deterministic local harnessGoogle Cloud CLI 585.0.0 · Enterprise MongoDB-compatible target optional/billableLast reviewed: 17 September 2026

1. AtlasMart problem: snapshot copy plus live writes creates a moving target

After compatibility testing, AtlasMart must load existing data while production continues to change. A one-time dump can be correct at its own instant and already stale before it finishes. The documented migration architecture handles this by combining a bulk/backfill path with change capture, then freezing source writes long enough to drain the tail.

Learning outcomes
  • Distinguish Firestore managed export/import from MongoDB-source migration tooling.
  • Explain the documented Datastream → Cloud Storage → Dataflow → Firestore flow.
  • Rehearse snapshot plus ordered/idempotent CDC locally with durable checkpoints.
  • Preserve sidelined/dead-letter records for repair and replay.
  • Bound optional cloud migration by IAM, billing, location, monitoring and cleanup controls.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 20 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps project identity demo-atlasmart-firestore. The mandatory lab is local/no-cost and uses Node.js 22.23.2, the MongoDB Node driver 6.21.0 from Chapter 19, and deterministic Extended-JSON-like fixtures plus an append-only change log. No real Firestore Enterprise MongoDB-compatible database, Datastream stream, Dataflow job, Cloud Storage bucket, service-account key, or billable migration resource is required. The optional managed path uses an isolated Enterprise MongoDB-compatible database, an explicitly selected region, Google Cloud CLI 585.0.0, IAM/ADC or SCRAM as appropriate, and a hard operation/budget/cleanup plan. The local harness is a semantic rehearsal: it does not prove managed service throughput, billing, network latency, Datastream/Dataflow behavior, IAM, or production cutover timing.

2. Edition, mode, client, security, and billing boundary

This chapter targets Firestore Enterprise edition with MongoDB compatibility. That is not Firestore Standard Native mode and it is not Enterprise Native mode. Enterprise Native exposes Core and Pipeline operations; MongoDB compatibility exposes a MongoDB-compatible endpoint and MQL/BSON surface. Do not move a Standard/Core/Pipeline assumption into this migration unless the current target documentation explicitly supports it.

Boundary Chapter 20 assumption
Application client MongoDB-compatible server driver/tool path; Chapter 19 pins Node driver 6.21.0.
End-user auth Firebase Authentication/App Check are not treated as the MongoDB driver authorization layer; backend application authorization remains explicit.
Database identity Use current MongoDB-compatible authentication/IAM/SCRAM/OIDC guidance; never ship service-account credentials to browsers/mobile clients.
Security Rules Do not assume mobile/web Firestore Security Rules authorize MongoDB-compatible driver operations.
Local execution Deterministic migration harness only; it simulates contracts and CDC evidence, not a managed MongoDB-compatible service.
Managed execution Enterprise target plus Datastream/Dataflow/Cloud Storage are optional and potentially billable; record region, IAM principal, budget and cleanup.
Performance/cost Local operation counts and latency are not production Read/Write Unit, network, Datastream or Dataflow evidence.
Rollback Source remains authoritative until the cutover state machine says otherwise; source retirement is a separate approved step.

3. Two different meanings of “export/import”

Firestore with MongoDB compatibility has a managed export/import service for Firestore database data and recovery/offline processing. That service is not a generic MongoDB-source ingestion format. For migration from a MongoDB-compatible source, current Google guidance uses Datastream to Cloud Storage and a Dataflow pipeline to the Firestore MongoDB-compatible destination. Do not substitute one for the other because both happen to use the words “export” and “import.”

Mechanism Source Destination Use
Managed Firestore export/import Firestore MongoDB-compatible database Firestore MongoDB-compatible database / offline processing backup-like copy, recovery, offline use; billing required
mongoexport/mongodump MongoDB-compatible source files bounded tool-based loads/tests; verify compatibility
Datastream + Cloud Storage + Dataflow MongoDB-compatible source Firestore MongoDB-compatible database documented minimal-downtime migration

4. Documented managed pipeline

  1. Create the MongoDB source Datastream connection profile.
  2. Create the Cloud Storage destination/bucket and connection profile.
  3. Start a Datastream stream that captures data at rest plus change events.
  4. Start the Dataflow migration pipeline reading the bucket and writing the Firestore MongoDB-compatible database.
  5. Monitor errors, sidelined records, throughput/backlog, and data freshness.
  6. Freeze source writes, drain remaining changes, move reads, then enable target writes.
Source prerequisites

The current migration guide requires a source that supports MongoDB Change Streams and therefore a replica set or sharded cluster. It documents MongoDB 4.0+ with minimum patch requirements for older major/minor lines. Validate the current source-version matrix before execution; do not infer it from the application driver version.

5. Local no-cost rehearsal: snapshot + append-only change log

The mandatory lab models exactly the mechanism that can be proven without cloud resources. source-fixture.json is the initial snapshot; changes.json is an ordered source change log; target.json is the destination adapter; checkpoint.json records the last committed event.

changes.mjs•••
const changes = [
  {seq:1, op:"insert", ns:"orders", id:"o-9003", doc:{_id:"o-9003",tenantId:"tenant-a",total:89,status:"OPEN"}},
  {seq:2, op:"update", ns:"products", id:"p-1001", patch:{$set:{stock:7}}},
  {seq:3, op:"update", ns:"orders", id:"o-9003", patch:{$set:{status:"PAID"}}},
  {seq:4, op:"delete", ns:"orders", id:"o-9002"}
];

// The checkpoint is durable migration state, not an in-memory loop index.
for (const event of changes) {
  if (event.seq <= checkpoint.lastAppliedSeq) continue;
  await applyIdempotently(event);
  checkpoint.lastAppliedSeq = event.seq;
  await saveCheckpoint(checkpoint);
}

6. Why the checkpoint must advance after the durable target write

If AtlasMart saves the checkpoint first and crashes before the target mutation, the event is skipped forever. If it writes the target first and crashes before checkpoint persistence, the event can be replayed. Therefore every change application must be idempotent or otherwise detect duplicates.

safe application order•••
async function applyAndCheckpoint(event) {
  if (await alreadyApplied(event.seq)) return;

  await target.transaction(async tx => {
    await applyEvent(tx, event);
    await tx.upsert("_migrationEvents", String(event.seq), {
      seq: event.seq,
      appliedAt: new Date().toISOString()
    });
  });

  await saveCheckpoint({lastAppliedSeq:event.seq});
}

The local JSON adapter can emulate this with an atomic temp-file rename. A real database implementation must use the supported transaction/idempotency semantics of the target instead of pretending filesystem atomicity is equivalent.

7. Sidelined records are evidence, not garbage

The official migration guide warns that unsupported data types, invalid _id values, or oversized documents can fail writes and be sidelined. AtlasMart’s local migrator emits a dead-letter record containing the source namespace, canonical source ID, reason code, transform version, source hash, event sequence, and retry count—without copying unnecessary sensitive fields.

dead-letter.json•••
{
  "namespace":"products",
  "sourceId":"legacy-77",
  "reason":"UNSUPPORTED_BSON_CODE",
  "sourceHash":"sha256:...",
  "transformVersion":null,
  "lastEventSeq":218,
  "attempts":1,
  "status":"NEEDS_OWNER_DECISION"
}

8. Initial load reconciliation before live replay

After the snapshot load, compare collection counts and canonical sampled hashes before declaring the backfill complete. Then replay change events from the defined capture boundary. A count mismatch is diagnostic; a count match is necessary but not sufficient.

reconcile.mjs•••
const report = {
  sourceCount: source.length,
  targetCount: target.length,
  missingIds: sourceIds.filter(id => !targetIds.has(id)),
  extraIds: [...targetIds].filter(id => !sourceIds.has(id)),
  sampleHashMismatches: compareSamples(source, target, 50)
};
console.log(JSON.stringify(report, null, 2));
if (report.missingIds.length || report.sampleHashMismatches.length) process.exitCode = 1;

9. Optional managed path: configuration, IAM, billing, and location

The cloud migration is not a Firebase Emulator Suite exercise. It uses billed Google Cloud services and production IAM. Use an isolated project/database, explicit source and destination regions, a dedicated migration principal, budgets/alerts, bounded worker settings, and a documented deletion/stop plan. The current guide calls out Datastream, Dataflow, Cloud Storage, and Datastore/Firestore IAM roles as part of the migration setup.

managed-path shape (illustrative placeholders)•••
# Use the current official migration guide to fill exact resource arguments.
gcloud datastream connection-profiles create mongodb SOURCE_PROFILE ...
gcloud datastream connection-profiles create cloud-storage DEST_PROFILE ...
gcloud datastream streams create STREAM ...
# Start the documented Firestore MongoDB compatibility Dataflow template.
# Monitor Datastream data freshness + Dataflow backlog/sidelined outputs.
Do not paste production credentials into lesson commands

Prefer ADC/workload identity/service accounts with least privilege. Connection secrets belong in approved secret-management paths, not shell history or repository files.

10. Controlled failure: duplicate the third change event

Append event sequence 3 twice and crash the local process after applying it but before writing the checkpoint. Restart. The repaired runner must converge to one business effect because event application is idempotent and the migration-event marker is durable.

What this proves

The local design tolerates one replay pattern. It does not prove Datastream ordering, Dataflow retries, managed target transactions, or production throughput. Those need bounded cloud evidence and service metrics.

11. Performance and cost evidence

Record source read rate, destination write rate, backlog, data freshness/lag, error/sidelined count, bytes transferred, p50/p95/p99 apply latency, and cloud cost dimensions. The local harness reports operation counts and measured process latency only; it must not fabricate managed billable units.

latency-summary.mjs•••
function percentile(sorted, p) {
  if (!sorted.length) return null;
  return sorted[Math.min(sorted.length-1, Math.ceil(p*sorted.length)-1)];
}
latencies.sort((a,b)=>a-b);
console.log({
  count:latencies.length,
  p50_ms:percentile(latencies,0.50),
  p95_ms:percentile(latencies,0.95),
  p99_ms:percentile(latencies,0.99)
});

Production judgment

Choose a data-movement mechanism only after matching it to downtime/RPO, source topology, data volume, transform needs, supported types, cost, and rollback. For minimal downtime from a MongoDB-compatible source, the current documented Datastream/Dataflow pattern is the primary reference. For small controlled imports or recovery, supported database tools or Firestore managed export/import may serve different purposes; validate semantics and limits rather than mixing the workflows.

Verification checklist and cleanup

  • Snapshot boundary and CDC start point are recorded.
  • Checkpoint advances only after durable/idempotent application.
  • Duplicate replay test converges.
  • Dead-letter records remain inspectable/replayable.
  • Counts and sampled hashes reconcile after backfill and replay.
  • Optional cloud project has budget, IAM and stop/delete plan.
  • No managed service is represented by fake emulator metrics.

Bridge to Lesson 4

Lesson 4 turns a continuously converging target into a traffic migration. It defines dual-run evidence, source-write freeze, lag gates, endpoint switching, target-write enablement, and a rollback window with explicit authority.

Knowledge check

  1. Why is Firestore managed export/import not the documented MongoDB-source migration pipeline?
  2. Why save the CDC checkpoint after the target mutation?
  3. What should happen to an unsupported target write?
  4. What does data freshness represent during managed migration?
  5. Why no production throughput claim from the local harness?
Review the answers

1. It operates on Firestore exports; the MongoDB-source path uses Datastream to Cloud Storage and Dataflow to the target.

2. Saving it first can lose an event on crash; replay after target-first must be handled idempotently.

3. Preserve a dead-letter/sidelined record with reason and repair metadata, then reprocess after an approved fix.

4. How far captured/applied change data trails the source, a key cutover gate.

5. It does not reproduce Datastream, Dataflow, network, Firestore infrastructure, IAM or billing behavior.

Summary and next step

This lesson established the working contract for Bulk Migration with Export/Import or Datastream→Cloud Storage→Dataflow Patterns and Change Capture. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Dual-Run/Cutover Planning, Validation, Lag Monitoring, Read/Write Freeze, and Rollback.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.