Chapter 20 · Migrating MongoDB Workloads to Firestore MongoDB Compatibility

Execute a Migration Rehearsal with Counts, Checksums / Samples, Query Regression, Performance, and Failure Recovery

Execute a complete AtlasMart migration rehearsal with counts, hashes, query regression, latency distributions, failure injection, recovery, cutover, rollback, and go/no-go evidence.

Advanced · 180–240 minutesMongoDB migration · compatibility · CDC · cutover · rollbackNode 22.23.2 · mongodb 6.21.0 · deterministic local harnessGoogle Cloud CLI 585.0.0 · Enterprise MongoDB-compatible target optional/billableLast reviewed: 17 September 2026

1. AtlasMart capstone: prove the migration plan can fail safely before production

The final lesson is a rehearsal, not a slide deck. AtlasMart runs the same synthetic dataset through inventory, compatibility translation, initial load, CDC replay, validation, query regression, latency measurement, a deliberate failure, cutover state transitions, rollback, and cleanup. Every step writes evidence that a reviewer can inspect after the run.

Learning outcomes
  • Execute a complete deterministic migration rehearsal end to end.
  • Produce counts, hashes, query-regression, lag, dead-letter and percentile evidence.
  • Inject and recover from a migration failure without deleting evidence.
  • Exercise reversible read/write cutover and rollback logic.
  • Generate a go/no-go report that clearly separates local proof from production-only checks.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 20 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps project identity demo-atlasmart-firestore. The mandatory lab is local/no-cost and uses Node.js 22.23.2, the MongoDB Node driver 6.21.0 from Chapter 19, and deterministic Extended-JSON-like fixtures plus an append-only change log. No real Firestore Enterprise MongoDB-compatible database, Datastream stream, Dataflow job, Cloud Storage bucket, service-account key, or billable migration resource is required. The optional managed path uses an isolated Enterprise MongoDB-compatible database, an explicitly selected region, Google Cloud CLI 585.0.0, IAM/ADC or SCRAM as appropriate, and a hard operation/budget/cleanup plan. The local harness is a semantic rehearsal: it does not prove managed service throughput, billing, network latency, Datastream/Dataflow behavior, IAM, or production cutover timing.

2. Edition, mode, client, security, and billing boundary

This chapter targets Firestore Enterprise edition with MongoDB compatibility. That is not Firestore Standard Native mode and it is not Enterprise Native mode. Enterprise Native exposes Core and Pipeline operations; MongoDB compatibility exposes a MongoDB-compatible endpoint and MQL/BSON surface. Do not move a Standard/Core/Pipeline assumption into this migration unless the current target documentation explicitly supports it.

Boundary Chapter 20 assumption
Application client MongoDB-compatible server driver/tool path; Chapter 19 pins Node driver 6.21.0.
End-user auth Firebase Authentication/App Check are not treated as the MongoDB driver authorization layer; backend application authorization remains explicit.
Database identity Use current MongoDB-compatible authentication/IAM/SCRAM/OIDC guidance; never ship service-account credentials to browsers/mobile clients.
Security Rules Do not assume mobile/web Firestore Security Rules authorize MongoDB-compatible driver operations.
Local execution Deterministic migration harness only; it simulates contracts and CDC evidence, not a managed MongoDB-compatible service.
Managed execution Enterprise target plus Datastream/Dataflow/Cloud Storage are optional and potentially billable; record region, IAM principal, budget and cleanup.
Performance/cost Local operation counts and latency are not production Read/Write Unit, network, Datastream or Dataflow evidence.
Rollback Source remains authoritative until the cutover state machine says otherwise; source retirement is a separate approved step.

3. Rehearsal directory

directory•••
chapter20/
  source-fixture.json
  changes.json
  compatibility-matrix.json
  inventory.mjs
  transform.mjs
  bulk-load.mjs
  replay.mjs
  reconcile.mjs
  golden-queries.mjs
  cutover.mjs
  inject-failure.mjs
  evidence/
    inventory.json
    counts.json
    hashes.json
    queries.json
    cdc.json
    dead-letters.ndjson
    latency.json
    cutover.json
    final-report.json

4. Phase 1 — deterministic reset and inventory

Delete only the local rehearsal state, reseed the fixture, then run the Lesson 1 scanner. Expected state: the clean baseline has zero blockers; the explicit failure fixture produces one blocker and prevents advancement.

run phase 1•••
rm -rf evidence target-state.json checkpoint.json
node seed.mjs
node inventory.mjs
test -f evidence/inventory.json

5. Phase 2 — transform and bulk load

Apply only approved transforms, retaining transform version/provenance. Bulk-load into the local target adapter. Record source/target counts and canonical hash samples.

bulk-load assertion•••
const source = await loadSource();
const target = await bulkLoad(source.map(transformLegacy));
assert.equal(await target.count(), source.length);
await writeJson("evidence/counts.json", {
  source: source.length,
  target: await target.count()
});

6. Phase 3 — replay CDC and measure convergence

Replay the deterministic change log with idempotent event markers and a durable checkpoint. Measure local apply latency but label it local-only. After replay, source truth and target truth should reflect the same logical event sequence.

CDC evidence•••
{
  "firstSeq":1,
  "lastSeq":4,
  "lastAppliedSeq":4,
  "duplicatesObserved":1,
  "duplicatesWithBusinessEffect":0,
  "deadLetters":0,
  "lagModel":"deterministic sequence distance",
  "managedDataFreshness":"NOT_MEASURED_LOCALLY"
}

7. Phase 4 — counts are necessary; hashes and queries test meaning

Reconcile counts, then compare canonical sampled document hashes, then run the golden query suite. The report keeps each layer separate so a matching count cannot mask a value/query mismatch.

reconciliation gate•••
assert.equal(report.countMismatch, 0);
assert.equal(report.sampleHashMismatch.length, 0);
assert.equal(queryReport.failures.length, 0);
assert.equal(cdcReport.deadLetters, 0);

8. Phase 5 — performance baseline without invented numbers

Run the same bounded local operation set multiple times and compute p50/p95/p99 from measured samples. Store hardware/runtime context. The report explicitly says these percentiles are harness latency, not Firestore or Dataflow SLO evidence.

latency evidence schema•••
{
  "environment":"local semantic harness",
  "node":"22.23.2",
  "samples":200,
  "p50_ms":"MEASURED",
  "p95_ms":"MEASURED",
  "p99_ms":"MEASURED",
  "errors":"MEASURED",
  "claim":"does not predict managed Firestore/Datastream/Dataflow latency"
}

9. Phase 6 — controlled failure and repair

Choose exactly one failure per run. The default injection corrupts the target copy of p-1001.stock after bulk load. Count remains unchanged, but the sampled hash and a golden inventory query fail. Repair by replaying the authoritative snapshot/change sequence or the named correction; never edit the evidence file to green.

inject-failure.mjs•••
const target = await openTarget();
await target.patch("products","p-1001", {stock:999});
console.log("Injected TARGET_VALUE_CORRUPTION for p-1001");
Expected failure

The rehearsal is successful when the gate turns red for the correct reason, blocks cutover, preserves the mismatch evidence, then returns green only after deterministic repair and revalidation.

10. Phase 7 — rehearse cutover and rollback states

cutover.mjs•••
const sequence = [
  "SOURCE_ONLY",
  "MIGRATING",
  "FREEZE_SOURCE_WRITES",
  "TARGET_READS",
  "TARGET_PRIMARY",
  "OBSERVE"
];
for (const next of sequence) {
  assertTransitionAllowed(state, next, evidence);
  state = next;
  await recordState(state);
}

// Rollback drill (local): freeze target writes, identify target-only writes,
// reconcile them, then return authority to source.

In the local harness this is routing state. In production, it maps to actual application/service configuration, write barriers, Datastream/Dataflow completion evidence, connection credentials, and target smoke tests.

11. Final go/no-go report

evidence/final-report.json•••
{
  "chapter":"20",
  "decision":"NO_GO_UNTIL_PRODUCTION_ONLY_CHECKS_PASS",
  "local":{
    "inventory":"PASS",
    "compatibilityMatrix":"PASS",
    "bulkLoad":"PASS",
    "cdcReplay":"PASS",
    "countReconciliation":"PASS",
    "sampleHashes":"PASS",
    "goldenQueries":"PASS",
    "failureDetectionAndRepair":"PASS",
    "cutoverStateMachine":"PASS",
    "rollbackDrill":"PASS"
  },
  "productionRequired":[
    "exact target driver/tool compatibility",
    "IAM/OIDC/SCRAM path",
    "Datastream source connection/change capture",
    "Dataflow target writes and sidelined output",
    "managed data freshness/backlog",
    "query/index performance and p95/p99",
    "billing/cost evidence",
    "real service endpoint cutover"
  ]
}

12. Optional managed rehearsal acceptance gates

Gate Evidence Failure action
Compatibility same pinned driver + matrix against isolated target stop; refactor or reject
Data quality counts + sampled/full hashes + query regression hold; repair/replay
CDC freshness/backlog acceptable; sidelined count resolved hold source as authority
Security least-privilege IAM/auth/application authorization stop; security owner fixes
Performance measured p50/p95/p99 + error/deadline rate tune model/index/region or reject
Cost bounded observed usage vs forecast revise plan/budget
Cutover freeze/drain/read switch/write switch runbook rehearsed do not schedule production
Rollback source retention + target-only write reconciliation tested extend rollback design

13. Backup/recovery and long-term portability

Migration success includes the day after cutover. Define target backup/export policy, recovery tests, source retirement date, schema/index ownership, driver upgrade cadence, release-note review, portability boundaries, and exit-data format. Firestore managed export/import is useful in the target lifecycle, but it is not a substitute for the MongoDB-source migration and rollback design.

14. Cost and observability record

Store the estimate assumptions and observed cloud usage next to the migration evidence: source egress, Datastream, Cloud Storage, Dataflow workers, Firestore operations/storage/indexes, logs, and any dual-run period. Keep alerts for dead letters, lag/data freshness, write errors, latency tails, and cutover epoch mismatches. The exact price changes by region and time; fetch current pricing before each rehearsal rather than embedding a permanent number in the runbook.

Production judgment

The migration is ready to schedule only when local semantic evidence and all production-only gates are green, rollback authority is named, the source retention window is approved, and the business accepts the measured downtime/RPO/cost envelope. If a required behavior remains unsupported or economically unreasonable, remaining on MongoDB or choosing a different target is a valid engineering outcome.

Verification checklist and cleanup

  • Rehearsal starts from a deterministic reset.
  • Every phase writes machine-readable evidence.
  • Injected corruption blocks cutover and is repaired reproducibly.
  • Counts, hashes and query regression all pass independently.
  • Latency report is labeled local vs managed.
  • Production-only gates cannot be auto-promoted by simulation.
  • Rollback drill includes target-only write handling.
  • Optional cloud resources have explicit stop/delete commands and retained evidence.

15. Bridge to Chapter 21

After migration, data lifecycle becomes an operational requirement. Chapter 21 focuses on TTL, retention, expiration, backfills, and automation—especially the fact that TTL deletion is asynchronous and cannot be treated as an exact scheduler or a substitute for legal-retention design.

Knowledge check

  1. Why preserve separate count, hash and query gates?
  2. What does the default failure injection test?
  3. Why does the final report default to NO_GO for production?
  4. When can rollback become harder?
  5. What is a valid outcome if a required feature remains unsupported?
Review the answers

1. Each detects a different failure class; equal counts can coexist with corrupt values or semantic query differences.

2. That value corruption is detected, blocks cutover, preserves evidence and can be repaired deterministically.

3. A local harness cannot prove managed IAM, Datastream/Dataflow behavior, production latency, billing or endpoint cutover.

4. After target writes begin, because target-only mutations must be reconciled before source authority can safely resume.

5. Do not cut over; refactor, change target, or remain on MongoDB based on evidence.

Summary and next step

This lesson established the working contract for Execute a Migration Rehearsal with Counts, Checksums/Samples, Query Regression, Performance, and Failure Recovery. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to TTL Policies and Expiration Fields: Asynchronous Deletion Semantics and Non-Instant Guarantees.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.