Chapter 20 · Migrating MongoDB Workloads to Firestore MongoDB Compatibility
Execute a Migration Rehearsal with Counts, Checksums / Samples, Query Regression, Performance, and Failure Recovery
Execute a complete AtlasMart migration rehearsal with counts, hashes, query regression, latency distributions, failure injection, recovery, cutover, rollback, and go/no-go evidence.
1. AtlasMart capstone: prove the migration plan can fail safely before production
The final lesson is a rehearsal, not a slide deck. AtlasMart runs the same synthetic dataset through inventory, compatibility translation, initial load, CDC replay, validation, query regression, latency measurement, a deliberate failure, cutover state transitions, rollback, and cleanup. Every step writes evidence that a reviewer can inspect after the run.
- Execute a complete deterministic migration rehearsal end to end.
- Produce counts, hashes, query-regression, lag, dead-letter and percentile evidence.
- Inject and recover from a migration failure without deleting evidence.
- Exercise reversible read/write cutover and rollback logic.
- Generate a go/no-go report that clearly separates local proof from production-only checks.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project identity
demo-atlasmart-firestore. The mandatory lab is
local/no-cost and uses Node.js 22.23.2, the
MongoDB Node driver 6.21.0 from Chapter 19, and
deterministic Extended-JSON-like fixtures plus an append-only
change log. No real Firestore Enterprise MongoDB-compatible
database, Datastream stream, Dataflow job, Cloud Storage
bucket, service-account key, or billable migration resource is
required. The optional managed path uses an isolated
Enterprise MongoDB-compatible database, an explicitly selected
region, Google Cloud CLI 585.0.0, IAM/ADC or
SCRAM as appropriate, and a hard operation/budget/cleanup
plan. The local harness is a semantic rehearsal: it does not
prove managed service throughput, billing, network latency,
Datastream/Dataflow behavior, IAM, or production cutover
timing.
2. Edition, mode, client, security, and billing boundary
This chapter targets Firestore Enterprise edition with MongoDB compatibility. That is not Firestore Standard Native mode and it is not Enterprise Native mode. Enterprise Native exposes Core and Pipeline operations; MongoDB compatibility exposes a MongoDB-compatible endpoint and MQL/BSON surface. Do not move a Standard/Core/Pipeline assumption into this migration unless the current target documentation explicitly supports it.
| Boundary | Chapter 20 assumption |
|---|---|
| Application client | MongoDB-compatible server driver/tool path; Chapter 19 pins Node driver 6.21.0. |
| End-user auth | Firebase Authentication/App Check are not treated as the MongoDB driver authorization layer; backend application authorization remains explicit. |
| Database identity | Use current MongoDB-compatible authentication/IAM/SCRAM/OIDC guidance; never ship service-account credentials to browsers/mobile clients. |
| Security Rules | Do not assume mobile/web Firestore Security Rules authorize MongoDB-compatible driver operations. |
| Local execution | Deterministic migration harness only; it simulates contracts and CDC evidence, not a managed MongoDB-compatible service. |
| Managed execution | Enterprise target plus Datastream/Dataflow/Cloud Storage are optional and potentially billable; record region, IAM principal, budget and cleanup. |
| Performance/cost | Local operation counts and latency are not production Read/Write Unit, network, Datastream or Dataflow evidence. |
| Rollback | Source remains authoritative until the cutover state machine says otherwise; source retirement is a separate approved step. |
3. Rehearsal directory
chapter20/
source-fixture.json
changes.json
compatibility-matrix.json
inventory.mjs
transform.mjs
bulk-load.mjs
replay.mjs
reconcile.mjs
golden-queries.mjs
cutover.mjs
inject-failure.mjs
evidence/
inventory.json
counts.json
hashes.json
queries.json
cdc.json
dead-letters.ndjson
latency.json
cutover.json
final-report.json
4. Phase 1 — deterministic reset and inventory
Delete only the local rehearsal state, reseed the fixture, then run the Lesson 1 scanner. Expected state: the clean baseline has zero blockers; the explicit failure fixture produces one blocker and prevents advancement.
rm -rf evidence target-state.json checkpoint.json
node seed.mjs
node inventory.mjs
test -f evidence/inventory.json
5. Phase 2 — transform and bulk load
Apply only approved transforms, retaining transform version/provenance. Bulk-load into the local target adapter. Record source/target counts and canonical hash samples.
const source = await loadSource();
const target = await bulkLoad(source.map(transformLegacy));
assert.equal(await target.count(), source.length);
await writeJson("evidence/counts.json", {
source: source.length,
target: await target.count()
});
6. Phase 3 — replay CDC and measure convergence
Replay the deterministic change log with idempotent event markers and a durable checkpoint. Measure local apply latency but label it local-only. After replay, source truth and target truth should reflect the same logical event sequence.
{
"firstSeq":1,
"lastSeq":4,
"lastAppliedSeq":4,
"duplicatesObserved":1,
"duplicatesWithBusinessEffect":0,
"deadLetters":0,
"lagModel":"deterministic sequence distance",
"managedDataFreshness":"NOT_MEASURED_LOCALLY"
}
7. Phase 4 — counts are necessary; hashes and queries test meaning
Reconcile counts, then compare canonical sampled document hashes, then run the golden query suite. The report keeps each layer separate so a matching count cannot mask a value/query mismatch.
assert.equal(report.countMismatch, 0);
assert.equal(report.sampleHashMismatch.length, 0);
assert.equal(queryReport.failures.length, 0);
assert.equal(cdcReport.deadLetters, 0);
8. Phase 5 — performance baseline without invented numbers
Run the same bounded local operation set multiple times and compute p50/p95/p99 from measured samples. Store hardware/runtime context. The report explicitly says these percentiles are harness latency, not Firestore or Dataflow SLO evidence.
{
"environment":"local semantic harness",
"node":"22.23.2",
"samples":200,
"p50_ms":"MEASURED",
"p95_ms":"MEASURED",
"p99_ms":"MEASURED",
"errors":"MEASURED",
"claim":"does not predict managed Firestore/Datastream/Dataflow latency"
}
9. Phase 6 — controlled failure and repair
Choose exactly one failure per run. The default injection
corrupts the target copy of p-1001.stock after bulk
load. Count remains unchanged, but the sampled hash and a golden
inventory query fail. Repair by replaying the authoritative
snapshot/change sequence or the named correction; never edit the
evidence file to green.
const target = await openTarget();
await target.patch("products","p-1001", {stock:999});
console.log("Injected TARGET_VALUE_CORRUPTION for p-1001");
The rehearsal is successful when the gate turns red for the correct reason, blocks cutover, preserves the mismatch evidence, then returns green only after deterministic repair and revalidation.
10. Phase 7 — rehearse cutover and rollback states
const sequence = [
"SOURCE_ONLY",
"MIGRATING",
"FREEZE_SOURCE_WRITES",
"TARGET_READS",
"TARGET_PRIMARY",
"OBSERVE"
];
for (const next of sequence) {
assertTransitionAllowed(state, next, evidence);
state = next;
await recordState(state);
}
// Rollback drill (local): freeze target writes, identify target-only writes,
// reconcile them, then return authority to source.
In the local harness this is routing state. In production, it maps to actual application/service configuration, write barriers, Datastream/Dataflow completion evidence, connection credentials, and target smoke tests.
11. Final go/no-go report
{
"chapter":"20",
"decision":"NO_GO_UNTIL_PRODUCTION_ONLY_CHECKS_PASS",
"local":{
"inventory":"PASS",
"compatibilityMatrix":"PASS",
"bulkLoad":"PASS",
"cdcReplay":"PASS",
"countReconciliation":"PASS",
"sampleHashes":"PASS",
"goldenQueries":"PASS",
"failureDetectionAndRepair":"PASS",
"cutoverStateMachine":"PASS",
"rollbackDrill":"PASS"
},
"productionRequired":[
"exact target driver/tool compatibility",
"IAM/OIDC/SCRAM path",
"Datastream source connection/change capture",
"Dataflow target writes and sidelined output",
"managed data freshness/backlog",
"query/index performance and p95/p99",
"billing/cost evidence",
"real service endpoint cutover"
]
}
12. Optional managed rehearsal acceptance gates
| Gate | Evidence | Failure action |
|---|---|---|
| Compatibility | same pinned driver + matrix against isolated target | stop; refactor or reject |
| Data quality | counts + sampled/full hashes + query regression | hold; repair/replay |
| CDC | freshness/backlog acceptable; sidelined count resolved | hold source as authority |
| Security | least-privilege IAM/auth/application authorization | stop; security owner fixes |
| Performance | measured p50/p95/p99 + error/deadline rate | tune model/index/region or reject |
| Cost | bounded observed usage vs forecast | revise plan/budget |
| Cutover | freeze/drain/read switch/write switch runbook rehearsed | do not schedule production |
| Rollback | source retention + target-only write reconciliation tested | extend rollback design |
13. Backup/recovery and long-term portability
Migration success includes the day after cutover. Define target backup/export policy, recovery tests, source retirement date, schema/index ownership, driver upgrade cadence, release-note review, portability boundaries, and exit-data format. Firestore managed export/import is useful in the target lifecycle, but it is not a substitute for the MongoDB-source migration and rollback design.
14. Cost and observability record
Store the estimate assumptions and observed cloud usage next to the migration evidence: source egress, Datastream, Cloud Storage, Dataflow workers, Firestore operations/storage/indexes, logs, and any dual-run period. Keep alerts for dead letters, lag/data freshness, write errors, latency tails, and cutover epoch mismatches. The exact price changes by region and time; fetch current pricing before each rehearsal rather than embedding a permanent number in the runbook.
Production judgment
The migration is ready to schedule only when local semantic evidence and all production-only gates are green, rollback authority is named, the source retention window is approved, and the business accepts the measured downtime/RPO/cost envelope. If a required behavior remains unsupported or economically unreasonable, remaining on MongoDB or choosing a different target is a valid engineering outcome.
Verification checklist and cleanup
- Rehearsal starts from a deterministic reset.
- Every phase writes machine-readable evidence.
- Injected corruption blocks cutover and is repaired reproducibly.
- Counts, hashes and query regression all pass independently.
- Latency report is labeled local vs managed.
- Production-only gates cannot be auto-promoted by simulation.
- Rollback drill includes target-only write handling.
- Optional cloud resources have explicit stop/delete commands and retained evidence.
15. Bridge to Chapter 21
After migration, data lifecycle becomes an operational requirement. Chapter 21 focuses on TTL, retention, expiration, backfills, and automation—especially the fact that TTL deletion is asynchronous and cannot be treated as an exact scheduler or a substitute for legal-retention design.
Knowledge check
- Why preserve separate count, hash and query gates?
- What does the default failure injection test?
- Why does the final report default to NO_GO for production?
- When can rollback become harder?
- What is a valid outcome if a required feature remains unsupported?
Review the answers
1. Each detects a different failure class; equal counts can coexist with corrupt values or semantic query differences.
2. That value corruption is detected, blocks cutover, preserves evidence and can be repaired deterministically.
3. A local harness cannot prove managed IAM, Datastream/Dataflow behavior, production latency, billing or endpoint cutover.
4. After target writes begin, because target-only mutations must be reconciled before source authority can safely resume.
5. Do not cut over; refactor, change target, or remain on MongoDB based on evidence.
Summary and next step
This lesson established the working contract for Execute a Migration Rehearsal with Counts, Checksums/Samples, Query Regression, Performance, and Failure Recovery. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to TTL Policies and Expiration Fields: Asynchronous Deletion Semantics and Non-Instant Guarantees.
Authoritative references
- Firebase · Migrate to Firestore with MongoDB compatibility
- Google Cloud · MongoDB-compatible migration process
- Google Cloud · Configure migration resources and IAM
- Google Cloud · Datastream import from MongoDB source
- Google Cloud · Dataflow write to Firestore with MongoDB compatibility
- Google Cloud · Traffic migration, freeze and cutover
- Google Cloud · Migration troubleshooting
- Google Cloud · Supported BSON types, drivers and tools
- Google Cloud · Behavior differences from MongoDB
- Google Cloud · Managed export/import for Firestore with MongoDB compatibility
- Google Cloud · Firestore with MongoDB compatibility release notes