Chapter 20 · Migrating MongoDB Workloads to Firestore MongoDB Compatibility
Dual-Run / Cutover Planning, Validation, Lag Monitoring, Read / Write Freeze, and Rollback
Design an evidence-gated AtlasMart dual-run and cutover with source write freeze, lag validation, read switch, target write enablement, and rollback authority.
1. AtlasMart problem: a synchronized target is not yet the source of truth
AtlasMart now has an initial load and change replay. The dangerous final step is to change one connection string and hope. Cutover is a state machine with evidence gates: dual-run, validation, write freeze, tail drain, read switch, write switch, observation, and rollback authority.
- Define dual-run and shadow-read patterns without creating uncontrolled dual writes.
- Use lag/freshness, error, hash and query-regression gates before cutover.
- Orchestrate the documented source-write freeze and tail drain.
- Separate read cutover from target-write enablement.
- Design a rollback window that preserves source data and decision authority.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project identity
demo-atlasmart-firestore. The mandatory lab is
local/no-cost and uses Node.js 22.23.2, the
MongoDB Node driver 6.21.0 from Chapter 19, and
deterministic Extended-JSON-like fixtures plus an append-only
change log. No real Firestore Enterprise MongoDB-compatible
database, Datastream stream, Dataflow job, Cloud Storage
bucket, service-account key, or billable migration resource is
required. The optional managed path uses an isolated
Enterprise MongoDB-compatible database, an explicitly selected
region, Google Cloud CLI 585.0.0, IAM/ADC or
SCRAM as appropriate, and a hard operation/budget/cleanup
plan. The local harness is a semantic rehearsal: it does not
prove managed service throughput, billing, network latency,
Datastream/Dataflow behavior, IAM, or production cutover
timing.
2. Edition, mode, client, security, and billing boundary
This chapter targets Firestore Enterprise edition with MongoDB compatibility. That is not Firestore Standard Native mode and it is not Enterprise Native mode. Enterprise Native exposes Core and Pipeline operations; MongoDB compatibility exposes a MongoDB-compatible endpoint and MQL/BSON surface. Do not move a Standard/Core/Pipeline assumption into this migration unless the current target documentation explicitly supports it.
| Boundary | Chapter 20 assumption |
|---|---|
| Application client | MongoDB-compatible server driver/tool path; Chapter 19 pins Node driver 6.21.0. |
| End-user auth | Firebase Authentication/App Check are not treated as the MongoDB driver authorization layer; backend application authorization remains explicit. |
| Database identity | Use current MongoDB-compatible authentication/IAM/SCRAM/OIDC guidance; never ship service-account credentials to browsers/mobile clients. |
| Security Rules | Do not assume mobile/web Firestore Security Rules authorize MongoDB-compatible driver operations. |
| Local execution | Deterministic migration harness only; it simulates contracts and CDC evidence, not a managed MongoDB-compatible service. |
| Managed execution | Enterprise target plus Datastream/Dataflow/Cloud Storage are optional and potentially billable; record region, IAM principal, budget and cleanup. |
| Performance/cost | Local operation counts and latency are not production Read/Write Unit, network, Datastream or Dataflow evidence. |
| Rollback | Source remains authoritative until the cutover state machine says otherwise; source retirement is a separate approved step. |
3. Cutover state machine
| State | Reads | Writes | Required evidence to advance |
|---|---|---|---|
| SOURCE_ONLY | source | source | baseline complete |
| MIGRATING | source + shadow target | source | backfill/CDC healthy; blocker count zero |
| FREEZE_SOURCE_WRITES | source/target verification | blocked at source | backlog drains; data freshness minimal |
| TARGET_READS | target | still blocked | query/hash checks pass across all services |
| TARGET_PRIMARY | target | target | write smoke + business invariants pass |
| OBSERVE | target | target | SLO/error/cost/security stable for rollback window |
| ROLLBACK | source | source after explicit reconciliation | incident commander approves |
4. Dual-run does not automatically mean dual-write
Uncoordinated dual writes create two primaries and divergent conflict resolution. A safer migration usually keeps the source as the write authority while replication feeds the target, then performs shadow reads or sampled query comparisons against the target. If business requirements require application-level dual writes, define an ordering, idempotency key, compensation policy, and reconciliation mechanism; do not improvise it at cutover.
5. Endpoint router with a reversible switch
export function databaseFor(request, config) {
switch (config.mode) {
case "SOURCE_ONLY":
case "MIGRATING":
return sourceDb;
case "TARGET_READS":
if (request.method !== "GET") throw new Error("WRITE_FREEZE");
return targetDb;
case "TARGET_PRIMARY":
return targetDb;
default:
throw new Error(`unknown migration mode ${config.mode}`);
}
}
The switch must be centrally auditable and versioned. A feature flag hidden in multiple services is not a safe cutover mechanism unless every service proves it observed the same epoch.
6. Lag and validation gates
For the documented managed path, Datastream data freshness and Dataflow backlog/transactional-write activity are primary migration signals. AtlasMart adds application evidence: source/target counts, sampled canonical hashes, golden-query regression, dead-letter count, cross-tenant authorization checks, and latency/error metrics.
{
"blockers":0,
"deadLettersUnresolved":0,
"countMismatch":0,
"sampleHashMismatch":0,
"goldenQueryFailures":0,
"cdcLagSeconds":"MEASURED",
"dataFreshness":"MEASURED",
"sourceWritesFrozen":false,
"allServicesOnTargetReads":false
}
Do not hard-code a universal lag threshold. The acceptable window derives from AtlasMart’s RPO, write volume, and business invariants. The gate stores the measured value and the approved threshold separately.
7. Freeze source writes before the final drain
The current migration guide explicitly requires stopping source write traffic so remaining captured changes can reach the destination. Implement the freeze at the application/API layer and verify it with both positive and negative probes. A banner saying “maintenance mode” while background workers still write is not a freeze.
await assert.rejects(
() => sourceOrderService.create({externalOrderId:"freeze-test"}),
/WRITE_FREEZE/
);
assert.equal(await backgroundWorkerWriteCountSince(freezeEpoch), 0);
8. Read cutover precedes target writes
Once the tail is drained and all validation gates pass, move read traffic to the target. Confirm every service—web API, worker, cron, admin tool, reporting job—uses the target. Only then enable writes on the target. This sequencing avoids a period where different services treat different databases as primary writers.
9. Rollback is a data problem, not only a DNS/config problem
Before target writes begin, rollback can often mean restoring source-only traffic after re-enabling source writes. After target writes begin, rollback requires deciding how those target-only mutations return to the source or are otherwise preserved. Therefore define a rollback window, change-capture direction, write freeze, and authority in advance.
Do not delete or irreversibly mutate the source immediately after cutover. Retain it for the approved rollback window subject to security, retention, and cost policy. Backups help recovery but do not replace a tested reverse/reconciliation procedure.
10. Controlled failure: stale worker during read cutover
Simulate one background worker that still reads the source after
the API has moved to target. The gate must remain red because
allServicesOnTargetReads=false. This catches a
common migration failure where a forgotten worker produces
actions from stale source state.
[
{"service":"api","databaseEpoch":"target-2026-09-17","status":"OK"},
{"service":"worker-payments","databaseEpoch":"source","status":"FAIL"},
{"service":"worker-email","databaseEpoch":"target-2026-09-17","status":"OK"}
]
11. Failure matrix and authority
| Failure | Immediate action | Data action | Decision owner |
|---|---|---|---|
| dead-letter spike | stop advancement | repair/replay; keep source writes on | migration lead |
| hash/query mismatch | hold read cutover | classify semantic/data cause | application owner |
| target auth/IAM denial | hold cutover | fix least privilege; rerun test | security owner |
| target p99/error regression | hold or rollback | profile query/index/region/cost evidence | SRE + app owner |
| post-target-write incident | freeze target writes | reconcile target-only writes before rollback | incident commander |
12. Observe latency and cost without average-only reporting
Capture p50/p95/p99, timeouts, retry counts, target operation/byte billing evidence, and query explain/index behavior on the exact target workload. Compare against the source baseline gathered before migration. Averages can hide long tails that break API deadlines. Do not load-test the production customer project without an explicit bounded plan.
Production judgment
A good cutover is boring because every transition is gated by measured evidence and reversible authority. “Minimal downtime” still includes a deliberate write freeze in the current documented path. The correct freeze duration and rollback window come from measured backlog/data freshness plus AtlasMart’s RPO/RTO and business constraints, not a generic recipe.
Verification checklist and cleanup
- Single write authority is explicit in every phase.
- All services report the same database epoch before target writes.
- Source freeze includes background workers.
- Lag/freshness and validation gates are captured.
- Rollback after target writes has a reconciliation plan.
- Source retention window is approved.
- Managed migration resources are stopped only after evidence retention requirements are met.
Bridge to Lesson 5
Lesson 5 combines the entire chapter into one repeatable rehearsal: inventory, transform, backfill, CDC replay, validation, latency measurement, failure injection, cutover, rollback, and a machine-readable go/no-go report.
Knowledge check
- Why avoid uncontrolled dual writes?
- What is the purpose of the source write freeze?
- Why move reads before enabling target writes?
- What changes about rollback after target writes start?
- Why include every worker in the endpoint gate?
Review the answers
1. They create two write authorities and undefined conflict/reconciliation semantics.
2. To stop new source mutations so the final captured tail can drain and the target can converge.
3. It proves every service uses the target while preventing simultaneous primary writers.
4. Target-only mutations must be preserved/reconciled back to the source or otherwise handled.
5. A forgotten source reader can act on stale data even if the main API has cut over.
Summary and next step
This lesson established the working contract for Dual-Run/Cutover Planning, Validation, Lag Monitoring, Read/Write Freeze, and Rollback. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Execute a Migration Rehearsal with Counts, Checksums/Samples, Query Regression, Performance, and Failure Recovery.
Authoritative references
- Firebase · Migrate to Firestore with MongoDB compatibility
- Google Cloud · MongoDB-compatible migration process
- Google Cloud · Configure migration resources and IAM
- Google Cloud · Datastream import from MongoDB source
- Google Cloud · Dataflow write to Firestore with MongoDB compatibility
- Google Cloud · Traffic migration, freeze and cutover
- Google Cloud · Migration troubleshooting
- Google Cloud · Supported BSON types, drivers and tools
- Google Cloud · Behavior differences from MongoDB
- Google Cloud · Managed export/import for Firestore with MongoDB compatibility
- Google Cloud · Firestore with MongoDB compatibility release notes