Chapter 20 · Migrating MongoDB Workloads to Firestore MongoDB Compatibility

Dual-Run / Cutover Planning, Validation, Lag Monitoring, Read / Write Freeze, and Rollback

Design an evidence-gated AtlasMart dual-run and cutover with source write freeze, lag validation, read switch, target write enablement, and rollback authority.

Advanced · 180–240 minutesMongoDB migration · compatibility · CDC · cutover · rollbackNode 22.23.2 · mongodb 6.21.0 · deterministic local harnessGoogle Cloud CLI 585.0.0 · Enterprise MongoDB-compatible target optional/billableLast reviewed: 17 September 2026

1. AtlasMart problem: a synchronized target is not yet the source of truth

AtlasMart now has an initial load and change replay. The dangerous final step is to change one connection string and hope. Cutover is a state machine with evidence gates: dual-run, validation, write freeze, tail drain, read switch, write switch, observation, and rollback authority.

Learning outcomes
  • Define dual-run and shadow-read patterns without creating uncontrolled dual writes.
  • Use lag/freshness, error, hash and query-regression gates before cutover.
  • Orchestrate the documented source-write freeze and tail drain.
  • Separate read cutover from target-write enablement.
  • Design a rollback window that preserves source data and decision authority.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 20 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps project identity demo-atlasmart-firestore. The mandatory lab is local/no-cost and uses Node.js 22.23.2, the MongoDB Node driver 6.21.0 from Chapter 19, and deterministic Extended-JSON-like fixtures plus an append-only change log. No real Firestore Enterprise MongoDB-compatible database, Datastream stream, Dataflow job, Cloud Storage bucket, service-account key, or billable migration resource is required. The optional managed path uses an isolated Enterprise MongoDB-compatible database, an explicitly selected region, Google Cloud CLI 585.0.0, IAM/ADC or SCRAM as appropriate, and a hard operation/budget/cleanup plan. The local harness is a semantic rehearsal: it does not prove managed service throughput, billing, network latency, Datastream/Dataflow behavior, IAM, or production cutover timing.

2. Edition, mode, client, security, and billing boundary

This chapter targets Firestore Enterprise edition with MongoDB compatibility. That is not Firestore Standard Native mode and it is not Enterprise Native mode. Enterprise Native exposes Core and Pipeline operations; MongoDB compatibility exposes a MongoDB-compatible endpoint and MQL/BSON surface. Do not move a Standard/Core/Pipeline assumption into this migration unless the current target documentation explicitly supports it.

Boundary Chapter 20 assumption
Application client MongoDB-compatible server driver/tool path; Chapter 19 pins Node driver 6.21.0.
End-user auth Firebase Authentication/App Check are not treated as the MongoDB driver authorization layer; backend application authorization remains explicit.
Database identity Use current MongoDB-compatible authentication/IAM/SCRAM/OIDC guidance; never ship service-account credentials to browsers/mobile clients.
Security Rules Do not assume mobile/web Firestore Security Rules authorize MongoDB-compatible driver operations.
Local execution Deterministic migration harness only; it simulates contracts and CDC evidence, not a managed MongoDB-compatible service.
Managed execution Enterprise target plus Datastream/Dataflow/Cloud Storage are optional and potentially billable; record region, IAM principal, budget and cleanup.
Performance/cost Local operation counts and latency are not production Read/Write Unit, network, Datastream or Dataflow evidence.
Rollback Source remains authoritative until the cutover state machine says otherwise; source retirement is a separate approved step.

3. Cutover state machine

State Reads Writes Required evidence to advance
SOURCE_ONLY source source baseline complete
MIGRATING source + shadow target source backfill/CDC healthy; blocker count zero
FREEZE_SOURCE_WRITES source/target verification blocked at source backlog drains; data freshness minimal
TARGET_READS target still blocked query/hash checks pass across all services
TARGET_PRIMARY target target write smoke + business invariants pass
OBSERVE target target SLO/error/cost/security stable for rollback window
ROLLBACK source source after explicit reconciliation incident commander approves

4. Dual-run does not automatically mean dual-write

Uncoordinated dual writes create two primaries and divergent conflict resolution. A safer migration usually keeps the source as the write authority while replication feeds the target, then performs shadow reads or sampled query comparisons against the target. If business requirements require application-level dual writes, define an ordering, idempotency key, compensation policy, and reconciliation mechanism; do not improvise it at cutover.

5. Endpoint router with a reversible switch

router.mjs•••
export function databaseFor(request, config) {
  switch (config.mode) {
    case "SOURCE_ONLY":
    case "MIGRATING":
      return sourceDb;
    case "TARGET_READS":
      if (request.method !== "GET") throw new Error("WRITE_FREEZE");
      return targetDb;
    case "TARGET_PRIMARY":
      return targetDb;
    default:
      throw new Error(`unknown migration mode ${config.mode}`);
  }
}

The switch must be centrally auditable and versioned. A feature flag hidden in multiple services is not a safe cutover mechanism unless every service proves it observed the same epoch.

6. Lag and validation gates

For the documented managed path, Datastream data freshness and Dataflow backlog/transactional-write activity are primary migration signals. AtlasMart adds application evidence: source/target counts, sampled canonical hashes, golden-query regression, dead-letter count, cross-tenant authorization checks, and latency/error metrics.

gate.json•••
{
  "blockers":0,
  "deadLettersUnresolved":0,
  "countMismatch":0,
  "sampleHashMismatch":0,
  "goldenQueryFailures":0,
  "cdcLagSeconds":"MEASURED",
  "dataFreshness":"MEASURED",
  "sourceWritesFrozen":false,
  "allServicesOnTargetReads":false
}

Do not hard-code a universal lag threshold. The acceptable window derives from AtlasMart’s RPO, write volume, and business invariants. The gate stores the measured value and the approved threshold separately.

7. Freeze source writes before the final drain

The current migration guide explicitly requires stopping source write traffic so remaining captured changes can reach the destination. Implement the freeze at the application/API layer and verify it with both positive and negative probes. A banner saying “maintenance mode” while background workers still write is not a freeze.

freeze verification•••
await assert.rejects(
  () => sourceOrderService.create({externalOrderId:"freeze-test"}),
  /WRITE_FREEZE/
);
assert.equal(await backgroundWorkerWriteCountSince(freezeEpoch), 0);

8. Read cutover precedes target writes

Once the tail is drained and all validation gates pass, move read traffic to the target. Confirm every service—web API, worker, cron, admin tool, reporting job—uses the target. Only then enable writes on the target. This sequencing avoids a period where different services treat different databases as primary writers.

9. Rollback is a data problem, not only a DNS/config problem

Before target writes begin, rollback can often mean restoring source-only traffic after re-enabling source writes. After target writes begin, rollback requires deciding how those target-only mutations return to the source or are otherwise preserved. Therefore define a rollback window, change-capture direction, write freeze, and authority in advance.

Keep the source

Do not delete or irreversibly mutate the source immediately after cutover. Retain it for the approved rollback window subject to security, retention, and cost policy. Backups help recovery but do not replace a tested reverse/reconciliation procedure.

10. Controlled failure: stale worker during read cutover

Simulate one background worker that still reads the source after the API has moved to target. The gate must remain red because allServicesOnTargetReads=false. This catches a common migration failure where a forgotten worker produces actions from stale source state.

service endpoint evidence•••
[
  {"service":"api","databaseEpoch":"target-2026-09-17","status":"OK"},
  {"service":"worker-payments","databaseEpoch":"source","status":"FAIL"},
  {"service":"worker-email","databaseEpoch":"target-2026-09-17","status":"OK"}
]

11. Failure matrix and authority

Failure Immediate action Data action Decision owner
dead-letter spike stop advancement repair/replay; keep source writes on migration lead
hash/query mismatch hold read cutover classify semantic/data cause application owner
target auth/IAM denial hold cutover fix least privilege; rerun test security owner
target p99/error regression hold or rollback profile query/index/region/cost evidence SRE + app owner
post-target-write incident freeze target writes reconcile target-only writes before rollback incident commander

12. Observe latency and cost without average-only reporting

Capture p50/p95/p99, timeouts, retry counts, target operation/byte billing evidence, and query explain/index behavior on the exact target workload. Compare against the source baseline gathered before migration. Averages can hide long tails that break API deadlines. Do not load-test the production customer project without an explicit bounded plan.

Production judgment

A good cutover is boring because every transition is gated by measured evidence and reversible authority. “Minimal downtime” still includes a deliberate write freeze in the current documented path. The correct freeze duration and rollback window come from measured backlog/data freshness plus AtlasMart’s RPO/RTO and business constraints, not a generic recipe.

Verification checklist and cleanup

  • Single write authority is explicit in every phase.
  • All services report the same database epoch before target writes.
  • Source freeze includes background workers.
  • Lag/freshness and validation gates are captured.
  • Rollback after target writes has a reconciliation plan.
  • Source retention window is approved.
  • Managed migration resources are stopped only after evidence retention requirements are met.

Bridge to Lesson 5

Lesson 5 combines the entire chapter into one repeatable rehearsal: inventory, transform, backfill, CDC replay, validation, latency measurement, failure injection, cutover, rollback, and a machine-readable go/no-go report.

Knowledge check

  1. Why avoid uncontrolled dual writes?
  2. What is the purpose of the source write freeze?
  3. Why move reads before enabling target writes?
  4. What changes about rollback after target writes start?
  5. Why include every worker in the endpoint gate?
Review the answers

1. They create two write authorities and undefined conflict/reconciliation semantics.

2. To stop new source mutations so the final captured tail can drain and the target can converge.

3. It proves every service uses the target while preventing simultaneous primary writers.

4. Target-only mutations must be preserved/reconciled back to the source or otherwise handled.

5. A forgotten source reader can act on stale data even if the main API has cut over.

Summary and next step

This lesson established the working contract for Dual-Run/Cutover Planning, Validation, Lag Monitoring, Read/Write Freeze, and Rollback. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Execute a Migration Rehearsal with Counts, Checksums/Samples, Query Regression, Performance, and Failure Recovery.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.