A migration is a temporary distributed system and must be designed for partial failure.
Migrate Safely with Dual Reads/Writes, Backfills, CDC, Verification, Cutover, and Rollback
Build a migration as a reversible synchronization system with an explicit source of truth, backfill, CDC, shadow verification, progressive cutover, and rollback instead of a risky one-shot copy.
Define the source of truth at every migration phase and treat dual writes as a failure-prone transitional mechanism rather than a correctness guarantee.
Combine snapshot/backfill with CDC, stable identities, idempotent application, reconciliation, and source-version markers.
Use shadow reads and progressive cutover gates to detect semantic, freshness, performance, and completeness mismatches before broad traffic movement.
Keep rollback possible until the old system is no longer needed for correctness or recovery, and verify decommission prerequisites explicitly.
1. A database migration is not a file copy
AtlasMart wants to move catalog serving from its current relational representation to a document-oriented aggregate model. Production traffic cannot stop for a long export/import window. That turns the migration into a synchronization problem: historical rows must be transformed, concurrent changes must be captured, duplicates and ordering must be handled, the target must be verified, and traffic must move progressively while rollback remains possible.
Write down the phases before changing code: source-only → snapshot/backfill + CDC → shadow reads → limited target reads → target-primary with source fallback → target-only → decommission. Every phase names the authoritative writer/read source, permitted divergence, observability gates and rollback action.
2. Backfill, CDC and versions solve different parts of the problem
A snapshot gives a consistent historical starting point. Change Data Capture (CDC) carries later mutations. Stable record identities and source versions allow idempotent application, so replay or duplicate delivery does not regress state. Reconciliation compares source and target counts, hashes, key samples and domain invariants; it catches transformation bugs that “lag = zero” cannot prove.
| Phase | Source of truth | Required evidence | Rollback |
|---|---|---|---|
| Backfill | Old/source | snapshot position, transformed counts | discard target and retry |
| Shadow | Old/source | semantic diff, target p99, freshness | stop shadowing |
| Progressive reads | Old writer; target read subset | error/mismatch SLO by cohort | route cohort to source |
| Target primary | Explicit new authority | CDC/reconciliation clean, restore tested | predefined reverse route/write policy |
| Decommission | New target only | rollback window closed, archived evidence | requires new migration/restore |
3. Deliberately wrong approach: independent dual writes
A request writes the old database and then the new database directly. The first succeeds; the second times out. Retrying can duplicate side effects or overwrite a newer target value. Calling this “temporary” does not make it safe. The transitional system has all the partial-failure problems of any distributed workflow.
Prefer one authoritative commit plus an outbox/log/CDC feed where possible. If business constraints force dual writes, make them idempotent, versioned and observable, and run continuous reconciliation. Most importantly, define which side wins when the two disagree.
4. AtlasMart lab: snapshot, injected divergence, CDC, shadow verification
Python 3.13+ standard library only. CDC is represented by an in-memory ordered change log; no broker, connector or cloud service is required.
from copy import deepcopy
source={
"p1":{"version":1,"name":"Blue Mug","price":10},
"p2":{"version":1,"name":"Red Bowl","price":20},
}
target={}
change_log=[]
print("PHASE 1: source authoritative; snapshot/backfill")
snapshot=deepcopy(source)
for k,v in snapshot.items():
target[k]=deepcopy(v)
print("target after snapshot=", target)
print("\nPHASE 2: changes happen during/after backfill")
source['p1']={"version":2,"name":"Blue Mug","price":11}
change_log.append((1,'p1',deepcopy(source['p1'])))
source['p3']={"version":1,"name":"Green Plate","price":30}
change_log.append((2,'p3',deepcopy(source['p3'])))
print("broken dual write: p3 source succeeds, target write is lost")
print("source keys=", sorted(source), "target keys=", sorted(target))
print("\nCDC REPLAY WITH VERSION GUARD")
for offset,key,value in change_log + [change_log[0]]: # duplicate delivery too
current=target.get(key,{"version":0})
if value['version'] >= current['version']:
target[key]=deepcopy(value)
print("offset",offset,"key",key,"target_version",target[key]['version'])
print("\nSHADOW VERIFICATION")
def digest(d):
return sorted((k,v['version'],v['price']) for k,v in d.items())
print("source=",digest(source))
print("target=",digest(target))
print("match=",digest(source)==digest(target))
phase="target-primary-10-percent"
print("\nCUTOVER PHASE=",phase)
print("rollback rule: if mismatch/error/p99 gate fails, route reads back to source; keep CDC running")
print("decommission allowed only after rollback window + backup/restore + reconciliation evidence close")
The broken dual write leaves product p3 absent
from the target. CDC replay repairs it, duplicate delivery
does not regress the version, and the final source/target
digest matches before progressive cutover.
5. Production judgment
Migration acceptance requires more than row counts. Compare null/default semantics, ordering, timestamps, numeric precision, TTL, indexes, authorization filters, query results, p99 latency, failure behavior and restores. Backfills can overload the source, pollute caches and create hot partitions; rate-limit them and observe the source as a production workload.
Keep rollback honest. If target-only features begin writing data the source cannot represent, rollback has already become a new migration. Record that point explicitly. Once the target and source-of-truth boundaries are stable, Lesson 4 composes multiple stores without reintroducing ambiguous ownership.
Check your understanding
- Why are application dual writes unsafe by default?
- What is the source of truth during backfill?
- Why combine snapshot/backfill with CDC?
- What should a shadow read do?
- When is decommission safe?
Review the answers
1. Two independent commits can fail separately, creating divergence unless a stronger atomic mechanism, outbox/CDC strategy, or reconciliation path is used.
2. The migration plan must state it explicitly; commonly the old/source system remains authoritative while the target is populated and verified.
3. The snapshot covers historical state while CDC propagates changes that occur during and after the copy so the target can converge.
4. Read the candidate target without serving its result, compare semantics/freshness/performance with the authoritative path, and record mismatches.
5. After cutover stability, rollback criteria close, reconciliation is clean, backups/restores are proven, consumers are migrated, and the organization accepts the loss of the old fallback.
References
Foundational and current implementation references used for this lesson:
- PostgreSQL 18 — Transactions — Current official transaction semantics used as one relational implementation anchor.
- PostgreSQL 18 — Logical Replication — Current official snapshot-plus-change replication model relevant to migration and synchronization.
- Debezium 3.6 release series — Latest stable 3.6 line; 3.6.1.Final was released 2026-08-04. Optional CDC implementation anchor only.
- Apache Cassandra downloads — Current official release page listing Cassandra 5.0.9 as latest GA on 2026-08-07.
- Transactional Outbox pattern — A concise pattern reference for coupling application state and publication intent in one local transaction.
- PostgreSQL 18 — Logical Replication Restrictions — Current official reminder that replication/migration mechanisms have concrete scope and compatibility constraints.