Evolve AtlasMart customer documents with schemaVersion markers, mixed-version readers, lazy migration, idempotent eager backfills, checkpoints, and rollback-compatible expansion.
Schema Version Fields, Lazy Migration, Eager Backfills, and Mixed-Version Readers
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
AtlasMart is moving customer contact data from top-level
email/phone fields into a
contact subdocument. A
schema version marker is application metadata
that tells readers which interpretation a document follows. It
is not MongoDB's Feature Compatibility Version (FCV), and
MongoDB does not automatically migrate documents because a field
is named schemaVersion.
Design a schemaVersion contract that distinguishes application document generations without confusing it with server FCV.
Implement mixed-version readers that normalize old and new documents into one application view.
Compare lazy migration on read/write with eager background backfill and state the workload tradeoffs.
Make migration retries safe by filtering source state and deterministically setting target state.
Use a resumable high-water-mark checkpoint without turning the checkpoint itself into an incorrect migrated-count source of truth.
This lesson pins MongoDB Community Server
8.3.8 with image
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, a disposable standalone mongod bound only to
127.0.0.1:27049, mongosh 2.10.0, and
PyMongo 4.17.0 where driver behavior is shown.
The database is atlasmart; the collection is
customers_ch07_l3. Authentication and TLS are
disabled only because this is an isolated loopback learning
container. The lab reads FCV but never changes it. Default
read/write concern and primary read preference are used on
this standalone. Atlas, Search, KMS, Enterprise Advanced, and
paid services are not required. Product commands below were
reviewed against current official documentation; they were not
executed in this generation environment because
Docker/mongod/mongosh/PyMongo are unavailable here.
Result fragments labeled “expected” are contract-oriented shapes derived from the documented command semantics, not copied from a generation-time MongoDB run. Field order, error wording, log attributes, object identifiers, timing, and some diagnostic detail are version/environment dependent. Verify the invariant or count being taught rather than comparing output byte-for-byte.
1. Version the interpretation, not every deploy
A schema version should change when the stored representation or interpretation changes in a way readers/writers must understand. Do not increment it for an unrelated application deployment. Documents with no marker need an explicit convention; in this lesson, absence means version 1. That convention must be encoded in every mixed-version reader until the migration is complete.
MongoDB's Schema Versioning pattern explicitly supports keeping
multiple shapes in one collection. That can avoid a
stop-the-world migration, but compatibility code and indexes can
become more complex. If the same logical field moves from
email to contact.email, queries may
need to cover both paths and may require indexes for both during
the transition.
| Strategy | Read path | Write/backfill cost | Best fit / risk |
|---|---|---|---|
| Lazy read only | Reader maps v1/v2 to one view; stored v1 remains. | Minimal migration I/O. | Useful for cold historical data; mixed shapes can live indefinitely and complicate indexes/analytics. |
| Lazy read + write-back | A read may opportunistically convert the document. | Traffic-shaped writes and possible latency spikes on reads. | Good when hot records naturally converge; must make retries and concurrent writers safe. |
| Eager backfill | Background job scans source versions and writes target version. | Predictable extra read/write/replication load. | Good when a deadline requires convergence; throttle from measured system pressure. |
| Dual write during rollback window | New writer populates both old and new fields. | Write amplification and duplicate mutable state. | Useful only with explicit ownership/freshness rules; remove after old readers are retired. |
2. Preserve rollback compatibility while moving the shape
The first safe v2 write in this lesson adds
contact and schemaVersion:2 but keeps
email and phone. That duplication is
temporary and deliberate: an older application can still read
the top-level fields if the new release must be rolled back.
Once telemetry proves no old readers remain, a later contract
phase can remove the old fields and related indexes.
This is a classic expand/contract migration. The dangerous shortcut is to delete the old fields as soon as the first new writer exists. A deployment rollback would then restore an application binary that cannot interpret data already produced by the new binary.
3. Run an idempotent eager transformation and prove reruns are no-ops
docker rm -f atlasmart-mongo-ch07-l3 2>/dev/null || truedocker volume rm atlasmart-mongo-ch07-l3-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch07-l3 \ -p 127.0.0.1:27049:27017 \ -v atlasmart-mongo-ch07-l3-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27049/atlasmart?directConnection=true" --quiet --eval \'printjson({server:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1}))'
const customers = db.getCollection("customers_ch07_l3");customers.drop();customers.insertMany([ { _id:"c-7301", name:"Ayla", email:"ayla@example.test", phone:"+994-1001" }, { _id:"c-7302", schemaVersion:NumberInt(1), name:"Mina", email:"mina@example.test", phone:"+994-1002" }, { _id:"c-7303", schemaVersion:NumberInt(2), name:"Noah", email:"noah@example.test", phone:"+994-1003", contact:{email:"noah@example.test",phone:"+994-1003"} }]);function distribution() { return customers.aggregate([ { $group:{ _id:{ $ifNull:["$schemaVersion", NumberInt(1)] }, count:{ $sum:1 } } }, { $sort:{ _id:1 } } ]).toArray();}print("before:"); printjson(distribution());// Idempotent eager transformation for v1/missing-version documents.const r1 = customers.updateMany( { $or:[ {schemaVersion:{$exists:false}}, {schemaVersion:NumberInt(1)} ] }, [ { $set:{ contact:{ email:{ $ifNull:["$contact.email","$email"] }, phone:{ $ifNull:["$contact.phone","$phone"] } }, schemaVersion:NumberInt(2) } } // Old email/phone stay during the rollback window. ]);printjson(r1);print("after first pass:"); printjson(distribution());const r2 = customers.updateMany( { $or:[ {schemaVersion:{$exists:false}}, {schemaVersion:NumberInt(1)} ] }, [ { $set:{ contact:{email:"$email",phone:"$phone"}, schemaVersion:NumberInt(2) } } ]);print("rerun modified:", r2.modifiedCount);printjson(customers.find({}, {_id:1,schemaVersion:1,email:1,contact:1}).sort({_id:1}).toArray());
The first pass should modify the two v1/missing-version records; the second pass has no source-version records left and therefore modifies zero. That zero is important evidence: a retry after a timeout or job restart does not advance a document to a fictitious version 3.
// Wrong migration: rerunning changes already-migrated documents again.db.customers_ch07_l3.updateMany({}, { $inc:{ schemaVersion:NumberInt(1) } });db.customers_ch07_l3.updateMany({}, { $inc:{ schemaVersion:NumberInt(1) } });// v1 -> v3 after two runs, even though no v3 reader contract exists.// Safer migration: filter the source state and SET the target state deterministically.db.customers_ch07_l3.updateMany( { $or:[ {schemaVersion:{$exists:false}}, {schemaVersion:NumberInt(1)} ] }, [ { $set:{ schemaVersion:NumberInt(2), contact:{email:"$email",phone:"$phone"} } } ]);
4. Application readers and resumable checkpoints
A migration is not finished merely because the database job can transform records. During rollout, readers must tolerate the version set that can exist at each point. The PyMongo example below normalizes both versions, optionally performs a lazy compare-and-set write-back, and shows an eager batch loop with a durable high-water mark.
from pymongo import MongoClient, UpdateOneURI = "mongodb://127.0.0.1:27049/atlasmart?directConnection=true"client = MongoClient(URI, serverSelectionTimeoutMS=3000)coll = client.atlasmart.customers_ch07_l3progress = client.atlasmart.migration_progress_ch07def normalize(raw): version = raw.get("schemaVersion", 1) if version == 1: return { "_id": raw["_id"], "name": raw["name"], "contact": {"email": raw.get("email"), "phone": raw.get("phone")}, "schemaVersion": 2, } if version == 2: return {"_id": raw["_id"], "name": raw["name"], "contact": raw["contact"], "schemaVersion": 2} raise ValueError(f"unsupported schemaVersion={version}")def lazy_read(customer_id, write_back=False): raw = coll.find_one({"_id": customer_id}) view = normalize(raw) if write_back and raw.get("schemaVersion", 1) == 1: # Compare-and-set style filter: a second/retried migration becomes a no-op. coll.update_one( {"_id": customer_id, "$or": [{"schemaVersion": {"$exists": False}}, {"schemaVersion": 1}]}, {"$set": {"contact": view["contact"], "schemaVersion": 2}}, ) return viewprint("mixed-reader view:", lazy_read("c-7301", write_back=True))# Small lab batch size to make progress visible; not a production tuning recommendation.BATCH = 2last_id = progress.find_one({"_id": "customers-v1-v2"}) or {}last_id = last_id.get("lastId", "")while True: batch = list(coll.find({ "_id": {"$gt": last_id}, "$or": [{"schemaVersion": {"$exists": False}}, {"schemaVersion": 1}], }).sort("_id", 1).limit(BATCH)) if not batch: break ops = [] for raw in batch: view = normalize(raw) ops.append(UpdateOne( {"_id": raw["_id"], "$or": [{"schemaVersion": {"$exists": False}}, {"schemaVersion": 1}]}, {"$set": {"contact": view["contact"], "schemaVersion": 2}}, )) result = coll.bulk_write(ops, ordered=True) last_id = batch[-1]["_id"] # Store a high-water mark, not a blindly incremented migrated total. progress.update_one( {"_id": "customers-v1-v2"}, {"$set": {"lastId": last_id, "lastBatchModified": result.modified_count}}, upsert=True, ) print("checkpoint", last_id, "modified", result.modified_count)client.close()
The progress collection records lastId and the last
batch's modification count. It deliberately does not blindly
increment a total because a crash can occur between data writes
and checkpoint writes. The canonical convergence evidence is the
actual remaining source-version count/distribution. On
production workloads with ObjectId or other nontrivial ordering,
choose a stable checkpoint key that matches your access/index
design and test pagination behavior explicitly.
Each eager update is a real write: it consumes CPU, storage-engine cache, journal/checkpoint work, replica network/oplog capacity, indexes, and potentially shard-routing/balancer capacity. A lab batch of two only makes state visible. Production batch size and pacing must come from measured latency, replication lag, cache/disk pressure, and workload SLOs—not from a tutorial constant.
5. Verification, cleanup, and production judgment
Verification checklist
- The initial distribution contains two logical v1 documents and one v2 document.
- The eager migration reaches a 100% v2 distribution while preserving old fields for rollback.
- The same migration rerun reports zero modifications.
- The mixed-version reader can normalize a v1 document before migration and a v2 document after migration.
- The checkpoint stores a high-water mark and does not claim to be the authoritative migrated total.
- You can state when it is safe to remove the old fields and old indexes: only after old readers/writers and rollback dependence are gone.
Schema versioning is appropriate when rolling compatibility or a long migration horizon matters. Its guarantee is only the contract your application implements: MongoDB does not enforce that version 2 has a particular shape unless you also use validators. Extra versions increase query branches, index budget, test matrix, and operational complexity. Keep the supported-version window intentionally small and measure distribution until old versions are extinct.
For sharded collections, batch filters and checkpoint keys
interact with shard targeting; cross-shard updates can fan out.
For replica sets, backfill consumes oplog and can increase lag.
Durability and consistency remain governed by write/read
concern, not by schemaVersion. Security and tenant
filters must remain in every migration path; a maintenance
script that omits tenantId can become a
cross-tenant incident.
The next lesson focuses on what happens when old documents are missing fields or contain stale derived values. Defaults and repairs must preserve semantic meaning rather than merely make validators green.
docker rm -f atlasmart-mongo-ch07-l3docker volume rm atlasmart-mongo-ch07-l3-data
Check your understanding
-
Why is an application
schemaVersionfield unrelated to MongoDB FCV? - Why does the lesson retain top-level email/phone after writing v2 contact data?
- What makes the eager v1→v2 update idempotent?
- Why is an incrementing migrated counter unsafe as the only progress truth?
- When can lazy migration be better than an eager backfill?
Review the answers
schemaVersion is application-defined document
metadata. FCV is a server compatibility control governing
MongoDB feature behavior during upgrades. They solve
different problems.
It preserves a rollback/read-compatibility window for older application versions. The duplication is temporary and should be removed only after old readers are retired.
The filter selects only v1/missing-version source documents and the update deterministically sets version 2. Once converted, the same document no longer matches the source filter.
A crash can happen after data writes but before the counter checkpoint, or retries can reprocess a batch. Actual schema-version distribution is the authoritative convergence evidence.
When many records are cold and may never need transformation, lazy normalization avoids a large write storm. The tradeoff is longer-lived mixed schemas and more complicated queries/indexes.
Authoritative references
- MongoDB 8.3 release notes — Current server-series and patch-release status; re-check before reproducing the lab.
- MongoDB versioning — Explains major/minor release series and why the latest stable patch should be used.
- mongosh release notes — Current mongosh release and version-sensitive shell behavior.
- PyMongo release notes — Current Python driver line used when application-side migration behavior is demonstrated.
- Document and schema versioning — Official overview of retaining historical document/schema versions.
-
Schema Versioning pattern
— Official
schemaVersionpattern and mixed-shape query implications. - Updates with aggregation pipeline — Official update-pipeline stages used for deterministic field-to-field transformation.
- updateMany — Documents idempotent-update guidance and sharded-collection considerations.
- Bulk write operations with PyMongo — Driver-side batching primitives used for resumable eager migration.
- Atomicity and transactions — Single-document atomicity context for compare-and-set style migration writes.