Follow one owned shard-key range as MongoDB redistributes it and expose the resource cost around that move.
Chunks/Ranges, Balancer, Migration, Critical Sections, and Operational Headroom
Inspect chunks/ranges, balancer status, metadata changes, migration semantics, and client tail-latency measurement.
Learning objectives
Define chunks/ranges as shard-key intervals and inspect their current owners.
Explain what the balancer monitors and why a tiny lab may never trigger an automatic balancing round.
Run one explicit range migration and verify config metadata before and after.
Explain synchronization, the brief migration critical section, and asynchronous orphan cleanup without inventing a fixed pause duration.
Build an operational-headroom checklist for bandwidth, cache/I/O, replication, and tail latency during redistribution.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where the driver is used. The mandatory
topology is a disposable single-host sharded cluster with one
single-member config server replica set, two
single-member shard replica sets, and one
mongos router, exposed only on loopback diagnostic
ports 27127–27130. Single-member replica sets satisfy the
mechanism requirement but provide no production redundancy;
production deployments require properly sized multi-member
replica sets and independent failure domains. Authentication and
TLS are disabled only for this isolated local lab. Feature
Compatibility Version (FCV) is observed and never changed.
Reads/writes go through mongos; direct shard
connections are used only for explicitly labeled diagnostics.
Automatic balancing is observed but not forced to trigger from
an artificial tiny threshold. The lab performs one explicit
moveRange/moveChunk-style migration on
disposable data so ownership changes are deterministic. Atlas,
Search, Vector Search, KMS, and Enterprise Advanced are not
mandatory. Product commands were not executed in this generation
environment because Docker, mongod, mongos, mongosh, and PyMongo
are unavailable here; expected invariants are
documentation-derived and measured timing/output must be
recorded on the learner machine.
1. A chunk/range is ownership metadata, not a file boundary
MongoDB partitions a ranged shard-key space into chunks (ranges). Each range has an inclusive lower bound and exclusive upper bound and is assigned to one shard. The range record is cluster metadata; storage pages/files on the shard do not need to align one-for-one with that logical boundary.
docker rm -f atlasmart-ch16-l4-cfg atlasmart-ch16-l4-s1 atlasmart-ch16-l4-s2 atlasmart-ch16-l4-mongos 2>/dev/null || truedocker network rm atlasmart-ch16-l4-net 2>/dev/null || truedocker volume rm atlasmart-ch16-l4-cfg-data atlasmart-ch16-l4-s1-data atlasmart-ch16-l4-s2-data 2>/dev/null || truedocker network create atlasmart-ch16-l4-netdocker run -d --name atlasmart-ch16-l4-cfg --network atlasmart-ch16-l4-net -p 127.0.0.1:27127:27017 -v atlasmart-ch16-l4-cfg-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --configsvr --replSet atlasmart-cfg16-l4 --bind_ip_alldocker run -d --name atlasmart-ch16-l4-s1 --network atlasmart-ch16-l4-net -p 127.0.0.1:27128:27017 -v atlasmart-ch16-l4-s1-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --shardsvr --replSet atlasmart-shard16-l4-a --bind_ip_alldocker run -d --name atlasmart-ch16-l4-s2 --network atlasmart-ch16-l4-net -p 127.0.0.1:27129:27017 -v atlasmart-ch16-l4-s2-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --shardsvr --replSet atlasmart-shard16-l4-b --bind_ip_allfor PORT in 27127 27128 27129; do until mongosh "mongodb://127.0.0.1:$PORT/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donedonemongosh "mongodb://127.0.0.1:27127/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-cfg16-l4",configsvr:true,members:[{_id:0,host:"atlasmart-ch16-l4-cfg:27017"}]})'mongosh "mongodb://127.0.0.1:27128/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-shard16-l4-a",members:[{_id:0,host:"atlasmart-ch16-l4-s1:27017"}]})'mongosh "mongodb://127.0.0.1:27129/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-shard16-l4-b",members:[{_id:0,host:"atlasmart-ch16-l4-s2:27017"}]})'for PORT in 27127 27128 27129; do until mongosh "mongodb://127.0.0.1:$PORT/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donedonedocker run -d --name atlasmart-ch16-l4-mongos --network atlasmart-ch16-l4-net -p 127.0.0.1:27130:27017 --entrypoint mongos mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --configdb atlasmart-cfg16-l4/atlasmart-ch16-l4-cfg:27017 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27130/admin" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27130/admin" --quiet --eval 'printjson(sh.addShard("atlasmart-shard16-l4-a/atlasmart-ch16-l4-s1:27017"));printjson(sh.addShard("atlasmart-shard16-l4-b/atlasmart-ch16-l4-s2:27017"));printjson({hello:db.hello(),shards:db.adminCommand({listShards:1}).shards});'
const admin = db.getSiblingDB("admin");const app = db.getSiblingDB("atlasmart");printjson(sh.enableSharding("atlasmart", "atlasmart-shard16-l4-a"));app.orders.drop();app.orders.createIndex({tenantId:1,orderId:1});printjson(sh.shardCollection("atlasmart.orders", {tenantId:1}));const docs=[];for (const t of ["a","b","c","n","o","p"]) { for (let i=1;i<=4;i++) docs.push({tenantId:t,orderId:`${t}-${i}`,status:i%2?"open":"closed",amount:i*25});}app.orders.insertMany(docs);printjson(sh.splitAt("atlasmart.orders", {tenantId:"m"}));printjson(sh.moveChunk("atlasmart.orders", {tenantId:"z"}, "atlasmart-shard16-l4-b"));sh.status(true);
2. Snapshot ownership and balancer state
const cfg=db.getSiblingDB("config");const c=cfg.collections.findOne({_id:"atlasmart.orders"});printjson(cfg.chunks.find({uuid:c.uuid},{min:1,max:1,shard:1,history:1}).sort({min:1}).toArray());printjson(sh.balancerCollectionStatus("atlasmart.orders"));printjson(sh.getShardedDataDistribution());printjson({balancerEnabled:sh.getBalancerState(),balancerRunning:sh.isBalancerRunning()});
The balancer is a background control-plane process. A 24-document fixture is far below normal migration thresholds, so absence of automatic movement is expected and is not evidence that balancing is broken.
3. Move one range explicitly and verify ownership
const cfg=db.getSiblingDB("config");const coll=cfg.collections.findOne({_id:"atlasmart.orders"});const before=cfg.chunks.find({uuid:coll.uuid},{min:1,max:1,shard:1}).sort({min:1}).toArray();printjson({before});// Move the lower range [MinKey, "m") to shard B for an observable ownership change.printjson(db.getSiblingDB("admin").runCommand({ moveRange:"atlasmart.orders", toShard:"atlasmart-shard16-l4-b", min:{tenantId:MinKey}, max:{tenantId:"m"}}));const after=cfg.chunks.find({uuid:coll.uuid},{min:1,max:1,shard:1,history:1}).sort({min:1}).toArray();printjson({after});printjson(sh.getShardedDataDistribution());
Re-running the lab can make the requested source/target relationship differ. Use the cleanup/reset block first so the migration starts from the deterministic initial ownership. Never “repair” a production cluster by blindly issuing range moves from copied lesson commands.
4. What happens during migration
The recipient copies the range and catches up writes that occurred while copying. Near commit, MongoDB enters a brief critical section for the migrating range/collection operations needed to make the metadata ownership transition safe. After metadata commit, old copies become orphans until cleanup removes them. Current MongoDB can overlap later migration work with asynchronous cleanup, which is why migration completion and disk-space recovery are not identical timestamps.
The critical-section duration and client-visible tail-latency effect depend on range size, write rate, replication health, network, storage, metadata latency, and concurrent work. Record it in your environment; do not copy a folklore millisecond value.
5. Optional latency observation around a manual move
import statistics, timefrom pymongo import MongoClientclient=MongoClient("mongodb://127.0.0.1:27130/?directConnection=true",serverSelectionTimeoutMS=3000)coll=client.atlasmart.orderssamples=[]for _ in range(200): t=time.perf_counter() list(coll.find({"tenantId":"b"}).limit(10)) samples.append((time.perf_counter()-t)*1000) time.sleep(0.02)s=sorted(samples)def pct(p): return s[min(len(s)-1,int((len(s)-1)*p))]print({"n":len(s),"p50_ms":pct(.50),"p95_ms":pct(.95),"p99_ms":pct(.99),"max_ms":max(s)})# Run this loop concurrently with the manual migration and compare with a baseline run.
The script measures client latency only. A tiny local range may migrate too quickly to produce a visible difference. That outcome is legitimate evidence; do not enlarge destructive scope merely to manufacture a spike.
Redistribution consumes bandwidth, cache, disk I/O, replication capacity, and metadata work. Keep headroom before adding shards or moving large ranges. Monitor migrations, orphan cleanup, replication lag, cache pressure, disk space, and application tail latency. Schedule or constrain balancing when the workload requires it, but do not leave balancing disabled indefinitely without a distribution plan.
Bridge. Lesson 5 rebuilds the topology from scratch and traces one application query end-to-end from client to router to target shard.
docker rm -f atlasmart-ch16-l4-cfg atlasmart-ch16-l4-s1 atlasmart-ch16-l4-s2 atlasmart-ch16-l4-mongos 2>/dev/null || truedocker volume rm atlasmart-ch16-l4-cfg-data atlasmart-ch16-l4-s1-data atlasmart-ch16-l4-s2-data 2>/dev/null || truedocker network rm atlasmart-ch16-l4-net 2>/dev/null || true
Check your understanding
- Does a range correspond to one physical storage file?
- Why might the balancer do nothing in this lab?
- What is the migration critical section for?
- Why can disk cleanup lag behind metadata ownership change?
Review the answers
1. No. It is logical ownership metadata for shard-key intervals.
2. The fixture is far below normal imbalance/migration thresholds.
3. To safely finalize ownership/metadata transition after the recipient is synchronized.
4. Old copies can remain as orphans until asynchronous range cleanup completes.
Authoritative references
- MongoDB Sharding — Sharded-cluster purpose, chunks/ranges, targeted operations, and architecture.
- Routing with mongos — Router metadata cache, targeted versus broadcast operations, and aggregation routing evidence.
- Config Servers — Config server replica sets, metadata responsibilities, and config-shard alternatives.
-
config Database
— Internal metadata including
config.collections,config.chunks,config.shards, and settings. - sh.status() — Shards, databases, ranges, sharded data distribution, and migration summaries.
- sh.shardCollection() — Shard a collection and define its shard key.
-
sh.addShard()
— Add replica-set shards through
mongos. - Split Chunks/Ranges — Controlled manual split examples and operational cautions.
-
moveRange
— Explicit range migration through
mongos. - Sharded Cluster Balancer — Range migration procedure, thresholds, cleanup, and resource impact.
- MongoDB 8.3 Release Notes — Current 8.3 behavior including mongos-only DDL on sharded clusters.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17 baseline.