Prompt 17 · Lesson 02 · Distribution strategy
Ranged vs Hashed Sharding, Compound Keys, and Prefix-Aware Targeting
Ranged and hashed sharding trade natural locality for write distribution in different ways; compound prefixes let routing preserve selected business locality.
Learning objectives
Distinguish ranged from hashed sharding by the ordering of the distributed key space and the queries each favors.
Explain exact-equality targeting on a hashed key and why range predicates on the unhashed value lose locality.
Use compound ranged keys to preserve a useful prefix for tenant-scoped targeting.
Predict which AtlasMart queries target one/subset/all shards before reading explain output.
Reject the misconception that hashed sharding automatically fixes skew, hot application values, or every write bottleneck.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where the driver is used. The local topology
is a disposable single-host sharded cluster with one
single-member config server replica set, two
single-member shard replica sets, and one
mongos, exposed only on loopback diagnostic ports
27139–27142. This is sufficient for routing/distribution
mechanics but is not production high availability.
Authentication and TLS are disabled only for the isolated lab.
Feature Compatibility Version (FCV) is inspected in the lab. FCV
is observed and never changed. Application/DDL operations go
through mongos; direct shard connections are
diagnostics only. The lab creates separate ranged and hashed
collections so explain evidence is not confused by changing one
collection mid-lesson. Small one-host distributions are
mechanism evidence only. Atlas, Search, Vector Search, KMS, and
Enterprise Advanced are not mandatory. Product commands were not
executed in this generation environment because Docker, mongod,
mongos, mongosh, and PyMongo are unavailable here, so all
timing/load outputs must be measured on the learner machine
rather than copied as invented results.
1. Ranged and hashed sharding optimize different locality
Ranged sharding preserves the sort order of shard-key values in the range map, so nearby values can share ranges and range predicates can be targeted. Hashed sharding stores ranges of the server-computed hash instead; logically adjacent source values are intentionally scattered through hash space. That usually improves write distribution for a high-cardinality monotonic key, but it sacrifices natural range locality.
| Strategy | Good fit | Tradeoff |
|---|---|---|
Ranged {tenantId:1, seq:1} |
Tenant prefix and within-tenant ordered/range access. | Monotonic suffix can still create a hot edge within a very busy tenant prefix. |
Hashed {eventId:"hashed"} |
High-cardinality monotonically increasing IDs with equality lookups. | Range predicates on original eventId cannot use hashed ordering and tend to scatter. |
| Compound hashed | Need a retained prefix plus hashed distribution in another component. | Targeting/distribution become less intuitive; analyze real workload before adopting. |
docker rm -f atlasmart-ch17-l2-cfg atlasmart-ch17-l2-s1 atlasmart-ch17-l2-s2 atlasmart-ch17-l2-mongos 2>/dev/null || truedocker network rm atlasmart-ch17-l2-net 2>/dev/null || truedocker volume rm atlasmart-ch17-l2-cfg-data atlasmart-ch17-l2-s1-data atlasmart-ch17-l2-s2-data 2>/dev/null || truedocker network create atlasmart-ch17-l2-netdocker run -d --name atlasmart-ch17-l2-cfg --network atlasmart-ch17-l2-net -p 127.0.0.1:27139:27017 -v atlasmart-ch17-l2-cfg-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --configsvr --replSet atlasmart-cfg17-l2 --bind_ip_alldocker run -d --name atlasmart-ch17-l2-s1 --network atlasmart-ch17-l2-net -p 127.0.0.1:27140:27017 -v atlasmart-ch17-l2-s1-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --shardsvr --replSet atlasmart-shard17-l2-a --bind_ip_alldocker run -d --name atlasmart-ch17-l2-s2 --network atlasmart-ch17-l2-net -p 127.0.0.1:27141:27017 -v atlasmart-ch17-l2-s2-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --shardsvr --replSet atlasmart-shard17-l2-b --bind_ip_allfor PORT in 27139 27140 27141; do until mongosh "mongodb://127.0.0.1:$PORT/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donedonemongosh "mongodb://127.0.0.1:27139/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-cfg17-l2",configsvr:true,members:[{_id:0,host:"atlasmart-ch17-l2-cfg:27017"}]})'mongosh "mongodb://127.0.0.1:27140/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-shard17-l2-a",members:[{_id:0,host:"atlasmart-ch17-l2-s1:27017"}]})'mongosh "mongodb://127.0.0.1:27141/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-shard17-l2-b",members:[{_id:0,host:"atlasmart-ch17-l2-s2:27017"}]})'for PORT in 27139 27140 27141; do until mongosh "mongodb://127.0.0.1:$PORT/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donedonedocker run -d --name atlasmart-ch17-l2-mongos --network atlasmart-ch17-l2-net -p 127.0.0.1:27142:27017 --entrypoint mongos mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --configdb atlasmart-cfg17-l2/atlasmart-ch17-l2-cfg:27017 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27142/admin" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27142/admin" --quiet --eval 'printjson(sh.addShard("atlasmart-shard17-l2-a/atlasmart-ch17-l2-s1:27017"));printjson(sh.addShard("atlasmart-shard17-l2-b/atlasmart-ch17-l2-s2:27017"));printjson({hello:db.hello(),fcv:db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion,shards:db.adminCommand({listShards:1}).shards});'
2. Build ranged and hashed AtlasMart collections
const app=db.getSiblingDB("atlasmart");sh.enableSharding("atlasmart", "atlasmart-shard17-l2-a");app.events_range.drop();app.events_hash.drop();app.events_range.createIndex({tenantId:1,seq:1});printjson(sh.shardCollection("atlasmart.events_range", {tenantId:1,seq:1}));printjson(sh.splitAt("atlasmart.events_range", {tenantId:"m",seq:MinKey}));printjson(sh.moveChunk("atlasmart.events_range", {tenantId:"z",seq:1}, "atlasmart-shard17-l2-b"));app.events_hash.createIndex({eventId:"hashed"});printjson(sh.shardCollection("atlasmart.events_hash", {eventId:"hashed"}));const ranged=[], hashed=[];for (let i=1;i<=400;i++) { const tenant = i%2 ? "b" : "p"; ranged.push({tenantId:tenant,seq:i,eventId:i,kind:"order"}); hashed.push({tenantId:tenant,seq:i,eventId:i,kind:"order"});}app.events_range.insertMany(ranged);app.events_hash.insertMany(hashed);sh.status(true);
The hashed collection's exact per-shard count is not a fixed
expected number in this lesson. What matters is that the hash
breaks the source-value ordering; verify the actual distribution
with sh.getShardedDataDistribution().
3. Explain targeting, not just result correctness
function shardNames(explain) { const found = new Set(); function walk(v) { if (!v || typeof v !== "object") return; if (typeof v.shardName === "string") found.add(v.shardName); if (Array.isArray(v.shards)) { for (const s of v.shards) { if (s && typeof s.shardName === "string") found.add(s.shardName); walk(s); } } for (const k of Object.keys(v)) walk(v[k]); } walk(explain); return Array.from(found).sort();}const app=db.getSiblingDB("atlasmart");for (const [label,coll,query] of [ ["range-prefix", app.events_range, {tenantId:"b"}], ["range-full", app.events_range, {tenantId:"b",seq:42}], ["range-no-prefix", app.events_range, {seq:{$gte:390}}], ["hash-equality", app.events_hash, {eventId:42}], ["hash-range", app.events_hash, {eventId:{$gte:390,$lt:400}}],]) { const ex=coll.find(query).explain("executionStats"); printjson({label,query,shards:shardNames(ex),nReturned:ex.executionStats?.nReturned});}
For the ranged compound key, a query on the leading
tenantId prefix can narrow routing even if it omits
seq. A query on seq alone does not
contain that prefix. For the hashed collection, an equality
predicate lets mongos hash the value and target the
corresponding range, but a source-value range such as 390–399 is
not a contiguous hashed range.
4. Boundary case: “hashed means balanced” is too strong
from collections import Countervalues = ["mega"] * 700 + [f"t{i:03d}" for i in range(300)]counts = Counter(values)print("distinct source values:", len(counts))print("hottest source value frequency:", counts["mega"])print("Hashing maps each distinct value to a hash; it does not split 700 identical mega values into 700 independent values.")
Do not choose a hashed key only because inserts are increasing. Verify equality-versus-range query mix, per-value frequency, future zone requirements, and whether a compound prefix is needed for targeting. Hashed sharding solves an ordering/distribution problem, not every workload problem.
5. Production judgment
Prefer ranged keys when business locality and range targeting are valuable and you can avoid hot edge ranges. Prefer hashed components when write distribution for a high-cardinality monotonic field dominates and equality targeting is sufficient. Compound keys let you blend concerns, but every additional field becomes part of the routing/index contract and future zone design. Measure explain shard fan-out and per-shard workload; do not infer production performance from a one-host lab.
Bridge. Lesson 3 replaces hand-calculated
intuition with MongoDB's analyzeShardKey and
sampled workload metrics.
docker rm -f atlasmart-ch17-l2-cfg atlasmart-ch17-l2-s1 atlasmart-ch17-l2-s2 atlasmart-ch17-l2-mongos 2>/dev/null || truedocker volume rm atlasmart-ch17-l2-cfg-data atlasmart-ch17-l2-s1-data atlasmart-ch17-l2-s2-data 2>/dev/null || truedocker network rm atlasmart-ch17-l2-net 2>/dev/null || true
Check your understanding
- Why can ranged sharding target a source-value interval efficiently?
- Why is eventId equality compatible with a hashed shard key?
- Why does a source-value range perform poorly on a hashed key?
- What does the leading field of a compound shard key enable?
- Does hashing split one very frequent value across many hashes?
Review the answers
1. Because the range map preserves shard-key ordering, so mongos can map the interval to the overlapping owned ranges.
2. mongos can compute the same hash for the equality value and target the matching hashed range.
3. Adjacent source values are not adjacent in hash space, so the range generally cannot map to one contiguous shard-key range.
4. Queries containing that prefix can often narrow routing even when later shard-key fields are omitted.
5. No. Identical source values produce the same hash, so frequency skew remains a design concern.
Authoritative references
- Choose a Shard Key — Cardinality, frequency, monotonicity, query patterns, and shard-key tradeoffs.
- Troubleshoot Shard Keys — Hot ranges, uneven load, jumbo/indivisible ranges, and scatter/gather symptoms.
- analyzeShardKey — Key-characteristic and sampled read/write-distribution metrics.
- configureQueryAnalyzer — Query sampling for prospective shard-key analysis.
- Hashed Indexes for Sharding — Write distribution and range-query limitations.
- Zones — Zone membership, non-overlapping zone ranges, prefix requirements, and balancing behavior.
- addShardToZone — Associate shards with zone labels.
- Change a Shard Key — Refine versus reshard decision boundary.
- Refine a Shard Key — Append suffix fields without changing existing shard-key field types.
- reshardCollection — Full shard-key change, redistribution phases, demo mode, and abort/commit boundaries.
- unshardCollection — MongoDB 8.0+ consolidation of a sharded collection onto one shard.
- moveCollection — MongoDB 8.0+ relocation of an unsharded collection between shards.
- MongoDB Sharding — Routing, ranges, balancer behavior, and mongos-only client access.
- MongoDB 8.3 Release Notes — Current 8.3 release and sharding/DDL changes.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17 baseline.