Prompt 17 · Lesson 03 · Server-side shard-key analysis

analyzeShardKey and Workload Evidence for Prospective Keys

analyzeShardKey combines candidate-key properties with sampled query behavior, but its answer is only as representative as the indexes and workload window you feed it.

Intermediate–Advanced130–210 minutesShard-key/distribution engineering labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Configure query sampling safely and explain why sampled workload quality determines analyzeShardKey usefulness.

02

Run analyzeShardKey against multiple prospective keys and interpret cardinality, frequency, monotonicity, and routing-distribution fields.

03

Explain the supporting-index requirement for key-characteristic metrics and the difference from the final shardCollection index requirement.

04

Monitor query sampling with currentOp and sampled-query evidence without assuming deterministic sample counts.

05

Use analyzer output as evidence in a decision record rather than as an automatic shard-key selector.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where the driver is used. The local topology is a disposable single-host sharded cluster with one single-member config server replica set, two single-member shard replica sets, and one mongos, exposed only on loopback diagnostic ports 27143–27146. This is sufficient for routing/distribution mechanics but is not production high availability. Authentication and TLS are disabled only for the isolated lab. Feature Compatibility Version (FCV) is inspected in the lab. FCV is observed and never changed. Application/DDL operations go through mongos; direct shard connections are diagnostics only. analyzeShardKey cannot run on a standalone and must run through mongos for a sharded cluster. Query sampling is probabilistic/traffic-dependent; the lesson specifies expected fields and directional evidence, not fixed sample counts. Atlas, Search, Vector Search, KMS, and Enterprise Advanced are not mandatory. Product commands were not executed in this generation environment because Docker, mongod, mongos, mongosh, and PyMongo are unavailable here, so all timing/load outputs must be measured on the learner machine rather than copied as invented results.

1. analyzeShardKey turns key choice into measurable hypotheses

MongoDB 7.0+ provides analyzeShardKey for both sharded and unsharded collections in a replica-set/sharded deployment. It can return keyCharacteristics (cardinality, most-common-value frequency, uniqueness, monotonicity) and readWriteDistribution (how sampled reads/writes would target ranges under the candidate). The second family depends on representative query sampling; a synthetic five-second sample cannot stand in for a real production cycle.

Important index nuance

Key-characteristic metrics require a supporting simple-collation index that is not multikey, sparse, or partial. The analyzer accepts a broader set of supporting index patterns than the final shardCollection command, so “analyzable” does not automatically mean “ready to shard.”

start compact sharded topology (l3)
docker rm -f atlasmart-ch17-l3-cfg atlasmart-ch17-l3-s1 atlasmart-ch17-l3-s2 atlasmart-ch17-l3-mongos 2>/dev/null || truedocker network rm atlasmart-ch17-l3-net 2>/dev/null || truedocker volume rm atlasmart-ch17-l3-cfg-data atlasmart-ch17-l3-s1-data atlasmart-ch17-l3-s2-data 2>/dev/null || truedocker network create atlasmart-ch17-l3-netdocker run -d --name atlasmart-ch17-l3-cfg --network atlasmart-ch17-l3-net -p 127.0.0.1:27143:27017 -v atlasmart-ch17-l3-cfg-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --configsvr --replSet atlasmart-cfg17-l3 --bind_ip_alldocker run -d --name atlasmart-ch17-l3-s1 --network atlasmart-ch17-l3-net -p 127.0.0.1:27144:27017 -v atlasmart-ch17-l3-s1-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --shardsvr --replSet atlasmart-shard17-l3-a --bind_ip_alldocker run -d --name atlasmart-ch17-l3-s2 --network atlasmart-ch17-l3-net -p 127.0.0.1:27145:27017 -v atlasmart-ch17-l3-s2-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --shardsvr --replSet atlasmart-shard17-l3-b --bind_ip_allfor PORT in 27143 27144 27145; do  until mongosh "mongodb://127.0.0.1:$PORT/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donedonemongosh "mongodb://127.0.0.1:27143/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-cfg17-l3",configsvr:true,members:[{_id:0,host:"atlasmart-ch17-l3-cfg:27017"}]})'mongosh "mongodb://127.0.0.1:27144/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-shard17-l3-a",members:[{_id:0,host:"atlasmart-ch17-l3-s1:27017"}]})'mongosh "mongodb://127.0.0.1:27145/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-shard17-l3-b",members:[{_id:0,host:"atlasmart-ch17-l3-s2:27017"}]})'for PORT in 27143 27144 27145; do  until mongosh "mongodb://127.0.0.1:$PORT/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donedonedocker run -d --name atlasmart-ch17-l3-mongos --network atlasmart-ch17-l3-net -p 127.0.0.1:27146:27017 --entrypoint mongos mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --configdb atlasmart-cfg17-l3/atlasmart-ch17-l3-cfg:27017 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27146/admin" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27146/admin" --quiet --eval 'printjson(sh.addShard("atlasmart-shard17-l3-a/atlasmart-ch17-l3-s1:27017"));printjson(sh.addShard("atlasmart-shard17-l3-b/atlasmart-ch17-l3-s2:27017"));printjson({hello:db.hello(),fcv:db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion,shards:db.adminCommand({listShards:1}).shards});' 

2. Create candidate indexes and a skewed workload fixture

prepare unsharded AtlasMart orders for prospective analysis
const app=db.getSiblingDB("atlasmart");sh.enableSharding("atlasmart", "atlasmart-shard17-l3-a");app.orders_analysis.drop();const docs=[];for (let i=1;i<=1200;i++) {  const tenant = i<=420 ? "mega" : `t${String(((i-421)%39)+1).padStart(2,"0")}`;  docs.push({orderId:`o-${String(i).padStart(5,"0")}`,tenantId:tenant,region:["EU","US","APAC"][i%3],createdSeq:i,status:i%4?"paid":"open"});}app.orders_analysis.insertMany(docs);app.orders_analysis.createIndex({region:1});app.orders_analysis.createIndex({tenantId:1});app.orders_analysis.createIndex({createdSeq:1});app.orders_analysis.createIndex({tenantId:1,orderId:1});printjson(app.orders_analysis.getIndexes());

3. Turn on sampling, run representative query shapes, and observe sampling state

sample a controlled mixed workload
const app=db.getSiblingDB("atlasmart");printjson(app.orders_analysis.configureQueryAnalyzer({mode:"full",samplesPerSecond:20}));for (let i=0;i<120;i++) {  app.orders_analysis.findOne({tenantId:i%3===0?"mega":"t07",status:"paid"});  if (i%4===0) app.orders_analysis.find({region:"EU"}).limit(5).toArray();  if (i%5===0) app.orders_analysis.findOne({orderId:`o-${String((i%100)+1).padStart(5,"0")}`});  if (i%6===0) app.orders_analysis.updateOne({tenantId:"mega",orderId:`o-${String((i%100)+1).padStart(5,"0")}`},{$set:{touched:true}});}sleep(1500);printjson(db.getSiblingDB("admin").aggregate([  {$currentOp:{allUsers:true,localOps:false}},  {$match:{desc:/query analyzer/i}},  {$project:{desc:1,ns:1,samplesPerSecond:1,sampledReadsCount:1,sampledWritesCount:1,startTime:1}}]).toArray());

The exact number of sampled operations can vary. The evidence you need is that sampling is enabled for the intended namespace and that the workload window contains the query shapes you intend to evaluate. For production, sample long enough to capture normal, peak, batch, and tenant-skew patterns.

4. Compare prospective keys with one consistent summary

run analyzeShardKey for four candidates
const app=db.getSiblingDB("atlasmart");function summarize(name, result) {  const k=result.keyCharacteristics || {};  const r=result.readDistribution || {};  const w=result.writeDistribution || {};  printjson({    candidate:name,    distinct:k.numDistinctValues,    hottest:k.mostCommonValues?.[0],    monotonicity:k.monotonicity,    readSamples:r.sampleSize?.total,    singleShardReads:r.percentageOfSingleShardReads,    scatterReads:r.percentageOfScatterGatherReads,    writeSamples:w.sampleSize?.total,    singleShardWrites:w.percentageOfSingleShardWrites,    scatterWrites:w.percentageOfScatterGatherWrites  });}for (const [name,key] of [  ["region",{region:1}],  ["tenant",{tenantId:1}],  ["createdSeq",{createdSeq:1}],  ["tenant+order",{tenantId:1,orderId:1}],]) {  const result=app.orders_analysis.analyzeShardKey(key,{keyCharacteristics:true,readWriteDistribution:true,sampleRate:1});  summarize(name,result);}

Directional expectations: region has low cardinality; tenantId exposes the hot mega value; createdSeq should show monotonic behavior for this insertion pattern; and {tenantId,orderId} offers many distinct values while aligning with tenant-scoped order lookups. The exact sampled read/write percentages are workload evidence, not constants.

5. Controlled failure: remove the supporting index

make key-characteristic analysis fail safely, then repair
const app=db.getSiblingDB("atlasmart");app.orders_analysis.dropIndex({region:1});try {  printjson(app.orders_analysis.analyzeShardKey({region:1},{keyCharacteristics:true,readWriteDistribution:false,sampleRate:1}));} catch (e) {  print(`Expected supporting-index error: ${e.codeName || e.message}`);}app.orders_analysis.createIndex({region:1});printjson(app.orders_analysis.configureQueryAnalyzer({mode:"off"}));
Wrong approach

Do not run analyzeShardKey once during a quiet hour and declare the highest-cardinality key the winner. Query-distribution metrics are only as representative as the sampled workload, and monotonicity can also be misleading after migrations because donor deletion/recipient insertion changes record ordering.

6. Production judgment

Store analyzer output beside workload metadata: sampling window, traffic mix, collection size, candidate indexes, tenant skew, read/write ratio, and deployment version. Re-run when workload or distribution changes. Analyzer metrics narrow the decision space; they do not estimate your future p99, operational skill, zone design, or resharding cost.

Bridge. Lesson 4 takes one key with a geographic prefix and uses zones to turn logical ranges into placement policy.

cleanup / full reset
docker rm -f atlasmart-ch17-l3-cfg atlasmart-ch17-l3-s1 atlasmart-ch17-l3-s2 atlasmart-ch17-l3-mongos 2>/dev/null || truedocker volume rm atlasmart-ch17-l3-cfg-data atlasmart-ch17-l3-s1-data atlasmart-ch17-l3-s2-data 2>/dev/null || truedocker network rm atlasmart-ch17-l3-net 2>/dev/null || true

Check your understanding

  1. What does keyCharacteristics measure?
  2. What does readWriteDistribution depend on?
  3. Can analyzeShardKey run on a standalone?
  4. Why can sample counts differ between runs?
  5. Does analyzeShardKey automatically choose the production key?
Review the answers

1. Candidate-key properties such as distinct values, most-common frequencies, uniqueness, and monotonicity.

2. Representative sampled queries collected by the query analyzer.

3. No. It requires a replica-set or sharded deployment; on a sharded cluster run it through mongos.

4. Query sampling is workload- and timing-dependent; the lesson does not promise deterministic sample counts.

5. No. It provides evidence that must be combined with workload, topology, zone, operational, and migration constraints.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.