Treat vector retrieval as an evaluated geometric index with fixed dimensions, similarity semantics, and measurable ANN recall.

Embeddings and Vector Search: Dimensions, Similarity Metrics, ANN Recall, and Index Choices

Model coordinates correctly before asking the database a spatial question.

Advanced120–190 minutesVector recall labMongoDB 8.3.8 · mongot 1.70.1+ · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Explain embeddings, dimensions, cosine/dot-product/euclidean similarity, and the need to use a compatible query embedding space.

02

Create a vectorSearch index with vector and metadata filter fields and verify its asynchronous lifecycle.

03

Compare exact nearest-neighbor (ENN) results with approximate nearest-neighbor (ANN) results to compute recall@k/overlap.

04

Measure the numCandidates latency/recall tradeoff instead of accepting a universal tuning value.

05

Diagnose dimension mismatch, similarity mismatch, weak filtering, and memory/cost consequences.

Reproducible lab baseline

This lesson pins MongoDB Server 8.3.8 through the evaluation image mongodb/mongodb-atlas-local:8.3.8-20260827T164726Z, mongosh 2.10.0, and PyMongo 4.17.0 where client measurement is useful. The current self-managed Search GA line is mongot 1.70.1 or later for MongoDB 8.3; production Community and Enterprise deployments operate mongot as a separate process, while Atlas manages it. The mandatory lab uses a disposable single-node replica-set development image on loopback port 27172 named atlasmart-ch21-l3. It is an evaluation topology, not high availability. The local credential is synthetic and must never be reused outside this lab. FCV is observed and never changed. Atlas cloud, KMS, external embedding APIs, and paid search nodes are not required. Automated Embedding is not used; vectors are deterministic teaching data so no model key or network call is needed. Product labs were not executed in this generation environment; index build time, search scores, recall, freshness lag, and latency must be measured locally rather than copied as invented output. The 4-dimensional vectors are deliberately hand-authored teaching vectors. They make retrieval mechanics deterministic but are not a benchmark of any embedding model, RAG system, or recommendation quality.

1. A vector index preserves geometry from an embedding model

An embedding maps an item into a numeric vector whose geometry is intended to encode useful similarity. Dimensions are the number of numeric components. The index enforces that query vectors match the configured dimensionality. Production embeddings must be created with a stable model/version/preprocessing contract; changing that contract is a data migration, not a harmless client upgrade.

MongoDB Vector Search supports cosine, dotProduct, and euclidean similarity. The right function depends on the embedding model and normalization. Do not select one because its score “looks larger.”

2. Build vector and filter fields

start the pinned local Search/Vector Search evaluation lab
docker rm -f atlasmart-ch21-l3 2>/dev/null || truedocker volume rm atlasmart-ch21-l3-db atlasmart-ch21-l3-config atlasmart-ch21-l3-search 2>/dev/null || truedocker run -d --name atlasmart-ch21-l3 \  -p 127.0.0.1:27172:27017 \  -e MONGODB_INITDB_ROOT_USERNAME=atlaslab \  -e MONGODB_INITDB_ROOT_PASSWORD=local-only-change-me \  -e DO_NOT_TRACK=1 \  -v atlasmart-ch21-l3-db:/data/db \  -v atlasmart-ch21-l3-config:/data/configdb \  -v atlasmart-ch21-l3-search:/data/mongot \  mongodb/mongodb-atlas-local:8.3.8-20260827T164726Zuntil [ "$(docker inspect -f '{{.State.Health.Status}}' atlasmart-ch21-l3 2>/dev/null)" = "healthy" ]; do sleep 2; donemongosh "mongodb://atlaslab:local-only-change-me@127.0.0.1:27172/admin?directConnection=true" --eval 'db.runCommand({ping:1})' 
seed deterministic vectors and build text/vector indexes
db = db.getSiblingDB("atlasmart");const c = db.catalog_search_ch21;c.drop();c.insertMany([ {_id:"P01",tenantId:"tenant-a",active:true, category:"power", title:"Portable USB-C Power Bank", description:"Compact travel battery with fast USB-C charging", embedding:[0.96,0.10,0.05,0.02], priceCents:4900}, {_id:"P02",tenantId:"tenant-a",active:true, category:"power", title:"65W USB-C Travel Charger", description:"GaN wall charger for laptops phones and travel", embedding:[0.92,0.16,0.03,0.02], priceCents:5900}, {_id:"P03",tenantId:"tenant-a",active:true, category:"cables", title:"Braided USB-C Cable", description:"Two meter durable charging and data cable", embedding:[0.82,0.20,0.08,0.04], priceCents:1800}, {_id:"P04",tenantId:"tenant-a",active:true, category:"audio", title:"Noise Cancelling Travel Headphones", description:"Over ear headphones for flights and commuting", embedding:[0.08,0.92,0.12,0.03], priceCents:12900}, {_id:"P05",tenantId:"tenant-a",active:true, category:"travel", title:"Universal Travel Adapter", description:"International plug adapter with USB-C ports", embedding:[0.72,0.14,0.55,0.06], priceCents:3900}, {_id:"P06",tenantId:"tenant-a",active:false,category:"power", title:"Legacy Power Brick", description:"Discontinued high capacity portable battery", embedding:[0.90,0.06,0.03,0.02], priceCents:3500}, {_id:"P07",tenantId:"tenant-b",active:true, category:"power", title:"Tenant B Private Charger", description:"Private catalog USB-C charging device", embedding:[0.95,0.11,0.03,0.02], priceCents:5100}, {_id:"P08",tenantId:"tenant-b",active:true, category:"security", title:"Tenant B Security Token", description:"Private authentication hardware token", embedding:[0.05,0.04,0.09,0.97], priceCents:7600}, {_id:"P09",tenantId:"tenant-a",active:true, category:"bags", title:"Laptop Travel Backpack", description:"Carry-on backpack with laptop compartment", embedding:[0.10,0.22,0.91,0.05], priceCents:8900}, {_id:"P10",tenantId:"tenant-a",active:true, category:"power", title:"Wireless Charging Pad", description:"Desk charger for Qi compatible phones", embedding:[0.78,0.14,0.05,0.03], priceCents:3200}, {_id:"P11",tenantId:"tenant-a",active:true, category:"audio", title:"USB-C Earbuds", description:"Wired earbuds with USB-C connector", embedding:[0.35,0.82,0.04,0.03], priceCents:2900}, {_id:"P12",tenantId:"tenant-a",active:true, category:"travel", title:"Packing Cube Set", description:"Lightweight organizers for carry-on travel", embedding:[0.08,0.13,0.96,0.02], priceCents:2600}]);print("documents", c.countDocuments({}));c.createSearchIndex("catalog_text", {  mappings:{dynamic:false,fields:{    title:{type:"string",analyzer:"lucene.english"},    description:{type:"string",analyzer:"lucene.english"},    tenantId:{type:"token",normalizer:"lowercase"},    active:{type:"boolean"},    category:{type:"token",normalizer:"lowercase"},    priceCents:{type:"number"}  }}});c.createSearchIndex("catalog_vector", "vectorSearch", {  fields:[    {type:"vector", path:"embedding", numDimensions:4, similarity:"cosine"},    {type:"filter", path:"tenantId"},    {type:"filter", path:"active"},    {type:"filter", path:"category"}  ]});printjson(c.getSearchIndexes());
wait for both indexes
const wanted = ["catalog_text", "catalog_vector"];for (let attempt=0; attempt<120; attempt++) {  const m = new Map(c.getSearchIndexes().map(x => [x.name,x]));  const ready = wanted.every(n => m.get(n) && m.get(n).status === "READY" && m.get(n).queryable === true);  if (ready) { print("READY", wanted.join(",")); break; }  sleep(1000);}printjson(c.getSearchIndexes());

3. ENN creates a small-lab ground truth; ANN trades work for speed

Exact nearest-neighbor (ENN) evaluates exact vector similarity over the eligible set and is useful as a judgement baseline when the collection is small enough. Approximate nearest-neighbor (ANN) searches an approximate graph/index structure and can miss neighbors. numCandidates controls how many candidates ANN considers before returning limit results; more candidates usually improve recall but consume more work.

compare exact and approximate neighbors under the same tenant filter
const q=[0.94,0.12,0.04,0.02];const exact=c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,exact:true,limit:5,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1,title:1,score:{$meta:"vectorSearchScore"}}}]).toArray();const approx=c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,numCandidates:7,limit:5,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1,title:1,score:{$meta:"vectorSearchScore"}}}]).toArray();print("ENN"); printjson(exact);print("ANN"); printjson(approx);const E=new Set(exact.map(x=>x._id));const overlap=approx.filter(x=>E.has(x._id)).length;print("recall@5 against ENN",overlap/exact.length);

On this tiny dataset the results may be identical. That does not prove production ANN recall. The point is the method: preserve an exact or curated judgement list, then measure overlap/recall and latency across representative queries.

4. Tune numCandidates with a distribution, not one lucky query

small repeated experiment; record p50/p95 and recall yourself
const candidates=[5,7,10];for (const n of candidates) {  const ms=[]; let recallSum=0;  for (let r=0;r<20;r++) {    const t=Date.now();    const a=c.aggregate([      {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:[0.94,0.12,0.04,0.02],numCandidates:n,limit:5,filter:{tenantId:"tenant-a",active:true}}},      {$project:{_id:1}}    ]).toArray();    ms.push(Date.now()-t);    const ids=new Set(a.map(x=>x._id)); recallSum += exact.filter(x=>ids.has(x._id)).length/exact.length;  }  ms.sort((a,b)=>a-b);  printjson({numCandidates:n,avgRecall:recallSum/20,p50:ms[Math.floor(ms.length*.50)],p95:ms[Math.floor(ms.length*.95)]});}

The documentation offers over-request guidance, but no universal value can substitute for your corpus size, filters, concurrency, vector dimensions, quantization, hardware, and quality target. Measure tails, not only means.

5. Controlled failures: dimensions and security filtering

dimension mismatch should fail rather than silently truncate
try {  printjson(c.aggregate([{$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:[0.9,0.1,0.0],exact:true,limit:3}}]).toArray());} catch (e) { print("expected dimension failure",e.code,e.codeName,e.message); }print("correct 4D query with tenant pre-filter");printjson(c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:[0.94,0.12,0.04,0.02],exact:true,limit:5,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1,tenantId:1,title:1,score:{$meta:"vectorSearchScore"}}}]).toArray());

A dimension mismatch is a model/index contract error. A missing tenant filter is worse: it can return a semantically close document from another tenant. Add security/business filter fields to the vector index and pre-filter inside vector retrieval.

6. Index choices, quantization, and capacity

Vector dimensions scale memory/disk linearly. Quantization can greatly reduce the in-memory representation but can affect accuracy and may require rescoring. Treat float/scalar/binary choices as an evaluated quality/cost decision. Production sizing must include vector count, dimensions, quantization, filters, concurrency, search-node memory, and index rebuild headroom.

7. Production judgment

Pin the embedding model, dimensions, normalization, similarity function, and index definition as one contract. Keep an evaluation set, measure ANN against ENN or human judgements, track latency distributions, and treat model changes as migrations requiring dual-write/backfill/cutover planning.

Bridge. Lesson 4 fuses lexical and vector retrieval so exact product names and semantic similarity can complement rather than replace one another.

cleanup only this lesson lab
docker rm -f atlasmart-ch21-l3 2>/dev/null || truedocker volume rm atlasmart-ch21-l3-db atlasmart-ch21-l3-config atlasmart-ch21-l3-search 2>/dev/null || true

Check your understanding

  1. Why must queryVector dimensionality match numDimensions?
  2. What is ENN useful for?
  3. What does numCandidates trade?
  4. Why include tenantId as a vector filter field?
  5. Is a 4D teaching vector a production embedding benchmark?
Review the answers

1. The index geometry is defined in a fixed-dimensional space; a different length is a contract error, not a value that MongoDB should silently truncate.

2. Ground truth on a manageable eligible set, allowing ANN overlap/recall evaluation.

3. More candidate exploration usually improves ANN recall but increases work/latency; tune empirically.

4. So tenant eligibility is enforced before semantic ranking rather than after cross-tenant candidates are retrieved.

5. No. It is deterministic data for mechanics only.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.