Treat vector retrieval as an evaluated geometric index with fixed dimensions, similarity semantics, and measurable ANN recall.
Embeddings and Vector Search: Dimensions, Similarity Metrics, ANN Recall, and Index Choices
Model coordinates correctly before asking the database a spatial question.
Learning objectives
Explain embeddings, dimensions, cosine/dot-product/euclidean similarity, and the need to use a compatible query embedding space.
Create a vectorSearch index with vector and metadata filter fields and verify its asynchronous lifecycle.
Compare exact nearest-neighbor (ENN) results with approximate nearest-neighbor (ANN) results to compute recall@k/overlap.
Measure the numCandidates latency/recall tradeoff instead of accepting a universal tuning value.
Diagnose dimension mismatch, similarity mismatch, weak filtering, and memory/cost consequences.
This lesson pins MongoDB Server 8.3.8 through the
evaluation image
mongodb/mongodb-atlas-local:8.3.8-20260827T164726Z,
mongosh 2.10.0, and PyMongo
4.17.0 where client measurement is useful. The
current self-managed Search GA line is
mongot 1.70.1 or later for MongoDB 8.3; production
Community and Enterprise deployments operate
mongot as a separate process, while Atlas manages
it. The mandatory lab uses a disposable single-node replica-set
development image on loopback port 27172 named
atlasmart-ch21-l3. It is an evaluation topology,
not high availability. The local credential is synthetic and
must never be reused outside this lab.
FCV is observed and never changed. Atlas cloud,
KMS, external embedding APIs, and paid search nodes are not
required. Automated Embedding is not used; vectors are
deterministic teaching data so no model key or network call is
needed. Product labs were not executed in this generation
environment; index build time, search scores, recall, freshness
lag, and latency must be measured locally rather than copied as
invented output. The 4-dimensional vectors are deliberately
hand-authored teaching vectors. They make retrieval mechanics
deterministic but are not a benchmark of any embedding model,
RAG system, or recommendation quality.
1. A vector index preserves geometry from an embedding model
An embedding maps an item into a numeric vector whose geometry is intended to encode useful similarity. Dimensions are the number of numeric components. The index enforces that query vectors match the configured dimensionality. Production embeddings must be created with a stable model/version/preprocessing contract; changing that contract is a data migration, not a harmless client upgrade.
MongoDB Vector Search supports cosine, dotProduct, and euclidean similarity. The right function depends on the embedding model and normalization. Do not select one because its score “looks larger.”
2. Build vector and filter fields
docker rm -f atlasmart-ch21-l3 2>/dev/null || truedocker volume rm atlasmart-ch21-l3-db atlasmart-ch21-l3-config atlasmart-ch21-l3-search 2>/dev/null || truedocker run -d --name atlasmart-ch21-l3 \ -p 127.0.0.1:27172:27017 \ -e MONGODB_INITDB_ROOT_USERNAME=atlaslab \ -e MONGODB_INITDB_ROOT_PASSWORD=local-only-change-me \ -e DO_NOT_TRACK=1 \ -v atlasmart-ch21-l3-db:/data/db \ -v atlasmart-ch21-l3-config:/data/configdb \ -v atlasmart-ch21-l3-search:/data/mongot \ mongodb/mongodb-atlas-local:8.3.8-20260827T164726Zuntil [ "$(docker inspect -f '{{.State.Health.Status}}' atlasmart-ch21-l3 2>/dev/null)" = "healthy" ]; do sleep 2; donemongosh "mongodb://atlaslab:local-only-change-me@127.0.0.1:27172/admin?directConnection=true" --eval 'db.runCommand({ping:1})'
db = db.getSiblingDB("atlasmart");const c = db.catalog_search_ch21;c.drop();c.insertMany([ {_id:"P01",tenantId:"tenant-a",active:true, category:"power", title:"Portable USB-C Power Bank", description:"Compact travel battery with fast USB-C charging", embedding:[0.96,0.10,0.05,0.02], priceCents:4900}, {_id:"P02",tenantId:"tenant-a",active:true, category:"power", title:"65W USB-C Travel Charger", description:"GaN wall charger for laptops phones and travel", embedding:[0.92,0.16,0.03,0.02], priceCents:5900}, {_id:"P03",tenantId:"tenant-a",active:true, category:"cables", title:"Braided USB-C Cable", description:"Two meter durable charging and data cable", embedding:[0.82,0.20,0.08,0.04], priceCents:1800}, {_id:"P04",tenantId:"tenant-a",active:true, category:"audio", title:"Noise Cancelling Travel Headphones", description:"Over ear headphones for flights and commuting", embedding:[0.08,0.92,0.12,0.03], priceCents:12900}, {_id:"P05",tenantId:"tenant-a",active:true, category:"travel", title:"Universal Travel Adapter", description:"International plug adapter with USB-C ports", embedding:[0.72,0.14,0.55,0.06], priceCents:3900}, {_id:"P06",tenantId:"tenant-a",active:false,category:"power", title:"Legacy Power Brick", description:"Discontinued high capacity portable battery", embedding:[0.90,0.06,0.03,0.02], priceCents:3500}, {_id:"P07",tenantId:"tenant-b",active:true, category:"power", title:"Tenant B Private Charger", description:"Private catalog USB-C charging device", embedding:[0.95,0.11,0.03,0.02], priceCents:5100}, {_id:"P08",tenantId:"tenant-b",active:true, category:"security", title:"Tenant B Security Token", description:"Private authentication hardware token", embedding:[0.05,0.04,0.09,0.97], priceCents:7600}, {_id:"P09",tenantId:"tenant-a",active:true, category:"bags", title:"Laptop Travel Backpack", description:"Carry-on backpack with laptop compartment", embedding:[0.10,0.22,0.91,0.05], priceCents:8900}, {_id:"P10",tenantId:"tenant-a",active:true, category:"power", title:"Wireless Charging Pad", description:"Desk charger for Qi compatible phones", embedding:[0.78,0.14,0.05,0.03], priceCents:3200}, {_id:"P11",tenantId:"tenant-a",active:true, category:"audio", title:"USB-C Earbuds", description:"Wired earbuds with USB-C connector", embedding:[0.35,0.82,0.04,0.03], priceCents:2900}, {_id:"P12",tenantId:"tenant-a",active:true, category:"travel", title:"Packing Cube Set", description:"Lightweight organizers for carry-on travel", embedding:[0.08,0.13,0.96,0.02], priceCents:2600}]);print("documents", c.countDocuments({}));c.createSearchIndex("catalog_text", { mappings:{dynamic:false,fields:{ title:{type:"string",analyzer:"lucene.english"}, description:{type:"string",analyzer:"lucene.english"}, tenantId:{type:"token",normalizer:"lowercase"}, active:{type:"boolean"}, category:{type:"token",normalizer:"lowercase"}, priceCents:{type:"number"} }}});c.createSearchIndex("catalog_vector", "vectorSearch", { fields:[ {type:"vector", path:"embedding", numDimensions:4, similarity:"cosine"}, {type:"filter", path:"tenantId"}, {type:"filter", path:"active"}, {type:"filter", path:"category"} ]});printjson(c.getSearchIndexes());
const wanted = ["catalog_text", "catalog_vector"];for (let attempt=0; attempt<120; attempt++) { const m = new Map(c.getSearchIndexes().map(x => [x.name,x])); const ready = wanted.every(n => m.get(n) && m.get(n).status === "READY" && m.get(n).queryable === true); if (ready) { print("READY", wanted.join(",")); break; } sleep(1000);}printjson(c.getSearchIndexes());
3. ENN creates a small-lab ground truth; ANN trades work for speed
Exact nearest-neighbor (ENN) evaluates exact
vector similarity over the eligible set and is useful as a
judgement baseline when the collection is small enough.
Approximate nearest-neighbor (ANN) searches an
approximate graph/index structure and can miss neighbors.
numCandidates controls how many candidates ANN
considers before returning limit results; more
candidates usually improve recall but consume more work.
const q=[0.94,0.12,0.04,0.02];const exact=c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,exact:true,limit:5,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1,title:1,score:{$meta:"vectorSearchScore"}}}]).toArray();const approx=c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,numCandidates:7,limit:5,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1,title:1,score:{$meta:"vectorSearchScore"}}}]).toArray();print("ENN"); printjson(exact);print("ANN"); printjson(approx);const E=new Set(exact.map(x=>x._id));const overlap=approx.filter(x=>E.has(x._id)).length;print("recall@5 against ENN",overlap/exact.length);
On this tiny dataset the results may be identical. That does not prove production ANN recall. The point is the method: preserve an exact or curated judgement list, then measure overlap/recall and latency across representative queries.
4. Tune numCandidates with a distribution, not one lucky query
const candidates=[5,7,10];for (const n of candidates) { const ms=[]; let recallSum=0; for (let r=0;r<20;r++) { const t=Date.now(); const a=c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:[0.94,0.12,0.04,0.02],numCandidates:n,limit:5,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1}} ]).toArray(); ms.push(Date.now()-t); const ids=new Set(a.map(x=>x._id)); recallSum += exact.filter(x=>ids.has(x._id)).length/exact.length; } ms.sort((a,b)=>a-b); printjson({numCandidates:n,avgRecall:recallSum/20,p50:ms[Math.floor(ms.length*.50)],p95:ms[Math.floor(ms.length*.95)]});}
The documentation offers over-request guidance, but no universal value can substitute for your corpus size, filters, concurrency, vector dimensions, quantization, hardware, and quality target. Measure tails, not only means.
5. Controlled failures: dimensions and security filtering
try { printjson(c.aggregate([{$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:[0.9,0.1,0.0],exact:true,limit:3}}]).toArray());} catch (e) { print("expected dimension failure",e.code,e.codeName,e.message); }print("correct 4D query with tenant pre-filter");printjson(c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:[0.94,0.12,0.04,0.02],exact:true,limit:5,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1,tenantId:1,title:1,score:{$meta:"vectorSearchScore"}}}]).toArray());
A dimension mismatch is a model/index contract error. A missing tenant filter is worse: it can return a semantically close document from another tenant. Add security/business filter fields to the vector index and pre-filter inside vector retrieval.
6. Index choices, quantization, and capacity
Vector dimensions scale memory/disk linearly. Quantization can greatly reduce the in-memory representation but can affect accuracy and may require rescoring. Treat float/scalar/binary choices as an evaluated quality/cost decision. Production sizing must include vector count, dimensions, quantization, filters, concurrency, search-node memory, and index rebuild headroom.
7. Production judgment
Pin the embedding model, dimensions, normalization, similarity function, and index definition as one contract. Keep an evaluation set, measure ANN against ENN or human judgements, track latency distributions, and treat model changes as migrations requiring dual-write/backfill/cutover planning.
Bridge. Lesson 4 fuses lexical and vector retrieval so exact product names and semantic similarity can complement rather than replace one another.
docker rm -f atlasmart-ch21-l3 2>/dev/null || truedocker volume rm atlasmart-ch21-l3-db atlasmart-ch21-l3-config atlasmart-ch21-l3-search 2>/dev/null || true
Check your understanding
- Why must queryVector dimensionality match numDimensions?
- What is ENN useful for?
- What does numCandidates trade?
- Why include tenantId as a vector filter field?
- Is a 4D teaching vector a production embedding benchmark?
Review the answers
1. The index geometry is defined in a fixed-dimensional space; a different length is a contract error, not a value that MongoDB should silently truncate.
2. Ground truth on a manageable eligible set, allowing ANN overlap/recall evaluation.
3. More candidate exploration usually improves ANN recall but increases work/latency; tune empirically.
4. So tenant eligibility is enforced before semantic ranking rather than after cross-tenant candidates are retrieved.
5. No. It is deterministic data for mechanics only.
Authoritative references
- Self-Managed MongoDB Search and Vector Search — mongot architecture, Community/Enterprise deployment paths, and Atlas-managed responsibility boundary.
- mongot Compatibility and Requirements — MongoDB 8.3 / mongot 1.70.1 compatibility, platforms, and topology requirements.
- Local Development Quickstart — Atlas Local evaluation topology and search/vector index workflow.
- Self-Managed mongot Release Notes — mongot 1.70.1 GA baseline for Community and Enterprise self-managed search.
- Verify mongot Connection — health/readiness endpoints, index state, and end-to-end verification.
- Troubleshoot Self-Managed mongot — replication lag, stale indexes, rebuilds, disk pressure, and query failures.
- mongot Metrics Reference — per-index size, status, replication lag, and resource metrics.
- MongoDB Search Queries and Indexes — dynamic/static mappings, analyzers, query/index relationship.
- MongoDB Search text Operator — analyzed text matching, options, and scoring semantics.
- MongoDB Search compound Operator — must/should/filter/mustNot clauses and score-neutral filtering.
- Search Highlighting — highlight metadata and index requirements.
- Search Faceting — facet collector and current token-oriented string faceting guidance.
- Vector Search Index Fields — vector/filter fields, dimensions, similarity, and quantization options.
- MongoDB Search vectorSearch Operator — ANN/ENN, numCandidates, filters, dimensions, and vector scoring.
- Measure Vector Search Accuracy — ENN judgement lists, ANN recall/overlap, numCandidates tuning, and reranking.
- $rankFusion — reciprocal-rank fusion for multi-pipeline retrieval on MongoDB 8.0+.
- $scoreFusion — score normalization and weighted fusion, GA in MongoDB 8.3.
- getSearchIndexes() — search/vector index lifecycle states including READY, BUILDING, and STALE.
- MongoDB 8.3 Release Notes — current server line, 8.3.8 patch baseline, and scoreFusion GA.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17 driver baseline.