Fuse lexical and semantic signals only after preserving filters and evaluating ranking, recall, latency, and leakage.
Hybrid Text + Vector Retrieval, Metadata Filtering, Reranking, and Evaluation
Model coordinates correctly before asking the database a spatial question.
Learning objectives
Explain why lexical and semantic retrieval fail on different query classes and why hybrid search can improve robustness.
Build text and vector retrieval branches with identical tenant/visibility filtering.
Use rank fusion and the MongoDB 8.3 scoreFusion option while understanding their different score assumptions.
Evaluate precision-style judgement lists, recall/overlap, latency, and filter leakage rather than relying on a few screenshots.
Diagnose fusion weight/ranking mistakes and preserve a reranking/evaluation loop.
This lesson pins MongoDB Server 8.3.8 through the
evaluation image
mongodb/mongodb-atlas-local:8.3.8-20260827T164726Z,
mongosh 2.10.0, and PyMongo
4.17.0 where client measurement is useful. The
current self-managed Search GA line is
mongot 1.70.1 or later for MongoDB 8.3; production
Community and Enterprise deployments operate
mongot as a separate process, while Atlas manages
it. The mandatory lab uses a disposable single-node replica-set
development image on loopback port 27173 named
atlasmart-ch21-l4. It is an evaluation topology,
not high availability. The local credential is synthetic and
must never be reused outside this lab.
FCV is observed and never changed. Atlas cloud,
KMS, external embedding APIs, and paid search nodes are not
required. Automated Embedding is not used; vectors are
deterministic teaching data so no model key or network call is
needed. Product labs were not executed in this generation
environment; index build time, search scores, recall, freshness
lag, and latency must be measured locally rather than copied as
invented output. MongoDB 8.3 makes
$scoreFusion generally available;
$rankFusion is GA from MongoDB 8.0. The lesson uses
rank-based fusion as the more score-scale-agnostic baseline and
shows score fusion as an 8.3-specific alternative.
1. Lexical and semantic signals fail differently
Exact model names, SKUs, acronyms, and freshly coined product terms often favor lexical search because the words themselves matter. Paraphrases such as “battery for charging a laptop on flights” may favor semantic vectors even when the exact terms differ. Hybrid retrieval keeps both signals and combines their ranked evidence.
A hybrid system is not automatically better. If either branch leaks unauthorized documents, has poor recall, or dominates with badly calibrated weights, fusion can amplify the defect.
2. Build both indexes and verify readiness
docker rm -f atlasmart-ch21-l4 2>/dev/null || truedocker volume rm atlasmart-ch21-l4-db atlasmart-ch21-l4-config atlasmart-ch21-l4-search 2>/dev/null || truedocker run -d --name atlasmart-ch21-l4 \ -p 127.0.0.1:27173:27017 \ -e MONGODB_INITDB_ROOT_USERNAME=atlaslab \ -e MONGODB_INITDB_ROOT_PASSWORD=local-only-change-me \ -e DO_NOT_TRACK=1 \ -v atlasmart-ch21-l4-db:/data/db \ -v atlasmart-ch21-l4-config:/data/configdb \ -v atlasmart-ch21-l4-search:/data/mongot \ mongodb/mongodb-atlas-local:8.3.8-20260827T164726Zuntil [ "$(docker inspect -f '{{.State.Health.Status}}' atlasmart-ch21-l4 2>/dev/null)" = "healthy" ]; do sleep 2; donemongosh "mongodb://atlaslab:local-only-change-me@127.0.0.1:27173/admin?directConnection=true" --eval 'db.runCommand({ping:1})'
db = db.getSiblingDB("atlasmart");const c = db.catalog_search_ch21;c.drop();c.insertMany([ {_id:"P01",tenantId:"tenant-a",active:true, category:"power", title:"Portable USB-C Power Bank", description:"Compact travel battery with fast USB-C charging", embedding:[0.96,0.10,0.05,0.02], priceCents:4900}, {_id:"P02",tenantId:"tenant-a",active:true, category:"power", title:"65W USB-C Travel Charger", description:"GaN wall charger for laptops phones and travel", embedding:[0.92,0.16,0.03,0.02], priceCents:5900}, {_id:"P03",tenantId:"tenant-a",active:true, category:"cables", title:"Braided USB-C Cable", description:"Two meter durable charging and data cable", embedding:[0.82,0.20,0.08,0.04], priceCents:1800}, {_id:"P04",tenantId:"tenant-a",active:true, category:"audio", title:"Noise Cancelling Travel Headphones", description:"Over ear headphones for flights and commuting", embedding:[0.08,0.92,0.12,0.03], priceCents:12900}, {_id:"P05",tenantId:"tenant-a",active:true, category:"travel", title:"Universal Travel Adapter", description:"International plug adapter with USB-C ports", embedding:[0.72,0.14,0.55,0.06], priceCents:3900}, {_id:"P06",tenantId:"tenant-a",active:false,category:"power", title:"Legacy Power Brick", description:"Discontinued high capacity portable battery", embedding:[0.90,0.06,0.03,0.02], priceCents:3500}, {_id:"P07",tenantId:"tenant-b",active:true, category:"power", title:"Tenant B Private Charger", description:"Private catalog USB-C charging device", embedding:[0.95,0.11,0.03,0.02], priceCents:5100}, {_id:"P08",tenantId:"tenant-b",active:true, category:"security", title:"Tenant B Security Token", description:"Private authentication hardware token", embedding:[0.05,0.04,0.09,0.97], priceCents:7600}, {_id:"P09",tenantId:"tenant-a",active:true, category:"bags", title:"Laptop Travel Backpack", description:"Carry-on backpack with laptop compartment", embedding:[0.10,0.22,0.91,0.05], priceCents:8900}, {_id:"P10",tenantId:"tenant-a",active:true, category:"power", title:"Wireless Charging Pad", description:"Desk charger for Qi compatible phones", embedding:[0.78,0.14,0.05,0.03], priceCents:3200}, {_id:"P11",tenantId:"tenant-a",active:true, category:"audio", title:"USB-C Earbuds", description:"Wired earbuds with USB-C connector", embedding:[0.35,0.82,0.04,0.03], priceCents:2900}, {_id:"P12",tenantId:"tenant-a",active:true, category:"travel", title:"Packing Cube Set", description:"Lightweight organizers for carry-on travel", embedding:[0.08,0.13,0.96,0.02], priceCents:2600}]);print("documents", c.countDocuments({}));c.createSearchIndex("catalog_text", { mappings:{dynamic:false,fields:{ title:{type:"string",analyzer:"lucene.english"}, description:{type:"string",analyzer:"lucene.english"}, tenantId:{type:"token",normalizer:"lowercase"}, active:{type:"boolean"}, category:{type:"token",normalizer:"lowercase"}, priceCents:{type:"number"} }}});c.createSearchIndex("catalog_vector", "vectorSearch", { fields:[ {type:"vector", path:"embedding", numDimensions:4, similarity:"cosine"}, {type:"filter", path:"tenantId"}, {type:"filter", path:"active"}, {type:"filter", path:"category"} ]});printjson(c.getSearchIndexes());
const wanted = ["catalog_text", "catalog_vector"];for (let attempt=0; attempt<120; attempt++) { const m = new Map(c.getSearchIndexes().map(x => [x.name,x])); const ready = wanted.every(n => m.get(n) && m.get(n).status === "READY" && m.get(n).queryable === true); if (ready) { print("READY", wanted.join(",")); break; } sleep(1000);}printjson(c.getSearchIndexes());
3. Reciprocal-rank fusion combines order without assuming score scale
$rankFusion combines the position of each document
in named input pipelines. This is useful when lexical and vector
score magnitudes are not directly comparable. The input
pipelines must remain selection/scoring pipelines over the same
collection.
const q=[0.94,0.12,0.04,0.02];printjson(c.aggregate([ {$rankFusion:{ input:{pipelines:{ lexical:[ {$search:{index:"catalog_text",compound:{ must:[{text:{query:"travel charger",path:["title","description"]}}], filter:[{equals:{path:"tenantId",value:"tenant-a"}},{equals:{path:"active",value:true}}] }}}, {$limit:8} ], semantic:[ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,numCandidates:10,limit:8,filter:{tenantId:"tenant-a",active:true}}} ] }}, combination:{weights:{lexical:1.2,semantic:1.0}}, scoreDetails:true }}, {$project:{_id:1,title:1,tenantId:1,scoreDetails:{$meta:"scoreDetails"}}}, {$limit:6}]).toArray());
Inspect which products come from one or both branches. Do not
couple application logic to the exact internal
scoreDetails shape; MongoDB documents that detailed
formatting is not a stable contract.
4. MongoDB 8.3 score fusion is useful only after normalization choices are justified
$scoreFusion normalizes scores and combines
weighted input scores. It is GA in MongoDB 8.3. This is more
expressive than rank fusion, but it also requires you to reason
about score distributions and normalization.
printjson(c.aggregate([ {$scoreFusion:{ input:{pipelines:{ lexical:[{$search:{index:"catalog_text",compound:{must:[{text:{query:"travel charger",path:["title","description"]}}],filter:[{equals:{path:"tenantId",value:"tenant-a"}},{equals:{path:"active",value:true}}]}}},{$limit:8}], semantic:[{$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,numCandidates:10,limit:8,filter:{tenantId:"tenant-a",active:true}}}] },normalization:"sigmoid"}, combination:{weights:{lexical:1.2,semantic:1.0},method:"avg"}, scoreDetails:true }}, {$project:{_id:1,title:1,scoreDetails:{$meta:"scoreDetails"}}}, {$limit:6}]).toArray());
Do not choose rank fusion versus score fusion from a single query. Evaluate a labelled set of query intents: exact-name queries, descriptive queries, ambiguous queries, category queries, and adversarial tenant/security cases.
5. Evaluate relevance and leakage together
from collections import defaultdictjudgements={ "travel charger":{"P02","P05","P01"}, "portable battery":{"P01","P06"}, "usb c audio":{"P11"}}def precision_at_k(returned,relevant,k): top=returned[:k] return sum(x in relevant for x in top)/max(1,len(top))def recall_at_k(returned,relevant,k): return sum(x in relevant for x in returned[:k])/max(1,len(relevant))# Replace these fixture rankings with product-lab output.fixture={"travel charger":["P02","P01","P05","P10"],"portable battery":["P01","P10","P06"],"usb c audio":["P11","P03"]}for q,rel in judgements.items(): r=fixture[q] print(q,"P@3",precision_at_k(r,rel,3),"R@3",recall_at_k(r,rel,3))
For real production evaluation, add latency percentiles, zero-result rate, click/engagement proxies, freshness probes, and filter leakage tests that assert every returned document belongs to the authorized tenant and visibility state.
6. Wrong hybrid pattern: filter after retrieval
Run broad text/vector retrieval across all tenants,
fuse/rerank it, then apply
$match:{tenantId:"tenant-a"} at the end. Even if
final output appears filtered, cross-tenant candidates have
already influenced candidate selection, scores, logs, traces,
caches, or downstream rerankers. Repair the design by applying
the security filter inside
every retrieval branch.
Security filtering is an eligibility rule, not a relevance
preference. Test it with synthetic “highly relevant but
forbidden” documents such as P07; the test should
fail if that document is ever visible to the authorized branch.
7. Production judgment
Choose fusion weights from an evaluation suite and segment results by query class. Track score/rank drift after analyzer or embedding changes. If you add an external reranker, bound the candidate set, latency, cost, and data sent outside the database; never send unauthorized candidates to the reranker.
Bridge. Lesson 5 turns retrieval into a production RAG/recommendation service with freshness, security, cost, and observability contracts.
docker rm -f atlasmart-ch21-l4 2>/dev/null || truedocker volume rm atlasmart-ch21-l4-db atlasmart-ch21-l4-config atlasmart-ch21-l4-search 2>/dev/null || true
Check your understanding
- Why can lexical and semantic retrieval complement each other?
- Why is rank fusion attractive across heterogeneous retrievers?
- What changed in MongoDB 8.3?
- Where must tenant filtering occur?
- What should relevance evaluation include besides P@k/recall?
Review the answers
1. They fail differently: lexical search preserves exact tokens/proper nouns, while vector retrieval can recover paraphrases and semantic similarity.
2. It combines rank positions without assuming the underlying score magnitudes are directly comparable.
3. $scoreFusion became generally available, adding normalized weighted score fusion.
4. Inside every retrieval branch before fusion/reranking.
5. Latency distributions, freshness, zero-result behavior, security leakage, cost, and query-class segmentation.
Authoritative references
- Self-Managed MongoDB Search and Vector Search — mongot architecture, Community/Enterprise deployment paths, and Atlas-managed responsibility boundary.
- mongot Compatibility and Requirements — MongoDB 8.3 / mongot 1.70.1 compatibility, platforms, and topology requirements.
- Local Development Quickstart — Atlas Local evaluation topology and search/vector index workflow.
- Self-Managed mongot Release Notes — mongot 1.70.1 GA baseline for Community and Enterprise self-managed search.
- Verify mongot Connection — health/readiness endpoints, index state, and end-to-end verification.
- Troubleshoot Self-Managed mongot — replication lag, stale indexes, rebuilds, disk pressure, and query failures.
- mongot Metrics Reference — per-index size, status, replication lag, and resource metrics.
- MongoDB Search Queries and Indexes — dynamic/static mappings, analyzers, query/index relationship.
- MongoDB Search text Operator — analyzed text matching, options, and scoring semantics.
- MongoDB Search compound Operator — must/should/filter/mustNot clauses and score-neutral filtering.
- Search Highlighting — highlight metadata and index requirements.
- Search Faceting — facet collector and current token-oriented string faceting guidance.
- Vector Search Index Fields — vector/filter fields, dimensions, similarity, and quantization options.
- MongoDB Search vectorSearch Operator — ANN/ENN, numCandidates, filters, dimensions, and vector scoring.
- Measure Vector Search Accuracy — ENN judgement lists, ANN recall/overlap, numCandidates tuning, and reranking.
- $rankFusion — reciprocal-rank fusion for multi-pipeline retrieval on MongoDB 8.0+.
- $scoreFusion — score normalization and weighted fusion, GA in MongoDB 8.3.
- getSearchIndexes() — search/vector index lifecycle states including READY, BUILDING, and STALE.
- MongoDB 8.3 Release Notes — current server line, 8.3.8 patch baseline, and scoreFusion GA.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17 driver baseline.