Fuse lexical and semantic signals only after preserving filters and evaluating ranking, recall, latency, and leakage.

Hybrid Text + Vector Retrieval, Metadata Filtering, Reranking, and Evaluation

Model coordinates correctly before asking the database a spatial question.

Advanced120–190 minutesHybrid retrieval labMongoDB 8.3.8 · mongot 1.70.1+ · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Explain why lexical and semantic retrieval fail on different query classes and why hybrid search can improve robustness.

02

Build text and vector retrieval branches with identical tenant/visibility filtering.

03

Use rank fusion and the MongoDB 8.3 scoreFusion option while understanding their different score assumptions.

04

Evaluate precision-style judgement lists, recall/overlap, latency, and filter leakage rather than relying on a few screenshots.

05

Diagnose fusion weight/ranking mistakes and preserve a reranking/evaluation loop.

Reproducible lab baseline

This lesson pins MongoDB Server 8.3.8 through the evaluation image mongodb/mongodb-atlas-local:8.3.8-20260827T164726Z, mongosh 2.10.0, and PyMongo 4.17.0 where client measurement is useful. The current self-managed Search GA line is mongot 1.70.1 or later for MongoDB 8.3; production Community and Enterprise deployments operate mongot as a separate process, while Atlas manages it. The mandatory lab uses a disposable single-node replica-set development image on loopback port 27173 named atlasmart-ch21-l4. It is an evaluation topology, not high availability. The local credential is synthetic and must never be reused outside this lab. FCV is observed and never changed. Atlas cloud, KMS, external embedding APIs, and paid search nodes are not required. Automated Embedding is not used; vectors are deterministic teaching data so no model key or network call is needed. Product labs were not executed in this generation environment; index build time, search scores, recall, freshness lag, and latency must be measured locally rather than copied as invented output. MongoDB 8.3 makes $scoreFusion generally available; $rankFusion is GA from MongoDB 8.0. The lesson uses rank-based fusion as the more score-scale-agnostic baseline and shows score fusion as an 8.3-specific alternative.

1. Lexical and semantic signals fail differently

Exact model names, SKUs, acronyms, and freshly coined product terms often favor lexical search because the words themselves matter. Paraphrases such as “battery for charging a laptop on flights” may favor semantic vectors even when the exact terms differ. Hybrid retrieval keeps both signals and combines their ranked evidence.

A hybrid system is not automatically better. If either branch leaks unauthorized documents, has poor recall, or dominates with badly calibrated weights, fusion can amplify the defect.

2. Build both indexes and verify readiness

start the pinned local Search/Vector Search evaluation lab
docker rm -f atlasmart-ch21-l4 2>/dev/null || truedocker volume rm atlasmart-ch21-l4-db atlasmart-ch21-l4-config atlasmart-ch21-l4-search 2>/dev/null || truedocker run -d --name atlasmart-ch21-l4 \  -p 127.0.0.1:27173:27017 \  -e MONGODB_INITDB_ROOT_USERNAME=atlaslab \  -e MONGODB_INITDB_ROOT_PASSWORD=local-only-change-me \  -e DO_NOT_TRACK=1 \  -v atlasmart-ch21-l4-db:/data/db \  -v atlasmart-ch21-l4-config:/data/configdb \  -v atlasmart-ch21-l4-search:/data/mongot \  mongodb/mongodb-atlas-local:8.3.8-20260827T164726Zuntil [ "$(docker inspect -f '{{.State.Health.Status}}' atlasmart-ch21-l4 2>/dev/null)" = "healthy" ]; do sleep 2; donemongosh "mongodb://atlaslab:local-only-change-me@127.0.0.1:27173/admin?directConnection=true" --eval 'db.runCommand({ping:1})' 
seed the shared corpus and create lexical/vector indexes
db = db.getSiblingDB("atlasmart");const c = db.catalog_search_ch21;c.drop();c.insertMany([ {_id:"P01",tenantId:"tenant-a",active:true, category:"power", title:"Portable USB-C Power Bank", description:"Compact travel battery with fast USB-C charging", embedding:[0.96,0.10,0.05,0.02], priceCents:4900}, {_id:"P02",tenantId:"tenant-a",active:true, category:"power", title:"65W USB-C Travel Charger", description:"GaN wall charger for laptops phones and travel", embedding:[0.92,0.16,0.03,0.02], priceCents:5900}, {_id:"P03",tenantId:"tenant-a",active:true, category:"cables", title:"Braided USB-C Cable", description:"Two meter durable charging and data cable", embedding:[0.82,0.20,0.08,0.04], priceCents:1800}, {_id:"P04",tenantId:"tenant-a",active:true, category:"audio", title:"Noise Cancelling Travel Headphones", description:"Over ear headphones for flights and commuting", embedding:[0.08,0.92,0.12,0.03], priceCents:12900}, {_id:"P05",tenantId:"tenant-a",active:true, category:"travel", title:"Universal Travel Adapter", description:"International plug adapter with USB-C ports", embedding:[0.72,0.14,0.55,0.06], priceCents:3900}, {_id:"P06",tenantId:"tenant-a",active:false,category:"power", title:"Legacy Power Brick", description:"Discontinued high capacity portable battery", embedding:[0.90,0.06,0.03,0.02], priceCents:3500}, {_id:"P07",tenantId:"tenant-b",active:true, category:"power", title:"Tenant B Private Charger", description:"Private catalog USB-C charging device", embedding:[0.95,0.11,0.03,0.02], priceCents:5100}, {_id:"P08",tenantId:"tenant-b",active:true, category:"security", title:"Tenant B Security Token", description:"Private authentication hardware token", embedding:[0.05,0.04,0.09,0.97], priceCents:7600}, {_id:"P09",tenantId:"tenant-a",active:true, category:"bags", title:"Laptop Travel Backpack", description:"Carry-on backpack with laptop compartment", embedding:[0.10,0.22,0.91,0.05], priceCents:8900}, {_id:"P10",tenantId:"tenant-a",active:true, category:"power", title:"Wireless Charging Pad", description:"Desk charger for Qi compatible phones", embedding:[0.78,0.14,0.05,0.03], priceCents:3200}, {_id:"P11",tenantId:"tenant-a",active:true, category:"audio", title:"USB-C Earbuds", description:"Wired earbuds with USB-C connector", embedding:[0.35,0.82,0.04,0.03], priceCents:2900}, {_id:"P12",tenantId:"tenant-a",active:true, category:"travel", title:"Packing Cube Set", description:"Lightweight organizers for carry-on travel", embedding:[0.08,0.13,0.96,0.02], priceCents:2600}]);print("documents", c.countDocuments({}));c.createSearchIndex("catalog_text", {  mappings:{dynamic:false,fields:{    title:{type:"string",analyzer:"lucene.english"},    description:{type:"string",analyzer:"lucene.english"},    tenantId:{type:"token",normalizer:"lowercase"},    active:{type:"boolean"},    category:{type:"token",normalizer:"lowercase"},    priceCents:{type:"number"}  }}});c.createSearchIndex("catalog_vector", "vectorSearch", {  fields:[    {type:"vector", path:"embedding", numDimensions:4, similarity:"cosine"},    {type:"filter", path:"tenantId"},    {type:"filter", path:"active"},    {type:"filter", path:"category"}  ]});printjson(c.getSearchIndexes());
wait for READY
const wanted = ["catalog_text", "catalog_vector"];for (let attempt=0; attempt<120; attempt++) {  const m = new Map(c.getSearchIndexes().map(x => [x.name,x]));  const ready = wanted.every(n => m.get(n) && m.get(n).status === "READY" && m.get(n).queryable === true);  if (ready) { print("READY", wanted.join(",")); break; }  sleep(1000);}printjson(c.getSearchIndexes());

3. Reciprocal-rank fusion combines order without assuming score scale

$rankFusion combines the position of each document in named input pipelines. This is useful when lexical and vector score magnitudes are not directly comparable. The input pipelines must remain selection/scoring pipelines over the same collection.

hybrid retrieval with tenant filters inside both branches
const q=[0.94,0.12,0.04,0.02];printjson(c.aggregate([ {$rankFusion:{   input:{pipelines:{     lexical:[       {$search:{index:"catalog_text",compound:{         must:[{text:{query:"travel charger",path:["title","description"]}}],         filter:[{equals:{path:"tenantId",value:"tenant-a"}},{equals:{path:"active",value:true}}]       }}}, {$limit:8}     ],     semantic:[       {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,numCandidates:10,limit:8,filter:{tenantId:"tenant-a",active:true}}}     ]   }},   combination:{weights:{lexical:1.2,semantic:1.0}},   scoreDetails:true }}, {$project:{_id:1,title:1,tenantId:1,scoreDetails:{$meta:"scoreDetails"}}}, {$limit:6}]).toArray());

Inspect which products come from one or both branches. Do not couple application logic to the exact internal scoreDetails shape; MongoDB documents that detailed formatting is not a stable contract.

4. MongoDB 8.3 score fusion is useful only after normalization choices are justified

$scoreFusion normalizes scores and combines weighted input scores. It is GA in MongoDB 8.3. This is more expressive than rank fusion, but it also requires you to reason about score distributions and normalization.

8.3 scoreFusion alternative
printjson(c.aggregate([ {$scoreFusion:{   input:{pipelines:{     lexical:[{$search:{index:"catalog_text",compound:{must:[{text:{query:"travel charger",path:["title","description"]}}],filter:[{equals:{path:"tenantId",value:"tenant-a"}},{equals:{path:"active",value:true}}]}}},{$limit:8}],     semantic:[{$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,numCandidates:10,limit:8,filter:{tenantId:"tenant-a",active:true}}}]   },normalization:"sigmoid"},   combination:{weights:{lexical:1.2,semantic:1.0},method:"avg"},   scoreDetails:true }}, {$project:{_id:1,title:1,scoreDetails:{$meta:"scoreDetails"}}}, {$limit:6}]).toArray());

Do not choose rank fusion versus score fusion from a single query. Evaluate a labelled set of query intents: exact-name queries, descriptive queries, ambiguous queries, category queries, and adversarial tenant/security cases.

5. Evaluate relevance and leakage together

deterministic local evaluation harness for judgement lists
from collections import defaultdictjudgements={  "travel charger":{"P02","P05","P01"},  "portable battery":{"P01","P06"},  "usb c audio":{"P11"}}def precision_at_k(returned,relevant,k):    top=returned[:k]    return sum(x in relevant for x in top)/max(1,len(top))def recall_at_k(returned,relevant,k):    return sum(x in relevant for x in returned[:k])/max(1,len(relevant))# Replace these fixture rankings with product-lab output.fixture={"travel charger":["P02","P01","P05","P10"],"portable battery":["P01","P10","P06"],"usb c audio":["P11","P03"]}for q,rel in judgements.items():    r=fixture[q]    print(q,"P@3",precision_at_k(r,rel,3),"R@3",recall_at_k(r,rel,3))

For real production evaluation, add latency percentiles, zero-result rate, click/engagement proxies, freshness probes, and filter leakage tests that assert every returned document belongs to the authorized tenant and visibility state.

6. Wrong hybrid pattern: filter after retrieval

Unsafe design:

Run broad text/vector retrieval across all tenants, fuse/rerank it, then apply $match:{tenantId:"tenant-a"} at the end. Even if final output appears filtered, cross-tenant candidates have already influenced candidate selection, scores, logs, traces, caches, or downstream rerankers. Repair the design by applying the security filter inside every retrieval branch.

Security filtering is an eligibility rule, not a relevance preference. Test it with synthetic “highly relevant but forbidden” documents such as P07; the test should fail if that document is ever visible to the authorized branch.

7. Production judgment

Choose fusion weights from an evaluation suite and segment results by query class. Track score/rank drift after analyzer or embedding changes. If you add an external reranker, bound the candidate set, latency, cost, and data sent outside the database; never send unauthorized candidates to the reranker.

Bridge. Lesson 5 turns retrieval into a production RAG/recommendation service with freshness, security, cost, and observability contracts.

cleanup only this lesson lab
docker rm -f atlasmart-ch21-l4 2>/dev/null || truedocker volume rm atlasmart-ch21-l4-db atlasmart-ch21-l4-config atlasmart-ch21-l4-search 2>/dev/null || true

Check your understanding

  1. Why can lexical and semantic retrieval complement each other?
  2. Why is rank fusion attractive across heterogeneous retrievers?
  3. What changed in MongoDB 8.3?
  4. Where must tenant filtering occur?
  5. What should relevance evaluation include besides P@k/recall?
Review the answers

1. They fail differently: lexical search preserves exact tokens/proper nouns, while vector retrieval can recover paraphrases and semantic similarity.

2. It combines rank positions without assuming the underlying score magnitudes are directly comparable.

3. $scoreFusion became generally available, adding normalized weighted score fusion.

4. Inside every retrieval branch before fusion/reranking.

5. Latency distributions, freshness, zero-result behavior, security leakage, cost, and query-class segmentation.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.