Operate RAG/recommendation retrieval with explicit freshness, security, cost, migration, and observability contracts.
Design RAG/Recommendation Retrieval with Freshness, Security Filters, Cost, and Observability
Model coordinates correctly before asking the database a spatial question.
Learning objectives
Design RAG and recommendation retrieval as an end-to-end production contract rather than a demo query.
Enforce tenant/visibility filters before ranking and test explicitly for forbidden-document leakage.
Measure Search freshness lag and distinguish database commit time from derived-index availability.
Instrument latency, recall/relevance, index status/size, replication lag, cost drivers, and fallback behavior.
Plan idempotent embedding/index migrations, reconciliation, rollback, and graceful degradation.
This lesson pins MongoDB Server 8.3.8 through the
evaluation image
mongodb/mongodb-atlas-local:8.3.8-20260827T164726Z,
mongosh 2.10.0, and PyMongo
4.17.0 where client measurement is useful. The
current self-managed Search GA line is
mongot 1.70.1 or later for MongoDB 8.3; production
Community and Enterprise deployments operate
mongot as a separate process, while Atlas manages
it. The mandatory lab uses a disposable single-node replica-set
development image on loopback port 27174 named
atlasmart-ch21-l5. It is an evaluation topology,
not high availability. The local credential is synthetic and
must never be reused outside this lab.
FCV is observed and never changed. Atlas cloud,
KMS, external embedding APIs, and paid search nodes are not
required. Automated Embedding is not used; vectors are
deterministic teaching data so no model key or network call is
needed. Product labs were not executed in this generation
environment; index build time, search scores, recall, freshness
lag, and latency must be measured locally rather than copied as
invented output. No large language model (LLM),
retrieval-augmented generation (RAG) API, or external
embedding/reranking service is required. The retrieval layer can
be tested locally with deterministic vectors and judgement
lists; generation quality must be evaluated separately if an LLM
is later added.
1. RAG is retrieval plus an application contract, not a magic database mode
Retrieval-augmented generation (RAG) retrieves evidence and gives that evidence to a generative model. A recommendation service similarly retrieves candidates and ranks them for a user/context. MongoDB can maintain lexical and vector indexes, but it does not automatically guarantee that retrieved evidence is authorized, fresh enough, correct, or sufficient for a final model answer.
AtlasMart therefore defines explicit service invariants: only tenant-a/public products are eligible; inactive products must not appear; recently published safety notices must become searchable within an observed freshness objective; and semantic search must meet a measured recall target against a judgement set.
2. Build the retrieval layer and enforce eligibility at source
docker rm -f atlasmart-ch21-l5 2>/dev/null || truedocker volume rm atlasmart-ch21-l5-db atlasmart-ch21-l5-config atlasmart-ch21-l5-search 2>/dev/null || truedocker run -d --name atlasmart-ch21-l5 \ -p 127.0.0.1:27174:27017 \ -e MONGODB_INITDB_ROOT_USERNAME=atlaslab \ -e MONGODB_INITDB_ROOT_PASSWORD=local-only-change-me \ -e DO_NOT_TRACK=1 \ -v atlasmart-ch21-l5-db:/data/db \ -v atlasmart-ch21-l5-config:/data/configdb \ -v atlasmart-ch21-l5-search:/data/mongot \ mongodb/mongodb-atlas-local:8.3.8-20260827T164726Zuntil [ "$(docker inspect -f '{{.State.Health.Status}}' atlasmart-ch21-l5 2>/dev/null)" = "healthy" ]; do sleep 2; donemongosh "mongodb://atlaslab:local-only-change-me@127.0.0.1:27174/admin?directConnection=true" --eval 'db.runCommand({ping:1})'
db = db.getSiblingDB("atlasmart");const c = db.catalog_search_ch21;c.drop();c.insertMany([ {_id:"P01",tenantId:"tenant-a",active:true, category:"power", title:"Portable USB-C Power Bank", description:"Compact travel battery with fast USB-C charging", embedding:[0.96,0.10,0.05,0.02], priceCents:4900}, {_id:"P02",tenantId:"tenant-a",active:true, category:"power", title:"65W USB-C Travel Charger", description:"GaN wall charger for laptops phones and travel", embedding:[0.92,0.16,0.03,0.02], priceCents:5900}, {_id:"P03",tenantId:"tenant-a",active:true, category:"cables", title:"Braided USB-C Cable", description:"Two meter durable charging and data cable", embedding:[0.82,0.20,0.08,0.04], priceCents:1800}, {_id:"P04",tenantId:"tenant-a",active:true, category:"audio", title:"Noise Cancelling Travel Headphones", description:"Over ear headphones for flights and commuting", embedding:[0.08,0.92,0.12,0.03], priceCents:12900}, {_id:"P05",tenantId:"tenant-a",active:true, category:"travel", title:"Universal Travel Adapter", description:"International plug adapter with USB-C ports", embedding:[0.72,0.14,0.55,0.06], priceCents:3900}, {_id:"P06",tenantId:"tenant-a",active:false,category:"power", title:"Legacy Power Brick", description:"Discontinued high capacity portable battery", embedding:[0.90,0.06,0.03,0.02], priceCents:3500}, {_id:"P07",tenantId:"tenant-b",active:true, category:"power", title:"Tenant B Private Charger", description:"Private catalog USB-C charging device", embedding:[0.95,0.11,0.03,0.02], priceCents:5100}, {_id:"P08",tenantId:"tenant-b",active:true, category:"security", title:"Tenant B Security Token", description:"Private authentication hardware token", embedding:[0.05,0.04,0.09,0.97], priceCents:7600}, {_id:"P09",tenantId:"tenant-a",active:true, category:"bags", title:"Laptop Travel Backpack", description:"Carry-on backpack with laptop compartment", embedding:[0.10,0.22,0.91,0.05], priceCents:8900}, {_id:"P10",tenantId:"tenant-a",active:true, category:"power", title:"Wireless Charging Pad", description:"Desk charger for Qi compatible phones", embedding:[0.78,0.14,0.05,0.03], priceCents:3200}, {_id:"P11",tenantId:"tenant-a",active:true, category:"audio", title:"USB-C Earbuds", description:"Wired earbuds with USB-C connector", embedding:[0.35,0.82,0.04,0.03], priceCents:2900}, {_id:"P12",tenantId:"tenant-a",active:true, category:"travel", title:"Packing Cube Set", description:"Lightweight organizers for carry-on travel", embedding:[0.08,0.13,0.96,0.02], priceCents:2600}]);print("documents", c.countDocuments({}));c.createSearchIndex("catalog_text", { mappings:{dynamic:false,fields:{ title:{type:"string",analyzer:"lucene.english"}, description:{type:"string",analyzer:"lucene.english"}, tenantId:{type:"token",normalizer:"lowercase"}, active:{type:"boolean"}, category:{type:"token",normalizer:"lowercase"}, priceCents:{type:"number"} }}});c.createSearchIndex("catalog_vector", "vectorSearch", { fields:[ {type:"vector", path:"embedding", numDimensions:4, similarity:"cosine"}, {type:"filter", path:"tenantId"}, {type:"filter", path:"active"}, {type:"filter", path:"category"} ]});printjson(c.getSearchIndexes());
const wanted = ["catalog_text", "catalog_vector"];for (let attempt=0; attempt<120; attempt++) { const m = new Map(c.getSearchIndexes().map(x => [x.name,x])); const ready = wanted.every(n => m.get(n) && m.get(n).status === "READY" && m.get(n).queryable === true); if (ready) { print("READY", wanted.join(",")); break; } sleep(1000);}printjson(c.getSearchIndexes());
const q=[0.94,0.12,0.04,0.02];const safe=c.aggregate([ {$vectorSearch:{index:"catalog_vector",path:"embedding",queryVector:q,numCandidates:12,limit:8,filter:{tenantId:"tenant-a",active:true}}}, {$project:{_id:1,tenantId:1,active:1,title:1}}]).toArray();printjson(safe);if (safe.some(x=>x.tenantId!=="tenant-a" || x.active!==true)) throw new Error("SECURITY FILTER LEAK");if (safe.some(x=>x._id==="P07")) throw new Error("forbidden tenant-b candidate leaked");print("security filter test: PASS");
3. Freshness is a measured lag between source commit and search visibility
mongot synchronizes from MongoDB using change
streams. That means the source document can exist before a
search query sees it. Do not promise “read-your-writes through
Search” merely because the MongoDB insert succeeded. Define and
measure an acceptable freshness objective, and provide a
source-of-truth fallback for workflows that require immediate
visibility.
const marker="fresh-"+new Date().toISOString();const committedAt=Date.now();c.insertOne({_id:marker,tenantId:"tenant-a",active:true,category:"power",title:"Fresh Safety Charger Notice",description:marker,embedding:[0.91,0.12,0.04,0.02],priceCents:1});let seenAt=null;for (let i=0;i<120;i++) { const r=c.aggregate([ {$search:{index:"catalog_text",compound:{must:[{text:{query:marker,path:"description"}}],filter:[{equals:{path:"tenantId",value:"tenant-a"}}]}}}, {$limit:1} ]).toArray(); if (r.length) { seenAt=Date.now(); break; } sleep(250);}printjson({marker,committedAt,seenAt,freshnessLagMs:seenAt===null?null:seenAt-committedAt});if (seenAt===null) print("not visible within probe window; inspect index status/lag instead of inventing success");
Repeat this probe under normal and loaded conditions. Correlate
large lag with getSearchIndexes(),
STALE/rebuild state, and self-managed
mongot replication-lag metrics.
4. Production observability spans correctness, quality, latency, and cost
| Signal | Question answered | Failure if ignored |
|---|---|---|
| Index status / queryable | Can the selected generation serve? | Queries fail or serve stale generation |
| mongot replication lag | How far is derived Search behind source? | Freshness SLO silently violated |
| P50/P95/P99 query latency | What do users experience at the tail? | Average hides overload |
| Recall@k / judgement quality | Does ANN/retrieval find relevant evidence? | Fast but wrong RAG |
| Leakage tests | Can forbidden documents enter candidates? | Cross-tenant/security incident |
| Index bytes / memory / CPU | What does retrieval cost to operate? | OOM, disk pressure, rebuild failure |
| Embedding version coverage | Are old/new vector spaces mixed? | Nonsensical similarity during migration |
| Freshness probe | How quickly do committed changes become searchable? | Stale recommendations/safety content |
For self-managed Search, monitor mongot status,
Lucene index size, disk headroom, memory, query errors, and
change-stream replication. Atlas exposes managed Search metrics
instead, but the responsibility to define
relevance/security/freshness objectives remains with the
application team.
5. Cost model and graceful degradation
Cost grows with indexed fields, analyzer complexity, document volume, vector count, dimensions, quantization choice, candidate count, concurrency, replication/rebuild work, and dedicated Search-node capacity. External embedding and reranking APIs add request and egress costs as well as privacy boundaries. This lab intentionally avoids them.
A resilient application defines fallbacks. If semantic search is degraded, lexical search may still be acceptable for some queries. If Search is stale and an invariant requires immediate correctness, query the source collection by an ordinary indexed identifier. If both retrieval paths fail, return a controlled “temporarily unavailable” result rather than hallucinating evidence.
6. Migration and rollback: embedding changes are data migrations
When replacing an embedding model, write the model/version into each document or embedding record. Backfill a new vector field, build a separate vector index, evaluate old versus new, then cut traffic gradually. Do not mix embeddings from incompatible spaces in one index and hope scores remain meaningful. Keep the old field/index until rollback criteria expire.
The same applies to analyzer/mapping changes: new definition, build state, judgement suite, security tests, controlled cutover, rollback definition. Search indexes are derived and rebuildable, but rebuild time and source/oplog history must be part of the capacity plan.
7. Retrieval is not an authorization or truth engine
Retrieve across all tenants, send top candidates to an external reranker/LLM, then remove unauthorized documents before rendering. This has already disclosed data to a downstream system. Repair by applying authorization-compatible metadata filters inside every search/vector branch, validating the returned candidate set before any external call, and restricting indexed sensitive fields.
CDC-style freshness also does not transform Search into a domain-event bus. Search synchronization is for derived retrieval state, not for exactly-once business side effects.
8. Production judgment and next bridge
A production RAG/recommendation architecture needs four
independent test suites: retrieval relevance/recall, security
leakage, freshness/reconciliation, and latency/cost under
concurrency. Add generation evaluation only after retrieval
evidence is reliable. Record server/mongot/driver
versions and index definitions alongside every benchmark.
Bridge. Chapter 22 moves from retrieval correctness to the database security boundary itself: authentication, authorization, TLS, network exposure, and auditing.
docker rm -f atlasmart-ch21-l5 2>/dev/null || truedocker volume rm atlasmart-ch21-l5-db atlasmart-ch21-l5-config atlasmart-ch21-l5-search 2>/dev/null || true
Check your understanding
- Why is Search freshness not identical to MongoDB write acknowledgement?
- Where should tenant/visibility filtering happen in RAG retrieval?
- What must be versioned during an embedding migration?
- What four evaluation dimensions should exist before adding generation quality?
- What is a safe fallback when immediate correctness matters but Search is stale?
Review the answers
1. Search is a separately maintained derived index synchronized from changes, so visibility can lag the source commit.
2. Inside each retrieval branch before fusion, reranking, caching, logging, or any external model call.
3. At minimum the embedding model/version and vector field/index definition; incompatible spaces should not be mixed as if comparable.
4. Retrieval relevance/recall, security leakage, freshness/reconciliation, and latency/cost under representative load.
5. Use an ordinary source-of-truth indexed read for that invariant-sensitive path, or fail closed/controlled rather than fabricate evidence.
Authoritative references
- Self-Managed MongoDB Search and Vector Search — mongot architecture, Community/Enterprise deployment paths, and Atlas-managed responsibility boundary.
- mongot Compatibility and Requirements — MongoDB 8.3 / mongot 1.70.1 compatibility, platforms, and topology requirements.
- Local Development Quickstart — Atlas Local evaluation topology and search/vector index workflow.
- Self-Managed mongot Release Notes — mongot 1.70.1 GA baseline for Community and Enterprise self-managed search.
- Verify mongot Connection — health/readiness endpoints, index state, and end-to-end verification.
- Troubleshoot Self-Managed mongot — replication lag, stale indexes, rebuilds, disk pressure, and query failures.
- mongot Metrics Reference — per-index size, status, replication lag, and resource metrics.
- MongoDB Search Queries and Indexes — dynamic/static mappings, analyzers, query/index relationship.
- MongoDB Search text Operator — analyzed text matching, options, and scoring semantics.
- MongoDB Search compound Operator — must/should/filter/mustNot clauses and score-neutral filtering.
- Search Highlighting — highlight metadata and index requirements.
- Search Faceting — facet collector and current token-oriented string faceting guidance.
- Vector Search Index Fields — vector/filter fields, dimensions, similarity, and quantization options.
- MongoDB Search vectorSearch Operator — ANN/ENN, numCandidates, filters, dimensions, and vector scoring.
- Measure Vector Search Accuracy — ENN judgement lists, ANN recall/overlap, numCandidates tuning, and reranking.
- $rankFusion — reciprocal-rank fusion for multi-pipeline retrieval on MongoDB 8.0+.
- $scoreFusion — score normalization and weighted fusion, GA in MongoDB 8.3.
- getSearchIndexes() — search/vector index lifecycle states including READY, BUILDING, and STALE.
- MongoDB 8.3 Release Notes — current server line, 8.3.8 patch baseline, and scoreFusion GA.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17 driver baseline.