Add semantic/RAG capability without losing score meaning, provenance, or tenant isolation.
Add Vector/Hybrid/RAG Retrieval with Explicit Quality Benchmarks and Tenant/Security Filters
Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.
Learning outcomes
Extend AtlasMart with deterministic dense-vector, lexical, hybrid, and RAG retrieval without confusing score scales.
Benchmark exact ground truth, ANN recall, lexical quality, fused ranking, and RAG retrieval independently from answer generation.
Apply tenant/access filters before protected content can become candidates or model context.
Design deterministic local fallbacks for model/inference failure rather than making paid/external inference mandatory.
Produce retrieval traces and provenance evidence that can explain every chunk placed into a generated-answer context.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: semantic convenience can violate both relevance and tenancy
AtlasMart wants natural-language product help and RAG-assisted support. The unsafe shortcut is to retrieve global semantic candidates first and remove unauthorized tenants afterward. A vector hit has already crossed the security boundary by the time it reaches the application or model. The capstone therefore treats tenant/security filtering as part of candidate generation, not presentation filtering.
2. Deterministic embedding contract
The mandatory lab keeps the Chapter 25/26 teaching contract: small precomputed eight-dimensional vectors with an immutable model identifier. A real embedding model can be added only after its dimensions, normalization, similarity, version, provider/license, privacy boundary, and reindex policy are recorded. Changing models without versioning produces vectors whose geometric meaning is no longer comparable.
{
"sku": "P-1001",
"tenant_id": "tenant-a",
"chunk_id": "P-1001#overview",
"source_uri": "atlasmart://catalog/P-1001",
"source_version": "catalog-v1",
"embedding_model": "atlasmart-precomputed-8d-v1",
"embedding_dims": 8,
"embedding_normalized": true,
"embedding": [0.12,-0.05,0.44,0.18,0.03,-0.21,0.09,0.31]
}
3. Evaluate retrieval components before fusion
| Retriever | Evidence | Do not infer |
|---|---|---|
| BM25 / lexical | judged top-k, NDCG/MRR-like metric, p95/p99 | that a high BM25 score is comparable to cosine similarity |
| exact vector | ground-truth nearest neighbors on small fixture | that brute-force cost scales acceptably to production |
| ANN vector | Recall@k versus exact ground truth + latency/memory | that ANN results are “true” because the same ANN index agrees with itself |
| hybrid fusion | component ranks + declared fusion method + judged metric | that raw score addition is calibrated |
| reranker | candidate-set size, lift, incremental latency/cost | that reranking every document is acceptable |
| RAG context | chunk IDs, coverage/precision, citations | that fluent answer text proves factual grounding |
4. Portable hybrid baseline: fuse ranks at the application boundary
Both products have native hybrid capabilities, but their APIs and scoring/normalization features are not identical. For the mandatory cross-platform learning path, retrieve lexical and vector candidate lists separately, then use a deterministic rank-fusion function in the lab harness. Native Elastic/OpenSearch hybrid features can be compared as optional platform-specific implementations.
def rrf(lists, k=60):
score = {}
for ranked_ids in lists:
for rank, doc_id in enumerate(ranked_ids, start=1):
score[doc_id] = score.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(score.items(), key=lambda kv: (-kv[1], kv[0]))
# Never add BM25 and vector scores directly unless you have an explicit,
# validated calibration model. Rank fusion avoids pretending the units match.
5. Tenant/security filter is part of retrieval
{
"tenant": "tenant-a",
"query_text": "waterproof hiking footwear for winter",
"retrieval": {
"lexical_filter": {"tenant_id": "tenant-a"},
"vector_filter": {"tenant_id": "tenant-a"},
"top_k_each": 25,
"fusion": "rrf",
"rerank_top_n": 10,
"context_top_n": 4
}
}
Whether the filter is expressed as Elasticsearch DLS, an explicit Query DSL filter, OpenSearch DLS, or a k-NN/neural filter depends on the product and license/plugin surface. The invariant is stronger: unauthorized documents must not enter the candidate pool. Add negative tests using a document that would be semantically perfect for the query but belongs to another tenant.
6. RAG context packer and provenance
{
"query_id": "q-017",
"tenant_id": "tenant-a",
"retrieval_version": "atlasmart-retrieval-v1",
"chunks": [
{"chunk_id":"P-1001#overview", "source_uri":"atlasmart://catalog/P-1001", "rank":1},
{"chunk_id":"KB-0042#returns", "source_uri":"atlasmart://kb/KB-0042", "rank":2}
],
"generation": {
"provider": "optional-local-or-external",
"model": "<record-if-used>",
"fallback": "return retrieved passages with citations"
}
}
If a model is unavailable, rate-limited, or disallowed by data-governance policy, a retrieval-only answer with citations is an acceptable deterministic fallback. The mandatory lab therefore does not require an external paid model.
7. Wrong approach: judge only the final answer
8. Production judgment
Accept semantic/hybrid/RAG changes only when they beat or complement the lexical baseline on judged queries without violating latency/cost, tenant isolation, freshness, and failure behavior. Lesson 4 now treats the entire platform as an operational system: load, profile, security, restore, failover, upgrade, monitoring, and incident drills.
Check your understanding
- Why must tenant filtering occur before or during candidate generation?
- Why keep lexical and semantic baselines?
- Why use exact vector ground truth on a small fixture?
- Why is raw BM25 + vector score addition unsafe by default?
- What is the mandatory RAG fallback?
Review the answers
1. Post-filtering can allow protected documents to enter application/model context, so the confidentiality boundary has already failed.
2. They show whether hybrid complexity creates measurable quality lift and provide fallback paths when inference/vector features fail.
3. It provides an independent reference for ANN recall instead of benchmarking an approximate index against itself.
4. The scores have different semantics/scales and are not automatically calibrated.
5. Return authorized retrieved passages with provenance/citations when generation is unavailable or not permitted.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Download Elasticsearch 9.5.3
- Elasticsearch mappings
- Elasticsearch text analysis
- Elasticsearch index templates
- Elasticsearch ingest pipelines
- Elasticsearch data streams
- Elasticsearch ILM
- Elasticsearch Query DSL
- Elasticsearch aggregations
- Elasticsearch vector search
- Elasticsearch hybrid search
- Elasticsearch security
- Elasticsearch snapshot and restore
- Elasticsearch performance guidance
- Elasticsearch subscription feature matrix
- OpenSearch 3.8 version history
- OpenSearch downloads and Apache 2.0 licensing
- OpenSearch mappings and field types
- OpenSearch index templates
- OpenSearch data streams
- OpenSearch ingest pipelines
- OpenSearch Index State Management
- OpenSearch query DSL
- OpenSearch aggregations
- OpenSearch vector search
- OpenSearch hybrid search
- OpenSearch Security plugin
- OpenSearch snapshot and restore
- OpenSearch performance tuning
- OpenSearch Benchmark