Add semantic/RAG capability without losing score meaning, provenance, or tenant isolation.

Add Vector/Hybrid/RAG Retrieval with Explicit Quality Benchmarks and Tenant/Security Filters

Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.

Intermediate → Advanced190–260 minutesVector, hybrid, RAG & tenant security · Chapter 31 · Lesson 03Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · free/local mandatory pathLast reviewed: September 2026

Learning outcomes

01

Extend AtlasMart with deterministic dense-vector, lexical, hybrid, and RAG retrieval without confusing score scales.

02

Benchmark exact ground truth, ANN recall, lexical quality, fused ranking, and RAG retrieval independently from answer generation.

03

Apply tenant/access filters before protected content can become candidates or model context.

04

Design deterministic local fallbacks for model/inference failure rather than making paid/external inference mandatory.

05

Produce retrieval traces and provenance evidence that can explain every chunk placed into a generated-answer context.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Capstone baseline. Examples are frozen to Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), with their bundled JVMs. The established disposable local endpoints remain Elasticsearch at https://localhost:9200 and OpenSearch at https://localhost:9201 on the atlasmart-search Docker network. The generation environment did not execute live clusters, so no latency, throughput, relevance, restore-time, or cost number is presented as measured unless the learner records it.

1. AtlasMart problem: semantic convenience can violate both relevance and tenancy

AtlasMart wants natural-language product help and RAG-assisted support. The unsafe shortcut is to retrieve global semantic candidates first and remove unauthorized tenants afterward. A vector hit has already crossed the security boundary by the time it reaches the application or model. The capstone therefore treats tenant/security filtering as part of candidate generation, not presentation filtering.

2. Deterministic embedding contract

The mandatory lab keeps the Chapter 25/26 teaching contract: small precomputed eight-dimensional vectors with an immutable model identifier. A real embedding model can be added only after its dimensions, normalization, similarity, version, provider/license, privacy boundary, and reindex policy are recorded. Changing models without versioning produces vectors whose geometric meaning is no longer comparable.

Vector provenance fields
{
  "sku": "P-1001",
  "tenant_id": "tenant-a",
  "chunk_id": "P-1001#overview",
  "source_uri": "atlasmart://catalog/P-1001",
  "source_version": "catalog-v1",
  "embedding_model": "atlasmart-precomputed-8d-v1",
  "embedding_dims": 8,
  "embedding_normalized": true,
  "embedding": [0.12,-0.05,0.44,0.18,0.03,-0.21,0.09,0.31]
}

3. Evaluate retrieval components before fusion

Retriever Evidence Do not infer
BM25 / lexical judged top-k, NDCG/MRR-like metric, p95/p99 that a high BM25 score is comparable to cosine similarity
exact vector ground-truth nearest neighbors on small fixture that brute-force cost scales acceptably to production
ANN vector Recall@k versus exact ground truth + latency/memory that ANN results are “true” because the same ANN index agrees with itself
hybrid fusion component ranks + declared fusion method + judged metric that raw score addition is calibrated
reranker candidate-set size, lift, incremental latency/cost that reranking every document is acceptable
RAG context chunk IDs, coverage/precision, citations that fluent answer text proves factual grounding

4. Portable hybrid baseline: fuse ranks at the application boundary

Both products have native hybrid capabilities, but their APIs and scoring/normalization features are not identical. For the mandatory cross-platform learning path, retrieve lexical and vector candidate lists separately, then use a deterministic rank-fusion function in the lab harness. Native Elastic/OpenSearch hybrid features can be compared as optional platform-specific implementations.

Deterministic reciprocal-rank fusion
def rrf(lists, k=60):
    score = {}
    for ranked_ids in lists:
        for rank, doc_id in enumerate(ranked_ids, start=1):
            score[doc_id] = score.get(doc_id, 0.0) + 1.0 / (k + rank)
    return sorted(score.items(), key=lambda kv: (-kv[1], kv[0]))

# Never add BM25 and vector scores directly unless you have an explicit,
# validated calibration model. Rank fusion avoids pretending the units match.

5. Tenant/security filter is part of retrieval

Protected candidate request shape
{
  "tenant": "tenant-a",
  "query_text": "waterproof hiking footwear for winter",
  "retrieval": {
    "lexical_filter": {"tenant_id": "tenant-a"},
    "vector_filter":  {"tenant_id": "tenant-a"},
    "top_k_each": 25,
    "fusion": "rrf",
    "rerank_top_n": 10,
    "context_top_n": 4
  }
}

Whether the filter is expressed as Elasticsearch DLS, an explicit Query DSL filter, OpenSearch DLS, or a k-NN/neural filter depends on the product and license/plugin surface. The invariant is stronger: unauthorized documents must not enter the candidate pool. Add negative tests using a document that would be semantically perfect for the query but belongs to another tenant.

6. RAG context packer and provenance

Context packer output
{
  "query_id": "q-017",
  "tenant_id": "tenant-a",
  "retrieval_version": "atlasmart-retrieval-v1",
  "chunks": [
    {"chunk_id":"P-1001#overview", "source_uri":"atlasmart://catalog/P-1001", "rank":1},
    {"chunk_id":"KB-0042#returns", "source_uri":"atlasmart://kb/KB-0042", "rank":2}
  ],
  "generation": {
    "provider": "optional-local-or-external",
    "model": "<record-if-used>",
    "fallback": "return retrieved passages with citations"
  }
}

If a model is unavailable, rate-limited, or disallowed by data-governance policy, a retrieval-only answer with citations is an acceptable deterministic fallback. The mandatory lab therefore does not require an external paid model.

7. Wrong approach: judge only the final answer

Wrong approach: declare RAG successful because a generated answer sounds plausible. Failure: a weak retriever, leaked tenant chunk, missing citation, or hallucinated answer can be hidden by fluent prose. Repair: score retrieval independently (Recall/NDCG/MRR-like), record context precision/coverage and provenance, run security-negative tests, then evaluate answer faithfulness/citation use separately.

8. Production judgment

Accept semantic/hybrid/RAG changes only when they beat or complement the lexical baseline on judged queries without violating latency/cost, tenant isolation, freshness, and failure behavior. Lesson 4 now treats the entire platform as an operational system: load, profile, security, restore, failover, upgrade, monitoring, and incident drills.

Check your understanding

  1. Why must tenant filtering occur before or during candidate generation?
  2. Why keep lexical and semantic baselines?
  3. Why use exact vector ground truth on a small fixture?
  4. Why is raw BM25 + vector score addition unsafe by default?
  5. What is the mandatory RAG fallback?
Review the answers

1. Post-filtering can allow protected documents to enter application/model context, so the confidentiality boundary has already failed.

2. They show whether hybrid complexity creates measurable quality lift and provide fallback paths when inference/vector features fail.

3. It provides an independent reference for ANN recall instead of benchmarking an approximate index against itself.

4. The scores have different semantics/scales and are not automatically calibrated.

5. Return authorized retrieved passages with provenance/citations when generation is unavailable or not permitted.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.