Fuse and rerank only with declared math, candidate budgets, judged relevance, and end-to-end latency evidence.
Hybrid Score Normalization, Weighted Combination, Reciprocal-Rank-Like Fusion, Re-Ranking, and Search Quality Evaluation
Combine lexical, dense, sparse, semantic, neural, and reranking signals without losing score meaning, evaluation discipline, or operational control across the two platforms.
Learning outcomes
Explain why raw lexical, dense, and sparse scores cannot be summed safely without calibration.
Implement and compare score normalization/weighted fusion with RRF-style rank fusion.
Use reranking as a bounded second stage and preserve first-stage diagnostics.
Compute judged Recall/NDCG/MRR-like metrics independently of generated answers.
Treat latency/cost and relevance as a joint decision surface rather than optimizing one metric.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
semantic_text usually invokes the Inference API and
can therefore carry license/service requirements; OpenSearch ML
Commons can run supported local models or connectors, but
model/hardware support remains deployment-specific.
1. AtlasMart problem: the “best” hybrid system depends on what you preserve
RRF preserves rank agreement and ignores score magnitude. Normalized weighted fusion preserves some score-margin information but depends on normalization and candidate windows. Reranking can model query-document interaction more deeply, but it adds model cost and may reorder a good first-stage ranking badly if the reranker is mismatched. There is no universally correct fusion rule.
2. Three fusion families
| Method | Uses | Strength | Failure mode |
|---|---|---|---|
| raw weighted sum | un-normalized scores | simple | meaningless when component scales differ |
| normalized weighted combination | normalized component scores | retains score margins and supports business weights | normalizer/candidate outliers can change contributions |
| RRF / reciprocal-rank-like | rank positions | robust to score-scale mismatch | throws away score margins and depends on candidate recall |
3. Deterministic AtlasMart retrieval fixture
Carry forward Chapter 25's eight product vectors and
tenant/category metadata. These vectors are a transparent
mechanics fixture, not outputs claimed from a production
embedding model. Every semantic experiment records a model
contract such as atlasmart-product-v1, dimensions,
normalization, source text, and tenant metadata.
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]
Security invariant: tenant/category eligibility is applied before a result is allowed into evaluation or reranking. A semantically close document outside the caller's authorization scope is not a candidate.
4. Deterministic RRF example
def rrf(rank_lists, k=60):
out = {}
for ranking in rank_lists:
for rank, doc_id in enumerate(ranking, start=1):
out[doc_id] = out.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(out.items(), key=lambda x: (-x[1], x[0]))
lexical = ["P-1002", "P-1001", "P-1005", "P-1006"]
semantic = ["P-1005", "P-1001", "P-1002", "P-1006"]
print(rrf([lexical, semantic]))
Record the rank constant and each child ranking. Do not publish only the fused list: debugging requires the components. Elastic's RRF retriever and OpenSearch's score-ranker processor implement product-native rank-fusion surfaces, but the client-side function is a portable oracle for this tiny fixture.
5. Score normalization changes ranking
OpenSearch documents min-max, L2, and z-score normalization for
the normalization processor. Different techniques can produce
different winners from the same component scores. Elastic's
linear retriever similarly normalizes before combining.
Therefore weights are not portable numbers:
0.7 semantic in one product/configuration is not
automatically equivalent to 0.7 in the other.
6. Rerank only candidates that deserve the budget
Start with a first-stage candidate set large enough to contain relevant documents, then rerank a bounded top-N. Measure: first-stage Recall@N; reranker NDCG/MRR at the display cutoff; rerank p95/p99; error/fallback rate; and cost per query if an external service is involved. Candidate N is a capacity parameter, not merely a relevance knob.
7. Judged evaluation—not generated-answer vibes
For search quality, a judgment set maps
(query, document) to relevance grades.
Recall@k tests whether relevant candidates are
retrieved. MRR rewards the position of the
first relevant result. NDCG@k rewards graded
relevance near the top and normalizes against the ideal ranking.
These metrics evaluate retrieval/ranking directly; Chapter 27
will separately evaluate generated answers.
import math, statistics, time
def dcg(ids, qrels, k=5):
return sum((2**qrels.get(doc,0)-1)/math.log2(i+2) for i,doc in enumerate(ids[:k]))
def ndcg(ids, qrels, k=5):
ideal=[d for d,_ in sorted(qrels.items(), key=lambda x:(-x[1],x[0]))]
denom=dcg(ideal,qrels,k)
return dcg(ids,qrels,k)/denom if denom else 0.0
def mrr(ids, qrels):
for i,doc in enumerate(ids,1):
if qrels.get(doc,0)>0: return 1/i
return 0.0
# Replace rankings with captured IDs from each interface; never compare raw BM25/vector scores directly.
qrels={"P-1005":3,"P-1001":2,"P-1002":2,"P-1006":1}
for name,ids in {
"lexical":["P-1002","P-1001","P-1005","P-1006"],
"semantic":["P-1005","P-1001","P-1002","P-1006"],
"hybrid":["P-1005","P-1002","P-1001","P-1006"],
}.items():
print(name, "NDCG@4=", round(ndcg(ids,qrels,4),4), "MRR=", round(mrr(ids,qrels),4))
8. Latency evidence
Measure end-to-end latency at the client boundary and decompose it into query embedding/inference, first-stage search, fusion, fetch, reranking, and network where possible. Report p50/p95/p99 with sample count and concurrency. Do not compare Profile API timings to normal request latency: profiling adds overhead and is diagnostic, not a benchmark.
9. Wrong approach: optimize only hybrid NDCG
A 1% quality lift that doubles p99, triples external-model cost, or causes frequent fallbacks may be a poor production trade. Conversely, the lowest-latency path may violate relevance objectives. Define a release gate with both quality and operational budgets, then compare deltas against the same query/judgment set.
10. Decision record template
experiment: atlasmart-hybrid-2026-09-a
corpus_version: catalog-fixture-v1
judgments_version: qrels-v1
embedding_contract: atlasmart-product-v1
lexical_config: bm25-default
semantic_config: dense-precomputed-v1
fusion: rrf
rank_constant: 60
candidate_windows: {lexical: 50, semantic: 50, rerank: 20}
filters: tenant-before-fusion
metrics: [Recall@20, NDCG@10, MRR, p95_ms, p99_ms, fallback_rate]
notes: synthetic mechanics fixture; production numbers must be measured
11. Production judgment
Fusion is a versioned ranking component. Treat normalizer, weights, rank constant, candidate windows, reranker model, and filter placement as deployable configuration with A/B or offline regression tests. Keep a lexical-only baseline so semantic quality lift and failure behavior stay observable.
Bridge: Lesson 5 turns the concepts into equivalent AtlasMart pipelines and an explicit portability/complexity comparison.
Check your understanding
- What information does RRF intentionally discard?
- Why are fusion weights not portable?
- Which metric tests first-stage candidate coverage?
- Why measure reranker candidate N?
- What remains the most important baseline?
Review the answers
1. Raw score magnitudes; it uses rank positions.
2. Normalization, score distributions, candidate windows, and product semantics differ.
3. Recall@k/N over judged relevant documents.
4. It controls both quality ceiling and inference cost/latency.
5. A stable lexical-only path evaluated on the same judgments and filters.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
These links are the product documentation surfaces used to freeze examples for the September 2026 lesson baseline. Re-check them before production rollout because syntax, licensing, managed-service support, and model availability can change independently of this course.
- Elastic semantic search quickstart
- Elastic semantic_text setup/configuration
- Elastic search/retrieve semantic_text
- Elastic hybrid search
- Elastic retrievers overview
- Elastic linear retriever
- Elastic sparse_vector query
- Elastic ranking and reranking
- Elastic semantic reranking
- OpenSearch neural query
- OpenSearch neural sparse query
- OpenSearch hybrid search
- OpenSearch normalization processor
- OpenSearch score ranker processor
- OpenSearch rerank processor
- OpenSearch text embedding processor
- OpenSearch custom local models