Fuse and rerank only with declared math, candidate budgets, judged relevance, and end-to-end latency evidence.

Hybrid Score Normalization, Weighted Combination, Reciprocal-Rank-Like Fusion, Re-Ranking, and Search Quality Evaluation

Combine lexical, dense, sparse, semantic, neural, and reranking signals without losing score meaning, evaluation discipline, or operational control across the two platforms.

Intermediate → Advanced140–185 minutesFusion + judged-evaluation lab · Chapter 26 · Lesson 04Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Explain why raw lexical, dense, and sparse scores cannot be summed safely without calibration.

02

Implement and compare score normalization/weighted fusion with RRF-style rank fusion.

03

Use reranking as a bounded second stage and preserve first-stage diagnostics.

04

Compute judged Recall/NDCG/MRR-like metrics independently of generated answers.

05

Treat latency/cost and relevance as a joint decision surface rather than optimizing one metric.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned platform baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0 with their bundled JVMs. The mandatory AtlasMart path uses precomputed local vectors and client-side evaluation so a paid inference service is not required. Elastic semantic_text usually invokes the Inference API and can therefore carry license/service requirements; OpenSearch ML Commons can run supported local models or connectors, but model/hardware support remains deployment-specific.

1. AtlasMart problem: the “best” hybrid system depends on what you preserve

RRF preserves rank agreement and ignores score magnitude. Normalized weighted fusion preserves some score-margin information but depends on normalization and candidate windows. Reranking can model query-document interaction more deeply, but it adds model cost and may reorder a good first-stage ranking badly if the reranker is mismatched. There is no universally correct fusion rule.

2. Three fusion families

Method Uses Strength Failure mode
raw weighted sum un-normalized scores simple meaningless when component scales differ
normalized weighted combination normalized component scores retains score margins and supports business weights normalizer/candidate outliers can change contributions
RRF / reciprocal-rank-like rank positions robust to score-scale mismatch throws away score margins and depends on candidate recall

3. Deterministic AtlasMart retrieval fixture

Carry forward Chapter 25's eight product vectors and tenant/category metadata. These vectors are a transparent mechanics fixture, not outputs claimed from a production embedding model. Every semantic experiment records a model contract such as atlasmart-product-v1, dimensions, normalization, source text, and tenant metadata.

Products and precomputed vectors
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]

Security invariant: tenant/category eligibility is applied before a result is allowed into evaluation or reranking. A semantically close document outside the caller's authorization scope is not a candidate.

4. Deterministic RRF example

Rank fusion
def rrf(rank_lists, k=60):
    out = {}
    for ranking in rank_lists:
        for rank, doc_id in enumerate(ranking, start=1):
            out[doc_id] = out.get(doc_id, 0.0) + 1.0 / (k + rank)
    return sorted(out.items(), key=lambda x: (-x[1], x[0]))

lexical = ["P-1002", "P-1001", "P-1005", "P-1006"]
semantic = ["P-1005", "P-1001", "P-1002", "P-1006"]
print(rrf([lexical, semantic]))

Record the rank constant and each child ranking. Do not publish only the fused list: debugging requires the components. Elastic's RRF retriever and OpenSearch's score-ranker processor implement product-native rank-fusion surfaces, but the client-side function is a portable oracle for this tiny fixture.

5. Score normalization changes ranking

OpenSearch documents min-max, L2, and z-score normalization for the normalization processor. Different techniques can produce different winners from the same component scores. Elastic's linear retriever similarly normalizes before combining. Therefore weights are not portable numbers: 0.7 semantic in one product/configuration is not automatically equivalent to 0.7 in the other.

6. Rerank only candidates that deserve the budget

Start with a first-stage candidate set large enough to contain relevant documents, then rerank a bounded top-N. Measure: first-stage Recall@N; reranker NDCG/MRR at the display cutoff; rerank p95/p99; error/fallback rate; and cost per query if an external service is involved. Candidate N is a capacity parameter, not merely a relevance knob.

7. Judged evaluation—not generated-answer vibes

For search quality, a judgment set maps (query, document) to relevance grades. Recall@k tests whether relevant candidates are retrieved. MRR rewards the position of the first relevant result. NDCG@k rewards graded relevance near the top and normalizes against the ideal ranking. These metrics evaluate retrieval/ranking directly; Chapter 27 will separately evaluate generated answers.

Portable judged-metric harness
import math, statistics, time

def dcg(ids, qrels, k=5):
    return sum((2**qrels.get(doc,0)-1)/math.log2(i+2) for i,doc in enumerate(ids[:k]))

def ndcg(ids, qrels, k=5):
    ideal=[d for d,_ in sorted(qrels.items(), key=lambda x:(-x[1],x[0]))]
    denom=dcg(ideal,qrels,k)
    return dcg(ids,qrels,k)/denom if denom else 0.0

def mrr(ids, qrels):
    for i,doc in enumerate(ids,1):
        if qrels.get(doc,0)>0: return 1/i
    return 0.0

# Replace rankings with captured IDs from each interface; never compare raw BM25/vector scores directly.
qrels={"P-1005":3,"P-1001":2,"P-1002":2,"P-1006":1}
for name,ids in {
  "lexical":["P-1002","P-1001","P-1005","P-1006"],
  "semantic":["P-1005","P-1001","P-1002","P-1006"],
  "hybrid":["P-1005","P-1002","P-1001","P-1006"],
}.items():
    print(name, "NDCG@4=", round(ndcg(ids,qrels,4),4), "MRR=", round(mrr(ids,qrels),4))

8. Latency evidence

Measure end-to-end latency at the client boundary and decompose it into query embedding/inference, first-stage search, fusion, fetch, reranking, and network where possible. Report p50/p95/p99 with sample count and concurrency. Do not compare Profile API timings to normal request latency: profiling adds overhead and is diagnostic, not a benchmark.

9. Wrong approach: optimize only hybrid NDCG

A 1% quality lift that doubles p99, triples external-model cost, or causes frequent fallbacks may be a poor production trade. Conversely, the lowest-latency path may violate relevance objectives. Define a release gate with both quality and operational budgets, then compare deltas against the same query/judgment set.

10. Decision record template

Store this with each experiment
experiment: atlasmart-hybrid-2026-09-a
corpus_version: catalog-fixture-v1
judgments_version: qrels-v1
embedding_contract: atlasmart-product-v1
lexical_config: bm25-default
semantic_config: dense-precomputed-v1
fusion: rrf
rank_constant: 60
candidate_windows: {lexical: 50, semantic: 50, rerank: 20}
filters: tenant-before-fusion
metrics: [Recall@20, NDCG@10, MRR, p95_ms, p99_ms, fallback_rate]
notes: synthetic mechanics fixture; production numbers must be measured

11. Production judgment

Fusion is a versioned ranking component. Treat normalizer, weights, rank constant, candidate windows, reranker model, and filter placement as deployable configuration with A/B or offline regression tests. Keep a lexical-only baseline so semantic quality lift and failure behavior stay observable.

Bridge: Lesson 5 turns the concepts into equivalent AtlasMart pipelines and an explicit portability/complexity comparison.

Check your understanding

  1. What information does RRF intentionally discard?
  2. Why are fusion weights not portable?
  3. Which metric tests first-stage candidate coverage?
  4. Why measure reranker candidate N?
  5. What remains the most important baseline?
Review the answers

1. Raw score magnitudes; it uses rank positions.

2. Normalization, score distributions, candidate windows, and product semantics differ.

3. Recall@k/N over judged relevant documents.

4. It controls both quality ceiling and inference cost/latency.

5. A stable lexical-only path evaluated on the same judgments and filters.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

These links are the product documentation surfaces used to freeze examples for the September 2026 lesson baseline. Re-check them before production rollout because syntax, licensing, managed-service support, and model availability can change independently of this course.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.