Compose Elastic lexical, dense, sparse, semantic, and reranking signals without pretending their raw scores share units.
Elastic Semantic/Vector Retrieval and Hybrid Ranking Concepts Across Lexical and Dense/Sparse Signals
Combine lexical, dense, sparse, semantic, neural, and reranking signals without losing score meaning, evaluation discipline, or operational control across the two platforms.
Learning outcomes
Use Elastic lexical, dense, sparse, and semantic_text retrieval as separate first-stage signals.
Compose first-stage retrievers with RRF or normalized linear combination without adding incomparable raw scores.
Understand when semantic_text simplifies inference and when explicit dense_vector/sparse_vector control is required.
Bound reranking candidate count and include inference/rerank time in tail-latency budgets.
Preserve a free/local path based on precomputed vectors and application-side fusion.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
semantic_text usually invokes the Inference API and
can therefore carry license/service requirements; OpenSearch ML
Commons can run supported local models or connectors, but
model/hardware support remains deployment-specific.
1. AtlasMart problem: three useful signals disagree
For “waterproof footwear for wet trail,” BM25 may reward literal trail shoes, dense retrieval may favor waterproof shell because “waterproof” dominates the compact fixture, and sparse learned retrieval may promote a product carrying semantically expanded trail/waterproof tokens. None of the raw score scales has common units. Hybrid search is therefore a rank/normalization problem before it is a weighting problem.
2. Elastic retrieval surfaces in 9.5.3
| Surface | Best use | Important boundary |
|---|---|---|
match / Query DSL |
BM25 lexical retrieval | query-relative BM25 score; not comparable to vector score |
semantic_text + match |
managed semantic retrieval | inference endpoint/license/service availability is part of the path |
knn / dense_vector |
explicit dense-vector control | model/dimension/similarity contract stays your responsibility |
sparse_vector query |
learned sparse retrieval such as ELSER | query vectors/tokens must match the document model |
rrf retriever |
rank-based fusion | uses rank positions, not score calibration |
linear retriever |
normalized weighted score fusion | weights are meaningful only after declared normalization |
text_similarity_reranker |
cross-encoder-like second stage | rerank only a bounded candidate set; model adds latency/cost |
3. Deterministic AtlasMart retrieval fixture
Carry forward Chapter 25's eight product vectors and
tenant/category metadata. These vectors are a transparent
mechanics fixture, not outputs claimed from a production
embedding model. Every semantic experiment records a model
contract such as atlasmart-product-v1, dimensions,
normalization, source text, and tenant metadata.
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]
Security invariant: tenant/category eligibility is applied before a result is allowed into evaluation or reranking. A semantically close document outside the caller's authorization scope is not a candidate.
4. RRF: merge rankings without pretending the scores share units
RRF assigns each document a contribution based on rank, commonly
1/(rank_constant + rank). The constant and
candidate windows affect sensitivity, so they belong in the
experiment record. The method is robust to score-scale mismatch
but it can still fail if one retriever has poor recall.
def rrf(rank_lists, k=60):
out = {}
for ranking in rank_lists:
for rank, doc_id in enumerate(ranking, start=1):
out[doc_id] = out.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(out.items(), key=lambda x: (-x[1], x[0]))
lexical = ["P-1002", "P-1001", "P-1005", "P-1006"]
semantic = ["P-1005", "P-1001", "P-1002", "P-1006"]
print(rrf([lexical, semantic]))
This client-side calculation is the mandatory free/local baseline. In Elasticsearch, the RRF retriever can perform equivalent rank fusion server-side; validate the exact request syntax against your installed 9.5.3 docs.
5. Linear fusion: only after normalization
Elasticsearch's linear retriever normalizes and combines child
retriever scores. That is materially different from
bm25_score + cosine_score. A weight of 0.7 is
meaningful only with a declared normalizer and evaluated query
set. Min-max style normalization can be sensitive to
candidate-window extremes; rank fusion avoids score calibration
but discards score margins. Evaluate both against judgments.
6. Dense and sparse are complementary, not “old vs new”
Dense embeddings compress meaning into a fixed vector and work well for conceptual similarity. Learned sparse retrieval expands text into weighted dimensions/tokens and can preserve term-like interpretability and inverted-index execution properties. Lexical BM25 remains valuable for exact product codes, brands, model numbers, negation-sensitive phrases, and rare terms. A strong production design keeps baselines for all three.
7. Wrong approach: add BM25 and cosine directly
A BM25 score of 8 and a cosine-derived score of 0.82 do not imply BM25 is “ten times stronger.” Their scales are query-, corpus-, similarity-, and implementation-dependent. Repair by RRF or an explicitly normalized/validated score-combination method. Keep component ranks and scores in debug telemetry so a regression can be localized.
8. Elastic request sketches
POST atlasmart-products-v1/_search
{
"retriever": {
"rrf": {
"retrievers": [
{"standard":{"query":{"match":{"name":"waterproof trail layer"}}}},
{"standard":{"query":{"match":{"name_semantic":"waterproof trail layer"}}}}
]
}
}
}
POST atlasmart-products-v1/_search
{
"retriever": {
"rrf": {
"retrievers": [
{"standard":{"query":{"bool":{"must":{"match":{"name":"waterproof trail layer"}},"filter":{"term":{"tenant":"tenant-a"}}}}}},
{"standard":{"query":{"knn":{"field":"embedding_v1","query_vector":[0.66,0.43,0.08,0.07,0.09,0.57,0.10,0.20],"k":20,"num_candidates":80,"filter":{"term":{"tenant":"tenant-a"}}}}}}
]
}
}
}
The second sketch keeps authorization in both branches. Do not rely on post-fusion filtering to hide unauthorized candidates.
9. Reranking is a second-stage budget
A cross-encoder or text-similarity model examines query+document pairs more deeply than a bi-encoder embedding. That makes it suitable for tens of candidates, not an unbounded index. Measure candidate-count sensitivity separately from first-stage recall. If the first stage never retrieves a relevant product, reranking cannot rescue it.
10. Production judgment
Choose the simplest composition that meets judged quality and latency. Record candidate windows, fusion type, normalization, weights, model endpoint/version, filter placement, and fallback. The same search interface may behave differently across Elastic Cloud, Serverless, and self-managed deployments because inference endpoints, licenses, and hardware differ.
Bridge: Lesson 3 implements the same intent using OpenSearch neural, neural_sparse, k-NN, ingest/search pipelines, and ML Commons.
Check your understanding
- Why prefer RRF when component scores have unrelated scales?
- When is linear fusion appropriate?
- Can reranking repair missing first-stage recall?
- Why retain BM25 in a semantic system?
- Where should tenant filters be applied?
Review the answers
1. RRF operates on rank positions and does not require score calibration.
2. When score normalization and weights are explicit and validated against judgments.
3. No. It can only reorder candidates that were retrieved.
4. Exact identifiers, rare terms, and literal lexical intent remain important.
5. Inside every candidate-producing branch before fusion/reranking.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
These links are the product documentation surfaces used to freeze examples for the September 2026 lesson baseline. Re-check them before production rollout because syntax, licensing, managed-service support, and model availability can change independently of this course.
- Elastic semantic search quickstart
- Elastic semantic_text setup/configuration
- Elastic search/retrieve semantic_text
- Elastic hybrid search
- Elastic retrievers overview
- Elastic linear retriever
- Elastic sparse_vector query
- Elastic ranking and reranking
- Elastic semantic reranking
- OpenSearch neural query
- OpenSearch neural sparse query
- OpenSearch hybrid search
- OpenSearch normalization processor
- OpenSearch score ranker processor
- OpenSearch rerank processor
- OpenSearch text embedding processor
- OpenSearch custom local models