Build equivalent keyword+semantic intent in both platforms, compare it with one evaluation contract, and keep rollback visible.
Build Equivalent Keyword+Semantic Search Pipelines in Both Platforms and Compare Quality, Latency, and Operational Complexity
Combine lexical, dense, sparse, semantic, neural, and reranking signals without losing score meaning, evaluation discipline, or operational control across the two platforms.
Learning outcomes
Build equivalent AtlasMart keyword+semantic intent on Elasticsearch and OpenSearch without claiming API identity.
Use fixed vectors and client-side RRF as a common free/local comparison oracle.
Compare lexical-only, semantic-only, and hybrid rankings with the same judgments and authorization filters.
Measure operational complexity across model/pipeline state, fallback, observability, upgrade, and managed-service boundaries.
Produce a go/no-go decision record that separates quality evidence from latency/cost evidence.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
semantic_text usually invokes the Inference API and
can therefore carry license/service requirements; OpenSearch ML
Commons can run supported local models or connectors, but
model/hardware support remains deployment-specific.
1. Capstone scenario: “equivalent intent,” not “equivalent JSON”
AtlasMart requires: tenant-scoped keyword retrieval, tenant-scoped semantic retrieval over the same catalog text, one fused ranking, optional bounded reranking, lexical fallback when inference is unavailable, and judged metrics. The implementation should preserve those requirements even though Elastic retrievers and OpenSearch search pipelines use different APIs.
2. Deterministic AtlasMart retrieval fixture
Carry forward Chapter 25's eight product vectors and
tenant/category metadata. These vectors are a transparent
mechanics fixture, not outputs claimed from a production
embedding model. Every semantic experiment records a model
contract such as atlasmart-product-v1, dimensions,
normalization, source text, and tenant metadata.
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]
Security invariant: tenant/category eligibility is applied before a result is allowed into evaluation or reranking. A semantically close document outside the caller's authorization scope is not a candidate.
3. Common portable baseline
- Index the same eight documents with the same text, tenant, category, SKU, and precomputed 8D vector.
- Use the same query vector and tenant filter.
- Capture lexical top-N IDs and dense-vector top-N IDs independently.
- Fuse the IDs in the client with the RRF reference implementation.
- Evaluate all three rankings with the same qrels.
- Record end-to-end request timings when you execute a cluster; do not fill in invented milliseconds.
def rrf(rank_lists, k=60):
out = {}
for ranking in rank_lists:
for rank, doc_id in enumerate(ranking, start=1):
out[doc_id] = out.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(out.items(), key=lambda x: (-x[1], x[0]))
lexical = ["P-1002", "P-1001", "P-1005", "P-1006"]
semantic = ["P-1005", "P-1001", "P-1002", "P-1006"]
print(rrf([lexical, semantic]))
import math, statistics, time
def dcg(ids, qrels, k=5):
return sum((2**qrels.get(doc,0)-1)/math.log2(i+2) for i,doc in enumerate(ids[:k]))
def ndcg(ids, qrels, k=5):
ideal=[d for d,_ in sorted(qrels.items(), key=lambda x:(-x[1],x[0]))]
denom=dcg(ideal,qrels,k)
return dcg(ids,qrels,k)/denom if denom else 0.0
def mrr(ids, qrels):
for i,doc in enumerate(ids,1):
if qrels.get(doc,0)>0: return 1/i
return 0.0
# Replace rankings with captured IDs from each interface; never compare raw BM25/vector scores directly.
qrels={"P-1005":3,"P-1001":2,"P-1002":2,"P-1006":1}
for name,ids in {
"lexical":["P-1002","P-1001","P-1005","P-1006"],
"semantic":["P-1005","P-1001","P-1002","P-1006"],
"hybrid":["P-1005","P-1002","P-1001","P-1006"],
}.items():
print(name, "NDCG@4=", round(ndcg(ids,qrels,4),4), "MRR=", round(mrr(ids,qrels),4))
4. Elasticsearch closest supported implementation
Use a text field for BM25 and either
semantic_text or an explicit
dense_vector/sparse_vector field for
semantic retrieval. Combine with the RRF retriever when rank
fusion is desired, or the linear retriever when you have
validated normalization/weights. Add
text_similarity_reranker only if its model endpoint
and latency budget are part of the deployment contract.
POST atlasmart-products-v1/_search
{
"retriever": {
"rrf": {
"retrievers": [
{"standard":{"query":{"bool":{"must":{"match":{"name":"waterproof trail layer"}},"filter":{"term":{"tenant":"tenant-a"}}}}}},
{"standard":{"query":{"knn":{"field":"embedding_v1","query_vector":[0.66,0.43,0.08,0.07,0.09,0.57,0.10,0.20],"k":20,"num_candidates":80,"filter":{"term":{"tenant":"tenant-a"}}}}}}
]
}
}
}
5. OpenSearch closest supported implementation
Use the same lexical and vector fields. A
hybrid query feeds either a normalization processor
(score-based) or a score-ranker processor (RRF). If you use
neural instead of raw k-NN, the model/ML Commons
state enters the request path. The free/local baseline can still
use the precomputed query vector.
PUT /_search/pipeline/atlasmart-rrf-v1
{
"phase_results_processors": [
{
"score-ranker-processor": {
"combination": {"technique":"rrf", "rank_constant":60}
}
}
]
}
POST atlasmart-products-v1/_search?search_pipeline=atlasmart-rrf-v1
{
"query":{"hybrid":{"queries":[
{"bool":{"must":{"match":{"name":"waterproof trail layer"}},"filter":{"term":{"tenant":"tenant-a"}}}},
{"knn":{"embedding_v1":{"vector":[0.66,0.43,0.08,0.07,0.09,0.57,0.10,0.20],"k":20,"filter":{"term":{"tenant":"tenant-a"}}}}}
]}}
}
If your exact 3.8 build names/parameters the processor
differently, stop and check
/_nodes/search_pipelines plus the versioned docs.
Do not “fix” it by switching to an undocumented request.
6. Comparison matrix
| Decision surface | Elasticsearch 9.5.3 | OpenSearch 3.8.0 |
|---|---|---|
| managed semantic field |
semantic_text automates much of
inference/vector plumbing
|
neural/semantic-field and ML Commons workflows; verify exact field/model configuration |
| rank fusion | RRF retriever | score-ranker processor / RRF in search pipeline |
| score fusion | linear retriever with normalization | normalization processor + combination/weights |
| sparse semantic |
sparse_vector, ELSER/compatible endpoint
|
neural_sparse on rank_features or
sparse_vector modes
|
| rerank | text-similarity reranker retriever or ES|QL RERANK | rerank processor / ML inference + rerank |
| operational state | inference endpoints, retriever config, model/service credentials | ML Commons models/connectors, ingest/search pipelines, plugin state |
| free deterministic lab | precomputed dense vectors + lexical + client/server RRF where available | precomputed dense vectors + lexical + client/server fusion where available |
7. Safe failure injection
Do not take down a shared model endpoint. In the disposable lab, simulate inference failure in the client by refusing to supply the query vector, or reference a deliberately absent model only in an isolated cluster and immediately revert. Success criteria: semantic branch fails visibly; lexical branch remains authorized and functional; fallback metric increments; no retry storm; the request does not leak cross-tenant data.
8. Quality and latency release gate
required_evidence:
- lexical_only metrics on qrels-v1
- semantic_only metrics on qrels-v1
- hybrid metrics on qrels-v1
- p50/p95/p99 end-to-end timings with sample count and concurrency
- inference/fusion/rerank fallback and error rates
- tenant-filter tests for every retrieval branch
- model/pipeline version inventory
- rollback path verified
go_when:
- quality lift is statistically/operationally meaningful for priority queries
- p99 and throughput fit the service SLO under representative concurrency
- fallback behavior preserves minimum acceptable lexical service
- cost/capacity and model lifecycle are owned
- upgrade and snapshot/restore runbooks include ML/pipeline dependencies
9. What the lab proves—and does not prove
The fixed corpus proves API mechanics, filter placement, fusion math, and evaluation harness behavior. It does not prove production NDCG, p99, memory, or cost. Those require representative catalog text, embedding distributions, real judgments, realistic concurrency, and the actual inference/reranking deployment.
10. Production judgment and bridge to Chapter 27
Choose between Elastic and OpenSearch using evidence that includes relevance, inference/model operations, security, managed-service fit, licensing, pipelines, observability, upgrades, and exit strategy—not just one hybrid demo. Keep model version, qrels, fusion config, and fallback policy deployable and auditable.
Chapter 27 adds RAG and agentic search. The key boundary is that retrieval quality must remain independently measurable: a fluent generated answer can hide a bad retriever, while a strong retriever can be undermined by poor context packing or generation.
Check your understanding
- What makes the two platform implementations “equivalent”?
- Why keep client-side RRF in the lab?
- What must be identical for a fair comparison?
- What evidence must not be fabricated?
- What changes in Chapter 27?
Review the answers
1. They satisfy the same retrieval, security, fallback, and evaluation intent—not the same JSON.
2. It is a portable, deterministic reference independent of product-native fusion features.
3. Corpus, vectors/model contract, query set, qrels, tenant filters, candidate cutoffs, and benchmark conditions.
4. Latency, throughput, memory, index-build, model cost, and production quality measurements.
5. Generation is added after retrieval, and retrieval must still be evaluated separately from generated-answer quality.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
These links are the product documentation surfaces used to freeze examples for the September 2026 lesson baseline. Re-check them before production rollout because syntax, licensing, managed-service support, and model availability can change independently of this course.
- Elastic semantic search quickstart
- Elastic semantic_text setup/configuration
- Elastic search/retrieve semantic_text
- Elastic hybrid search
- Elastic retrievers overview
- Elastic linear retriever
- Elastic sparse_vector query
- Elastic ranking and reranking
- Elastic semantic reranking
- OpenSearch neural query
- OpenSearch neural sparse query
- OpenSearch hybrid search
- OpenSearch normalization processor
- OpenSearch score ranker processor
- OpenSearch rerank processor
- OpenSearch text embedding processor
- OpenSearch custom local models