Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation
Named Queries, Explain/Profile, Score Debugging, and Avoiding Score Comparisons Across Different Retrieval Sources
Debug ranking with named queries, _explain and Profile while keeping semantics, score attribution and latency diagnosis separate; treat raw scores from different retrieval sources as incomparable until a defined fusion or reranking method is applied.
Learning outcomes
When AtlasMart receives a ranking complaint, three questions are often mixed together: “which clause matched?”, “why did this document receive this lexical score?”, and “where did the query spend execution time?” Named queries, Explain and Profile answer different parts. A fourth question—“can I compare this BM25 score with a vector score?”—requires an explicit fusion/calibration strategy, not arithmetic on raw values.
Use named queries to make large compound queries observable at the hit level without treating names as score explanations.
Use _explain for one document/query pair and distinguish score attribution from retrieval-quality judgment.
Use Profile for search-execution timing while accounting for profiling overhead and important time it does not measure.
Diagnose semantic, scoring and performance regressions with a repeatable evidence sequence rather than random query edits.
Explain why BM25, vector, neural, reranker and business-feature scores belong to different scales until a defined fusion/reranking method combines them.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. This lesson uses lexical queries only for executable examples. Vector/neural scores are discussed as a boundary because later chapters introduce them with explicit normalization/fusion/evaluation.
The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.
1. Named queries answer “which clause matched?”
GET atlasmart-products-relevance-v1/_search
{
"query": {
"bool": {
"must": [
{
"multi_match": {
"query": "wireless headphones",
"fields": ["name^3", "description"],
"_name": "lexical_product_text"
}
}
],
"should": [
{
"rank_feature": {
"field": "popularity",
"saturation": { "pivot": 60 },
"boost": 0.20,
"_name": "popularity_prior"
}
}
]
}
}
}
Where supported, search hits can report
matched_queries. That is excellent for test
diagnostics and product analytics such as “did this result match
the lexical title clause or only a secondary feature?” A name
does not tell you how much a clause contributed or how expensive
it was; use Explain/Profile for those questions.
Treat query names as observability identifiers. If dashboards or tests depend on them, version/rename deliberately just as you would an API field.
2. Explain answers “why did this document score this way?”
GET atlasmart-products-relevance-v1/_explain/p7
{
"query": {
"multi_match": {
"query": "wireless headphones",
"type": "best_fields",
"fields": ["name^3", "description"],
"tie_breaker": 0.15
}
}
}
Trace the explanation from query structure to BM25 term contributions, boosts and combined clauses. Compare p1 and p7 when a “wireless headphones” query behaves unexpectedly. The explanation is per-document and can be expensive; it is not designed for every production request.
OpenSearch documents that the dedicated Explain API does not support the search_pipeline parameter. For a search using a pipeline, use the Search API with explain=true when supported for that path. Do not assume every Explain surface covers every post-processing stage.
3. Profile answers “where did shard-level search execution spend time?”
GET atlasmart-products-relevance-v1/_search
{
"profile": true,
"query": {
"multi_match": {
"query": "wireless headphones",
"fields": ["name^3", "description"]
}
}
}
Profile exposes query rewrite/execution/collector timing by shard and underlying Lucene query structure. Both product docs warn that profiling adds significant overhead. It also does not represent end-to-end latency: network time, queue waiting and some coordinating-node merge time are outside the profile coverage. Use it to identify expensive components, then benchmark without profiling.
| Tool | Primary question | What it does not prove |
|---|---|---|
| named queries | Which labeled clauses matched this hit? | Score attribution, quality, or latency. |
| _explain | Why did this document receive this score/match outcome? | System-wide latency or whether ranking is good for users. |
| profile | Which search components consumed execution time? | Normal production latency; profiling itself adds overhead. |
| _rank_eval / offline harness | Did a query configuration rank judged documents well? | Production tail latency, business impact, or unbiased judgments. |
4. Debug in a fixed order
1. Freeze: query body/template version, index generation, product/version, user segment, timestamp.
2. Reproduce: same query + filters + routing/PIT context where relevant.
3. Verify semantics: analyzed tokens, mapping, filters, nested path, eligible candidates.
4. Inspect ranked IDs + matched_queries.
5. Explain only surprising document/query pairs.
6. Profile only if the issue is latency/execution shape; never use profiled latency as benchmark latency.
7. Compare against the judged set and previous baseline.
8. Change one mechanism at a time; rerun quality + latency + failure checks.
9. Record the experiment and rollback condition.
5. Wrong approach: compare raw scores from different retrieval sources
BM25 returned 6.2 and vector search returned 0.82, therefore BM25 is “7.5× more confident”; add the raw numbers and sort.
These values are generated by different scoring functions with different ranges/distributions and semantics. Even two BM25 scores from different queries are not calibrated probabilities. Hybrid systems need a defined fusion or reranking method—such as normalized score fusion, reciprocal-rank-style fusion, or a learned reranker—plus judged-query evaluation. Preserve component provenance so you can tell which retriever contributed which candidates.
6. Score debugging must preserve privacy and authorization
Explain trees, query logs and relevance judgments can expose query text, document terms, tenant identifiers or business features. Apply the same access controls/redaction/retention policy as other search telemetry. Never weaken document/tenant authorization to make debugging easier; retrieve diagnostic evidence only for data the operator is authorized to inspect.
Check your understanding
- What does a named query tell you?
- When should _explain be used?
- Why is Profile not a latency benchmark?
- Why can raw BM25 and vector scores not simply be added?
- What should be frozen before debugging a ranking regression?
Review the answers
1. Which labeled query clauses matched a hit; it does not by itself explain numeric contribution or cost.
2. On selected document/query pairs when you need score/match attribution, not on every production hit.
3. It adds overhead and excludes parts of end-to-end latency such as network/queue time and some coordinating work.
4. They are not calibrated to a common scale or semantic meaning; fusion must define normalization/rank combination and be evaluated.
5. Query/template version, index generation/data, product/version, relevant context/segment and timestamp so the experiment is reproducible.
Production judgment
Operate relevance debugging as an engineering workflow: reproducible inputs, least-privilege evidence, named clauses, targeted Explain, targeted Profile, judged-query comparison, one change at a time and an explicit rollback condition. A compelling explanation for one document is not proof that the whole ranking improved.
Summary and next step
You now have tools to attribute score and execution behavior without conflating them. The final lesson turns judgments and ranked IDs into metrics and release thresholds so relevance changes are governed like code rather than approved from screenshots.
Authoritative references
- Elastic: BM25 and similarity settings — Default BM25 behavior, k1/b parameters and expert similarity configuration.
- Elastic: multi_match query — Field boosts and best_fields/most_fields/cross_fields semantics.
- Elastic: dis_max query — Best-clause scoring plus tie_breaker contribution.
- Elastic: function_score query — Business-signal scoring, decay functions, score_mode and boost_mode.
- Elastic: rank_feature query — Optimized numeric ranking features and scoring functions.
- Elastic: ranking evaluation — Judged-query evaluation with precision, recall, MRR and DCG-family metrics.
- OpenSearch: keyword search and BM25 — BM25 fundamentals and the OpenSearch 3.x LegacyBM25-to-BM25 score-scale change.
- OpenSearch: multi_match queries — Field boosts and multi-field matching modes.
- OpenSearch: function_score query — Function-score composition, decay and score combination.
- OpenSearch: rank_feature query — Rank-feature field/query behavior and saturation/log/sigmoid options.
- OpenSearch: Explain API — Per-document score explanation and diagnostic limitations.
- OpenSearch: Profile API — Search execution timing with explicit profiling overhead/coverage limits.
- OpenSearch: Ranking Evaluation API — Judged-query ranking-quality evaluation.