Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation

Named Queries, Explain/Profile, Score Debugging, and Avoiding Score Comparisons Across Different Retrieval Sources

Debug ranking with named queries, _explain and Profile while keeping semantics, score attribution and latency diagnosis separate; treat raw scores from different retrieval sources as incomparable until a defined fusion or reranking method is applied.

Intermediate100–120 minutesJudged-query relevance labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

When AtlasMart receives a ranking complaint, three questions are often mixed together: “which clause matched?”, “why did this document receive this lexical score?”, and “where did the query spend execution time?” Named queries, Explain and Profile answer different parts. A fourth question—“can I compare this BM25 score with a vector score?”—requires an explicit fusion/calibration strategy, not arithmetic on raw values.

01

Use named queries to make large compound queries observable at the hit level without treating names as score explanations.

02

Use _explain for one document/query pair and distinguish score attribution from retrieval-quality judgment.

03

Use Profile for search-execution timing while accounting for profiling overhead and important time it does not measure.

04

Diagnose semantic, scoring and performance regressions with a repeatable evidence sequence rather than random query edits.

05

Explain why BM25, vector, neural, reranker and business-feature scores belong to different scales until a defined fusion/reranking method combines them.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. This lesson uses lexical queries only for executable examples. Vector/neural scores are discussed as a boundary because later chapters introduce them with explicit normalization/fusion/evaluation.

Execution and safety note

The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.

1. Named queries answer “which clause matched?”

Dev Tools · label lexical and business clauses
GET atlasmart-products-relevance-v1/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "multi_match": {
            "query": "wireless headphones",
            "fields": ["name^3", "description"],
            "_name": "lexical_product_text"
          }
        }
      ],
      "should": [
        {
          "rank_feature": {
            "field": "popularity",
            "saturation": { "pivot": 60 },
            "boost": 0.20,
            "_name": "popularity_prior"
          }
        }
      ]
    }
  }
}

Where supported, search hits can report matched_queries. That is excellent for test diagnostics and product analytics such as “did this result match the lexical title clause or only a secondary feature?” A name does not tell you how much a clause contributed or how expensive it was; use Explain/Profile for those questions.

Name stability

Treat query names as observability identifiers. If dashboards or tests depend on them, version/rename deliberately just as you would an API field.

2. Explain answers “why did this document score this way?”

Dev Tools · explain a specific hit
GET atlasmart-products-relevance-v1/_explain/p7
{
  "query": {
    "multi_match": {
      "query": "wireless headphones",
      "type": "best_fields",
      "fields": ["name^3", "description"],
      "tie_breaker": 0.15
    }
  }
}

Trace the explanation from query structure to BM25 term contributions, boosts and combined clauses. Compare p1 and p7 when a “wireless headphones” query behaves unexpectedly. The explanation is per-document and can be expensive; it is not designed for every production request.

OpenSearch-specific boundary

OpenSearch documents that the dedicated Explain API does not support the search_pipeline parameter. For a search using a pipeline, use the Search API with explain=true when supported for that path. Do not assume every Explain surface covers every post-processing stage.

3. Profile answers “where did shard-level search execution spend time?”

Dev Tools · profile the search
GET atlasmart-products-relevance-v1/_search
{
  "profile": true,
  "query": {
    "multi_match": {
      "query": "wireless headphones",
      "fields": ["name^3", "description"]
    }
  }
}

Profile exposes query rewrite/execution/collector timing by shard and underlying Lucene query structure. Both product docs warn that profiling adds significant overhead. It also does not represent end-to-end latency: network time, queue waiting and some coordinating-node merge time are outside the profile coverage. Use it to identify expensive components, then benchmark without profiling.

Tool Primary question What it does not prove
named queries Which labeled clauses matched this hit? Score attribution, quality, or latency.
_explain Why did this document receive this score/match outcome? System-wide latency or whether ranking is good for users.
profile Which search components consumed execution time? Normal production latency; profiling itself adds overhead.
_rank_eval / offline harness Did a query configuration rank judged documents well? Production tail latency, business impact, or unbiased judgments.

4. Debug in a fixed order

ranking-debug playbook
1. Freeze: query body/template version, index generation, product/version, user segment, timestamp.
2. Reproduce: same query + filters + routing/PIT context where relevant.
3. Verify semantics: analyzed tokens, mapping, filters, nested path, eligible candidates.
4. Inspect ranked IDs + matched_queries.
5. Explain only surprising document/query pairs.
6. Profile only if the issue is latency/execution shape; never use profiled latency as benchmark latency.
7. Compare against the judged set and previous baseline.
8. Change one mechanism at a time; rerun quality + latency + failure checks.
9. Record the experiment and rollback condition.

5. Wrong approach: compare raw scores from different retrieval sources

Deliberately wrong fusion

BM25 returned 6.2 and vector search returned 0.82, therefore BM25 is “7.5× more confident”; add the raw numbers and sort.

These values are generated by different scoring functions with different ranges/distributions and semantics. Even two BM25 scores from different queries are not calibrated probabilities. Hybrid systems need a defined fusion or reranking method—such as normalized score fusion, reciprocal-rank-style fusion, or a learned reranker—plus judged-query evaluation. Preserve component provenance so you can tell which retriever contributed which candidates.

6. Score debugging must preserve privacy and authorization

Explain trees, query logs and relevance judgments can expose query text, document terms, tenant identifiers or business features. Apply the same access controls/redaction/retention policy as other search telemetry. Never weaken document/tenant authorization to make debugging easier; retrieve diagnostic evidence only for data the operator is authorized to inspect.

Check your understanding

  1. What does a named query tell you?
  2. When should _explain be used?
  3. Why is Profile not a latency benchmark?
  4. Why can raw BM25 and vector scores not simply be added?
  5. What should be frozen before debugging a ranking regression?
Review the answers

1. Which labeled query clauses matched a hit; it does not by itself explain numeric contribution or cost.

2. On selected document/query pairs when you need score/match attribution, not on every production hit.

3. It adds overhead and excludes parts of end-to-end latency such as network/queue time and some coordinating work.

4. They are not calibrated to a common scale or semantic meaning; fusion must define normalization/rank combination and be evaluated.

5. Query/template version, index generation/data, product/version, relevant context/segment and timestamp so the experiment is reproducible.

Production judgment

Operate relevance debugging as an engineering workflow: reproducible inputs, least-privilege evidence, named clauses, targeted Explain, targeted Profile, judged-query comparison, one change at a time and an explicit rollback condition. A compelling explanation for one document is not proof that the whole ranking improved.

Summary and next step

You now have tools to attribute score and execution behavior without conflating them. The final lesson turns judgments and ranked IDs into metrics and release thresholds so relevance changes are governed like code rather than approved from screenshots.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.