Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation

TF/IDF Intuition, BM25 Parameters, Document Length, Rare Terms, and Why Scores Are Query-Relative

Build an evidence-based BM25 mental model: connect term frequency, inverse document frequency and field-length normalization to AtlasMart rankings, then prove why _score is meaningful only within the exact query/corpus/scoring context that produced it.

Intermediate100–120 minutesJudged-query relevance labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart already returns the right broad set of products, but product managers now ask why one headset appears above another and whether a higher _score means “90% relevant.” That is the wrong mental model. Ranking starts with evidence produced by the analyzed query and indexed terms, then a similarity such as BM25 turns corpus statistics into a relative score for this query. The goal is to predict directional effects and validate them against judgments—not to memorize one numeric score.

01

Explain term frequency and inverse document frequency as evidence signals and connect them to BM25 saturation and field-length normalization.

02

Interpret BM25 k1 and b as expert controls rather than universal tuning knobs, and state what a change is expected to influence.

03

Use _explain to identify term-frequency, IDF, field-length and boost contributions for one document/query pair.

04

Explain why _score is query-, corpus-, shard/statistics-, version- and query-structure-relative rather than a probability or stable business KPI.

05

Establish a lexical baseline and judged-query habit before attempting field boosts or business signals.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. This lesson leaves BM25 at product defaults for the primary lab; any parameter experiment uses a separate disposable index and must be judged against the same query set.

Execution and safety note

The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.

1. Matching answers “eligible”; ranking answers “which first?”

A lexical query matches indexed terms derived from text analysis. Term frequency (TF) asks how often a query term occurs in a document field. Document frequency counts how many indexed documents contain that term; inverse document frequency (IDF) makes a term that appears in fewer documents more discriminative. BM25 combines these ideas with term-frequency saturation and document-length normalization.

Signal Direction Boundary to remember
Term frequency More occurrences can add evidence Gain saturates; ten repetitions are not ten times as relevant.
Inverse document frequency Rarer query terms usually contribute more IDF changes when the corpus/index statistics change.
Field length Longer fields can be normalized downward Effect is controlled by b and field norms; it is not a generic “shorter is better” rule.
Boosts Increase the contribution of a clause/field They encode policy; a large boost can overpower useful lexical evidence.
_score Orders hits for one scoring context It is not a probability and is not safely comparable across unrelated queries/retrievers/versions.
BM25 mental model · directional, not a frozen score oracle
Conceptual BM25 contribution for one query term t in document d:

score(t,d) ~= IDF(t) * [ tf(t,d) * (k1 + 1) ]
                         ---------------------------------
                         tf(t,d) + k1 * (1 - b + b * dl/avgdl)

Interpretation:
- IDF(t): rarer terms usually contribute more evidence.
- tf(t,d): repeated occurrences help, but k1 makes the gain saturate.
- dl/avgdl: longer fields are normalized according to b.
- field/query boosts multiply or otherwise combine with these lexical contributions.

Do NOT copy this as a cross-product exact-score oracle. Lucene/OpenSearch implementation details,
index statistics, boosts and query rewrites determine the actual explanation tree.

2. Build a corpus small enough to reason about

The ranking fixture deliberately reuses AtlasMart product names from earlier chapters and adds business fields used later. One primary shard removes cross-shard statistics as a confounder for the first experiments. The corpus is still synthetic and too small for production tuning; its purpose is to make cause and effect observable.

Dev Tools · create the disposable relevance fixture
DELETE atlasmart-products-relevance-v1
PUT atlasmart-products-relevance-v1
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0
  },
  "mappings": {
    "properties": {
      "sku":          { "type": "keyword" },
      "name":         { "type": "text", "fields": { "raw": { "type": "keyword" } } },
      "description":  { "type": "text" },
      "brand_text":   { "type": "text" },
      "category":     { "type": "keyword" },
      "available":    { "type": "boolean" },
      "price":        { "type": "double" },
      "rating":       { "type": "float" },
      "orders_30d":   { "type": "integer" },
      "launched_at":  { "type": "date" },
      "popularity":   { "type": "rank_feature" }
    }
  }
}
Dev Tools · load deterministic AtlasMart documents
POST atlasmart-products-relevance-v1/_bulk?refresh=true
{ "index": { "_id": "p1" } }
{ "sku":"AM-AU-100", "name":"Wireless Noise Cancelling Headphones", "description":"Over-ear Bluetooth headphones with active noise cancellation for travel", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":199.99, "rating":4.7, "orders_30d":920, "launched_at":"2026-08-10", "popularity":95 }
{ "index": { "_id": "p2" } }
{ "sku":"AM-AU-200", "name":"Wired Studio Headphones", "description":"Closed-back monitoring headphones for studio recording and mixing", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":89.99, "rating":4.4, "orders_30d":310, "launched_at":"2025-11-15", "popularity":55 }
{ "index": { "_id": "p3" } }
{ "sku":"AM-WB-300", "name":"Portable Bluetooth Speaker", "description":"Compact waterproof wireless speaker for travel and outdoor use", "brand_text":"WaveBox", "category":"electronics/audio", "available":false, "price":79.99, "rating":4.2, "orders_30d":760, "launched_at":"2026-06-01", "popularity":80 }
{ "index": { "_id": "p4" } }
{ "sku":"AM-AU-400", "name":"USB-C Noise Cancelling Earbuds", "description":"In-ear USB-C earbuds with active noise cancellation", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":129.99, "rating":4.5, "orders_30d":410, "launched_at":"2026-09-01", "popularity":60 }
{ "index": { "_id": "p5" } }
{ "sku":"AM-WN-500", "name":"Aluminum Laptop Stand", "description":"Adjustable desktop stand for laptops with ventilated aluminum construction", "brand_text":"WorkNest", "category":"office/accessories", "available":true, "price":49.00, "rating":4.6, "orders_30d":1200, "launched_at":"2026-01-20", "popularity":97 }
{ "index": { "_id": "p6" } }
{ "sku":"AM-ST-600", "name":"Lightweight Running Shoes", "description":"Breathable road running shoes for daily training", "brand_text":"Stride", "category":"sports/running", "available":true, "price":69.00, "rating":4.1, "orders_30d":1500, "launched_at":"2026-08-25", "popularity":99 }
{ "index": { "_id": "p7" } }
{ "sku":"AM-CA-700", "name":"Wireless Headphones Carrying Case", "description":"Hard protective case sized for over-ear headphones", "brand_text":"CarryAll", "category":"electronics/accessories", "available":true, "price":24.00, "rating":4.8, "orders_30d":1800, "launched_at":"2026-07-03", "popularity":100 }
{ "index": { "_id": "p8" } }
{ "sku":"AM-AU-800", "name":"Audiophile Studio Monitoring Headphones", "description":"Reference monitoring headphones with neutral tuning, replaceable cable, studio adapter, hard case, documentation, spare pads and accessories for long mixing sessions", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":249.00, "rating":4.9, "orders_30d":140, "launched_at":"2026-04-12", "popularity":45 }

3. Observe the ranking before explaining it

Dev Tools · lexical baseline
GET atlasmart-products-relevance-v1/_search
{
  "track_total_hits": true,
  "query": {
    "match": {
      "name": "wireless headphones"
    }
  },
  "_source": ["sku","name","category","orders_30d","launched_at"]
}

Record the ordered IDs and _score values as evidence for this exact run. The useful assertion is not “p1 has score 2.73.” The useful assertion is whether the judged-intent ordering is acceptable and why the explanation tree differs between hits.

Dev Tools · explain two candidates
GET atlasmart-products-relevance-v1/_explain/p1
{
  "query": {
    "match": {
      "name": "wireless headphones"
    }
  }
}

GET atlasmart-products-relevance-v1/_explain/p7
{
  "query": {
    "match": {
      "name": "wireless headphones"
    }
  }
}

4. Read an Explain tree without worshipping it

The Explain API decomposes a score for one document under one query. Look for the analyzed term, IDF/doc-frequency statistics, term frequency, norm/field-length contribution and boosts. It is a causal debugging trace for that score calculation, not a performance profiler and not a global statement about the index. Running Explain broadly is expensive; use it on carefully chosen query/document pairs.

OpenSearch 3.x score-scale caveat

OpenSearch 3.0 switched its default from LegacyBM25Similarity to Lucene-native BM25Similarity. The documented change can lower numeric score magnitude by a constant factor while preserving ranking order. This is a concrete reason never to build product thresholds around raw _score magnitude.

5. k1 and b are hypotheses, not tuning folklore

Parameter What it controls If increased, generally Why not tune blindly
k1 How quickly term-frequency contribution saturates Repeated terms can keep contributing for longer May reward repetition/spam-like text; value interacts with corpus and field semantics.
b Strength of field-length normalization Longer fields are penalized more relative to average length Can hurt legitimately long product descriptions; field-length distribution matters.
discount_overlaps Whether zero-position-increment overlap tokens count toward norm length Depends on analyzer/token graph Synonym/token-graph behavior makes this an analysis contract, not a casual switch.

If AtlasMart has a clear hypothesis—such as “long descriptions are over-penalized despite graded judgments”—create a separate experiment index, keep the same documents and judgments, and compare quality metrics. Do not change production similarity because one query “looks better.”

Dev Tools · isolated BM25 experiment index
PUT atlasmart-bm25-experiment
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0,
    "similarity": {
      "bm25_experiment": {
        "type": "BM25",
        "k1": 0.8,
        "b": 0.3
      }
    }
  },
  "mappings": {
    "properties": {
      "name": { "type": "text", "similarity": "bm25_experiment" }
    }
  }
}

6. Wrong approach: turn _score into a product probability

Deliberately wrong rule

Reject every hit with _score < 2.0 because scores below 2 mean low relevance.

This rule has no stable semantic basis. Query length, term rarity, field boosts, corpus growth, shard statistics, similarity implementation and query structure can move scores. Replace score thresholds with business rules grounded in explicit features (for example availability), judged ranking metrics, calibrated rerankers when appropriate, or query-specific acceptance logic whose semantics are tested.

7. First relevance fixture: judgments before tuning

judgments.csv · graded AtlasMart intent
query_id,query,document_id,grade
q1,wireless headphones,p1,3
q1,wireless headphones,p7,2
q1,wireless headphones,p2,1
q2,studio headphones,p2,3
q2,studio headphones,p8,3
q3,noise cancelling,p1,3
q3,noise cancelling,p4,3
q4,bluetooth speaker,p3,3
q4,bluetooth speaker,p1,1
q5,laptop stand,p5,3
q6,running shoes,p6,3

Grades encode information need, not “truth forever.” A grade of 3 means strongly relevant for this lab’s stated intent; 0 or omission does not automatically mean universally irrelevant. Store query text, locale/device/segment context when relevant, judgment provenance and date. The later evaluation lesson will turn this file into metrics and release gates.

Check your understanding

  1. Why can a rare term contribute more than a common term?
  2. What does increasing k1 generally change?
  3. Why is _score not a probability?
  4. What is Explain good for?
  5. When should BM25 parameters be changed?
Review the answers

1. Its inverse document frequency is higher, so matching that term provides more discriminatory evidence within the indexed corpus.

2. It changes term-frequency saturation so repeated occurrences can keep adding score for longer; it does not make relevance objectively better.

3. It is the output of the current scoring formula and query/corpus context, not a calibrated estimate bounded to a stable probability interpretation.

4. Diagnosing why one document did or did not receive its score under one query, including term statistics, norms and boosts.

5. Only from a clear retrieval hypothesis evaluated against representative judgments and operational constraints, preferably in an isolated index/experiment.

Production judgment

Most teams should start with default BM25 and spend early relevance effort on mappings, analyzers, query intent, field selection and representative judgments. Similarity tuning can matter, but it is a high-leverage expert control whose benefit must survive query segments, corpus growth and upgrades. Record version, index generation, analyzer, similarity parameters and evaluation set with every experiment.

Summary and next step

BM25 converts lexical evidence into a query-relative ranking. The next lesson controls how evidence from several fields is combined so “name match,” “description match” and “brand/category match” express AtlasMart’s intent deliberately rather than accidentally.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.