Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation
TF/IDF Intuition, BM25 Parameters, Document Length, Rare Terms, and Why Scores Are Query-Relative
Build an evidence-based BM25 mental model: connect term frequency, inverse document frequency and field-length normalization to AtlasMart rankings, then prove why _score is meaningful only within the exact query/corpus/scoring context that produced it.
Learning outcomes
AtlasMart already returns the right broad set of products, but
product managers now ask why one headset appears above another
and whether a higher _score means “90% relevant.”
That is the wrong mental model. Ranking starts with evidence
produced by the analyzed query and indexed terms, then a
similarity such as BM25 turns corpus statistics into a relative
score for this query. The goal is to predict directional effects
and validate them against judgments—not to memorize one numeric
score.
Explain term frequency and inverse document frequency as evidence signals and connect them to BM25 saturation and field-length normalization.
Interpret BM25 k1 and b as expert controls rather than universal tuning knobs, and state what a change is expected to influence.
Use _explain to identify term-frequency, IDF, field-length and boost contributions for one document/query pair.
Explain why _score is query-, corpus-, shard/statistics-, version- and query-structure-relative rather than a probability or stable business KPI.
Establish a lexical baseline and judged-query habit before attempting field boosts or business signals.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. This lesson leaves BM25 at product defaults for the primary lab; any parameter experiment uses a separate disposable index and must be judged against the same query set.
The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.
1. Matching answers “eligible”; ranking answers “which first?”
A lexical query matches indexed terms derived from text analysis. Term frequency (TF) asks how often a query term occurs in a document field. Document frequency counts how many indexed documents contain that term; inverse document frequency (IDF) makes a term that appears in fewer documents more discriminative. BM25 combines these ideas with term-frequency saturation and document-length normalization.
| Signal | Direction | Boundary to remember |
|---|---|---|
| Term frequency | More occurrences can add evidence | Gain saturates; ten repetitions are not ten times as relevant. |
| Inverse document frequency | Rarer query terms usually contribute more | IDF changes when the corpus/index statistics change. |
| Field length | Longer fields can be normalized downward | Effect is controlled by b and field norms; it is not a generic “shorter is better” rule. |
| Boosts | Increase the contribution of a clause/field | They encode policy; a large boost can overpower useful lexical evidence. |
| _score | Orders hits for one scoring context | It is not a probability and is not safely comparable across unrelated queries/retrievers/versions. |
Conceptual BM25 contribution for one query term t in document d:
score(t,d) ~= IDF(t) * [ tf(t,d) * (k1 + 1) ]
---------------------------------
tf(t,d) + k1 * (1 - b + b * dl/avgdl)
Interpretation:
- IDF(t): rarer terms usually contribute more evidence.
- tf(t,d): repeated occurrences help, but k1 makes the gain saturate.
- dl/avgdl: longer fields are normalized according to b.
- field/query boosts multiply or otherwise combine with these lexical contributions.
Do NOT copy this as a cross-product exact-score oracle. Lucene/OpenSearch implementation details,
index statistics, boosts and query rewrites determine the actual explanation tree.
2. Build a corpus small enough to reason about
The ranking fixture deliberately reuses AtlasMart product names from earlier chapters and adds business fields used later. One primary shard removes cross-shard statistics as a confounder for the first experiments. The corpus is still synthetic and too small for production tuning; its purpose is to make cause and effect observable.
DELETE atlasmart-products-relevance-v1
PUT atlasmart-products-relevance-v1
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0
},
"mappings": {
"properties": {
"sku": { "type": "keyword" },
"name": { "type": "text", "fields": { "raw": { "type": "keyword" } } },
"description": { "type": "text" },
"brand_text": { "type": "text" },
"category": { "type": "keyword" },
"available": { "type": "boolean" },
"price": { "type": "double" },
"rating": { "type": "float" },
"orders_30d": { "type": "integer" },
"launched_at": { "type": "date" },
"popularity": { "type": "rank_feature" }
}
}
}
POST atlasmart-products-relevance-v1/_bulk?refresh=true
{ "index": { "_id": "p1" } }
{ "sku":"AM-AU-100", "name":"Wireless Noise Cancelling Headphones", "description":"Over-ear Bluetooth headphones with active noise cancellation for travel", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":199.99, "rating":4.7, "orders_30d":920, "launched_at":"2026-08-10", "popularity":95 }
{ "index": { "_id": "p2" } }
{ "sku":"AM-AU-200", "name":"Wired Studio Headphones", "description":"Closed-back monitoring headphones for studio recording and mixing", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":89.99, "rating":4.4, "orders_30d":310, "launched_at":"2025-11-15", "popularity":55 }
{ "index": { "_id": "p3" } }
{ "sku":"AM-WB-300", "name":"Portable Bluetooth Speaker", "description":"Compact waterproof wireless speaker for travel and outdoor use", "brand_text":"WaveBox", "category":"electronics/audio", "available":false, "price":79.99, "rating":4.2, "orders_30d":760, "launched_at":"2026-06-01", "popularity":80 }
{ "index": { "_id": "p4" } }
{ "sku":"AM-AU-400", "name":"USB-C Noise Cancelling Earbuds", "description":"In-ear USB-C earbuds with active noise cancellation", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":129.99, "rating":4.5, "orders_30d":410, "launched_at":"2026-09-01", "popularity":60 }
{ "index": { "_id": "p5" } }
{ "sku":"AM-WN-500", "name":"Aluminum Laptop Stand", "description":"Adjustable desktop stand for laptops with ventilated aluminum construction", "brand_text":"WorkNest", "category":"office/accessories", "available":true, "price":49.00, "rating":4.6, "orders_30d":1200, "launched_at":"2026-01-20", "popularity":97 }
{ "index": { "_id": "p6" } }
{ "sku":"AM-ST-600", "name":"Lightweight Running Shoes", "description":"Breathable road running shoes for daily training", "brand_text":"Stride", "category":"sports/running", "available":true, "price":69.00, "rating":4.1, "orders_30d":1500, "launched_at":"2026-08-25", "popularity":99 }
{ "index": { "_id": "p7" } }
{ "sku":"AM-CA-700", "name":"Wireless Headphones Carrying Case", "description":"Hard protective case sized for over-ear headphones", "brand_text":"CarryAll", "category":"electronics/accessories", "available":true, "price":24.00, "rating":4.8, "orders_30d":1800, "launched_at":"2026-07-03", "popularity":100 }
{ "index": { "_id": "p8" } }
{ "sku":"AM-AU-800", "name":"Audiophile Studio Monitoring Headphones", "description":"Reference monitoring headphones with neutral tuning, replaceable cable, studio adapter, hard case, documentation, spare pads and accessories for long mixing sessions", "brand_text":"Auralux", "category":"electronics/audio", "available":true, "price":249.00, "rating":4.9, "orders_30d":140, "launched_at":"2026-04-12", "popularity":45 }
3. Observe the ranking before explaining it
GET atlasmart-products-relevance-v1/_search
{
"track_total_hits": true,
"query": {
"match": {
"name": "wireless headphones"
}
},
"_source": ["sku","name","category","orders_30d","launched_at"]
}
Record the ordered IDs and _score values as
evidence for this exact run. The useful assertion is not “p1 has
score 2.73.” The useful assertion is whether the judged-intent
ordering is acceptable and why the explanation tree differs
between hits.
GET atlasmart-products-relevance-v1/_explain/p1
{
"query": {
"match": {
"name": "wireless headphones"
}
}
}
GET atlasmart-products-relevance-v1/_explain/p7
{
"query": {
"match": {
"name": "wireless headphones"
}
}
}
4. Read an Explain tree without worshipping it
The Explain API decomposes a score for one document under one query. Look for the analyzed term, IDF/doc-frequency statistics, term frequency, norm/field-length contribution and boosts. It is a causal debugging trace for that score calculation, not a performance profiler and not a global statement about the index. Running Explain broadly is expensive; use it on carefully chosen query/document pairs.
OpenSearch 3.0 switched its default from LegacyBM25Similarity to Lucene-native BM25Similarity. The documented change can lower numeric score magnitude by a constant factor while preserving ranking order. This is a concrete reason never to build product thresholds around raw _score magnitude.
5. k1 and b are hypotheses, not tuning folklore
| Parameter | What it controls | If increased, generally | Why not tune blindly |
|---|---|---|---|
| k1 | How quickly term-frequency contribution saturates | Repeated terms can keep contributing for longer | May reward repetition/spam-like text; value interacts with corpus and field semantics. |
| b | Strength of field-length normalization | Longer fields are penalized more relative to average length | Can hurt legitimately long product descriptions; field-length distribution matters. |
| discount_overlaps | Whether zero-position-increment overlap tokens count toward norm length | Depends on analyzer/token graph | Synonym/token-graph behavior makes this an analysis contract, not a casual switch. |
If AtlasMart has a clear hypothesis—such as “long descriptions are over-penalized despite graded judgments”—create a separate experiment index, keep the same documents and judgments, and compare quality metrics. Do not change production similarity because one query “looks better.”
PUT atlasmart-bm25-experiment
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0,
"similarity": {
"bm25_experiment": {
"type": "BM25",
"k1": 0.8,
"b": 0.3
}
}
},
"mappings": {
"properties": {
"name": { "type": "text", "similarity": "bm25_experiment" }
}
}
}
6. Wrong approach: turn _score into a product probability
Reject every hit with _score < 2.0 because scores below 2 mean low relevance.
This rule has no stable semantic basis. Query length, term rarity, field boosts, corpus growth, shard statistics, similarity implementation and query structure can move scores. Replace score thresholds with business rules grounded in explicit features (for example availability), judged ranking metrics, calibrated rerankers when appropriate, or query-specific acceptance logic whose semantics are tested.
7. First relevance fixture: judgments before tuning
query_id,query,document_id,grade
q1,wireless headphones,p1,3
q1,wireless headphones,p7,2
q1,wireless headphones,p2,1
q2,studio headphones,p2,3
q2,studio headphones,p8,3
q3,noise cancelling,p1,3
q3,noise cancelling,p4,3
q4,bluetooth speaker,p3,3
q4,bluetooth speaker,p1,1
q5,laptop stand,p5,3
q6,running shoes,p6,3
Grades encode information need, not “truth forever.” A grade of 3 means strongly relevant for this lab’s stated intent; 0 or omission does not automatically mean universally irrelevant. Store query text, locale/device/segment context when relevant, judgment provenance and date. The later evaluation lesson will turn this file into metrics and release gates.
Check your understanding
- Why can a rare term contribute more than a common term?
- What does increasing k1 generally change?
- Why is _score not a probability?
- What is Explain good for?
- When should BM25 parameters be changed?
Review the answers
1. Its inverse document frequency is higher, so matching that term provides more discriminatory evidence within the indexed corpus.
2. It changes term-frequency saturation so repeated occurrences can keep adding score for longer; it does not make relevance objectively better.
3. It is the output of the current scoring formula and query/corpus context, not a calibrated estimate bounded to a stable probability interpretation.
4. Diagnosing why one document did or did not receive its score under one query, including term statistics, norms and boosts.
5. Only from a clear retrieval hypothesis evaluated against representative judgments and operational constraints, preferably in an isolated index/experiment.
Production judgment
Most teams should start with default BM25 and spend early relevance effort on mappings, analyzers, query intent, field selection and representative judgments. Similarity tuning can matter, but it is a high-leverage expert control whose benefit must survive query segments, corpus growth and upgrades. Record version, index generation, analyzer, similarity parameters and evaluation set with every experiment.
Summary and next step
BM25 converts lexical evidence into a query-relative ranking. The next lesson controls how evidence from several fields is combined so “name match,” “description match” and “brand/category match” express AtlasMart’s intent deliberately rather than accidentally.
Authoritative references
- Elastic: BM25 and similarity settings — Default BM25 behavior, k1/b parameters and expert similarity configuration.
- Elastic: multi_match query — Field boosts and best_fields/most_fields/cross_fields semantics.
- Elastic: dis_max query — Best-clause scoring plus tie_breaker contribution.
- Elastic: function_score query — Business-signal scoring, decay functions, score_mode and boost_mode.
- Elastic: rank_feature query — Optimized numeric ranking features and scoring functions.
- Elastic: ranking evaluation — Judged-query evaluation with precision, recall, MRR and DCG-family metrics.
- OpenSearch: keyword search and BM25 — BM25 fundamentals and the OpenSearch 3.x LegacyBM25-to-BM25 score-scale change.
- OpenSearch: multi_match queries — Field boosts and multi-field matching modes.
- OpenSearch: function_score query — Function-score composition, decay and score combination.
- OpenSearch: rank_feature query — Rank-feature field/query behavior and saturation/log/sigmoid options.
- OpenSearch: Explain API — Per-document score explanation and diagnostic limitations.
- OpenSearch: Profile API — Search execution timing with explicit profiling overhead/coverage limits.
- OpenSearch: Ranking Evaluation API — Judged-query ranking-quality evaluation.