Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation
Field Boosts, DisMax, Tie Breakers, Cross/Best/Most Fields, and Multi-Match Strategy
Control how lexical evidence from several fields is combined: compare field boosts, dis_max/tie_breaker, and multi_match best_fields, most_fields and cross_fields strategies against judged AtlasMart queries instead of choosing by intuition.
Learning outcomes
AtlasMart’s query wireless headphones can match a
product name, a long description, a brand phrase or several
fields at once. Relevance engineering begins by deciding whether
one best field should dominate, evidence should accumulate
across fields, or terms should be treated as one logical
person/name-like field set. multi_match and
dis_max express these different policies.
Use per-field boosts as explicit relevance policy and explain why boosts must be evaluated rather than chosen by round numbers.
Explain dis_max scoring and tie_breaker as “best clause plus bounded evidence from additional clauses.”
Choose best_fields, most_fields or cross_fields from field semantics and analyzer compatibility rather than trial-and-error.
Observe ranked IDs, matched clauses and Explain output for representative queries while avoiding frozen numeric-score assertions.
Create baseline-versus-candidate experiments whose only changed variable is field-combination policy.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. The portable lab uses full-text fields with the same standard analysis unless noted. Product-specific alternatives can be explored only after the common baseline is understood.
The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.
1. A field boost is a statement of product intent
If a shopper types a product phrase, AtlasMart usually wants a
name match to count more than the same words buried in a long
description. A boost such as name^3 says exactly
that. It does not mean the name is “three times more relevant”
in any user-calibrated sense; it multiplies that clause’s
scoring contribution within the query.
| Policy question | Likely mechanism | Failure mode |
|---|---|---|
| Should one strong field dominate? | best_fields / dis_max | A weak match in many fields may otherwise outrank a precise title match. |
| Should evidence from several analyzed representations accumulate? | most_fields | Correlated duplicate fields can double-count similar evidence. |
| Are several short same-analyzer fields one logical entity? | cross_fields | Different analyzers/boosts/statistics can make blending hard to reason about. |
| Should secondary field matches help a little? | dis_max tie_breaker | A large tie_breaker quietly turns “best field” toward accumulation. |
2. best_fields is a dis_max policy
With best_fields, the strongest matching field
supplies the primary score. A nonzero
tie_breaker adds a fraction of the other matching
field scores. That is useful when a product matching both name
and description should gain modest confidence without letting
verbose descriptions dominate.
GET atlasmart-products-relevance-v1/_search
{
"query": {
"dis_max": {
"queries": [
{ "match": { "name": { "query": "wireless headphones", "boost": 3.0 } } },
{ "match": { "description": { "query": "wireless headphones" } } }
],
"tie_breaker": 0.15
}
}
}
Compare this result to the equivalent multi_match
best_fields query. The acceptance test is ordered
IDs plus explanation shape—not equality of a remembered
floating-point score.
3. best_fields, most_fields and cross_fields solve different problems
# A. best_fields: strongest single field dominates; optional tie_breaker adds evidence from others
GET atlasmart-products-relevance-v1/_search
{
"query": {
"multi_match": {
"query": "wireless headphones",
"type": "best_fields",
"fields": ["name^3", "description", "brand_text^1.5"],
"tie_breaker": 0.15
}
}
}
# B. most_fields: contributions from multiple analyzed representations/fields accumulate
GET atlasmart-products-relevance-v1/_search
{
"query": {
"multi_match": {
"query": "wireless headphones",
"type": "most_fields",
"fields": ["name^3", "description", "brand_text^1.5"]
}
}
}
# C. cross_fields: term-centric treatment across compatible analyzed fields
GET atlasmart-products-relevance-v1/_search
{
"query": {
"multi_match": {
"query": "Auralux headphones",
"type": "cross_fields",
"fields": ["name", "brand_text"],
"operator": "and"
}
}
}
| Mode | Mental model | Good fit | Caution |
|---|---|---|---|
| best_fields | Use strongest field; optionally add tie-break evidence | title/name + description | Field-centric operator/minimum_should_match behavior can surprise. |
| most_fields | Sum evidence from matching fields | same content indexed with several analysis strategies | Can over-reward duplicated/correlated representations. |
| cross_fields | Blend compatible field statistics; terms can land across fields | first/last name or similarly analyzed short fields | Different analyzers split fields into groups; boosts and norms complicate interpretation. |
4. Controlled experiment: only change one variable
Run query q1 and q2 from the judgment
file with these candidates:
baseline: multi_match best_fields fields=[name,description,brand_text] tie_breaker=0
candidate A: best_fields fields=[name^3,description,brand_text^1.5] tie_breaker=0.15
candidate B: most_fields fields=[name^3,description,brand_text^1.5]
candidate C: cross_fields fields=[name,brand_text] operator=and # only for entity-like query q="Auralux headphones"
For each run capture:
- ordered hit IDs
- judged grades at each rank
- _shards.failed and timed_out
- product/version and index generation
- evaluation metric(s) from Lesson 5
Do not change analyzers, fixture data and field-combination mode in the same experiment.
5. Wrong approach: make every field important
fields=["name^10","description^8","brand_text^6","*"] because “more signals should improve recall.”
This confuses retrieval breadth with ranking quality, can include fields never intended for full-text search, increases clause/resource cost and makes explanations opaque. Restrict searchable fields, use boosts tied to an explicit hypothesis, and inspect failures by query segment. If “brand + product type” is a distinct intent, test it as a segment instead of forcing one giant query to solve every intent equally.
6. Cross-product compatibility discipline
Both products support the core multi_match modes
used here, but defaults and clause limits are operational
settings, not a portable performance guarantee. Elastic
currently documents a default Boolean-clause limit different
from the OpenSearch documentation. Therefore never teach “N
fields × M terms is always safe”; keep field lists bounded and
validate the exact deployment settings.
Check your understanding
- What does best_fields optimize for?
- What does tie_breaker do?
- When is most_fields appropriate?
- Why can cross_fields be difficult to reason about?
- Why avoid changing boosts and analyzers in one experiment?
Review the answers
1. A document with one strongly matching field; it is commonly implemented through dis_max semantics.
2. It adds a bounded fraction of scores from additional matching clauses on top of the best clause.
3. When several fields or analyzed representations contain complementary evidence whose scores are intended to accumulate.
4. It blends term statistics across compatible analyzer groups, while boosts, analyzers and norms can change how that blending behaves.
5. You lose causal attribution: an observed metric change cannot be tied to one mechanism.
Production judgment
Field-combination strategy belongs in versioned search templates/configurations with representative query judgments. Monitor head, torso and tail queries separately; a title-heavy boost that helps known-item searches may hurt descriptive discovery queries. Re-run evaluation after mapping/analyzer changes because those alter the evidence being combined.
Summary and next step
You can now combine lexical evidence deliberately. The next lesson introduces non-text business signals—popularity, freshness and operational priorities—while protecting the lexical relevance you just made measurable.
Authoritative references
- Elastic: BM25 and similarity settings — Default BM25 behavior, k1/b parameters and expert similarity configuration.
- Elastic: multi_match query — Field boosts and best_fields/most_fields/cross_fields semantics.
- Elastic: dis_max query — Best-clause scoring plus tie_breaker contribution.
- Elastic: function_score query — Business-signal scoring, decay functions, score_mode and boost_mode.
- Elastic: rank_feature query — Optimized numeric ranking features and scoring functions.
- Elastic: ranking evaluation — Judged-query evaluation with precision, recall, MRR and DCG-family metrics.
- OpenSearch: keyword search and BM25 — BM25 fundamentals and the OpenSearch 3.x LegacyBM25-to-BM25 score-scale change.
- OpenSearch: multi_match queries — Field boosts and multi-field matching modes.
- OpenSearch: function_score query — Function-score composition, decay and score combination.
- OpenSearch: rank_feature query — Rank-feature field/query behavior and saturation/log/sigmoid options.
- OpenSearch: Explain API — Per-document score explanation and diagnostic limitations.
- OpenSearch: Profile API — Search execution timing with explicit profiling overhead/coverage limits.
- OpenSearch: Ranking Evaluation API — Judged-query ranking-quality evaluation.