Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation

Field Boosts, DisMax, Tie Breakers, Cross/Best/Most Fields, and Multi-Match Strategy

Control how lexical evidence from several fields is combined: compare field boosts, dis_max/tie_breaker, and multi_match best_fields, most_fields and cross_fields strategies against judged AtlasMart queries instead of choosing by intuition.

Intermediate100–120 minutesJudged-query relevance labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart’s query wireless headphones can match a product name, a long description, a brand phrase or several fields at once. Relevance engineering begins by deciding whether one best field should dominate, evidence should accumulate across fields, or terms should be treated as one logical person/name-like field set. multi_match and dis_max express these different policies.

01

Use per-field boosts as explicit relevance policy and explain why boosts must be evaluated rather than chosen by round numbers.

02

Explain dis_max scoring and tie_breaker as “best clause plus bounded evidence from additional clauses.”

03

Choose best_fields, most_fields or cross_fields from field semantics and analyzer compatibility rather than trial-and-error.

04

Observe ranked IDs, matched clauses and Explain output for representative queries while avoiding frozen numeric-score assertions.

05

Create baseline-versus-candidate experiments whose only changed variable is field-combination policy.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. The portable lab uses full-text fields with the same standard analysis unless noted. Product-specific alternatives can be explored only after the common baseline is understood.

Execution and safety note

The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.

1. A field boost is a statement of product intent

If a shopper types a product phrase, AtlasMart usually wants a name match to count more than the same words buried in a long description. A boost such as name^3 says exactly that. It does not mean the name is “three times more relevant” in any user-calibrated sense; it multiplies that clause’s scoring contribution within the query.

Policy question Likely mechanism Failure mode
Should one strong field dominate? best_fields / dis_max A weak match in many fields may otherwise outrank a precise title match.
Should evidence from several analyzed representations accumulate? most_fields Correlated duplicate fields can double-count similar evidence.
Are several short same-analyzer fields one logical entity? cross_fields Different analyzers/boosts/statistics can make blending hard to reason about.
Should secondary field matches help a little? dis_max tie_breaker A large tie_breaker quietly turns “best field” toward accumulation.

2. best_fields is a dis_max policy

With best_fields, the strongest matching field supplies the primary score. A nonzero tie_breaker adds a fraction of the other matching field scores. That is useful when a product matching both name and description should gain modest confidence without letting verbose descriptions dominate.

Dev Tools · explicit dis_max
GET atlasmart-products-relevance-v1/_search
{
  "query": {
    "dis_max": {
      "queries": [
        { "match": { "name":        { "query": "wireless headphones", "boost": 3.0 } } },
        { "match": { "description": { "query": "wireless headphones" } } }
      ],
      "tie_breaker": 0.15
    }
  }
}

Compare this result to the equivalent multi_match best_fields query. The acceptance test is ordered IDs plus explanation shape—not equality of a remembered floating-point score.

3. best_fields, most_fields and cross_fields solve different problems

Dev Tools · three field-combination strategies
# A. best_fields: strongest single field dominates; optional tie_breaker adds evidence from others
GET atlasmart-products-relevance-v1/_search
{
  "query": {
    "multi_match": {
      "query": "wireless headphones",
      "type": "best_fields",
      "fields": ["name^3", "description", "brand_text^1.5"],
      "tie_breaker": 0.15
    }
  }
}

# B. most_fields: contributions from multiple analyzed representations/fields accumulate
GET atlasmart-products-relevance-v1/_search
{
  "query": {
    "multi_match": {
      "query": "wireless headphones",
      "type": "most_fields",
      "fields": ["name^3", "description", "brand_text^1.5"]
    }
  }
}

# C. cross_fields: term-centric treatment across compatible analyzed fields
GET atlasmart-products-relevance-v1/_search
{
  "query": {
    "multi_match": {
      "query": "Auralux headphones",
      "type": "cross_fields",
      "fields": ["name", "brand_text"],
      "operator": "and"
    }
  }
}
Mode Mental model Good fit Caution
best_fields Use strongest field; optionally add tie-break evidence title/name + description Field-centric operator/minimum_should_match behavior can surprise.
most_fields Sum evidence from matching fields same content indexed with several analysis strategies Can over-reward duplicated/correlated representations.
cross_fields Blend compatible field statistics; terms can land across fields first/last name or similarly analyzed short fields Different analyzers split fields into groups; boosts and norms complicate interpretation.

4. Controlled experiment: only change one variable

Run query q1 and q2 from the judgment file with these candidates:

experiment matrix
baseline:  multi_match best_fields fields=[name,description,brand_text] tie_breaker=0
candidate A: best_fields fields=[name^3,description,brand_text^1.5] tie_breaker=0.15
candidate B: most_fields fields=[name^3,description,brand_text^1.5]
candidate C: cross_fields fields=[name,brand_text] operator=and  # only for entity-like query q="Auralux headphones"

For each run capture:
- ordered hit IDs
- judged grades at each rank
- _shards.failed and timed_out
- product/version and index generation
- evaluation metric(s) from Lesson 5
Do not change analyzers, fixture data and field-combination mode in the same experiment.

5. Wrong approach: make every field important

Deliberately wrong candidate

fields=["name^10","description^8","brand_text^6","*"] because “more signals should improve recall.”

This confuses retrieval breadth with ranking quality, can include fields never intended for full-text search, increases clause/resource cost and makes explanations opaque. Restrict searchable fields, use boosts tied to an explicit hypothesis, and inspect failures by query segment. If “brand + product type” is a distinct intent, test it as a segment instead of forcing one giant query to solve every intent equally.

6. Cross-product compatibility discipline

Both products support the core multi_match modes used here, but defaults and clause limits are operational settings, not a portable performance guarantee. Elastic currently documents a default Boolean-clause limit different from the OpenSearch documentation. Therefore never teach “N fields × M terms is always safe”; keep field lists bounded and validate the exact deployment settings.

Check your understanding

  1. What does best_fields optimize for?
  2. What does tie_breaker do?
  3. When is most_fields appropriate?
  4. Why can cross_fields be difficult to reason about?
  5. Why avoid changing boosts and analyzers in one experiment?
Review the answers

1. A document with one strongly matching field; it is commonly implemented through dis_max semantics.

2. It adds a bounded fraction of scores from additional matching clauses on top of the best clause.

3. When several fields or analyzed representations contain complementary evidence whose scores are intended to accumulate.

4. It blends term statistics across compatible analyzer groups, while boosts, analyzers and norms can change how that blending behaves.

5. You lose causal attribution: an observed metric change cannot be tied to one mechanism.

Production judgment

Field-combination strategy belongs in versioned search templates/configurations with representative query judgments. Monitor head, torso and tail queries separately; a title-heavy boost that helps known-item searches may hurt descriptive discovery queries. Re-run evaluation after mapping/analyzer changes because those alter the evidence being combined.

Summary and next step

You can now combine lexical evidence deliberately. The next lesson introduces non-text business signals—popularity, freshness and operational priorities—while protecting the lexical relevance you just made measurable.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.