Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation

Function Score, Decay, Rank Feature, Business Signals, Freshness, and Popularity Without Drowning Text Relevance

Add bounded business signals without allowing popularity or freshness to erase lexical intent: compare function_score, decay functions and rank_feature, inspect their score contribution, and gate changes with relevance judgments.

Intermediate100–120 minutesJudged-query relevance labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart wants popular and recently launched products to get a modest advantage—but never wants a best-selling laptop stand to outrank headphones for the query wireless headphones. Business signals must therefore modify an already meaningful candidate set, use bounded transforms, and be judged against both relevance and bias/freshness risks.

01

Explain function_score as a post-retrieval scoring layer over documents already matched by a query, including score_mode and boost_mode.

02

Use field_value_factor and date decay with bounded transforms instead of adding raw orders or timestamps to BM25.

03

Use rank_feature/rank_feature query for purpose-built numeric ranking signals and understand saturation/log/sigmoid-style diminishing returns.

04

Detect and repair “popularity drowning lexical relevance” with judged queries, explicit caps/weights and query-segment analysis.

05

Treat feature freshness, missing values, manipulation/bias and update cost as ranking-system design constraints.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. The shared fixture stores both orders_30d (ordinary integer for function_score) and popularity (rank_feature) so the two mechanisms can be compared without changing product data.

Execution and safety note

The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.

1. Business signals are priors, not substitutes for query intent

Lexical relevance estimates how well document text matches the query. A business signal is separate evidence such as recent sales, rating, freshness or editorial quality. Some signals are useful priors: if two products match text similarly, stronger popularity can break the tie. They become dangerous when their scale is so large that text intent stops mattering.

Signal Potential benefit Risk requiring governance
orders_30d Current demand/popularity Self-reinforcing rich-get-richer loop; campaigns/bots can distort.
launched_at Freshness for fast-moving categories Older evergreen products can be unfairly buried.
rating Quality prior Small-sample ratings and selection bias.
availability Usually an eligibility/filter rule Treating out-of-stock as a small score penalty may violate product policy.
profit/editorial Business objectives Can conflict with user relevance; requires explicit policy and experiment review.

2. function_score: transform, weight, combine

Dev Tools · bounded popularity + freshness candidate
GET atlasmart-products-relevance-v1/_search
{
  "query": {
    "function_score": {
      "query": {
        "multi_match": {
          "query": "wireless headphones",
          "fields": ["name^3", "description", "brand_text^1.5"],
          "type": "best_fields",
          "tie_breaker": 0.15
        }
      },
      "functions": [
        {
          "field_value_factor": {
            "field": "orders_30d",
            "modifier": "log1p",
            "factor": 0.15,
            "missing": 0
          },
          "weight": 0.20
        },
        {
          "gauss": {
            "launched_at": {
              "origin": "2026-09-11",
              "scale": "60d",
              "offset": "7d",
              "decay": 0.5
            }
          },
          "weight": 0.15
        }
      ],
      "score_mode": "sum",
      "boost_mode": "sum",
      "max_boost": 1.5
    }
  }
}

The exact weights are intentionally experiment values, not recommendations. log1p compresses a large orders range so 1,800 orders do not contribute 1,800 times the score of one order. A Gaussian decay turns age into a smooth bounded signal. score_mode combines functions; boost_mode combines the function result with the lexical query. max_boost constrains the aggregate function contribution but does not by itself prove relevance safety.

3. rank_feature is purpose-built for numeric ranking evidence

Dev Tools · lexical must + popularity should
GET atlasmart-products-relevance-v1/_search
{
  "query": {
    "bool": {
      "must": {
        "multi_match": {
          "query": "wireless headphones",
          "fields": ["name^3", "description"]
        }
      },
      "should": [
        {
          "rank_feature": {
            "field": "popularity",
            "saturation": { "pivot": 60 },
            "boost": 0.25
          }
        }
      ]
    }
  }
}

rank_feature is mapped specifically for scoring rather than general aggregations/sorts. Saturation is useful when a business value should have diminishing returns: being moderately popular can help, while the difference between “very popular” and “extremely popular” should not dominate the text match. Both Elasticsearch and OpenSearch expose rank-feature concepts, but advanced parameters and performance behavior must be verified per target version.

4. Controlled failure: let popularity drown relevance

Deliberately wrong experiment

Replace the lexical score with a large raw popularity/order signal (for example boost_mode=replace with a high-scale field value) and then declare the result “better” because popular products rise.

Run wireless headphones. If p5 or p6—high-demand but lexically irrelevant—enters the top results, the experiment failed the basic intent constraint. The repair is not “reduce weight until it looks okay” on this one query. Restore lexical candidate semantics, use bounded transforms, add representative judged queries and choose thresholds from aggregate plus segment metrics.

5. Freshness is a time-dependent feature contract

A decay query has an origin (the reference point), scale, optional offset and target decay. Hard-coding “now” into an evaluation makes tomorrow’s result different even if code/data are unchanged. For reproducible offline tests, pin the origin date. In production, a dynamic current-time origin is valid, but your evaluation harness should record the evaluation timestamp and feature snapshot.

Freshness non-guarantee

A decay function is not lifecycle/retention. It changes ranking contribution; it does not delete old documents, guarantee freshness of indexed business fields, or replace an availability/filter policy.

6. Feature quality includes update mechanics and abuse resistance

  • Record where each feature comes from and how stale it may be.
  • Define missing-value behavior explicitly; missing should not accidentally get the strongest score.
  • Normalize/cap skewed features before they reach scoring.
  • Separate business goals from security/authorization. Tenant or inventory permissions are filters, never score bonuses.
  • Measure update frequency and indexing cost. A signal refreshed every second can dominate write load.
  • Audit manipulation: fake clicks/orders/reviews can become ranking attacks.

Check your understanding

  1. Why should business signals normally modify a lexical query instead of replace it?
  2. Why use log or saturation transforms for popularity?
  3. What does boost_mode control?
  4. Why pin the date origin in offline freshness evaluation?
  5. What is one non-relevance risk of a popularity signal?
Review the answers

1. The lexical query preserves user-intent eligibility; business signals then refine ordering among relevant candidates.

2. They bound/dampen large numeric ranges so extreme values have diminishing influence instead of overwhelming text relevance.

3. How the combined function score is combined with the original query score, such as sum, multiply or replace.

4. Otherwise rankings change simply because wall-clock time advances, making experiments irreproducible.

5. It can encode feedback loops or be manipulated by bot/fraud activity, so provenance and abuse resistance matter.

Production judgment

Every business signal needs an owner, freshness SLA, transform, missing-value rule, upper influence bound, manipulation model and evaluation evidence. Report metrics by query/category/locale/tenant segments where applicable; an aggregate improvement can conceal severe regressions for a smaller segment. Keep a pure lexical baseline available for diagnosis and rollback.

Summary and next step

Business signals can improve ranking only when their contribution is bounded and measurable. The next lesson builds a debugging discipline that separates “why did this document score this way?” from “why is this query slow?” and from “is this score comparable to another retrieval system?”

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.