Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation
Function Score, Decay, Rank Feature, Business Signals, Freshness, and Popularity Without Drowning Text Relevance
Add bounded business signals without allowing popularity or freshness to erase lexical intent: compare function_score, decay functions and rank_feature, inspect their score contribution, and gate changes with relevance judgments.
Learning outcomes
AtlasMart wants popular and recently launched products to get a modest advantage—but never wants a best-selling laptop stand to outrank headphones for the query wireless headphones. Business signals must therefore modify an already meaningful candidate set, use bounded transforms, and be judged against both relevance and bias/freshness risks.
Explain function_score as a post-retrieval scoring layer over documents already matched by a query, including score_mode and boost_mode.
Use field_value_factor and date decay with bounded transforms instead of adding raw orders or timestamps to BM25.
Use rank_feature/rank_feature query for purpose-built numeric ranking signals and understand saturation/log/sigmoid-style diminishing returns.
Detect and repair “popularity drowning lexical relevance” with judged queries, explicit caps/weights and query-segment analysis.
Treat feature freshness, missing values, manipulation/bias and update cost as ranking-system design constraints.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. The shared fixture stores both orders_30d (ordinary integer for function_score) and popularity (rank_feature) so the two mechanisms can be compared without changing product data.
The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.
1. Business signals are priors, not substitutes for query intent
Lexical relevance estimates how well document text matches the query. A business signal is separate evidence such as recent sales, rating, freshness or editorial quality. Some signals are useful priors: if two products match text similarly, stronger popularity can break the tie. They become dangerous when their scale is so large that text intent stops mattering.
| Signal | Potential benefit | Risk requiring governance |
|---|---|---|
| orders_30d | Current demand/popularity | Self-reinforcing rich-get-richer loop; campaigns/bots can distort. |
| launched_at | Freshness for fast-moving categories | Older evergreen products can be unfairly buried. |
| rating | Quality prior | Small-sample ratings and selection bias. |
| availability | Usually an eligibility/filter rule | Treating out-of-stock as a small score penalty may violate product policy. |
| profit/editorial | Business objectives | Can conflict with user relevance; requires explicit policy and experiment review. |
2. function_score: transform, weight, combine
GET atlasmart-products-relevance-v1/_search
{
"query": {
"function_score": {
"query": {
"multi_match": {
"query": "wireless headphones",
"fields": ["name^3", "description", "brand_text^1.5"],
"type": "best_fields",
"tie_breaker": 0.15
}
},
"functions": [
{
"field_value_factor": {
"field": "orders_30d",
"modifier": "log1p",
"factor": 0.15,
"missing": 0
},
"weight": 0.20
},
{
"gauss": {
"launched_at": {
"origin": "2026-09-11",
"scale": "60d",
"offset": "7d",
"decay": 0.5
}
},
"weight": 0.15
}
],
"score_mode": "sum",
"boost_mode": "sum",
"max_boost": 1.5
}
}
}
The exact weights are intentionally experiment values,
not recommendations. log1p compresses a large
orders range so 1,800 orders do not contribute 1,800 times the
score of one order. A Gaussian decay turns age into a smooth
bounded signal. score_mode combines functions;
boost_mode combines the function result with the
lexical query. max_boost constrains the aggregate
function contribution but does not by itself prove relevance
safety.
3. rank_feature is purpose-built for numeric ranking evidence
GET atlasmart-products-relevance-v1/_search
{
"query": {
"bool": {
"must": {
"multi_match": {
"query": "wireless headphones",
"fields": ["name^3", "description"]
}
},
"should": [
{
"rank_feature": {
"field": "popularity",
"saturation": { "pivot": 60 },
"boost": 0.25
}
}
]
}
}
}
rank_feature is mapped specifically for scoring
rather than general aggregations/sorts. Saturation is useful
when a business value should have diminishing returns: being
moderately popular can help, while the difference between “very
popular” and “extremely popular” should not dominate the text
match. Both Elasticsearch and OpenSearch expose rank-feature
concepts, but advanced parameters and performance behavior must
be verified per target version.
4. Controlled failure: let popularity drown relevance
Replace the lexical score with a large raw popularity/order signal (for example boost_mode=replace with a high-scale field value) and then declare the result “better” because popular products rise.
Run wireless headphones. If p5 or p6—high-demand
but lexically irrelevant—enters the top results, the experiment
failed the basic intent constraint. The repair is not “reduce
weight until it looks okay” on this one query. Restore lexical
candidate semantics, use bounded transforms, add representative
judged queries and choose thresholds from aggregate plus segment
metrics.
5. Freshness is a time-dependent feature contract
A decay query has an origin (the reference point), scale, optional offset and target decay. Hard-coding “now” into an evaluation makes tomorrow’s result different even if code/data are unchanged. For reproducible offline tests, pin the origin date. In production, a dynamic current-time origin is valid, but your evaluation harness should record the evaluation timestamp and feature snapshot.
A decay function is not lifecycle/retention. It changes ranking contribution; it does not delete old documents, guarantee freshness of indexed business fields, or replace an availability/filter policy.
6. Feature quality includes update mechanics and abuse resistance
- Record where each feature comes from and how stale it may be.
- Define missing-value behavior explicitly; missing should not accidentally get the strongest score.
- Normalize/cap skewed features before they reach scoring.
- Separate business goals from security/authorization. Tenant or inventory permissions are filters, never score bonuses.
- Measure update frequency and indexing cost. A signal refreshed every second can dominate write load.
- Audit manipulation: fake clicks/orders/reviews can become ranking attacks.
Check your understanding
- Why should business signals normally modify a lexical query instead of replace it?
- Why use log or saturation transforms for popularity?
- What does boost_mode control?
- Why pin the date origin in offline freshness evaluation?
- What is one non-relevance risk of a popularity signal?
Review the answers
1. The lexical query preserves user-intent eligibility; business signals then refine ordering among relevant candidates.
2. They bound/dampen large numeric ranges so extreme values have diminishing influence instead of overwhelming text relevance.
3. How the combined function score is combined with the original query score, such as sum, multiply or replace.
4. Otherwise rankings change simply because wall-clock time advances, making experiments irreproducible.
5. It can encode feedback loops or be manipulated by bot/fraud activity, so provenance and abuse resistance matter.
Production judgment
Every business signal needs an owner, freshness SLA, transform, missing-value rule, upper influence bound, manipulation model and evaluation evidence. Report metrics by query/category/locale/tenant segments where applicable; an aggregate improvement can conceal severe regressions for a smaller segment. Keep a pure lexical baseline available for diagnosis and rollback.
Summary and next step
Business signals can improve ranking only when their contribution is bounded and measurable. The next lesson builds a debugging discipline that separates “why did this document score this way?” from “why is this query slow?” and from “is this score comparable to another retrieval system?”
Authoritative references
- Elastic: BM25 and similarity settings — Default BM25 behavior, k1/b parameters and expert similarity configuration.
- Elastic: multi_match query — Field boosts and best_fields/most_fields/cross_fields semantics.
- Elastic: dis_max query — Best-clause scoring plus tie_breaker contribution.
- Elastic: function_score query — Business-signal scoring, decay functions, score_mode and boost_mode.
- Elastic: rank_feature query — Optimized numeric ranking features and scoring functions.
- Elastic: ranking evaluation — Judged-query evaluation with precision, recall, MRR and DCG-family metrics.
- OpenSearch: keyword search and BM25 — BM25 fundamentals and the OpenSearch 3.x LegacyBM25-to-BM25 score-scale change.
- OpenSearch: multi_match queries — Field boosts and multi-field matching modes.
- OpenSearch: function_score query — Function-score composition, decay and score combination.
- OpenSearch: rank_feature query — Rank-feature field/query behavior and saturation/log/sigmoid options.
- OpenSearch: Explain API — Per-document score explanation and diagnostic limitations.
- OpenSearch: Profile API — Search execution timing with explicit profiling overhead/coverage limits.
- OpenSearch: Ranking Evaluation API — Judged-query ranking-quality evaluation.