Chapter 07 · Relevance Engineering: BM25, Field Weighting, Function Score, Ranking, and Evaluation
Build a Relevance Evaluation Set with Queries, Judgments, Metrics, Baselines, and Regression Thresholds
Create a reproducible relevance-evaluation harness with representative queries, graded judgments, Precision/Recall/MRR/NDCG metrics, baseline-versus-candidate reports, segment checks and regression thresholds suitable for CI.
Learning outcomes
Relevance work becomes reliable when AtlasMart can answer “did this candidate improve search?” with a versioned query set, explicit judgments, defined metrics and a pre-agreed release rule. Screenshots and memorable demo queries are useful diagnostics, but they are not a regression suite. This lesson builds the smallest rigorous evaluation loop that can grow with production traffic.
Design representative query sets and graded relevance judgments with provenance, segment context and review dates.
Choose Precision@k, Recall@k, reciprocal-rank/MRR and DCG/NDCG-style metrics based on the user task rather than metric popularity.
Run portable _rank_eval requests where appropriate and maintain an independent local metric harness for transparent acceptance logic.
Compare baseline and candidate per query and per segment, not only by one aggregate number.
Define regression thresholds and release gates that include correctness, quality, latency/resources, failures and product-version compatibility.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with one primary shard and zero replicas unless a section says otherwise. Both use BM25 by default, but raw score magnitude is not a compatibility contract. OpenSearch 3.x changed its default implementation from LegacyBM25 to Lucene-native BM25, which can change numeric score scale without changing relative ranking. Both products expose a Ranking Evaluation API, but the portable course discipline is the judged query/document set itself. OpenSearch also provides Search Relevance Workbench in newer versions; it is optional tooling, not the only learning path.
The generation environment does not provide Docker or live Elasticsearch/OpenSearch clusters, so cluster commands and expected result ordering are specified as reproducible acceptance tests rather than represented as captured output. Do not copy illustrative score numbers into tests. Run the fixture against each supported product/version, record actual ranked IDs and metrics, and clean up only the dedicated AtlasMart indices.
1. Build the test collection before the metric dashboard
query_id,query,document_id,grade
q1,wireless headphones,p1,3
q1,wireless headphones,p7,2
q1,wireless headphones,p2,1
q2,studio headphones,p2,3
q2,studio headphones,p8,3
q3,noise cancelling,p1,3
q3,noise cancelling,p4,3
q4,bluetooth speaker,p3,3
q4,bluetooth speaker,p1,1
q5,laptop stand,p5,3
q6,running shoes,p6,3
A useful judgment record should eventually include query ID/text, locale, category/intent segment, document ID, grade scale definition, assessor/provenance and timestamp. Mix known-item, category, descriptive and ambiguous queries. Sample from real privacy-reviewed traffic when possible; synthetic queries are a bootstrap, not a permanent substitute.
| Metric | Answers | Good for | Blind spot |
|---|---|---|---|
| Precision@k | What fraction of top k is relevant? | Top-results cleanliness | Ignores relevant results below k; binary threshold loses grade nuance. |
| Recall@k | What fraction of known relevant documents appear in top k? | Coverage-sensitive tasks | Requires sufficiently complete relevance judgments. |
| Reciprocal rank / MRR | How early is the first relevant result? | Known-item / first-answer tasks | Ignores later relevant results once first is found. |
| DCG/NDCG@k | Are highly graded documents near the top? | Graded ranking quality | Depends on judgment quality and gain/discount assumptions. |
2. Use _rank_eval for repeatable server-side evaluation
GET atlasmart-products-relevance-v1/_rank_eval
{
"requests": [
{
"id": "q1_wireless_headphones",
"request": {
"query": {
"multi_match": {
"query": "wireless headphones",
"fields": ["name^3", "description", "brand_text^1.5"],
"type": "best_fields",
"tie_breaker": 0.15
}
}
},
"ratings": [
{ "_index":"atlasmart-products-relevance-v1", "_id":"p1", "rating":3 },
{ "_index":"atlasmart-products-relevance-v1", "_id":"p7", "rating":2 },
{ "_index":"atlasmart-products-relevance-v1", "_id":"p2", "rating":1 }
]
},
{
"id": "q3_noise_cancelling",
"request": {
"query": {
"multi_match": {
"query": "noise cancelling",
"fields": ["name^3", "description"]
}
}
},
"ratings": [
{ "_index":"atlasmart-products-relevance-v1", "_id":"p1", "rating":3 },
{ "_index":"atlasmart-products-relevance-v1", "_id":"p4", "rating":3 }
]
}
],
"metric": {
"dcg": {
"k": 5,
"normalize": true
}
}
}
Run the request against the baseline query, then against a
candidate that changes one ranking mechanism. Store the request
body with the code/configuration version. Examine both overall
metric_score and per-query details. A higher
aggregate score does not excuse a catastrophic protected-query
regression.
Elasticsearch and OpenSearch both expose _rank_eval, but request/response options and surrounding relevance tooling can diverge. Keep your source-of-truth judgments in a product-neutral file and verify each target product/version.
3. Keep a transparent local metric implementation
from math import log2
def dcg(grades, k):
return sum((2**g - 1) / log2(i + 2) for i, g in enumerate(grades[:k]))
def ndcg(ranked_ids, grades_by_id, k=5):
observed = [grades_by_id.get(doc_id, 0) for doc_id in ranked_ids[:k]]
ideal = sorted(grades_by_id.values(), reverse=True)[:k]
ideal_score = dcg(ideal, k)
return 0.0 if ideal_score == 0 else dcg(observed, k) / ideal_score
def reciprocal_rank(ranked_ids, relevant_ids, k=10):
for rank, doc_id in enumerate(ranked_ids[:k], start=1):
if doc_id in relevant_ids:
return 1.0 / rank
return 0.0
# Populate ranked_ids with ACTUAL IDs returned by your baseline/candidate search.
grades = {"p1": 3, "p7": 2, "p2": 1}
ranked_ids = ["p1", "p7", "p2"] # example ordering only; replace from real response
print("NDCG@3", ndcg(ranked_ids, grades, 3))
print("RR@3", reciprocal_rank(ranked_ids, {"p1", "p7"}, 3))
The example ordering in this code is intentionally labeled illustrative. Replace it with IDs captured from a real search response. A local evaluator is useful for CI, comparing several retrievers and auditing metric assumptions. Unit-test the metric code itself with hand-computable cases before it becomes a release gate.
4. Baseline-versus-candidate report
experiment_id: rel-2026-09-11-01
product: elasticsearch|opensearch
server_version: 9.5.3|3.8.0
index_generation: atlasmart-products-relevance-v1
query_config_baseline: lexical-v1
query_config_candidate: lexical-name3-tie015-v2
judgment_set: atlasmart-relevance-judgments-v1
metrics: [ndcg@5, mrr@10, precision@5]
segments: [known_item, descriptive, category, brand_plus_type]
for each query:
baseline_ranked_ids
candidate_ranked_ids
grades_at_rank
metric_delta
shard_failures
timed_out
aggregate:
mean/median metric delta
regression_count + worst regressions
performance (separate unprofiled run):
p50/p95/p99 latency, throughput, concurrency, CPU/heap/cache, shard count
release_decision + reviewer + rollback_config
Do not mix quality-run latency gathered with
profile:true into production performance gates.
Profile adds overhead. Run quality and performance evidence as
separate experiments on the same candidate configuration.
5. Regression thresholds are product policy
Example release policy (choose thresholds from your own established baseline, not from this text):
FAIL the candidate if any of these is true:
- a critical known-item query loses its grade-3 item from top 3;
- NDCG@5 for any protected segment regresses beyond its approved tolerance;
- aggregate NDCG@5 improves but more than the allowed fraction of queries regress materially;
- _shards.failed > 0 or timed_out=true during a quality run;
- p95/p99 latency or CPU/heap/resource budget exceeds the separately measured performance SLO;
- feature freshness/provenance is missing for a ranking signal;
- results differ between Elasticsearch/OpenSearch in a way not covered by the support matrix.
PASS only when quality, correctness, failure handling and performance gates all pass.
There is no universal acceptable NDCG regression. A medical, security or known-item search may protect specific queries more strongly than a discovery feed. Define critical-query and segment rules before reviewing the candidate. Otherwise teams are tempted to move thresholds after seeing a desired feature’s results.
6. Wrong approach: optimize the demo queries
Tune boosts until “wireless headphones” and “laptop stand” look excellent, then ship because the average of two hand-picked demos improved.
This creates selection bias and encourages overfitting. Expand the query set, stratify by intent/frequency/locale/category where relevant, freeze a holdout slice, and inspect regressions—not only aggregate wins. Production experiments can later validate behavioral impact, but offline relevance evaluation remains essential for fast, deterministic regression detection.
7. OpenSearch Search Relevance Workbench is optional tooling
OpenSearch 3.x includes Search Relevance Workbench capabilities for query sets, judgments and experiments. It can accelerate experimentation, but the academy’s mandatory path remains the free/local index, portable judgment file and transparent metrics. This keeps the learning objective reproducible on both products and prevents a UI/plugin from becoming the conceptual model.
Check your understanding
- Why can aggregate NDCG improve while the release should still fail?
- When is MRR especially useful?
- Why store judgments outside product-specific tooling?
- Why measure performance separately from Profile output?
- What should a relevance release gate include besides one metric?
Review the answers
1. A protected query or user segment may regress materially; aggregate metrics can hide localized harm.
2. When users mainly need the first relevant result quickly, such as known-item or direct-answer search.
3. It preserves portability, provenance and reproducibility across Elasticsearch/OpenSearch/tool changes.
4. Profile adds overhead and omits parts of end-to-end latency, so it is diagnostic rather than an SLO benchmark.
5. Critical-query/segment quality, correctness/failures, unprofiled performance/resources, feature freshness/provenance, compatibility and rollback conditions.
Production judgment
Relevance engineering is experimental discipline: representative information needs, explicit judgments, versioned query configurations, per-segment metrics, reproducible environments, protected regressions, separate performance evidence and rollback. Update the test set as catalog/user behavior changes, but preserve historical suites so “improvement” cannot be defined only by the newest traffic.
Summary and next step
Chapter 07 moved AtlasMart from matching to measured ranking quality: BM25 evidence, field combination, bounded business signals, score diagnostics and judged-query release gates. Chapter 08 will add aggregations, where quality questions shift from ranking order to bucket semantics, distributed approximation, memory and analytical correctness.
Authoritative references
- Elastic: BM25 and similarity settings — Default BM25 behavior, k1/b parameters and expert similarity configuration.
- Elastic: multi_match query — Field boosts and best_fields/most_fields/cross_fields semantics.
- Elastic: dis_max query — Best-clause scoring plus tie_breaker contribution.
- Elastic: function_score query — Business-signal scoring, decay functions, score_mode and boost_mode.
- Elastic: rank_feature query — Optimized numeric ranking features and scoring functions.
- Elastic: ranking evaluation — Judged-query evaluation with precision, recall, MRR and DCG-family metrics.
- OpenSearch: keyword search and BM25 — BM25 fundamentals and the OpenSearch 3.x LegacyBM25-to-BM25 score-scale change.
- OpenSearch: multi_match queries — Field boosts and multi-field matching modes.
- OpenSearch: function_score query — Function-score composition, decay and score combination.
- OpenSearch: rank_feature query — Rank-feature field/query behavior and saturation/log/sigmoid options.
- OpenSearch: Explain API — Per-document score explanation and diagnostic limitations.
- OpenSearch: Profile API — Search execution timing with explicit profiling overhead/coverage limits.
- OpenSearch: Ranking Evaluation API — Judged-query ranking-quality evaluation.
- OpenSearch: Search Relevance Workbench — Optional query-set/judgment/experiment UI and plugin workflow for OpenSearch.
- OpenSearch: Evaluating search quality — Pointwise relevance evaluation in Search Relevance Workbench.