Benchmark AtlasMart vector retrieval with exact oracles, representative filters, p99, build time, capacity evidence, and explicit production gates.

Benchmark Vector Recall, p99 Latency, Index Build Time, Memory, and Filter Selectivity with Representative Queries

Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.

Intermediate → Advanced150–210 minutesVector benchmark & decision-gate lab · Chapter 25 · Lesson 05Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Build a representative vector benchmark that uses independent exact ground truth, judged queries, filters, and controlled client concurrency.

02

Measure recall@k, p95/p99 latency, index build time, index bytes, memory evidence, and filter selectivity without fabricated results.

03

Separate cold/warm cache behavior, ingestion/build phases, merge/refresh effects, and steady-state query performance.

04

Create a reproducible comparison record for Elasticsearch and OpenSearch without comparing transformed scores as if they were identical.

05

Choose a production candidate only when it satisfies quality, tail-latency, capacity, security, and operational-recovery gates.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned vector-search baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released September 3, 2026) and OpenSearch/OpenSearch Dashboards 3.8.0 (released August 4, 2026), using the course's established local TLS/auth conventions and bundled JVMs. The mandatory lab uses precomputed deterministic vectors, so no hosted model, GPU, paid inference endpoint, or proprietary API is required. Elasticsearch examples explicitly map dense_vector rather than relying on changing defaults; OpenSearch examples explicitly map knn_vector and use the built-in Lucene engine for the portable local path. No live cluster is available in this generation environment, so performance cells are labeled MEASURED and must be filled from the learner's run rather than invented.

1. AtlasMart problem: the fastest demo loses the products users need

A developer benchmarks one query against warm cache, reports average latency, and picks the configuration with the smallest number. The chosen system misses relevant filtered products for a minority tenant and has unpredictable p99 during indexing. The benchmark optimized a demo, not a service.

Wrong approach: one query, warm cache, mean latency.

A defensible vector benchmark needs query populations, exact ground truth, filter distributions, tail percentiles, indexing/build evidence, and quality gates.

2. Benchmark contract

Dimension Required evidence
Quality recall@k against exact reference; judged NDCG/MRR when business relevance matters
Latency p50/p95/p99 by query class, filter class, warm/cold state, and concurrency
Build/ingest documents/s, wall-clock index build time, refresh/merge state
Capacity store bytes, vector/native memory evidence, JVM/OS cache, replica/shard count
Filtering eligible-doc count/selectivity + recall/latency under each filter class
Operations recovery/restart behavior, snapshot/reindex plan, model/version traceability

3. Deterministic larger fixture generator

The eight-product fixture proves semantics; it is too small to prove performance. Generate a larger synthetic corpus only for systems mechanics. It still is not a substitute for representative production embeddings.

Generate reproducible normalized vectors without external libraries
import json, math, random
R=random.Random(25092026)
DIMS=32; N=5000

def unit(v):
    n=math.sqrt(sum(x*x for x in v)); return [x/n for x in v]

with open("atlasmart-v25.ndjson","w",encoding="utf-8") as f:
    for i in range(N):
        # Four broad clusters; deterministic noise creates realistic-ish neighborhoods.
        c=i % 4
        center=[0.0]*DIMS
        for j in range(8): center[c*8+j]=1.0
        v=unit([center[j] + R.gauss(0,0.25) for j in range(DIMS)])
        doc={"product_id":f"S-{i:05d}","tenant_id":f"tenant-{i%5}",
             "category":f"c{c}","embedding_model":"synthetic-v25","embedding":v}
        f.write(json.dumps({"index":{"_id":doc["product_id"]}})+"\n")
        f.write(json.dumps(doc)+"\n")
print("generated",N,"documents; seed=25092026; dims=32")

Label every chart/table from this corpus synthetic mechanics. Production decisions require real embedding distributions, query vectors, filter selectivity, and judged relevance.

4. Exact ground truth harness

Compute exact neighbors outside both products
import heapq, math

def cosine(a,b):
    return sum(x*y for x,y in zip(a,b))/(math.sqrt(sum(x*x for x in a))*math.sqrt(sum(y*y for y in b)))

def exact_topk(docs, q, k, predicate=lambda d: True):
    scored=((cosine(q,d["embedding"]), d["product_id"]) for d in docs if predicate(d))
    return [pid for score,pid in heapq.nlargest(k, scored)]

# Persist exact IDs for every benchmark query/filter tuple.
# The exact file is the quality oracle used for both Elasticsearch and OpenSearch.

5. Query matrix

Representative query dimensions
query_id  k   filter_class           concurrency  cache_state
Q01       10  none                   1            cold+warm
Q02       10  tenant=20% corpus      4            warm
Q03       10  tenant+category=5%     8            warm
Q04       50  none                   4            warm
Q05       10  highly-selective <1%   4            warm
# Add real AtlasMart query classes before production approval.

6. Timing harness rules

  • Use a monotonic client timer around the full HTTP request; record status/error and returned IDs.
  • Run a documented warm-up phase but keep cold-start observations separate rather than discarding them silently.
  • Use fixed concurrency and connection settings per run.
  • Do not run indexing/build and query benchmarks together unless the test is intentionally mixed-load.
  • Capture cluster/node/index stats before and after so tail latency can be correlated with GC, CPU, disk, merge, cache, rejection, and recovery signals from Chapters 14–16 and 24.
  • Repeat enough samples to make p99 meaningful; state sample count.

7. Result schema: no invented numbers

Benchmark CSV
run_id,product,version,engine,index_options,model,dims,docs,shards,replicas,query_set,filter_class,k,candidates_or_ef,recall_at_k,p50_ms,p95_ms,p99_ms,index_build_s,index_bytes,memory_evidence,sample_count,notes
RUN-001,elasticsearch,9.5.3,elastic,int8_hnsw,synthetic-v25,32,5000,1,0,Q01,none,10,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,fill from run
RUN-002,opensearch,3.8.0,lucene,hnsw,synthetic-v25,32,5000,1,0,Q01,none,10,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,fill from run

8. Filter selectivity must be recorded

Define selectivity = eligible_docs / total_docs. Report it for every filtered query. Two filters with the same syntax can have radically different selectivity by tenant or category, which changes both exact cost and ANN behavior. If a filter leaves fewer than k eligible documents, the correct result can contain fewer than k hits; do not mark that as ANN recall loss.

9. Build-time and recovery evidence

Time bulk ingestion, refresh completion, and the period until the index is in the intended searchable state. Record store size after merges stabilize according to a documented policy; do not force-merge a hot benchmark index merely to make a chart smaller. For production readiness, also test restart/recovery and snapshot/reindex procedures because an index that searches quickly but cannot be recovered within the service objective is not acceptable.

10. Decision gates

Gate Example team-defined question
Quality Does every priority query/filter class meet the agreed recall/judged metric?
Tail latency Does p99 meet the SLO at expected concurrency?
Capacity Does the configuration fit memory/disk with merge/recovery headroom?
Security Are tenant/visibility constraints enforced independently of vector similarity?
Lifecycle Can the team re-embed/reindex, snapshot, restore, and roll forward safely?

11. Cleanup/reset

Disposable local cleanup
DELETE /atlasmart-vector-es-v25
DELETE /atlasmart-vector-os-v25
# Delete only the product/index that exists on the cluster you are connected to.
# Remove generated atlasmart-v25.ndjson and benchmark outputs after archiving the evidence bundle.
# Do not delete shared course indices such as atlasmart-products-v1.

12. Bridge to Chapter 26

Chapter 25 establishes the vector contract and measurement discipline. Chapter 26 adds model invocation, sparse/neural retrieval, lexical+dense/sparse hybrid fusion, normalization, reciprocal-rank-like fusion, and reranking. The same rule carries forward: every new retrieval signal needs a versioned contract, a judged evaluation set, operational limits, and measurable failure behavior.

Check your understanding

  1. Why generate a larger synthetic corpus if it cannot prove production relevance?
  2. What is the quality oracle for both products?
  3. Why report filter selectivity?
  4. What makes p99 credible?
  5. What is the next chapter's key addition?
Review the answers

1. It can exercise indexing/search mechanics and benchmark harness behavior while remaining deterministic; real production decisions still need representative embeddings and judgments.

2. An independently computed exact top-k result over the same raw vectors and eligible filter set.

3. It changes the eligible set and can change exact/ANN cost and recall behavior.

4. A stated sample count, representative query population, fixed concurrency, clear warm/cold conditions, and full-request timing.

5. Model-powered semantic/sparse/neural/hybrid retrieval and fusion/reranking, evaluated with the same discipline.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.