Benchmark AtlasMart vector retrieval with exact oracles, representative filters, p99, build time, capacity evidence, and explicit production gates.
Benchmark Vector Recall, p99 Latency, Index Build Time, Memory, and Filter Selectivity with Representative Queries
Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.
Learning outcomes
Build a representative vector benchmark that uses independent exact ground truth, judged queries, filters, and controlled client concurrency.
Measure recall@k, p95/p99 latency, index build time, index bytes, memory evidence, and filter selectivity without fabricated results.
Separate cold/warm cache behavior, ingestion/build phases, merge/refresh effects, and steady-state query performance.
Create a reproducible comparison record for Elasticsearch and OpenSearch without comparing transformed scores as if they were identical.
Choose a production candidate only when it satisfies quality, tail-latency, capacity, security, and operational-recovery gates.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 (released September
3, 2026) and
OpenSearch/OpenSearch Dashboards 3.8.0
(released August 4, 2026), using the course's established local
TLS/auth conventions and bundled JVMs. The mandatory lab uses
precomputed deterministic vectors, so no hosted
model, GPU, paid inference endpoint, or proprietary API is
required. Elasticsearch examples explicitly map
dense_vector rather than relying on changing
defaults; OpenSearch examples explicitly map
knn_vector and use the built-in Lucene engine for
the portable local path. No live cluster is available in this
generation environment, so performance cells are labeled
MEASURED and must be filled from the learner's run
rather than invented.
1. AtlasMart problem: the fastest demo loses the products users need
A developer benchmarks one query against warm cache, reports average latency, and picks the configuration with the smallest number. The chosen system misses relevant filtered products for a minority tenant and has unpredictable p99 during indexing. The benchmark optimized a demo, not a service.
A defensible vector benchmark needs query populations, exact ground truth, filter distributions, tail percentiles, indexing/build evidence, and quality gates.
2. Benchmark contract
| Dimension | Required evidence |
|---|---|
| Quality | recall@k against exact reference; judged NDCG/MRR when business relevance matters |
| Latency | p50/p95/p99 by query class, filter class, warm/cold state, and concurrency |
| Build/ingest | documents/s, wall-clock index build time, refresh/merge state |
| Capacity | store bytes, vector/native memory evidence, JVM/OS cache, replica/shard count |
| Filtering | eligible-doc count/selectivity + recall/latency under each filter class |
| Operations | recovery/restart behavior, snapshot/reindex plan, model/version traceability |
3. Deterministic larger fixture generator
The eight-product fixture proves semantics; it is too small to prove performance. Generate a larger synthetic corpus only for systems mechanics. It still is not a substitute for representative production embeddings.
import json, math, random
R=random.Random(25092026)
DIMS=32; N=5000
def unit(v):
n=math.sqrt(sum(x*x for x in v)); return [x/n for x in v]
with open("atlasmart-v25.ndjson","w",encoding="utf-8") as f:
for i in range(N):
# Four broad clusters; deterministic noise creates realistic-ish neighborhoods.
c=i % 4
center=[0.0]*DIMS
for j in range(8): center[c*8+j]=1.0
v=unit([center[j] + R.gauss(0,0.25) for j in range(DIMS)])
doc={"product_id":f"S-{i:05d}","tenant_id":f"tenant-{i%5}",
"category":f"c{c}","embedding_model":"synthetic-v25","embedding":v}
f.write(json.dumps({"index":{"_id":doc["product_id"]}})+"\n")
f.write(json.dumps(doc)+"\n")
print("generated",N,"documents; seed=25092026; dims=32")
Label every chart/table from this corpus synthetic mechanics. Production decisions require real embedding distributions, query vectors, filter selectivity, and judged relevance.
4. Exact ground truth harness
import heapq, math
def cosine(a,b):
return sum(x*y for x,y in zip(a,b))/(math.sqrt(sum(x*x for x in a))*math.sqrt(sum(y*y for y in b)))
def exact_topk(docs, q, k, predicate=lambda d: True):
scored=((cosine(q,d["embedding"]), d["product_id"]) for d in docs if predicate(d))
return [pid for score,pid in heapq.nlargest(k, scored)]
# Persist exact IDs for every benchmark query/filter tuple.
# The exact file is the quality oracle used for both Elasticsearch and OpenSearch.
5. Query matrix
query_id k filter_class concurrency cache_state
Q01 10 none 1 cold+warm
Q02 10 tenant=20% corpus 4 warm
Q03 10 tenant+category=5% 8 warm
Q04 50 none 4 warm
Q05 10 highly-selective <1% 4 warm
# Add real AtlasMart query classes before production approval.
6. Timing harness rules
- Use a monotonic client timer around the full HTTP request; record status/error and returned IDs.
- Run a documented warm-up phase but keep cold-start observations separate rather than discarding them silently.
- Use fixed concurrency and connection settings per run.
- Do not run indexing/build and query benchmarks together unless the test is intentionally mixed-load.
- Capture cluster/node/index stats before and after so tail latency can be correlated with GC, CPU, disk, merge, cache, rejection, and recovery signals from Chapters 14–16 and 24.
- Repeat enough samples to make p99 meaningful; state sample count.
7. Result schema: no invented numbers
run_id,product,version,engine,index_options,model,dims,docs,shards,replicas,query_set,filter_class,k,candidates_or_ef,recall_at_k,p50_ms,p95_ms,p99_ms,index_build_s,index_bytes,memory_evidence,sample_count,notes
RUN-001,elasticsearch,9.5.3,elastic,int8_hnsw,synthetic-v25,32,5000,1,0,Q01,none,10,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,fill from run
RUN-002,opensearch,3.8.0,lucene,hnsw,synthetic-v25,32,5000,1,0,Q01,none,10,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,fill from run
8. Filter selectivity must be recorded
Define selectivity = eligible_docs / total_docs.
Report it for every filtered query. Two filters with the same
syntax can have radically different selectivity by tenant or
category, which changes both exact cost and ANN behavior. If a
filter leaves fewer than k eligible documents, the correct
result can contain fewer than k hits; do not mark that as ANN
recall loss.
9. Build-time and recovery evidence
Time bulk ingestion, refresh completion, and the period until the index is in the intended searchable state. Record store size after merges stabilize according to a documented policy; do not force-merge a hot benchmark index merely to make a chart smaller. For production readiness, also test restart/recovery and snapshot/reindex procedures because an index that searches quickly but cannot be recovered within the service objective is not acceptable.
10. Decision gates
| Gate | Example team-defined question |
|---|---|
| Quality | Does every priority query/filter class meet the agreed recall/judged metric? |
| Tail latency | Does p99 meet the SLO at expected concurrency? |
| Capacity | Does the configuration fit memory/disk with merge/recovery headroom? |
| Security | Are tenant/visibility constraints enforced independently of vector similarity? |
| Lifecycle | Can the team re-embed/reindex, snapshot, restore, and roll forward safely? |
11. Cleanup/reset
DELETE /atlasmart-vector-es-v25
DELETE /atlasmart-vector-os-v25
# Delete only the product/index that exists on the cluster you are connected to.
# Remove generated atlasmart-v25.ndjson and benchmark outputs after archiving the evidence bundle.
# Do not delete shared course indices such as atlasmart-products-v1.
12. Bridge to Chapter 26
Chapter 25 establishes the vector contract and measurement discipline. Chapter 26 adds model invocation, sparse/neural retrieval, lexical+dense/sparse hybrid fusion, normalization, reciprocal-rank-like fusion, and reranking. The same rule carries forward: every new retrieval signal needs a versioned contract, a judged evaluation set, operational limits, and measurable failure behavior.
Check your understanding
- Why generate a larger synthetic corpus if it cannot prove production relevance?
- What is the quality oracle for both products?
- Why report filter selectivity?
- What makes p99 credible?
- What is the next chapter's key addition?
Review the answers
1. It can exercise indexing/search mechanics and benchmark harness behavior while remaining deterministic; real production decisions still need representative embeddings and judgments.
2. An independently computed exact top-k result over the same raw vectors and eligible filter set.
3. It changes the eligible set and can change exact/ANN cost and recall behavior.
4. A stated sample count, representative query population, fixed concurrency, clear warm/cold conditions, and full-request timing.
5. Model-powered semantic/sparse/neural/hybrid retrieval and fusion/reranking, evaluated with the same discipline.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic dense_vector field — dimensions, similarity, exact/script scoring, HNSW and quantized index options.
- Elastic kNN search — approximate versus brute-force kNN, filters, nested vector search, and query forms.
- Elastic sparse_vector field — weighted sparse features, pruning, and specialized query boundaries.
-
OpenSearch vector index creation
—
knn_vector, dimensions, workload modes, and storage choices. - OpenSearch methods and engines — HNSW/IVF and Lucene/Faiss/NMSLIB/JVector capability boundaries.
-
OpenSearch k-NN query
—
k, filters, and method parameters such asef_search. - OpenSearch efficient k-NN filtering — filter-aware exact/approximate behavior.
- OpenSearch vector quantization — byte/binary, scalar, product, and engine-specific compression choices.
- OpenSearch sparse_vector — neural sparse ANN support introduced in OpenSearch 3.3.
- Elastic Stack 9.5.3 release and OpenSearch version history — pinned September/August 2026 baselines.