Use exact ground truth to quantify HNSW/ANN quality, then trade recall against tail latency, build cost, and memory with product-correct controls.
Exact vs Approximate kNN, HNSW Concepts, ef/m Parameters, Recall/Latency/Memory Tradeoffs, and Ground-Truth Testing
Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.
Learning outcomes
Explain exact kNN versus approximate nearest-neighbor retrieval and why ANN must be evaluated against an independent exact baseline.
Explain HNSW layers, graph connectivity, construction/search candidate breadth, and the different roles of m, ef_construction, ef_search, and Elasticsearch num_candidates.
Compute recall@k and separate quality loss from latency, indexing, memory, and filter effects.
Run a parameter sweep without changing multiple variables at once or treating Profile output as a benchmark.
Choose a candidate ANN configuration from measured service objectives rather than a universal HNSW recipe.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 (released September
3, 2026) and
OpenSearch/OpenSearch Dashboards 3.8.0
(released August 4, 2026), using the course's established local
TLS/auth conventions and bundled JVMs. The mandatory lab uses
precomputed deterministic vectors, so no hosted
model, GPU, paid inference endpoint, or proprietary API is
required. Elasticsearch examples explicitly map
dense_vector rather than relying on changing
defaults; OpenSearch examples explicitly map
knn_vector and use the built-in Lucene engine for
the portable local path. No live cluster is available in this
generation environment, so performance cells are labeled
MEASURED and must be filled from the learner's run
rather than invented.
1. AtlasMart problem: “ANN says it is correct because ANN agrees with ANN”
A benchmark compares two HNSW configurations by asking each one for top-10 results, then declares the faster configuration “100% accurate” because repeated runs agree with themselves. No exact reference ranking exists. That test measures repeatability, not recall.
Approximate retrieval needs an independent reference: brute-force similarity over the same eligible document set or a sufficiently exact product mode. Then measure how many exact top-k neighbors ANN recovered.
2. Exact versus approximate kNN
| Dimension | Exact | Approximate (ANN) |
|---|---|---|
| Search work | score every eligible vector | explore a search structure/candidate subset |
| Recall | exact for the metric/eligible set | can miss true neighbors |
| Scaling | cost grows with eligible vectors | trades recall for lower query cost |
| Primary use here | ground truth and small/restrictive-filter cases | production candidate retrieval at scale |
3. HNSW mechanism
Hierarchical Navigable Small World (HNSW) organizes vectors in a graph. Upper sparse layers provide long jumps; lower dense layers refine the neighborhood. Search starts from entry points and explores promising graph neighbors rather than calculating distance to every vector.
| Control | Stage | General effect when increased |
|---|---|---|
m |
index construction | more graph links; often more memory/build cost and potentially better recall |
ef_construction |
index construction | broader candidate set while building; slower indexing, potentially better graph quality |
ef_search |
query, where exposed | broader search exploration; typically higher recall and latency |
num_candidates |
Elasticsearch query | more per-shard candidates considered for approximate kNN; quality/latency tradeoff |
Do not copy a numeric recipe between products. Elasticsearch
exposes HNSW construction controls in
index_options and query candidate breadth through
num_candidates; OpenSearch can expose
ef_search through method/query settings depending
on the chosen engine.
4. Recall@k
For a query with exact top-k set G and ANN returned
set A, recall@k = |G ∩ A| / k. Recall
ignores ordering within the set, so pair it with ranking metrics
when order matters. Also record filter condition, tenant, model
version, index version, and query population.
def recall_at_k(ground_truth, ann, k):
g=set(ground_truth[:k])
a=set(ann[:k])
return len(g & a)/k
# Example only: if exact=[A,B,C,D,E] and ANN=[A,C,B,X,E], recall@5=0.8
5. Elastic mapping: pin the ANN structure
PUT /atlasmart-vector-es-v25
{
"settings": {"number_of_shards": 1, "number_of_replicas": 0},
"mappings": {"properties": {
"product_id": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"category": {"type":"keyword"},
"embedding_model": {"type":"keyword"},
"embedding": {
"type":"dense_vector",
"dims":8,
"similarity":"cosine",
"index_options":{"type":"int8_hnsw","m":16,"ef_construction":100}
}
}}
}
The mapping deliberately specifies int8_hnsw so a
future default change does not silently change the experiment.
Quantization can reduce in-memory vector footprint but may
reduce recall; Elasticsearch retains raw float values on disk
for rescoring/reindex-related behavior, so memory and disk
effects must both be measured.
6. OpenSearch mapping: pin engine and method
PUT /atlasmart-vector-os-v25
{
"settings": {"index":{"knn":true,"number_of_shards":1,"number_of_replicas":0}},
"mappings": {"properties": {
"product_id": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"category": {"type":"keyword"},
"embedding_model": {"type":"keyword"},
"embedding": {
"type":"knn_vector",
"dimension":8,
"method": {
"name":"hnsw", "space_type":"cosinesimil", "engine":"lucene",
"parameters":{"m":16,"ef_construction":100}
}
}
}}
}
OpenSearch also supports other engine/method combinations such as Faiss HNSW/IVF; NMSLIB is deprecated. Optional engines/plugins change operational requirements. The mandatory lab stays on built-in Lucene to keep the local path reproducible.
7. Query sweep: change one variable
POST /atlasmart-vector-es-v25/_search
{
"size":4,
"knn":{
"field":"embedding",
"query_vector":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20],
"k":4,
"num_candidates":20
},
"_source":["product_id","category","tenant_id","embedding_model"]
}
# Repeat with controlled num_candidates values. Record MEASURED latency and recall@4.
POST /atlasmart-vector-os-v25/_search
{
"size":4,
"query":{"knn":{"embedding":{
"vector":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20],
"k":4,
"method_parameters":{"ef_search":20}
}}}
}
# Repeat with supported ef_search values for this exact engine/version. Record MEASURED evidence.
8. Benchmark discipline
Warm and cold behavior, client concurrency, refresh/merge work, shard count, JVM/OS cache, filter selectivity, and quantization can all move the result. Capture p50/p95/p99 across repeated representative queries; capture indexing throughput/build time separately. Profile APIs can explain operators but add instrumentation overhead and are not a latency benchmark.
query_set,config,k,filter,recall_at_k,p50_ms,p95_ms,p99_ms,index_build_s,index_bytes,rss_or_vector_memory
atlasmart-v1,es-int8-hnsw-c20,10,none,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED
atlasmart-v1,os-lucene-hnsw-ef20,10,none,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED
9. Production judgment
ANN configuration is a service-level decision. Start with a quality target on judged queries, then find the lowest-cost configuration that satisfies it under representative filters and concurrency. Re-run the sweep when model version, vector count/distribution, shard layout, engine, quantization, or hardware changes.
Check your understanding
- What does ANN recall compare against?
- What does increasing m usually buy and cost?
- Is OpenSearch ef_search the same API as Elasticsearch num_candidates?
- Why pin index options in the lab?
- Why can Profile output not substitute for p99 measurement?
Review the answers
1. An independent exact top-k reference over the same eligible document set.
2. Potentially better connectivity/recall at higher graph memory and build cost.
3. No. They express related search-breadth tradeoffs but are product/engine-specific controls.
4. To prevent product defaults from changing the experiment across versions.
5. Instrumentation changes execution and Profile is designed for diagnosis, not representative production benchmarking.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic dense_vector field — dimensions, similarity, exact/script scoring, HNSW and quantized index options.
- Elastic kNN search — approximate versus brute-force kNN, filters, nested vector search, and query forms.
- Elastic sparse_vector field — weighted sparse features, pruning, and specialized query boundaries.
-
OpenSearch vector index creation
—
knn_vector, dimensions, workload modes, and storage choices. - OpenSearch methods and engines — HNSW/IVF and Lucene/Faiss/NMSLIB/JVector capability boundaries.
-
OpenSearch k-NN query
—
k, filters, and method parameters such asef_search. - OpenSearch efficient k-NN filtering — filter-aware exact/approximate behavior.
- OpenSearch vector quantization — byte/binary, scalar, product, and engine-specific compression choices.
- OpenSearch sparse_vector — neural sparse ANN support introduced in OpenSearch 3.3.
- Elastic Stack 9.5.3 release and OpenSearch version history — pinned September/August 2026 baselines.