Use exact ground truth to quantify HNSW/ANN quality, then trade recall against tail latency, build cost, and memory with product-correct controls.

Exact vs Approximate kNN, HNSW Concepts, ef/m Parameters, Recall/Latency/Memory Tradeoffs, and Ground-Truth Testing

Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.

Intermediate → Advanced130–175 minutesANN/HNSW recall sweep lab · Chapter 25 · Lesson 02Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Explain exact kNN versus approximate nearest-neighbor retrieval and why ANN must be evaluated against an independent exact baseline.

02

Explain HNSW layers, graph connectivity, construction/search candidate breadth, and the different roles of m, ef_construction, ef_search, and Elasticsearch num_candidates.

03

Compute recall@k and separate quality loss from latency, indexing, memory, and filter effects.

04

Run a parameter sweep without changing multiple variables at once or treating Profile output as a benchmark.

05

Choose a candidate ANN configuration from measured service objectives rather than a universal HNSW recipe.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned vector-search baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released September 3, 2026) and OpenSearch/OpenSearch Dashboards 3.8.0 (released August 4, 2026), using the course's established local TLS/auth conventions and bundled JVMs. The mandatory lab uses precomputed deterministic vectors, so no hosted model, GPU, paid inference endpoint, or proprietary API is required. Elasticsearch examples explicitly map dense_vector rather than relying on changing defaults; OpenSearch examples explicitly map knn_vector and use the built-in Lucene engine for the portable local path. No live cluster is available in this generation environment, so performance cells are labeled MEASURED and must be filled from the learner's run rather than invented.

1. AtlasMart problem: “ANN says it is correct because ANN agrees with ANN”

A benchmark compares two HNSW configurations by asking each one for top-10 results, then declares the faster configuration “100% accurate” because repeated runs agree with themselves. No exact reference ranking exists. That test measures repeatability, not recall.

Wrong approach: benchmark ANN against itself.

Approximate retrieval needs an independent reference: brute-force similarity over the same eligible document set or a sufficiently exact product mode. Then measure how many exact top-k neighbors ANN recovered.

2. Exact versus approximate kNN

Dimension Exact Approximate (ANN)
Search work score every eligible vector explore a search structure/candidate subset
Recall exact for the metric/eligible set can miss true neighbors
Scaling cost grows with eligible vectors trades recall for lower query cost
Primary use here ground truth and small/restrictive-filter cases production candidate retrieval at scale

3. HNSW mechanism

Hierarchical Navigable Small World (HNSW) organizes vectors in a graph. Upper sparse layers provide long jumps; lower dense layers refine the neighborhood. Search starts from entry points and explores promising graph neighbors rather than calculating distance to every vector.

Control Stage General effect when increased
m index construction more graph links; often more memory/build cost and potentially better recall
ef_construction index construction broader candidate set while building; slower indexing, potentially better graph quality
ef_search query, where exposed broader search exploration; typically higher recall and latency
num_candidates Elasticsearch query more per-shard candidates considered for approximate kNN; quality/latency tradeoff

Do not copy a numeric recipe between products. Elasticsearch exposes HNSW construction controls in index_options and query candidate breadth through num_candidates; OpenSearch can expose ef_search through method/query settings depending on the chosen engine.

4. Recall@k

For a query with exact top-k set G and ANN returned set A, recall@k = |G ∩ A| / k. Recall ignores ordering within the set, so pair it with ranking metrics when order matters. Also record filter condition, tenant, model version, index version, and query population.

Recall function
def recall_at_k(ground_truth, ann, k):
    g=set(ground_truth[:k])
    a=set(ann[:k])
    return len(g & a)/k

# Example only: if exact=[A,B,C,D,E] and ANN=[A,C,B,X,E], recall@5=0.8

5. Elastic mapping: pin the ANN structure

Elasticsearch 9.5.3: explicit dense_vector mapping
PUT /atlasmart-vector-es-v25
{
  "settings": {"number_of_shards": 1, "number_of_replicas": 0},
  "mappings": {"properties": {
    "product_id": {"type":"keyword"},
    "tenant_id": {"type":"keyword"},
    "category": {"type":"keyword"},
    "embedding_model": {"type":"keyword"},
    "embedding": {
      "type":"dense_vector",
      "dims":8,
      "similarity":"cosine",
      "index_options":{"type":"int8_hnsw","m":16,"ef_construction":100}
    }
  }}
}

The mapping deliberately specifies int8_hnsw so a future default change does not silently change the experiment. Quantization can reduce in-memory vector footprint but may reduce recall; Elasticsearch retains raw float values on disk for rescoring/reindex-related behavior, so memory and disk effects must both be measured.

6. OpenSearch mapping: pin engine and method

OpenSearch 3.8.0: portable Lucene HNSW path
PUT /atlasmart-vector-os-v25
{
  "settings": {"index":{"knn":true,"number_of_shards":1,"number_of_replicas":0}},
  "mappings": {"properties": {
    "product_id": {"type":"keyword"},
    "tenant_id": {"type":"keyword"},
    "category": {"type":"keyword"},
    "embedding_model": {"type":"keyword"},
    "embedding": {
      "type":"knn_vector",
      "dimension":8,
      "method": {
        "name":"hnsw", "space_type":"cosinesimil", "engine":"lucene",
        "parameters":{"m":16,"ef_construction":100}
      }
    }
  }}
}

OpenSearch also supports other engine/method combinations such as Faiss HNSW/IVF; NMSLIB is deprecated. Optional engines/plugins change operational requirements. The mandatory lab stays on built-in Lucene to keep the local path reproducible.

7. Query sweep: change one variable

Elasticsearch approximate query sweep
POST /atlasmart-vector-es-v25/_search
{
  "size":4,
  "knn":{
    "field":"embedding",
    "query_vector":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20],
    "k":4,
    "num_candidates":20
  },
  "_source":["product_id","category","tenant_id","embedding_model"]
}
# Repeat with controlled num_candidates values. Record MEASURED latency and recall@4.
OpenSearch approximate query sweep
POST /atlasmart-vector-os-v25/_search
{
  "size":4,
  "query":{"knn":{"embedding":{
    "vector":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20],
    "k":4,
    "method_parameters":{"ef_search":20}
  }}}
}
# Repeat with supported ef_search values for this exact engine/version. Record MEASURED evidence.

8. Benchmark discipline

Warm and cold behavior, client concurrency, refresh/merge work, shard count, JVM/OS cache, filter selectivity, and quantization can all move the result. Capture p50/p95/p99 across repeated representative queries; capture indexing throughput/build time separately. Profile APIs can explain operators but add instrumentation overhead and are not a latency benchmark.

Result record
query_set,config,k,filter,recall_at_k,p50_ms,p95_ms,p99_ms,index_build_s,index_bytes,rss_or_vector_memory
atlasmart-v1,es-int8-hnsw-c20,10,none,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED
atlasmart-v1,os-lucene-hnsw-ef20,10,none,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED

9. Production judgment

ANN configuration is a service-level decision. Start with a quality target on judged queries, then find the lowest-cost configuration that satisfies it under representative filters and concurrency. Re-run the sweep when model version, vector count/distribution, shard layout, engine, quantization, or hardware changes.

Check your understanding

  1. What does ANN recall compare against?
  2. What does increasing m usually buy and cost?
  3. Is OpenSearch ef_search the same API as Elasticsearch num_candidates?
  4. Why pin index options in the lab?
  5. Why can Profile output not substitute for p99 measurement?
Review the answers

1. An independent exact top-k reference over the same eligible document set.

2. Potentially better connectivity/recall at higher graph memory and build cost.

3. No. They express related search-breadth tradeoffs but are product/engine-specific controls.

4. To prevent product defaults from changing the experiment across versions.

5. Instrumentation changes execution and Profile is designed for diagnosis, not representative production benchmarking.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.