Implement equivalent vector-retrieval intent through Elasticsearch dense_vector and OpenSearch knn_vector without pretending their APIs, engines, or scores are identical.

Elastic dense_vector/kNN Concepts and OpenSearch k-NN/vector Engine Options

Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.

Intermediate → Advanced120–165 minutesElastic/OpenSearch vector mapping lab · Chapter 25 · Lesson 03Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Configure equivalent-intent vector fields in Elasticsearch and OpenSearch without assuming mapping or query API identity.

02

Distinguish Elasticsearch dense_vector indexing/search choices from OpenSearch method/engine choices.

03

Use product-correct exact and approximate retrieval paths and interpret score transformations cautiously.

04

Explain dense versus sparse field support and identify the handoff from vector mechanics to semantic/neural workflows in Chapter 26.

05

Document engine, plugin, license, dimension, similarity, and migration assumptions in a portable decision record.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned vector-search baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released September 3, 2026) and OpenSearch/OpenSearch Dashboards 3.8.0 (released August 4, 2026), using the course's established local TLS/auth conventions and bundled JVMs. The mandatory lab uses precomputed deterministic vectors, so no hosted model, GPU, paid inference endpoint, or proprietary API is required. Elasticsearch examples explicitly map dense_vector rather than relying on changing defaults; OpenSearch examples explicitly map knn_vector and use the built-in Lucene engine for the portable local path. No live cluster is available in this generation environment, so performance cells are labeled MEASURED and must be filled from the learner's run rather than invented.

1. AtlasMart problem: “same JSON” portability

A platform team attempts to copy an Elasticsearch dense_vector mapping directly into OpenSearch. The products share Lucene ancestry, but the field names, method/engine abstraction, query parameters, quantization surface, defaults, and managed-service constraints differ. Equivalent intent does not imply interchangeable JSON.

Wrong approach: port syntax instead of semantics.

Write a product-neutral retrieval contract first—dimension, metric, exact/ANN requirement, quality target, filter semantics, memory budget—and then implement that contract with each product's supported APIs.

2. Equivalent intent matrix

Intent Elasticsearch 9.5.3 OpenSearch 3.8.0
Dense field dense_vector, dims knn_vector, dimension
Similarity cosine, dot_product, l2_norm, etc. engine-supported space_type, e.g. cosinesimil
Approximate index HNSW/quantized index options method + engine such as Lucene/Faiss HNSW
Exact path script_score or supported flat index option script/exact paths or supported flat method depending engine/version
Query breadth num_candidates / product query controls ef_search or engine-specific method parameters
Sparse field sparse_vector specialized retrieval sparse_vector neural sparse ANN plus other neural-sparse patterns

3. Elasticsearch dense_vector mechanics

Current Elasticsearch enables dense-vector indexing by default, but this course pins options. A dense_vector is single-valued; multiple passage vectors require nested/object patterns rather than multiple values in one field. Approximate kNN is supported through kNN search/query forms; brute-force vector functions in script_score provide a useful exact baseline on small or aggressively filtered sets.

Exact cosine reference over tenant-a
POST /atlasmart-vector-es-v25/_search
{
  "size":4,
  "query":{
    "script_score":{
      "query":{"term":{"tenant_id":"tenant-a"}},
      "script":{
        "source":"cosineSimilarity(params.q, 'embedding') + 1.0",
        "params":{"q":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20]}
      }
    }
  },
  "_source":["product_id","category","tenant_id"]
}
# Use this as an engine exact reference only after validating the eligible set and metric.

4. Elasticsearch quantization: memory is not free accuracy

Elasticsearch supports scalar and binary quantized vector index options such as int8_hnsw, int4_hnsw, and BBQ variants. The current reference notes that raw float values remain on disk for rescoring/reindex-related behavior, so quantization can reduce search memory while adding disk overhead. Some index options can be license-sensitive; the mandatory lab uses a free/local-supported path and labels Enterprise-only features instead of making them required.

Score boundary

Elasticsearch transforms raw vector relations into positive _score. Use rank/judgment metrics for portability; do not compare an Elastic _score numerically with an OpenSearch score and infer that one result is “more semantic.”

5. OpenSearch engine choices

OpenSearch's k-NN layer separates method (for example HNSW or IVF) from engine. Current documentation includes Lucene, Faiss, deprecated NMSLIB, and optional JVector support. Capability matrices differ for filtering, spaces, compression, training, and scale. The lab uses Lucene/HNSW because it is built in and filter-friendly; Faiss is a separate candidate when its scale/compression features match the workload.

OpenSearch engine-decision record
required_metric: cosine
required_filtering: tenant + category during vector retrieval
vector_count: MEASURED
updates_per_second: MEASURED
quality_target: recall@10 >= TEAM_TARGET
p99_target_ms: TEAM_TARGET
memory_budget: TEAM_TARGET
candidate_1: lucene/hnsw
candidate_2: faiss/hnsw
excluded: nmslib (deprecated)
optional_plugin_engine: evaluate only if exact 3.8 deployment/support policy permits

6. Sparse vectors: similar nouns, different systems

Elasticsearch sparse_vector stores weighted features for specialized sparse retrieval and is commonly paired with ELSER/semantic workflows. OpenSearch 3.8 supports both traditional neural sparse patterns and sparse_vector neural sparse ANN (introduced in 3.3). Sparse ANN parameters, model/pipeline APIs, pruning/quantization, and scoring are product-specific. Chapter 26 uses these systems in semantic and hybrid retrieval; Chapter 25 only establishes the representation and measurement contract.

7. Compatibility and migration checklist

  • Freeze model/version and preprocessing before cross-product comparison.
  • Recalculate exact ground truth from the raw fixture, not from either product's prior result.
  • Use the same eligible document set and filters.
  • Record engine/index options rather than “vector search enabled.”
  • Reindex when mapping/engine/model changes require it; do not expect in-place compatibility.
  • Snapshot source data/model metadata independently of vector replicas.
  • Re-run quality and resource tests after upgrade because defaults and implementations evolve.

8. Production judgment

A portable vector-search design is a contract plus test suite, not shared JSON. Preserve exact source IDs/content, model metadata, exact ground-truth tooling, and judged queries so either platform can be validated independently.

Check your understanding

  1. What is the OpenSearch abstraction that Elasticsearch does not expose in the same form?
  2. Why use a product-neutral contract before mappings?
  3. Can one dense_vector field hold multiple vectors in Elasticsearch?
  4. Why is NMSLIB not the default choice in this course?
  5. What must be regenerated after a model-space change?
Review the answers

1. An explicit method/engine choice such as HNSW with Lucene or Faiss.

2. It preserves intent while allowing different supported APIs and defaults.

3. No; dense_vector is single-valued, so multi-vector document patterns need separate/nested structures.

4. Current OpenSearch documentation marks it deprecated.

5. Vectors, exact ground truth, and the ANN index; then quality/resource evidence must be rerun.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.