Implement equivalent vector-retrieval intent through Elasticsearch dense_vector and OpenSearch knn_vector without pretending their APIs, engines, or scores are identical.
Elastic dense_vector/kNN Concepts and OpenSearch k-NN/vector Engine Options
Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.
Learning outcomes
Configure equivalent-intent vector fields in Elasticsearch and OpenSearch without assuming mapping or query API identity.
Distinguish Elasticsearch dense_vector indexing/search choices from OpenSearch method/engine choices.
Use product-correct exact and approximate retrieval paths and interpret score transformations cautiously.
Explain dense versus sparse field support and identify the handoff from vector mechanics to semantic/neural workflows in Chapter 26.
Document engine, plugin, license, dimension, similarity, and migration assumptions in a portable decision record.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 (released September
3, 2026) and
OpenSearch/OpenSearch Dashboards 3.8.0
(released August 4, 2026), using the course's established local
TLS/auth conventions and bundled JVMs. The mandatory lab uses
precomputed deterministic vectors, so no hosted
model, GPU, paid inference endpoint, or proprietary API is
required. Elasticsearch examples explicitly map
dense_vector rather than relying on changing
defaults; OpenSearch examples explicitly map
knn_vector and use the built-in Lucene engine for
the portable local path. No live cluster is available in this
generation environment, so performance cells are labeled
MEASURED and must be filled from the learner's run
rather than invented.
1. AtlasMart problem: “same JSON” portability
A platform team attempts to copy an Elasticsearch
dense_vector mapping directly into OpenSearch. The
products share Lucene ancestry, but the field names,
method/engine abstraction, query parameters, quantization
surface, defaults, and managed-service constraints differ.
Equivalent intent does not imply interchangeable JSON.
Write a product-neutral retrieval contract first—dimension, metric, exact/ANN requirement, quality target, filter semantics, memory budget—and then implement that contract with each product's supported APIs.
2. Equivalent intent matrix
| Intent | Elasticsearch 9.5.3 | OpenSearch 3.8.0 |
|---|---|---|
| Dense field | dense_vector, dims |
knn_vector, dimension |
| Similarity |
cosine, dot_product,
l2_norm, etc.
|
engine-supported space_type, e.g.
cosinesimil
|
| Approximate index | HNSW/quantized index options | method + engine such as Lucene/Faiss HNSW |
| Exact path |
script_score or supported flat index option
|
script/exact paths or supported flat method
depending engine/version
|
| Query breadth |
num_candidates / product query controls
|
ef_search or engine-specific method
parameters
|
| Sparse field | sparse_vector specialized retrieval |
sparse_vector neural sparse ANN plus other
neural-sparse patterns
|
3. Elasticsearch dense_vector mechanics
Current Elasticsearch enables dense-vector indexing by default,
but this course pins options. A dense_vector is
single-valued; multiple passage vectors require nested/object
patterns rather than multiple values in one field. Approximate
kNN is supported through kNN search/query forms; brute-force
vector functions in script_score provide a useful
exact baseline on small or aggressively filtered sets.
POST /atlasmart-vector-es-v25/_search
{
"size":4,
"query":{
"script_score":{
"query":{"term":{"tenant_id":"tenant-a"}},
"script":{
"source":"cosineSimilarity(params.q, 'embedding') + 1.0",
"params":{"q":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20]}
}
}
},
"_source":["product_id","category","tenant_id"]
}
# Use this as an engine exact reference only after validating the eligible set and metric.
4. Elasticsearch quantization: memory is not free accuracy
Elasticsearch supports scalar and binary quantized vector index
options such as int8_hnsw, int4_hnsw,
and BBQ variants. The current reference notes that raw float
values remain on disk for rescoring/reindex-related behavior, so
quantization can reduce search memory while adding disk
overhead. Some index options can be license-sensitive; the
mandatory lab uses a free/local-supported path and labels
Enterprise-only features instead of making them required.
Elasticsearch transforms raw vector relations into positive
_score. Use rank/judgment metrics for
portability; do not compare an Elastic
_score numerically with an OpenSearch score and
infer that one result is “more semantic.”
5. OpenSearch engine choices
OpenSearch's k-NN layer separates method (for example HNSW or IVF) from engine. Current documentation includes Lucene, Faiss, deprecated NMSLIB, and optional JVector support. Capability matrices differ for filtering, spaces, compression, training, and scale. The lab uses Lucene/HNSW because it is built in and filter-friendly; Faiss is a separate candidate when its scale/compression features match the workload.
required_metric: cosine
required_filtering: tenant + category during vector retrieval
vector_count: MEASURED
updates_per_second: MEASURED
quality_target: recall@10 >= TEAM_TARGET
p99_target_ms: TEAM_TARGET
memory_budget: TEAM_TARGET
candidate_1: lucene/hnsw
candidate_2: faiss/hnsw
excluded: nmslib (deprecated)
optional_plugin_engine: evaluate only if exact 3.8 deployment/support policy permits
6. Sparse vectors: similar nouns, different systems
Elasticsearch sparse_vector stores weighted
features for specialized sparse retrieval and is commonly paired
with ELSER/semantic workflows. OpenSearch 3.8 supports both
traditional neural sparse patterns and
sparse_vector neural sparse ANN (introduced in
3.3). Sparse ANN parameters, model/pipeline APIs,
pruning/quantization, and scoring are product-specific. Chapter
26 uses these systems in semantic and hybrid retrieval; Chapter
25 only establishes the representation and measurement contract.
7. Compatibility and migration checklist
- Freeze model/version and preprocessing before cross-product comparison.
- Recalculate exact ground truth from the raw fixture, not from either product's prior result.
- Use the same eligible document set and filters.
- Record engine/index options rather than “vector search enabled.”
- Reindex when mapping/engine/model changes require it; do not expect in-place compatibility.
- Snapshot source data/model metadata independently of vector replicas.
- Re-run quality and resource tests after upgrade because defaults and implementations evolve.
8. Production judgment
A portable vector-search design is a contract plus test suite, not shared JSON. Preserve exact source IDs/content, model metadata, exact ground-truth tooling, and judged queries so either platform can be validated independently.
Check your understanding
- What is the OpenSearch abstraction that Elasticsearch does not expose in the same form?
- Why use a product-neutral contract before mappings?
- Can one dense_vector field hold multiple vectors in Elasticsearch?
- Why is NMSLIB not the default choice in this course?
- What must be regenerated after a model-space change?
Review the answers
1. An explicit method/engine choice such as HNSW with Lucene or Faiss.
2. It preserves intent while allowing different supported APIs and defaults.
3. No; dense_vector is single-valued, so multi-vector document patterns need separate/nested structures.
4. Current OpenSearch documentation marks it deprecated.
5. Vectors, exact ground truth, and the ANN index; then quality/resource evidence must be rerun.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic dense_vector field — dimensions, similarity, exact/script scoring, HNSW and quantized index options.
- Elastic kNN search — approximate versus brute-force kNN, filters, nested vector search, and query forms.
- Elastic sparse_vector field — weighted sparse features, pruning, and specialized query boundaries.
-
OpenSearch vector index creation
—
knn_vector, dimensions, workload modes, and storage choices. - OpenSearch methods and engines — HNSW/IVF and Lucene/Faiss/NMSLIB/JVector capability boundaries.
-
OpenSearch k-NN query
—
k, filters, and method parameters such asef_search. - OpenSearch efficient k-NN filtering — filter-aware exact/approximate behavior.
- OpenSearch vector quantization — byte/binary, scalar, product, and engine-specific compression choices.
- OpenSearch sparse_vector — neural sparse ANN support introduced in OpenSearch 3.3.
- Elastic Stack 9.5.3 release and OpenSearch version history — pinned September/August 2026 baselines.