Govern every AtlasMart vector as a versioned model/data contract before using geometric proximity as retrieval evidence.
Embedding Semantics, Dimensions, Distance/Similarity, Normalization, Model Versioning, and Data Contracts
Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.
Learning outcomes
Define an embedding as a versioned model output contract rather than a generic semantic score.
Distinguish dense vectors from sparse token-weight vectors and choose an appropriate representation for a retrieval problem.
Explain dimensions, cosine, dot product, Euclidean/L2 distance, vector normalization, and why platform _score values are not directly portable.
Create a vector data contract that includes model/version, preprocessing, dimension, similarity, tenant/security metadata, and reindex rules.
Validate AtlasMart vectors deterministically before indexing and reject silent model or dimension drift.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 (released September
3, 2026) and
OpenSearch/OpenSearch Dashboards 3.8.0
(released August 4, 2026), using the course's established local
TLS/auth conventions and bundled JVMs. The mandatory lab uses
precomputed deterministic vectors, so no hosted
model, GPU, paid inference endpoint, or proprietary API is
required. Elasticsearch examples explicitly map
dense_vector rather than relying on changing
defaults; OpenSearch examples explicitly map
knn_vector and use the built-in Lucene engine for
the portable local path. No live cluster is available in this
generation environment, so performance cells are labeled
MEASURED and must be filled from the learner's run
rather than invented.
1. AtlasMart problem: “similar” changed after a model upgrade
AtlasMart's product API stores a vector next to every product. A
model team silently upgrades the encoder, keeps the field name
embedding, and starts writing vectors produced by
the new model into the existing index. Search still returns HTTP
200, but rankings become incoherent because old and new vectors
no longer share the same coordinate system.
An array of numbers does not identify its model, preprocessing, dimension semantics, normalization contract, or similarity function. A vector field must be governed like a schema plus model artifact.
2. Dense and sparse embeddings are different contracts
| Property | Dense vector | Sparse vector |
|---|---|---|
| Representation | fixed-length numeric array | mostly-zero dimensions, commonly token → weight |
| Typical mechanism | semantic geometry + kNN | weighted feature/posting-style retrieval |
| Interpretability | individual dimensions usually opaque | non-zero token/features can often be inspected |
| Primary chapter role | ANN/HNSW fundamentals | contrast representation; deeper sparse/hybrid search follows in Chapter 26 |
| Portability | only with the same model/preprocessing/metric contract | only with the same vocabulary/model/weighting contract |
Elasticsearch 9.5.3 has both dense_vector and
sparse_vector. OpenSearch 3.8.0 has
knn_vector for dense retrieval and a
sparse_vector path for neural sparse ANN. These
similarly named concepts do not imply identical storage, query
APIs, scoring, model management, or defaults.
3. Distance and similarity: rank relation, not semantic truth
For non-zero vectors a and b, cosine
similarity is (a·b)/(||a|| ||b||). Dot product is
a·b; its magnitude changes with vector norms unless
the model contract specifies unit normalization. L2 distance is
sqrt(sum((a_i-b_i)^2)); smaller means closer.
Search products transform these raw relations into positive
_score values differently. Compare
rank quality against judgments, not raw scores
across Elasticsearch and OpenSearch.
If the model emits unit vectors, dot product can represent
cosine efficiently. If it does not, silently switching to dot
product changes the ranking objective. Elasticsearch's
dot_product float vectors require unit-length
vectors; its cosine path normalizes internally.
OpenSearch engine behavior also varies by space/engine, so pin
the mapping and test it.
4. The AtlasMart embedding contract
embedding_contract:
id: atlasmart-product-v1
producer: deterministic-precomputed-fixture
model_version: fixture-2026-09
dimensions: 8
element_type: float32
preprocessing: none; vectors supplied directly
similarity_intent: cosine
required_metadata: [product_id, tenant_id, category, embedding_model]
missing_vector_policy: reject from vector index
model_change_policy: new field/index + measured reindex; never mix spaces
security_rule: tenant filter/authorization must be enforced independently of vector similarity
In production replace the fixture producer with the real model identifier, tokenizer/preprocessor version, truncation/chunking policy, normalization rule, and deployment checksum. If any contract item changes in a way that changes the vector space, treat it as a migration.
5. Deterministic AtlasMart fixture
The chapter uses eight 8-dimensional vectors. They are deliberately small for inspection; they are not claimed to be outputs of a real embedding model.
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]
[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20]
6. Ground truth before a search engine
import math
products = {
"P-1001":[.72,.50,.10,.10,.10,.40,.10,.15],
"P-1002":[.68,.55,.12,.10,.08,.35,.08,.20],
"P-1003":[.05,.05,.75,.45,.35,.10,.20,.05],
"P-1004":[.05,.05,.68,.50,.40,.05,.30,.05],
"P-1005":[.60,.40,.05,.05,.10,.65,.10,.25],
"P-1006":[.62,.32,.04,.04,.08,.58,.18,.18],
"P-1007":[.18,.12,.48,.32,.55,.20,.25,.12],
"P-1008":[.05,.05,.62,.52,.42,.08,.20,.06],
}
q=[.70,.42,.04,.05,.05,.55,.08,.20]
def cosine(a,b):
dot=sum(x*y for x,y in zip(a,b))
return dot/(math.sqrt(sum(x*x for x in a))*math.sqrt(sum(y*y for y in b)))
for pid,v in sorted(products.items(), key=lambda kv: cosine(q,kv[1]), reverse=True):
print(pid, f"{cosine(q,v):.6f}")
Deterministic invariant: the first four
products are P-1005, P-1006, P-1001, P-1002 for
this raw cosine fixture. The exact decimal scores are useful
only to verify the local arithmetic; they are not relevance
judgments.
7. Dimension/type validation is correctness, not convenience
Reject vectors with the wrong length, NaN/Infinity, unexpected
element type, missing model metadata, or an undefined cosine
vector such as all zeros. Elasticsearch
dense_vector supports at most 4096 dimensions in
the current reference; OpenSearch engine/method combinations
have their own limits. Never infer that a model is compatible
merely because two vectors have the same length.
def validate_vector(v, dims=8):
import math
if len(v) != dims: raise ValueError(f"expected {dims} dims, got {len(v)}")
if not all(isinstance(x,(int,float)) and math.isfinite(x) for x in v):
raise ValueError("non-finite vector value")
if math.sqrt(sum(float(x)*float(x) for x in v)) == 0:
raise ValueError("zero vector invalid for cosine")
return True
8. Production judgment
Vector search begins with data governance: model lifecycle, vector contract, tenant metadata, judged queries, and a migration plan. A high cosine value is evidence of geometric proximity under one contract—not a guarantee of relevance, safety, factuality, or user intent. Keep original source text/IDs available for audit and reranking; keep model/version metadata queryable; keep authorization separate from retrieval.
Check your understanding
- Why can two 768-dimensional vectors still be incompatible?
- When is dot product equivalent to cosine ranking?
- Why not compare Elasticsearch and OpenSearch _score directly?
- What should happen when an embedding model changes spaces?
- What proves the fixture is working?
Review the answers
1. Dimension count does not identify model, preprocessing, vocabulary, normalization, or coordinate semantics.
2. When the vector contract provides compatible unit-normalized vectors.
3. Products can transform distance/similarity into scores differently; compare ranked judgments or raw metric calculations under a shared contract.
4. Version the contract and re-embed/reindex into a new field or index; do not mix vectors from different spaces.
5. The independent cosine baseline produces the deterministic expected ordering before ANN is introduced.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic dense_vector field — dimensions, similarity, exact/script scoring, HNSW and quantized index options.
- Elastic kNN search — approximate versus brute-force kNN, filters, nested vector search, and query forms.
- Elastic sparse_vector field — weighted sparse features, pruning, and specialized query boundaries.
-
OpenSearch vector index creation
—
knn_vector, dimensions, workload modes, and storage choices. - OpenSearch methods and engines — HNSW/IVF and Lucene/Faiss/NMSLIB/JVector capability boundaries.
-
OpenSearch k-NN query
—
k, filters, and method parameters such asef_search. - OpenSearch efficient k-NN filtering — filter-aware exact/approximate behavior.
- OpenSearch vector quantization — byte/binary, scalar, product, and engine-specific compression choices.
- OpenSearch sparse_vector — neural sparse ANN support introduced in OpenSearch 3.3.
- Elastic Stack 9.5.3 release and OpenSearch version history — pinned September/August 2026 baselines.