Govern every AtlasMart vector as a versioned model/data contract before using geometric proximity as retrieval evidence.

Embedding Semantics, Dimensions, Distance/Similarity, Normalization, Model Versioning, and Data Contracts

Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.

Intermediate → Advanced125–165 minutesEmbedding-contract & exact-cosine lab · Chapter 25 · Lesson 01Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Define an embedding as a versioned model output contract rather than a generic semantic score.

02

Distinguish dense vectors from sparse token-weight vectors and choose an appropriate representation for a retrieval problem.

03

Explain dimensions, cosine, dot product, Euclidean/L2 distance, vector normalization, and why platform _score values are not directly portable.

04

Create a vector data contract that includes model/version, preprocessing, dimension, similarity, tenant/security metadata, and reindex rules.

05

Validate AtlasMart vectors deterministically before indexing and reject silent model or dimension drift.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned vector-search baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released September 3, 2026) and OpenSearch/OpenSearch Dashboards 3.8.0 (released August 4, 2026), using the course's established local TLS/auth conventions and bundled JVMs. The mandatory lab uses precomputed deterministic vectors, so no hosted model, GPU, paid inference endpoint, or proprietary API is required. Elasticsearch examples explicitly map dense_vector rather than relying on changing defaults; OpenSearch examples explicitly map knn_vector and use the built-in Lucene engine for the portable local path. No live cluster is available in this generation environment, so performance cells are labeled MEASURED and must be filled from the learner's run rather than invented.

1. AtlasMart problem: “similar” changed after a model upgrade

AtlasMart's product API stores a vector next to every product. A model team silently upgrades the encoder, keeps the field name embedding, and starts writing vectors produced by the new model into the existing index. Search still returns HTTP 200, but rankings become incoherent because old and new vectors no longer share the same coordinate system.

Wrong approach: treat a vector as self-describing.

An array of numbers does not identify its model, preprocessing, dimension semantics, normalization contract, or similarity function. A vector field must be governed like a schema plus model artifact.

2. Dense and sparse embeddings are different contracts

Property Dense vector Sparse vector
Representation fixed-length numeric array mostly-zero dimensions, commonly token → weight
Typical mechanism semantic geometry + kNN weighted feature/posting-style retrieval
Interpretability individual dimensions usually opaque non-zero token/features can often be inspected
Primary chapter role ANN/HNSW fundamentals contrast representation; deeper sparse/hybrid search follows in Chapter 26
Portability only with the same model/preprocessing/metric contract only with the same vocabulary/model/weighting contract

Elasticsearch 9.5.3 has both dense_vector and sparse_vector. OpenSearch 3.8.0 has knn_vector for dense retrieval and a sparse_vector path for neural sparse ANN. These similarly named concepts do not imply identical storage, query APIs, scoring, model management, or defaults.

3. Distance and similarity: rank relation, not semantic truth

For non-zero vectors a and b, cosine similarity is (a·b)/(||a|| ||b||). Dot product is a·b; its magnitude changes with vector norms unless the model contract specifies unit normalization. L2 distance is sqrt(sum((a_i-b_i)^2)); smaller means closer. Search products transform these raw relations into positive _score values differently. Compare rank quality against judgments, not raw scores across Elasticsearch and OpenSearch.

Normalization boundary

If the model emits unit vectors, dot product can represent cosine efficiently. If it does not, silently switching to dot product changes the ranking objective. Elasticsearch's dot_product float vectors require unit-length vectors; its cosine path normalizes internally. OpenSearch engine behavior also varies by space/engine, so pin the mapping and test it.

4. The AtlasMart embedding contract

Versioned contract
embedding_contract:
  id: atlasmart-product-v1
  producer: deterministic-precomputed-fixture
  model_version: fixture-2026-09
  dimensions: 8
  element_type: float32
  preprocessing: none; vectors supplied directly
  similarity_intent: cosine
  required_metadata: [product_id, tenant_id, category, embedding_model]
  missing_vector_policy: reject from vector index
  model_change_policy: new field/index + measured reindex; never mix spaces
  security_rule: tenant filter/authorization must be enforced independently of vector similarity

In production replace the fixture producer with the real model identifier, tokenizer/preprocessor version, truncation/chunking policy, normalization rule, and deployment checksum. If any contract item changes in a way that changes the vector space, treat it as a migration.

5. Deterministic AtlasMart fixture

The chapter uses eight 8-dimensional vectors. They are deliberately small for inspection; they are not claimed to be outputs of a real embedding model.

Products and precomputed vectors
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]
Query vector: waterproof hiking gear
[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20]

6. Ground truth before a search engine

Standard-library cosine baseline
import math

products = {
 "P-1001":[.72,.50,.10,.10,.10,.40,.10,.15],
 "P-1002":[.68,.55,.12,.10,.08,.35,.08,.20],
 "P-1003":[.05,.05,.75,.45,.35,.10,.20,.05],
 "P-1004":[.05,.05,.68,.50,.40,.05,.30,.05],
 "P-1005":[.60,.40,.05,.05,.10,.65,.10,.25],
 "P-1006":[.62,.32,.04,.04,.08,.58,.18,.18],
 "P-1007":[.18,.12,.48,.32,.55,.20,.25,.12],
 "P-1008":[.05,.05,.62,.52,.42,.08,.20,.06],
}
q=[.70,.42,.04,.05,.05,.55,.08,.20]
def cosine(a,b):
    dot=sum(x*y for x,y in zip(a,b))
    return dot/(math.sqrt(sum(x*x for x in a))*math.sqrt(sum(y*y for y in b)))
for pid,v in sorted(products.items(), key=lambda kv: cosine(q,kv[1]), reverse=True):
    print(pid, f"{cosine(q,v):.6f}")

Deterministic invariant: the first four products are P-1005, P-1006, P-1001, P-1002 for this raw cosine fixture. The exact decimal scores are useful only to verify the local arithmetic; they are not relevance judgments.

7. Dimension/type validation is correctness, not convenience

Reject vectors with the wrong length, NaN/Infinity, unexpected element type, missing model metadata, or an undefined cosine vector such as all zeros. Elasticsearch dense_vector supports at most 4096 dimensions in the current reference; OpenSearch engine/method combinations have their own limits. Never infer that a model is compatible merely because two vectors have the same length.

Preflight validator
def validate_vector(v, dims=8):
    import math
    if len(v) != dims: raise ValueError(f"expected {dims} dims, got {len(v)}")
    if not all(isinstance(x,(int,float)) and math.isfinite(x) for x in v):
        raise ValueError("non-finite vector value")
    if math.sqrt(sum(float(x)*float(x) for x in v)) == 0:
        raise ValueError("zero vector invalid for cosine")
    return True

8. Production judgment

Vector search begins with data governance: model lifecycle, vector contract, tenant metadata, judged queries, and a migration plan. A high cosine value is evidence of geometric proximity under one contract—not a guarantee of relevance, safety, factuality, or user intent. Keep original source text/IDs available for audit and reranking; keep model/version metadata queryable; keep authorization separate from retrieval.

Check your understanding

  1. Why can two 768-dimensional vectors still be incompatible?
  2. When is dot product equivalent to cosine ranking?
  3. Why not compare Elasticsearch and OpenSearch _score directly?
  4. What should happen when an embedding model changes spaces?
  5. What proves the fixture is working?
Review the answers

1. Dimension count does not identify model, preprocessing, vocabulary, normalization, or coordinate semantics.

2. When the vector contract provides compatible unit-normalized vectors.

3. Products can transform distance/similarity into scores differently; compare ranked judgments or raw metric calculations under a shared contract.

4. Version the contract and re-embed/reindex into a new field or index; do not mix vectors from different spaces.

5. The independent cosine baseline produces the deterministic expected ordering before ANN is introduced.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.