Turn semantic inference into an explicit, versioned dependency with safe lexical fallback and measurable failure behavior.
Semantic Search from Text to Embeddings: Ingest-Time vs Query-Time Model Invocation and Failure Boundaries
Combine lexical, dense, sparse, semantic, neural, and reranking signals without losing score meaning, evaluation discipline, or operational control across the two platforms.
Learning outcomes
Distinguish semantic search as a model+vector contract from lexical term matching and from generative answering.
Choose ingest-time versus query-time inference by freshness, latency, cost, and failure-domain requirements.
Identify model-version, chunking, token-limit, dimension, and tenant-metadata contracts that must be versioned together.
Design explicit fallback behavior for inference timeouts, unavailable endpoints, malformed vectors, and partial ingestion.
Prove semantic quality against lexical and exact-vector baselines rather than treating similarity score as truth.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
semantic_text usually invokes the Inference API and
can therefore carry license/service requirements; OpenSearch ML
Commons can run supported local models or connectors, but
model/hardware support remains deployment-specific.
1. AtlasMart problem: meaning improves recall, but inference becomes a dependency
A shopper searches for “rainproof hiking layer.” The catalog contains Waterproof shell and Hiking boots, but not the exact phrase. Lexical search can still work through analysis/synonyms, yet semantic retrieval may recover the intended product from meaning. The new capability introduces an inference contract: text must be converted into a dense or sparse representation by a specific model, version, preprocessing/chunking policy, and output shape.
Embedding means that model-produced representation. Dense means most dimensions are numeric and populated. Sparse means only a small set of dimensions/tokens carry non-zero weights. Inference endpoint is the service boundary that turns text into model output. Ingest-time inference runs when documents enter the index; query-time inference runs for user queries. Neither is “just another analyzer”: both can fail independently of Elasticsearch/OpenSearch indexing/search.
| Decision | Ingest-time inference | Query-time inference |
|---|---|---|
| Latency paid | during indexing/reindex | during each query |
| Failure effect | documents may be rejected, delayed, or indexed without semantic representation if your pipeline allows that | query may fail or require fallback |
| Model upgrade | usually requires new vectors/reindex or dual fields | query model can change quickly, but must remain compatible with stored representation |
| Freshness | semantic vector exists only after inference succeeds | documents can be current, but query semantics depend on live model |
| Cost control | batching/throughput oriented | tail-latency/concurrency oriented |
2. Elastic and OpenSearch do not expose the same semantic abstraction
Elasticsearch 9.5.3 recommends semantic_text for
managed text semantic workflows; a match on that
field performs semantic retrieval through the field's configured
inference endpoint. When you need explicit control, use
dense_vector, sparse_vector,
kNN/sparse-vector queries, and retrievers. The older
semantic query remains for compatibility but
current Elastic documentation recommends match for
new semantic_text projects.
OpenSearch 3.8.0 exposes neural for
dense/model-backed retrieval and neural_sparse for
sparse neural retrieval, while k-NN can consume precomputed
vectors directly. ML Commons and ingest/search pipelines can
host or invoke models. These names are not aliases for Elastic
APIs; migrate the retrieval intent, not the JSON.
3. Deterministic AtlasMart retrieval fixture
Carry forward Chapter 25's eight product vectors and
tenant/category metadata. These vectors are a transparent
mechanics fixture, not outputs claimed from a production
embedding model. Every semantic experiment records a model
contract such as atlasmart-product-v1, dimensions,
normalization, source text, and tenant metadata.
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]
Security invariant: tenant/category eligibility is applied before a result is allowed into evaluation or reranking. A semantically close document outside the caller's authorization scope is not a candidate.
4. Version the semantic contract, not just the model file
| Contract field | Why it changes results |
|---|---|
| model_id + version | different model weights can reorder neighbors even at the same dimension |
| task/output type | dense, sparse, rerank, and generative outputs are not interchangeable |
| dimensions / token space | mapping/query compatibility is structural |
| normalization / similarity | dot product, cosine, and L2-like comparisons have different assumptions |
| source fields + chunking | different text boundaries change what is embedded |
| language/preprocessing | case, truncation, HTML stripping, and token limits change representation |
| tenant/security metadata | retrieval must never expand authorization scope |
5. Failure boundaries and fallback ladder
Do not hide model failures behind an empty result. Record inference status separately from search status. A practical ladder for AtlasMart is: (1) semantic+lexical hybrid when semantic inference succeeds; (2) lexical-only when the query inference service is unavailable; (3) cached/reused query embedding only when its model/version/normalization key matches exactly; (4) explicit user-facing degradation if neither path is healthy.
try:
qvec = embed(query, model="atlasmart-product-v1", timeout_ms=150)
return hybrid_search(query, qvec, tenant=user.tenant)
except InferenceUnavailable:
metrics.increment("search.semantic_fallback")
return lexical_search(query, tenant=user.tenant)
The fallback is part of the SLO: measure its rate. A “successful” HTTP 200 that silently switches semantics on 30% of requests is an operational incident, not normal behavior.
6. Wrong approach: switch model IDs in place
Changing an embedding model at query time while documents still contain vectors from the old model can produce syntactically valid but semantically meaningless nearest neighbors. Repair with a dual-version migration: index a new vector field or new index, write model/version metadata, populate it, evaluate against judgments, cut search traffic gradually, and keep rollback until the old representation can be retired safely.
7. Local lab: prove the boundaries without any paid model
- Load the Chapter 25 fixture into plain keyword fields plus a precomputed 8D vector field on each product.
- Run lexical-only and vector-only candidate retrieval for a fixed query set.
- Simulate query inference failure by refusing to provide the precomputed query vector; verify lexical fallback and a dedicated fallback counter.
-
Attach
model_version=atlasmart-product-v1to requests/results. Reject a syntheticv2query vector against thev1document field in the harness. - Record result IDs/ranks; do not claim a real model latency because no model is being executed.
8. Production judgment
Semantic search is production-ready only when model availability, version migration, tenant filtering, fallbacks, and judged quality are observable. Similarity scores are query/model-relative signals, not probability or semantic truth. Backups must preserve mappings/model metadata and migration state; retries must not duplicate ingestion; and p95/p99 must include inference plus search when inference is on the request path.
Bridge: Lesson 2 keeps the model contract fixed and focuses on Elastic's lexical+dense+sparse retrieval composition and reranking surfaces.
Check your understanding
- Why can a model upgrade require reindexing?
- What is the safest free/local fallback for this chapter?
- Why is a semantic fallback rate an SLO signal?
- Does a high cosine score prove user relevance?
- What must tenant filtering protect?
Review the answers
1. Stored document vectors belong to the old model space; the new query vector is not automatically comparable.
2. Precomputed vectors plus lexical search and client-side fusion/evaluation.
3. HTTP success can hide a material change in retrieval behavior.
4. No. It is a model-relative geometric signal and must be validated against judgments.
5. Candidate eligibility before fusion or reranking, not merely fields shown after ranking.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
These links are the product documentation surfaces used to freeze examples for the September 2026 lesson baseline. Re-check them before production rollout because syntax, licensing, managed-service support, and model availability can change independently of this course.
- Elastic semantic search quickstart
- Elastic semantic_text setup/configuration
- Elastic search/retrieve semantic_text
- Elastic hybrid search
- Elastic retrievers overview
- Elastic linear retriever
- Elastic sparse_vector query
- Elastic ranking and reranking
- Elastic semantic reranking
- OpenSearch neural query
- OpenSearch neural sparse query
- OpenSearch hybrid search
- OpenSearch normalization processor
- OpenSearch score ranker processor
- OpenSearch rerank processor
- OpenSearch text embedding processor
- OpenSearch custom local models