Operate OpenSearch neural, neural_sparse, k-NN, ML Commons, and search pipelines as explicit retrieval and model state.
OpenSearch neural, neural_sparse, k-NN, Ingest/Search Pipelines, and ML Commons Model Integration
Combine lexical, dense, sparse, semantic, neural, and reranking signals without losing score meaning, evaluation discipline, or operational control across the two platforms.
Learning outcomes
Explain OpenSearch neural, neural_sparse, and k-NN as separate retrieval paths with different storage/model contracts.
Use ML Commons, ingest pipelines, and search pipelines as explicit operational dependencies.
Build hybrid search with normalization or rank-based fusion without assuming parity with Elastic retrievers.
Design safe local-model and precomputed-vector paths that do not require external paid inference.
Observe model deployment, pipeline failure, candidate ranks, and authorization separately.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
semantic_text usually invokes the Inference API and
can therefore carry license/service requirements; OpenSearch ML
Commons can run supported local models or connectors, but
model/hardware support remains deployment-specific.
1. AtlasMart problem: OpenSearch can own more of the ML path, but that creates more state
AtlasMart wants text-to-vector ingestion, text-to-vector query inference, hybrid fusion, and optional reranking. In OpenSearch 3.8.0, the Security plugin, ML Commons model registry/deployment state, ingest pipelines, search pipelines, k-NN mappings, and application credentials can all participate. That convenience expands the failure graph: a green index can coexist with an undeployed model or broken pipeline.
2. Four retrieval concepts to keep separate
| OpenSearch construct | Meaning | Typical dependency |
|---|---|---|
knn |
nearest-neighbor query over a vector field | precomputed query vector; no model required at search time |
neural |
text/image is converted to a dense representation then searched | deployed/associated model unless semantic field hides model id |
neural_sparse |
sparse learned token/dimension retrieval | rank_features or sparse_vector plus compatible model/analyzer/tokens |
hybrid |
multiple subqueries whose results are fused | search pipeline using normalization processor or score-ranker/RRF |
3. Deterministic AtlasMart retrieval fixture
Carry forward Chapter 25's eight product vectors and
tenant/category metadata. These vectors are a transparent
mechanics fixture, not outputs claimed from a production
embedding model. Every semantic experiment records a model
contract such as atlasmart-product-v1, dimensions,
normalization, source text, and tenant metadata.
P-1001,outdoor,tenant-a,Hiking boots,[0.72,0.50,0.10,0.10,0.10,0.40,0.10,0.15]
P-1002,outdoor,tenant-a,Trail shoes,[0.68,0.55,0.12,0.10,0.08,0.35,0.08,0.20]
P-1003,kitchen,tenant-a,Espresso machine,[0.05,0.05,0.75,0.45,0.35,0.10,0.20,0.05]
P-1004,kitchen,tenant-a,Coffee grinder,[0.05,0.05,0.68,0.50,0.40,0.05,0.30,0.05]
P-1005,outdoor,tenant-a,Waterproof shell,[0.60,0.40,0.05,0.05,0.10,0.65,0.10,0.25]
P-1006,outdoor,tenant-b,Backpacking tent,[0.62,0.32,0.04,0.04,0.08,0.58,0.18,0.18]
P-1007,kitchen,tenant-b,Travel mug,[0.18,0.12,0.48,0.32,0.55,0.20,0.25,0.12]
P-1008,kitchen,tenant-b,Blender,[0.05,0.05,0.62,0.52,0.42,0.08,0.20,0.06]
Security invariant: tenant/category eligibility is applied before a result is allowed into evaluation or reranking. A semantically close document outside the caller's authorization scope is not a candidate.
4. ML Commons and ingestion
The text_embedding ingest processor can call a
registered ML model and write a vector field. A separate
sparse-encoding workflow can write sparse representations. Model
token limits and truncation matter: if relevant text is outside
the model's effective input, the embedding never represents it.
Version chunking and source-field selection with the model.
PUT _ingest/pipeline/atlasmart-text-embedding-v1
{
"processors": [
{
"text_embedding": {
"model_id": "${ATLASMART_DENSE_MODEL_ID}",
"field_map": {"name": "embedding_v1"}
}
}
]
}
${ATLASMART_DENSE_MODEL_ID} is an
environment/substitution variable, not a magic course model ID.
The mandatory lab can bypass this processor and bulk-index the
fixed vectors.
5. neural and neural_sparse query boundaries
POST atlasmart-products-v1/_search
{
"query": {
"neural": {
"embedding_v1": {
"query_text": "waterproof trail layer",
"model_id": "${ATLASMART_DENSE_MODEL_ID}",
"k": 20
}
}
}
}
POST atlasmart-products-sparse-v1/_search
{
"query": {
"neural_sparse": {
"sparse_embedding_v1": {
"query_text": "waterproof trail layer",
"model_id": "${ATLASMART_SPARSE_MODEL_ID}"
}
}
}
}
The model IDs and field types are product/version-specific contracts. OpenSearch 3.8 also supports sparse-vector ANN patterns; do not reuse two-phase sparse pipelines or method parameters unless the selected sparse field/mode documents them.
6. Hybrid fusion with a search pipeline
The score-based normalization-processor normalizes
each subquery result set then combines scores. The rank-based
score-ranker-processor uses RRF and therefore
avoids requiring component scores to share a scale. Both run on
candidate results, so the number of candidates returned by each
subquery can affect quality and cost.
PUT /_search/pipeline/atlasmart-hybrid-v1
{
"phase_results_processors": [
{
"normalization-processor": {
"normalization": {"technique":"min_max"},
"combination": {
"technique":"arithmetic_mean",
"parameters":{"weights":[0.4,0.6]}
}
}
}
]
}
POST atlasmart-products-v1/_search?search_pipeline=atlasmart-hybrid-v1
{
"query":{"hybrid":{"queries":[
{"bool":{"must":{"match":{"name":"waterproof trail layer"}},"filter":{"term":{"tenant":"tenant-a"}}}},
{"knn":{"embedding_v1":{"vector":[0.66,0.43,0.08,0.07,0.09,0.57,0.10,0.20],"k":20,"filter":{"term":{"tenant":"tenant-a"}}}}}
]}}
}
Use hybrid score explanation only for troubleshooting because explanation work is expensive. Never turn it into a production benchmark.
7. Reranking and candidate budget
OpenSearch's rerank search response processor can
apply an OpenSearch-provided cross-encoder model or rerank by a
field. ML inference response processors can also invoke
registered models and feed a rerank step. The processing order
matters: in a hybrid pipeline, normalization/fusion can happen
before reranking. Capture the pre-rerank ranking in debug tests
if you need to attribute regressions.
8. Wrong approach: hide model failure inside pipeline complexity
If the model is undeployed, credentials expire, or the connector throttles, a pipeline may fail even though the underlying lexical index is healthy. Repair by instrumenting model state and pipeline errors, keeping lexical-only fallback, setting bounded retries/timeouts, and testing the failure path deliberately. Never store remote-model secrets in query bodies or lesson HTML.
9. Production judgment
OpenSearch gives you powerful in-cluster ML and search-pipeline composition, but every model, connector, pipeline, and plugin version becomes part of change management and snapshots/runbooks. Managed Amazon OpenSearch Service can expose a different subset/configuration; verify service-region/version support instead of assuming upstream parity.
Bridge: Lesson 4 compares normalization, weighted fusion, RRF, reranking, and judged evaluation in a product-neutral way.
Check your understanding
- What is the difference between knn and neural?
- Why can hybrid result quality depend on candidate size?
- Why is the search pipeline an operational dependency?
- What is the free/local escape hatch?
- Should upstream OpenSearch behavior be assumed identical to a managed service?
Review the answers
1. kNN can search a vector you already have; neural can invoke/associate a model to produce the query representation.
2. Fusion only sees documents each branch returned.
3. Its processors can fail or change ranking even when indices are healthy.
4. Index and query fixed precomputed vectors and fuse results without a model service.
5. No. Service version, plugin, IAM, network, and feature restrictions must be checked.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
These links are the product documentation surfaces used to freeze examples for the September 2026 lesson baseline. Re-check them before production rollout because syntax, licensing, managed-service support, and model availability can change independently of this course.
- Elastic semantic search quickstart
- Elastic semantic_text setup/configuration
- Elastic search/retrieve semantic_text
- Elastic hybrid search
- Elastic retrievers overview
- Elastic linear retriever
- Elastic sparse_vector query
- Elastic ranking and reranking
- Elastic semantic reranking
- OpenSearch neural query
- OpenSearch neural sparse query
- OpenSearch hybrid search
- OpenSearch normalization processor
- OpenSearch score ranker processor
- OpenSearch rerank processor
- OpenSearch text embedding processor
- OpenSearch custom local models