Implement schema, ingest, lifecycle, and core retrieval as versioned deployable artifacts.
Create Mappings/Analyzers/Templates, Ingest Pipelines, Data Streams/Indices, Lifecycle Policies, and Core Search/Aggregation APIs
Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.
Learning outcomes
Implement the AtlasMart schema as explicit mappings/analyzers/templates instead of relying on dynamic behavior.
Separate catalog indices from append-oriented telemetry data streams and map lifecycle intent to ILM versus ISM correctly.
Use ingest pipelines for deterministic normalization while keeping upstream source-of-truth semantics recoverable.
Prove core lexical search, filters, aggregations, pagination, aliases/data streams, and schema tests on both platforms.
Package schema and pipeline definitions as versioned deployment artifacts with rollback evidence.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: schema is a search contract
Lesson 1 approved requirements; now each field exists because a
query, aggregation, sort, retention rule, security filter, or
operational workflow needs it. text and
keyword are not cosmetic choices,
object and nested are not
interchangeable, and dynamic mapping is not “schema-free.” The
capstone keeps product catalog and telemetry separate because
their update shape, retention, query mix, and lifecycle differ.
2. Catalog mapping and analyzer contract
PUT atlasmart-products-v1
{
"settings": {"number_of_shards": 1, "number_of_replicas": 0},
"mappings": {
"dynamic": "strict",
"properties": {
"sku": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"name": {"type":"text", "fields":{"raw":{"type":"keyword"}}},
"category": {"type":"keyword"},
"brand": {"type":"keyword"},
"price": {"type":"scaled_float", "scaling_factor":100},
"available": {"type":"boolean"},
"updated_at":{"type":"date"},
"popularity":{"type":"rank_feature"}
}
},
"aliases": {"atlasmart-products-read": {}}
}
Before indexing, use analyze APIs to prove case folding, token boundaries, synonym behavior, and search-time analysis. Store analyzer tests as expected token arrays. If a synonym or analyzer change alters existing indexed terms, plan a new index plus reindex/alias cutover rather than assuming a live mapping edit rewrites old terms.
3. Template + data-stream contract for telemetry
Append-oriented events use a timestamp field and a data-stream/backing-index model. The same lifecycle intent—roll over bounded backing indices, retain a verified recovery window, then delete according to policy—maps differently to Elastic ILM/data tiers and OpenSearch ISM. Do not paste an ILM policy into OpenSearch or describe ISM state transitions as ILM phases.
telemetry_intent:
stream_name: atlasmart-telemetry
timestamp_field: "@timestamp"
rollover:
condition: <measured size/age policy>
retention:
delete_after: <approved retention>
recovery_gate:
snapshot_verified_before_delete: true
elastic_implementation:
template: composable index template
stream: Elasticsearch data stream
lifecycle: ILM policy + data tiers where topology supports them
opensearch_implementation:
template: OpenSearch index template
stream: OpenSearch data stream
lifecycle: ISM policy/states/transitions/actions
4. Ingest pipeline: normalize, do not hide source ambiguity
Use ingest processors for deterministic operations such as timestamp parsing, field rename/copy, simple enrichment, and dead-letter/failure routing. Preserve raw or replayable source where governance allows, because a pipeline bug otherwise becomes irreversible historical corruption. Expensive business joins or transformations that are easier to test upstream can stay upstream.
PUT _ingest/pipeline/atlasmart-catalog-normalize-v1
{
"processors": [
{"set": {"field":"schema_version", "value":"catalog-v1"}},
{"lowercase": {"field":"category"}},
{"date": {"field":"updated_at", "formats":["ISO8601"]}}
],
"on_failure": [
{"set": {"field":"ingest_error", "value":"{{ _ingest.on_failure_message }}"}}
]
}
Verify exact processor syntax against the pinned product before promotion; similar processor names do not guarantee identical edge behavior.
5. Core query and aggregation contract
POST atlasmart-products-read/_search
{
"size": 10,
"query": {
"bool": {
"must": [{"multi_match": {"query":"waterproof trail shoe", "fields":["name^3","brand","category"]}}],
"filter": [
{"term": {"tenant_id":"tenant-a"}},
{"term": {"available": true}}
]
}
},
"aggs": {
"by_brand": {"terms": {"field":"brand", "size":10}},
"price": {"percentiles": {"field":"price", "percents":[50,95]}}
},
"sort": [{"_score":"desc"},{"sku":"asc"}]
}
The stable secondary sort key matters when pagination extends beyond one page. Validate counts/facets against a known fixture, not only HTTP 200. Profile/Explain are diagnostic tools, not production latency benchmarks.
6. Deployment order and rollback
| Step | Evidence before proceeding | Rollback |
|---|---|---|
| 1. simulate/validate templates | resolved mapping/settings match expectation | do not create index/stream |
| 2. create versioned index/data stream | health/mapping/analyzer tests pass | delete only disposable target |
| 3. ingest fixture through pipeline | item failures zero or classified | replay source into prior target |
| 4. run query/agg tests | expected IDs/counts/buckets | keep old read alias/application target |
| 5. switch alias/application target | dual-read smoke tests pass | switch endpoint/alias back |
7. Wrong approach: one mega-index for the whole capstone
8. Production judgment
Lesson 2 is complete only when schemas and analyzers are versioned artifacts with deterministic tests, lifecycle intent is explicit, ingest failure behavior is visible, and an alias/stream strategy permits safe replacement. Lesson 3 adds vector, hybrid, and RAG retrieval without weakening the tenant or relevance contracts.
Check your understanding
- Why is dynamic mapping not a schema strategy?
- Why separate catalog indices from telemetry data streams?
- Why are ILM and ISM not interchangeable?
- What proves an ingest pipeline is correct?
- Why use versioned aliases/streams instead of editing everything in place?
Review the answers
1. It can infer types and create fields that do not match query, aggregation, security, or long-term compatibility requirements.
2. They have different mutation patterns, retention, freshness, query shapes, lifecycle, and access-control requirements.
3. They use different policy models, APIs, state/phase semantics, actions, and operational behavior.
4. Known-input tests, expected transformed documents, classified failure output, replayability, and downstream query/schema checks.
5. They provide a controlled cutover and rollback surface for changes that require a new index or backing generation.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Download Elasticsearch 9.5.3
- Elasticsearch mappings
- Elasticsearch text analysis
- Elasticsearch index templates
- Elasticsearch ingest pipelines
- Elasticsearch data streams
- Elasticsearch ILM
- Elasticsearch Query DSL
- Elasticsearch aggregations
- Elasticsearch vector search
- Elasticsearch hybrid search
- Elasticsearch security
- Elasticsearch snapshot and restore
- Elasticsearch performance guidance
- Elasticsearch subscription feature matrix
- OpenSearch 3.8 version history
- OpenSearch downloads and Apache 2.0 licensing
- OpenSearch mappings and field types
- OpenSearch index templates
- OpenSearch data streams
- OpenSearch ingest pipelines
- OpenSearch Index State Management
- OpenSearch query DSL
- OpenSearch aggregations
- OpenSearch vector search
- OpenSearch hybrid search
- OpenSearch Security plugin
- OpenSearch snapshot and restore
- OpenSearch performance tuning
- OpenSearch Benchmark