Implement schema, ingest, lifecycle, and core retrieval as versioned deployable artifacts.

Create Mappings/Analyzers/Templates, Ingest Pipelines, Data Streams/Indices, Lifecycle Policies, and Core Search/Aggregation APIs

Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.

Intermediate → Advanced190–260 minutesSchema, ingest, lifecycle & core APIs · Chapter 31 · Lesson 02Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · free/local mandatory pathLast reviewed: September 2026

Learning outcomes

01

Implement the AtlasMart schema as explicit mappings/analyzers/templates instead of relying on dynamic behavior.

02

Separate catalog indices from append-oriented telemetry data streams and map lifecycle intent to ILM versus ISM correctly.

03

Use ingest pipelines for deterministic normalization while keeping upstream source-of-truth semantics recoverable.

04

Prove core lexical search, filters, aggregations, pagination, aliases/data streams, and schema tests on both platforms.

05

Package schema and pipeline definitions as versioned deployment artifacts with rollback evidence.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Capstone baseline. Examples are frozen to Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), with their bundled JVMs. The established disposable local endpoints remain Elasticsearch at https://localhost:9200 and OpenSearch at https://localhost:9201 on the atlasmart-search Docker network. The generation environment did not execute live clusters, so no latency, throughput, relevance, restore-time, or cost number is presented as measured unless the learner records it.

1. AtlasMart problem: schema is a search contract

Lesson 1 approved requirements; now each field exists because a query, aggregation, sort, retention rule, security filter, or operational workflow needs it. text and keyword are not cosmetic choices, object and nested are not interchangeable, and dynamic mapping is not “schema-free.” The capstone keeps product catalog and telemetry separate because their update shape, retention, query mix, and lifecycle differ.

2. Catalog mapping and analyzer contract

Elasticsearch catalog mapping skeleton
PUT atlasmart-products-v1
{
  "settings": {"number_of_shards": 1, "number_of_replicas": 0},
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "sku":       {"type":"keyword"},
      "tenant_id": {"type":"keyword"},
      "name":      {"type":"text", "fields":{"raw":{"type":"keyword"}}},
      "category":  {"type":"keyword"},
      "brand":     {"type":"keyword"},
      "price":     {"type":"scaled_float", "scaling_factor":100},
      "available": {"type":"boolean"},
      "updated_at":{"type":"date"},
      "popularity":{"type":"rank_feature"}
    }
  },
  "aliases": {"atlasmart-products-read": {}}
}
Portability boundary. Treat the JSON as an Elasticsearch example, not a promise that every specialized field or option has identical OpenSearch behavior. Build an OpenSearch mapping from the same query/aggregation requirements and run the same contract tests.

Before indexing, use analyze APIs to prove case folding, token boundaries, synonym behavior, and search-time analysis. Store analyzer tests as expected token arrays. If a synonym or analyzer change alters existing indexed terms, plan a new index plus reindex/alias cutover rather than assuming a live mapping edit rewrites old terms.

3. Template + data-stream contract for telemetry

Append-oriented events use a timestamp field and a data-stream/backing-index model. The same lifecycle intent—roll over bounded backing indices, retain a verified recovery window, then delete according to policy—maps differently to Elastic ILM/data tiers and OpenSearch ISM. Do not paste an ILM policy into OpenSearch or describe ISM state transitions as ILM phases.

Shared intent, product-specific lifecycle artifacts
telemetry_intent:
  stream_name: atlasmart-telemetry
  timestamp_field: "@timestamp"
  rollover:
    condition: <measured size/age policy>
  retention:
    delete_after: <approved retention>
  recovery_gate:
    snapshot_verified_before_delete: true
elastic_implementation:
  template: composable index template
  stream: Elasticsearch data stream
  lifecycle: ILM policy + data tiers where topology supports them
opensearch_implementation:
  template: OpenSearch index template
  stream: OpenSearch data stream
  lifecycle: ISM policy/states/transitions/actions

4. Ingest pipeline: normalize, do not hide source ambiguity

Use ingest processors for deterministic operations such as timestamp parsing, field rename/copy, simple enrichment, and dead-letter/failure routing. Preserve raw or replayable source where governance allows, because a pipeline bug otherwise becomes irreversible historical corruption. Expensive business joins or transformations that are easier to test upstream can stay upstream.

Illustrative ingest pipeline contract
PUT _ingest/pipeline/atlasmart-catalog-normalize-v1
{
  "processors": [
    {"set": {"field":"schema_version", "value":"catalog-v1"}},
    {"lowercase": {"field":"category"}},
    {"date": {"field":"updated_at", "formats":["ISO8601"]}}
  ],
  "on_failure": [
    {"set": {"field":"ingest_error", "value":"{{ _ingest.on_failure_message }}"}}
  ]
}

Verify exact processor syntax against the pinned product before promotion; similar processor names do not guarantee identical edge behavior.

5. Core query and aggregation contract

Lexical + tenant filter + facets
POST atlasmart-products-read/_search
{
  "size": 10,
  "query": {
    "bool": {
      "must": [{"multi_match": {"query":"waterproof trail shoe", "fields":["name^3","brand","category"]}}],
      "filter": [
        {"term": {"tenant_id":"tenant-a"}},
        {"term": {"available": true}}
      ]
    }
  },
  "aggs": {
    "by_brand": {"terms": {"field":"brand", "size":10}},
    "price": {"percentiles": {"field":"price", "percents":[50,95]}}
  },
  "sort": [{"_score":"desc"},{"sku":"asc"}]
}

The stable secondary sort key matters when pagination extends beyond one page. Validate counts/facets against a known fixture, not only HTTP 200. Profile/Explain are diagnostic tools, not production latency benchmarks.

6. Deployment order and rollback

Step Evidence before proceeding Rollback
1. simulate/validate templates resolved mapping/settings match expectation do not create index/stream
2. create versioned index/data stream health/mapping/analyzer tests pass delete only disposable target
3. ingest fixture through pipeline item failures zero or classified replay source into prior target
4. run query/agg tests expected IDs/counts/buckets keep old read alias/application target
5. switch alias/application target dual-read smoke tests pass switch endpoint/alias back

7. Wrong approach: one mega-index for the whole capstone

Wrong approach: put products, telemetry, security events, and RAG chunks in one dynamic index so deployment is “simple.” Failure: incompatible retention, permissions, query shapes, mappings, update rates, and scaling units become coupled. Repair: keep workload-specific schemas and lifecycle while sharing infrastructure only where the measured noisy-neighbor and governance evidence allows it.

8. Production judgment

Lesson 2 is complete only when schemas and analyzers are versioned artifacts with deterministic tests, lifecycle intent is explicit, ingest failure behavior is visible, and an alias/stream strategy permits safe replacement. Lesson 3 adds vector, hybrid, and RAG retrieval without weakening the tenant or relevance contracts.

Check your understanding

  1. Why is dynamic mapping not a schema strategy?
  2. Why separate catalog indices from telemetry data streams?
  3. Why are ILM and ISM not interchangeable?
  4. What proves an ingest pipeline is correct?
  5. Why use versioned aliases/streams instead of editing everything in place?
Review the answers

1. It can infer types and create fields that do not match query, aggregation, security, or long-term compatibility requirements.

2. They have different mutation patterns, retention, freshness, query shapes, lifecycle, and access-control requirements.

3. They use different policy models, APIs, state/phase semantics, actions, and operational behavior.

4. Known-input tests, expected transformed documents, classified failure output, replayability, and downstream query/schema checks.

5. They provide a controlled cutover and rollback surface for changes that require a new index or backing generation.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.