Chapter 09 · Dynamic Templates, Index Templates, Component Templates, and Schema Evolution

Create a Versioned Schema Deployment Workflow with Template Tests and Zero-Downtime Reindexing

Turn schema evolution into a repeatable delivery workflow with template simulation tests, immutable versions, migration gates, alias cutover/rollback, drift checks, and evidence captured before production promotion.

Intermediate100–120 minutesSchema evolution labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

A reliable schema migration should be boring enough to repeat. AtlasMart therefore treats mappings, component templates, index templates, simulation results, migration code, acceptance fixtures, alias state and rollback instructions as one versioned release artifact—not a collection of commands remembered by an operator.

01

Define a schema manifest that makes ownership, version, compatibility, template names, priorities and expected aliases explicit.

02

Automate component/index-template simulation and resolved-contract assertions for representative index names.

03

Gate migration on deterministic mapping, query, aggregation, relevance and freshness evidence rather than HTTP success alone.

04

Coordinate bulk copy and live-write reconciliation before an atomic alias cutover.

05

Preserve rollback, audit evidence and post-cutover observability until the migration is formally closed.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 using the existing AtlasMart local-course conventions: Elasticsearch on https://localhost:9200 with its copied CA certificate, OpenSearch on https://localhost:9201 with the disposable demo certificate explicitly treated as local-only, and no moving latest tags. Template and mapping syntax is verified against current official documentation. The workflow has a shared conceptual core, but product-specific schema controls remain separate branches. Do not force OpenSearch-only dynamic allow-template modes or Elasticsearch runtime-template behavior into a single JSON file merely to claim portability.

Execution and safety note

The generation environment does not run the two search servers, so commands are specified as deterministic labs and expected invariants rather than represented as captured live output. Run them only against the disposable AtlasMart course indices/templates, record the actual API responses from your pinned versions, and never test alias cutover or destructive cleanup against production names.

1. Define the schema release artifact

schema-manifest-v2.yaml
schema: atlasmart-products
owner: search-platform
release: 2
compatible_servers:
  elasticsearch: 9.5.3
  opensearch: 3.8.0
components:
  settings: atlasmart-products-settings-v1
  mappings: atlasmart-products-mappings-v2
index_template: atlasmart-products-template-v2
index_pattern: atlasmart-products-v2-*
priority: 210
physical_target: atlasmart-products-v2-000001
read_alias: atlasmart-products-read
write_alias: atlasmart-products-write
rollback_target: atlasmart-products-v1-000001
write_reconciliation: bounded-pause-for-course-lab
acceptance:
  mapping_contract: required
  fixture_count: exact
  search_cases: required
  aggregation_cases: required
  relevance_regression: required
  shard_failures: zero

Store a checksum of the deployed component/index-template JSON if your release system supports it. The manifest separates logical release identity from whatever timestamp the cluster reports.

2. Preflight: inventory and simulate before mutation

Dev Tools · evidence capture
GET _component_template/atlasmart-products-*
GET _index_template/atlasmart-products-*
GET _alias/atlasmart-products-read
GET _alias/atlasmart-products-write

POST _index_template/_simulate_index/atlasmart-products-v2-canary

Fail the deployment on unexpected higher-priority overlap, missing components, alias pointing to an unknown generation, or a resolved mapping/settings diff that is not part of the release. This is where environment drift becomes visible while rollback is still trivial: nothing has moved.

3. Canary the resolved contract

Deployment job sketch
# Pseudocode for a product-aware schema deployment job
for target in ["elasticsearch-9.5.3", "opensearch-3.8.0"]:
    authenticate(target)
    assert_server_version(target)

    inventory = capture_templates_components_aliases(target)
    assert_no_manual_drift(inventory, expected_manifest)

    put_versioned_components(candidate)
    simulate_template(candidate)
    resolved = simulate_index("atlasmart-products-v2-canary")
    assert_resolved_contract(resolved)

    create_canary_index()
    index_boundary_fixtures()
    assert_mapping_field_caps_search_and_facets()
    delete_canary_index()

# Production migration is a separate controlled stage:
# build destination -> bulk copy -> reconcile writes -> validate -> alias cutover
# -> smoke/observe -> either keep v2 or execute tested rollback.

Boundary fixtures should include values that previously caused incidents: date-like identifiers, numeric-looking strings, maximum keyword lengths, null/empty arrays, merchant extension paths, wrong-type values expected to fail, and the query/facet/relevance cases used by the application. The canary proves behavior of the resolved template, not just validity of each component in isolation.

4. Migration state machine

State Entry evidence Exit condition
PREPARED v2 templates installed + simulation passes Canary tests pass.
COPYING destination created and write strategy selected Bulk copy succeeds or failures are classified/resolved.
RECONCILING bulk copy complete Destination contains every accepted write through a defined boundary.
VALIDATING freshness boundary met Counts/IDs/content/search/facets/relevance/shard health pass.
CUTOVER rollback target retained Atomic alias switch verified through alias-based smoke tests.
OBSERVING traffic on v2 SLO/error/relevance signals stable for agreed window.
ROLLED_BACK or CLOSED decision recorded If closed, old generation removal follows backup/retention policy.

A state machine prevents “somebody already switched the alias” ambiguity. Persist timestamps, task IDs, source/destination counts, failure samples, alias snapshots and decision owner for each transition.

5. Zero-downtime requires a live-write strategy

Strategy Strength Risk/cost
Short write pause at final cutover Simple and deterministic Only acceptable when bounded downtime/write buffering fits SLO.
Dual-write old + new generations Low cutover gap Application complexity; must handle partial dual-write failure/idempotency.
Authoritative log/CDC replay Strong reconciliation model Needs durable ordered/change source and lag observability.
Reindex only, no reconciliation Not a valid live-write strategy Can miss or stale writes accepted after source snapshot.

For the course lab, the write pause is explicit and small because the dataset is tiny. Calling that “zero downtime” for a real high-write production system would be misleading. In production, zero read downtime can be easy with aliases; zero lost/stale writes requires write coordination.

6. Regression gates: schema, relevance and operations

Example gate report
schema-contract: PASS
  required fields/types: PASS
  dynamic-extension cases: PASS
  incompatible payload rejected: PASS
migration-integrity: PASS
  source/destination IDs reconciled: PASS
  final write boundary reconciled: PASS
search-contract: PASS
  filters/facets: PASS
  lexical judged-query thresholds: PASS
operations: PASS
  shard failures: 0
  p95/p99 latency: measured against environment SLO
  indexing rejection/error rate: within SLO
rollback: READY
  old generation retained
  alias rollback request tested in staging

Do not hard-code universal p95/p99 numbers into the course. Establish environment-specific SLOs from representative load, then fail deployment when the candidate materially regresses them. Keep quality thresholds segmented by query class so an average metric cannot hide a damaging tail segment.

7. Rollback is a tested workflow, not a sentence in a runbook

Before cutover, execute the alias-reversal sequence in staging and prove the old generation still serves the previous contract. After production cutover, retain the old generation through the rollback window and document write reconciliation if v2 receives new writes. A snapshot can provide disaster recovery, but restoring a snapshot is slower and operationally different from an immediate alias rollback.

Check your understanding

  1. What is the main purpose of simulate-index in CI?
  2. Why can a component-template test pass while the deployment is still wrong?
  3. What makes reindex migration safe under ongoing writes?
  4. Why keep relevance tests in a schema migration?
  5. When is the old generation safe to delete?
Review the answers

1. To assert the final configuration a concrete future index name would receive after pattern, priority and component resolution.

2. Another index template may win by priority, composition order may override values, or environment drift may change resolution.

3. An explicit reconciliation strategy such as dual-write, CDC/replay, or a bounded write fence for the final delta.

4. Mapping/analyzer/type changes can alter ranking/search behavior even when all documents copied successfully.

5. After the observation/rollback window, write reconciliation, backup/retention requirements and closure decision are complete.

Production judgment

Treat schema deployment as a privileged release process with separation of duties, immutable artifacts, target-version compatibility tests and complete evidence. Watch mapping growth, template drift, reindex task health, indexing/search latency, rejection rates, alias state and query-quality regressions. Managed services may restrict settings, plugins or privileges, so promotion must validate the actual target control plane rather than assuming self-managed parity.

Summary and next step

Chapter 09 has moved schema from an incidental side effect of ingestion to a versioned deployable contract: controlled dynamic behavior, ordered dynamic templates, composable template resolution, simulation, new-generation migration, atomic alias cutover and rollback. Chapter 10 applies those template and lifecycle ideas to append-heavy data streams, logs and metrics.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.