Chapter 17 · Index Lifecycle: Elastic ILM/Data Tiers and OpenSearch ISM
Elastic ILM Phases/Actions, Data Tiers, Rollover Aliases/Data Streams, and Explain/Troubleshooting
Implement and diagnose an AtlasMart Elasticsearch lifecycle with ILM phases, actions, data-tier migration, rollover aliases/data streams, explain output, and safe troubleshooting.
Learning outcomes
AtlasMart now wants the lifecycle intent from Lesson 1 implemented in Elasticsearch. The target is not merely “create an ILM policy”; it is to understand phase execution, data-tier placement, alias/data-stream rollover, cached phase definitions, explain output and safe error recovery.
Model hot/warm/cold/frozen/delete ILM phases and distinguish phase age from rollover-relative age.
Use rollover aliases or data streams correctly and know the prerequisites that make rollover deterministic.
Explain data-tier migration and why tier preference is not equivalent to arbitrary node attributes.
Read ILM explain output as phase/action/step state and diagnose ERROR steps before retrying.
Recognize Serverless and subscription/topology boundaries without presenting ILM as universal Elastic behavior.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. AtlasMart keeps https://localhost:9200 for
Elasticsearch with CA verification and
https://localhost:9201 for the disposable
OpenSearch demo certificate using
OPENSEARCH_INITIAL_ADMIN_PASSWORD. Existing
containers remain atlasmart-es and
atlasmart-os. The lifecycle lab uses isolated
aliases atlasmart-life-es and
atlasmart-life-os, one primary shard and zero
replicas so a single-node local lab can run; that topology is
not production guidance. OpenSearch demo -k is
local-only. No moving latest tags are used.
Elastic ILM is a phase/action engine integrated with Elastic data tiers. OpenSearch ISM is a plugin state machine with states, ordered actions and transitions. Similar goals such as rollover and deletion do not make policy JSON, policy-update timing, tier semantics, troubleshooting APIs, permissions, managed-service behavior, or feature availability portable.
The generation environment does not run the AtlasMart Docker containers. Commands below are reproducible lab instructions, while response snippets are explicitly labeled expected shapes/invariants rather than fabricated measurements. Measure your own transition timing, I/O, CPU, merge time, p95/p99 search latency and storage consumption.
1. ILM phase execution is a state machine with cached phase definitions
ILM uses five phase names: hot, warm, cold, frozen and delete.
An index can progress only when the current phase actions
complete and the next phase age requirement is met. If rollover
occurs, later min_age calculations are relative to
rollover time rather than the original creation time.
When an index enters a phase, Elasticsearch caches that phase definition in index metadata. A later policy edit increments policy version, but ILM only applies changes to the cached phase if they are safe. Otherwise the index finishes the phase using the cached definition and sees the newer policy when it enters a later phase. This is why “I edited the JSON” does not mean every managed index instantly follows the new behavior.
| Phase | Typical intent | Important constraint |
|---|---|---|
| hot | active writes/search; rollover | shrink/force merge in hot require rollover; do not compact the active write index |
| warm | read-mostly; reduce shard/segment cost | requires capacity and placement that actually exists |
| cold | infrequent search; cheaper recovery-oriented storage | query/recovery SLO must tolerate colder placement |
| frozen | rare access; very low local footprint patterns | feature/topology/subscription specifics matter |
| delete | retention expiry | delete only after compliance/recoverability gates |
2. Build the disposable AtlasMart ILM policy
PUT _ilm/policy/atlasmart-life-es-v1
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": { "max_docs": 3 }
}
},
"warm": {
"min_age": "0ms",
"actions": {
"forcemerge": { "max_num_segments": 1 }
}
},
"delete": {
"min_age": "30d",
"actions": { "delete": {} }
}
}
}
}
max_docs: 3 and immediate warm eligibility exist
solely to make state observable. Production values must come
from shard growth, recovery, query and compliance
requirements. Force merge is resource intensive; here it runs
only after rollover on a tiny disposable generation.
PUT _index_template/atlasmart-life-es-template-v1
{
"index_patterns": ["atlasmart-life-es-*"],
"template": {
"settings": {
"index.number_of_shards": 1,
"index.number_of_replicas": 0,
"index.lifecycle.name": "atlasmart-life-es-v1",
"index.lifecycle.rollover_alias": "atlasmart-life-es"
},
"mappings": {
"properties": {
"@timestamp": {"type": "date"},
"service.name": {"type": "keyword"},
"log.level": {"type": "keyword"},
"message": {"type": "text"}
}
}
}
}
PUT atlasmart-life-es-000001
{
"aliases": {
"atlasmart-life-es": {"is_write_index": true}
}
}
3. Rollover alias versus data stream
An ILM rollover target may be a data stream or a compliant index
alias. For an alias, the backing index name must have a numeric
suffix such as -000001, the alias must point to it,
and it must be the write index. A data stream carries its own
backing-index/write-index model and is often the cleaner choice
for append-heavy telemetry.
Do not copy alias rollover settings into a data stream blindly. Chapter 10 established data-stream semantics; this chapter keeps an alias deliberately so the same lifecycle intent can be compared with OpenSearch ISM.
POST atlasmart-life-es/_doc
{"@timestamp":"2026-09-11T12:00:00Z","service.name":"checkout","log.level":"INFO","message":"order accepted"}
POST atlasmart-life-es/_doc
{"@timestamp":"2026-09-11T12:00:01Z","service.name":"checkout","log.level":"WARN","message":"inventory retry"}
POST atlasmart-life-es/_doc
{"@timestamp":"2026-09-11T12:00:02Z","service.name":"checkout","log.level":"INFO","message":"order confirmed"}
GET atlasmart-life-es-000001/_ilm/explain?human
GET _cat/aliases/atlasmart-life-es?v
GET _cat/indices/atlasmart-life-es-*?v
ILM polls periodically. The expected invariant is eventual
creation of a later generation and removal of write status from
-000001 once rollover completes. Exact wall-clock
timing is not guaranteed by the document ingest request.
4. Data tiers and migration
Elastic data tiers are node roles such as data_hot,
data_warm, data_cold and
data_frozen. ILM’s migrate action updates tier
preference for warm/cold placement and is automatically injected
in supported phases unless disabled. On a one-node lab there is
no meaningful cost tier to demonstrate—the same machine owns all
work—so the correct lesson is to inspect settings, not to invent
a tier-performance improvement.
GET atlasmart-life-es-000001/_settings?filter_path=*.settings.index.routing.allocation.include._tier_preference
GET _cluster/allocation/explain
{
"index": "atlasmart-life-es-000001",
"shard": 0,
"primary": true
}
In production, a tier migration is complete only if shard allocation actually moves to eligible capacity and query/recovery SLOs remain acceptable. A phase label alone does not prove placement.
5. Explain, error, retry, and dangerous manual movement
GET atlasmart-life-es-*/_ilm/explain?human
GET atlasmart-life-es-*/_ilm/explain?only_errors=true&human
GET _ilm/status
_ilm/explain identifies the policy, phase, action,
step, timing and failure metadata. If an index enters
ERROR, repair the root cause first, then
POST /index/_ilm/retry. Do not treat retry as a
root-cause fix.
_ilm/move.
The move-to-step API is explicitly an expert/destructive operation: it runs the selected step and can cause data loss if used incorrectly. Prefer explaining the state, fixing the allocation/alias/shrink/permission problem and retrying the failed step. Manual movement is an exceptional recovery action with change control and rollback evidence.
6. Security and managed-service boundaries
ILM executes with the privileges associated with the user who last updated the policy, so lifecycle automation can fail after role/permission changes. Elastic Cloud Serverless does not expose ILM in the same cluster-management form; data stream lifecycle is the simpler serverless lifecycle mechanism. Do not present a self-managed ILM runbook as a Serverless runbook.
The next lesson implements the same retention intent with OpenSearch ISM and shows why the JSON and state semantics differ.
Check your understanding
- When does ILM cache a phase definition?
- What does rollover change about age calculations?
- Does entering warm prove the shard moved to cheaper hardware?
- What should happen before an ILM retry?
- Why is _ilm/move not a normal troubleshooting shortcut?
Review the answers
1. When an index enters a phase; safe policy changes may update it, while unsafe retroactive changes wait for later phase progression.
2. Later ILM min_age values are relative to rollover time when rollover was used.
3. No. Inspect tier preference/allocation and actual node placement/capacity.
4. Diagnose and repair the underlying cause of the ERROR step.
5. It is an expert-level destructive API that can execute lifecycle steps out of the normal safe progression.
Summary
Elastic ILM combines phase/action execution with rollover and data-tier placement. Correct operation is demonstrated through explain, settings and allocation evidence—not policy presence alone. The next lesson maps the same business intent to OpenSearch ISM.
Authoritative references
- Elastic index lifecycle management — ILM scope, availability, and lifecycle concepts for indices and data streams.
- Elastic ILM phases and actions — Hot/warm/cold/frozen/delete phases, cached phase execution, and transition rules.
- Elastic Explain lifecycle API — Current phase/action/step, failures, and phase execution evidence.
- Elastic data tiers — Content/hot/warm/cold/frozen roles and lifecycle placement concepts.
- Elastic policy updates — Policy versions and cached phase-definition behavior.
- Elastic ILM troubleshooting — ERROR steps, retry behavior, and safe diagnosis.
- OpenSearch Index State Management — ISM model, job cadence, policy attachment, and managed-index workflow.
- OpenSearch ISM policies — States, actions, transitions, rollover, force merge, allocation, and templates.
- OpenSearch ISM API — Policy OCC, explain, retry, change-policy, and simulation APIs.
- OpenSearch ISM error prevention — Pre-action validation and explain diagnostics.
- OpenSearch artifacts by version — Current OpenSearch release baseline.
- Elasticsearch downloads — Current Elasticsearch release baseline.