Chapter 17 · Index Lifecycle: Elastic ILM/Data Tiers and OpenSearch ISM
Lifecycle Goals: Rollover, Retention, Tiering, Shrink/Force-Merge-Like Actions, and Cost/Recovery Tradeoffs
Translate AtlasMart retention and cost goals into lifecycle invariants before choosing Elastic ILM or OpenSearch ISM, with explicit rollover, tiering, merge/shrink, recovery, and compliance tradeoffs.
Learning outcomes
AtlasMart stores application logs for incident response and audit. The expensive mistake is to begin with vendor JSON instead of a retention contract. This lesson starts with lifecycle intent and turns it into observable invariants that can later be mapped to ILM or ISM.
Express rollover, retention and tiering as workload/recovery requirements rather than calendar folklore.
Distinguish rollover from retention, tier movement, shrink, force merge and backup.
Explain why shrink/force-merge-like work belongs after writes stop and must be resource-budgeted.
Compare Elastic ILM/data-tier and OpenSearch ISM state-machine control models without claiming equivalence.
Define acceptance evidence for lifecycle correctness, cost, recovery, compliance and rollback.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. AtlasMart keeps https://localhost:9200 for
Elasticsearch with CA verification and
https://localhost:9201 for the disposable
OpenSearch demo certificate using
OPENSEARCH_INITIAL_ADMIN_PASSWORD. Existing
containers remain atlasmart-es and
atlasmart-os. The lifecycle lab uses isolated
aliases atlasmart-life-es and
atlasmart-life-os, one primary shard and zero
replicas so a single-node local lab can run; that topology is
not production guidance. OpenSearch demo -k is
local-only. No moving latest tags are used.
Elastic ILM is a phase/action engine integrated with Elastic data tiers. OpenSearch ISM is a plugin state machine with states, ordered actions and transitions. Similar goals such as rollover and deletion do not make policy JSON, policy-update timing, tier semantics, troubleshooting APIs, permissions, managed-service behavior, or feature availability portable.
The generation environment does not run the AtlasMart Docker containers. Commands below are reproducible lab instructions, while response snippets are explicitly labeled expected shapes/invariants rather than fabricated measurements. Measure your own transition timing, I/O, CPU, merge time, p95/p99 search latency and storage consumption.
1. Lifecycle intent before policy syntax
A lifecycle policy is automation around the changing value and access pattern of an index. It does not make data “old” by itself. The design inputs are ingest rate, shard growth, recovery time objective, search frequency, query latency target, retention/compliance dates, snapshot/restore posture and storage economics.
Rollover changes the active write generation. Retention decides how long generations remain. Tiering places data on hardware/storage appropriate to expected access. Shrink creates a new index with fewer primary shards. Force merge compacts immutable Lucene segments. None of these actions is a backup.
| Decision | Mechanism question | Observable evidence |
|---|---|---|
| Rollover | When is the current write shard no longer an acceptable recovery/search unit? | primary-shard bytes/docs, indexing rate, recovery drill, p95/p99 |
| Tiering | Can older queries tolerate slower/cheaper resources? | age-segmented query latency, CPU/I/O, cache warmness, capacity headroom |
| Shrink | Are old generations over-sharded for their now-lower concurrency? | shard count, shard size, recovery time, allocation feasibility |
| Force merge | Is a read-only generation carrying merge/deleted-doc debt worth compacting? | segment count, deleted docs, merge I/O/time, free-disk headroom |
| Delete | Has retention expired and recoverability/compliance evidence passed? | snapshot/restore proof, legal hold state, policy explain state, deletion audit |
2. The two lifecycle control planes
Elasticsearch ILM organizes indices into named phases—hot, warm, cold, frozen and delete—and executes product-defined actions within those phases. Data tiers are node roles and placement preferences, not generic strings called “hot” or “warm.” ILM can automatically inject migration behavior for warm/cold phases when appropriate.
OpenSearch ISM instead manages an index through exactly one
named state at a time. A state contains ordered
actions and transition conditions. The state names are yours;
naming a state warm does not magically create an
Elastic-style warm data tier. Placement is a separate allocation
decision using OpenSearch-supported attributes/capabilities.
| Intent | Elastic ILM | OpenSearch ISM |
|---|---|---|
| Rollover | hot-phase rollover action on data stream or compliant alias | rollover action on managed index; alias/name prerequisites |
| Lifecycle progression | fixed phase vocabulary + action ordering | user-defined states + ordered actions + transitions |
| Tiering | data tiers / tier preference integrated with ILM | allocation/search-only/remote-storage capabilities depending topology; not same model |
| Explain | GET index/_ilm/explain |
GET _plugins/_ism/explain/index |
| Policy preview | reason from policy + explain; test on disposable indices | ISM simulate API can preview transition/action evaluation without mutation |
| Serverless caveat | Elastic Cloud Serverless uses data stream lifecycle rather than ILM | Amazon OpenSearch Serverless lifecycle/storage controls differ from upstream self-managed ISM |
3. Cost and recovery are coupled
Moving older data to cheaper storage can lower cost but increase recovery/search latency. Shrinking can reduce shard overhead but concentrates recovery into larger units. Force merge can reduce segment count but generates heavy I/O and temporary disk amplification. Deletion lowers cost only after it destroys the local copy. Therefore a lifecycle design is a sequence of bounded risks, not a cleanup script.
Those numbers contain no AtlasMart ingest volume, recovery objective, compliance requirement or measured query distribution. The repair is to define measurable thresholds and derive times from business/data behavior. Calendar ages can still be part of the policy, but they must encode a requirement, not folklore.
AtlasMart lifecycle lab contract
The chapter uses a small log-like fixture so lifecycle transitions are observable without consuming substantial resources. The application writes through aliases rather than physical index names. That isolates the lifecycle contract from generation numbers and lets the write target change atomically.
| Concern | Elasticsearch lab | OpenSearch lab |
|---|---|---|
| Write alias | atlasmart-life-es |
atlasmart-life-os |
| Bootstrap index | atlasmart-life-es-000001 |
atlasmart-life-os-000001 |
| Policy | atlasmart-life-es-v1 |
atlasmart-life-os-v1 |
| Template | atlasmart-life-es-template-v1 |
atlasmart-life-os-template-v1 |
| Mapping |
@timestamp date;
service.name/log.level
keyword; message text
|
same logical mapping |
| Local topology | 1 primary, 0 replicas | 1 primary, 0 replicas |
| Production deletion | 30-day example is policy intent, not a legal recommendation | same |
The production requirement used for comparison is: roll active logs before any primary shard violates the platform’s measured recovery/latency envelope; keep recent data on the performance tier that meets the SLO; make older data cheaper only when recovery and query latency permit; and delete only after the retention/compliance date and required recoverability evidence are satisfied.
4. Tiny cross-product observation plan
# Before lifecycle actions
GET _cat/aliases/atlasmart-life-*?v
GET atlasmart-life-es-000001/_settings
GET atlasmart-life-os-000001/_settings
# Elasticsearch
GET atlasmart-life-es-000001/_ilm/explain?human
# OpenSearch
GET _plugins/_ism/explain/atlasmart-life-os-000001?show_policy=true
# Always pair policy state with index/storage evidence
GET _cat/indices/atlasmart-life-*?v&h=index,docs.count,store.size,pri,rep
GET _cat/segments/atlasmart-life-*?v
An explain response proves what the lifecycle controller believes it is doing; it does not prove backups exist, that an older tier meets p99, or that a force merge is harmless. Those require separate evidence.
5. Production judgment
Choose lifecycle automation only after defining the data owner, retention authority, backup owner, acceptable recovery time, failure domain, tier capacity and operational escalation path. Never let lifecycle deletion outrun snapshot verification or legal/compliance state. Never run shrink/force merge on active write generations simply because the action exists.
The next lesson maps this intent specifically to Elastic ILM and data tiers.
Check your understanding
- Why is rollover not the same as retention?
- Why is force merge not a durability action?
- Does an ISM state named warm equal an Elastic warm tier?
- What must be proven before lifecycle deletion?
- Why avoid universal rollover sizes?
Review the answers
1. Rollover changes the active write generation; retention determines how long older generations remain.
2. It rewrites Lucene segments for storage/search characteristics; durability and recoverability require separate write/commit/replica/snapshot mechanisms.
3. No. ISM state names are policy states; Elastic data tiers are specific node roles and allocation semantics.
4. Retention/compliance eligibility plus the required backup/restore or other recoverability evidence.
5. Optimal rollover depends on ingest rate, shard topology, recovery/search SLOs, storage and operational headroom.
Summary
Lifecycle design begins with retention, recovery, query and cost requirements. ILM and ISM can implement overlapping goals, but their state models, tiering semantics, policy updates and APIs differ. Chapter 17 now applies that intent to Elastic ILM.
Authoritative references
- Elastic index lifecycle management — ILM scope, availability, and lifecycle concepts for indices and data streams.
- Elastic ILM phases and actions — Hot/warm/cold/frozen/delete phases, cached phase execution, and transition rules.
- Elastic Explain lifecycle API — Current phase/action/step, failures, and phase execution evidence.
- Elastic data tiers — Content/hot/warm/cold/frozen roles and lifecycle placement concepts.
- Elastic policy updates — Policy versions and cached phase-definition behavior.
- Elastic ILM troubleshooting — ERROR steps, retry behavior, and safe diagnosis.
- OpenSearch Index State Management — ISM model, job cadence, policy attachment, and managed-index workflow.
- OpenSearch ISM policies — States, actions, transitions, rollover, force merge, allocation, and templates.
- OpenSearch ISM API — Policy OCC, explain, retry, change-policy, and simulation APIs.
- OpenSearch ISM error prevention — Pre-action validation and explain diagnostics.
- OpenSearch artifacts by version — Current OpenSearch release baseline.
- Elasticsearch downloads — Current Elasticsearch release baseline.