Chapter 12 · Aliases, Rollover, Reindex, Update/Delete by Query, and Zero-Downtime Index Changes
Rollover Conditions by Age/Size/Docs and Coordinating with Data Streams/Lifecycle Policies
Design rollover from measurable age, primary-size and document-count conditions, then distinguish manual alias rollover from data-stream and lifecycle-policy automation on Elasticsearch and OpenSearch.
Learning outcomes
AtlasMart telemetry and high-write indexes cannot grow forever as one physical index. Rollover creates a new write index when a chosen condition is met. The condition is an operational trigger, not a universal shard-size recommendation. This lesson connects alias rollover, data-stream rollover and lifecycle automation without treating ILM and OpenSearch ISM as the same subsystem.
Explain how rollover changes the current write target for an alias or data stream.
Use age, primary-shard size and document-count conditions as observable triggers rather than folklore.
Distinguish manual rollover from Elastic ILM and OpenSearch Index State Management/lifecycle mechanisms.
Use dry-run style evidence and post-rollover metadata to prove which generation receives writes.
Choose rollover conditions from recovery, latency, storage and retention requirements instead of fixed internet rules.
Examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 using the copied CA certificate, OpenSearch on https://localhost:9201 using the disposable demo certificate only for local learning, pinned server versions, and no moving latest tags. Data streams are preferred for append-heavy timestamped data. Alias rollover remains useful for migration and non-data-stream designs. Elasticsearch ILM and OpenSearch ISM/lifecycle controls have different policy syntax and capabilities; do not copy one policy body into the other product.
The generation environment does not run both search servers. Requests below are deterministic lab specifications reviewed against current product documentation. Expected output describes invariants, not fabricated captured results. Execute against disposable AtlasMart resources and record your own task IDs, counts, conflicts, throttling time, p95/p99 latency, CPU, disk growth and rollback evidence.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Alias rollover requires a recognizable write target
PUT atlasmart-orders-000001
{
"settings":{"number_of_shards":1,"number_of_replicas":0},
"mappings":{"properties":{
"order_id":{"type":"keyword"},
"created_at":{"type":"date"},
"total":{"type":"scaled_float","scaling_factor":100}
}},
"aliases":{"atlasmart-orders":{"is_write_index":true}}
}
POST atlasmart-orders/_doc?refresh=true
{"order_id":"O-001","created_at":"2026-09-11T08:00:00Z","total":42.50}
GET _alias/atlasmart-orders
The alias names the current write index. A rollover can create
atlasmart-orders-000002, set it as the write target
and demote the old generation. Reads through the alias can still
span the indices attached to it.
2. Conditions describe when to consider a new generation
| Condition | What it observes | What it does not guarantee |
|---|---|---|
max_age |
Elapsed age of the current write index | A fixed data volume or stable query latency. |
max_docs |
Document count used by the rollover implementation | Equal storage across documents or shards. |
max_size |
Primary-store size threshold | Equal per-shard recovery time or replica footprint. |
| min/max variants where supported | Lower/upper constraints around rollover decision | A universal best-practice number. |
POST atlasmart-orders/_rollover
{
"conditions":{
"max_age":"7d",
"max_docs":1000000,
"max_size":"20gb"
}
}
GET _alias/atlasmart-orders
GET _cat/indices/atlasmart-orders-*?v
On a tiny lab fixture the conditions will probably be false. That is useful: capture the response fields that report which conditions matched instead of fabricating a rollover. For a deterministic learning exercise, use a very small document condition in a disposable index or perform an unconditional rollover after showing the conditional request.
3. Data-stream rollover changes backing-index generation
GET _data_stream/atlasmart-telemetry
POST atlasmart-telemetry/_rollover
GET _data_stream/atlasmart-telemetry
GET _cat/indices/.ds-atlasmart-telemetry-*?v
The previous write backing index becomes a normal backing index and a new generation receives new documents. Do not treat the backing index naming pattern as an application contract. Read and write through the data-stream name unless an administrative operation explicitly requires a backing index.
4. Lifecycle automation is product-specific
| Concern | Elasticsearch 9.5.x | OpenSearch 3.8.x |
|---|---|---|
| Automated lifecycle | ILM for supported self-managed/hosted deployments; Serverless has different lifecycle controls | Index State Management and related OpenSearch lifecycle features use OpenSearch-specific policy APIs. |
| Rollover primitive | Rollover API for alias/data stream | Rollover API for alias/data stream |
| Policy portability | Do not assume ILM policy JSON works elsewhere | Do not assume ISM policy JSON is accepted by Elasticsearch |
| Decision input | Age/docs/size plus product policy features | Age/docs/size plus OpenSearch policy features |
Keep the application-level acceptance criteria portable—maximum recovery time, maximum shard growth, retention/compliance obligations—even when the policy implementation differs.
5. Deliberately wrong approach: “50 GB is always correct”
A fixed rollover number copied from a blog ignores document density, query fan-out, storage hardware, snapshot speed, replica count, merge workload and recovery objectives. Repair the decision by measuring shard recovery throughput and representative query/indexing performance, then derive a rollover band that maintains headroom. Treat it as a hypothesis to revalidate after version, hardware or workload changes.
rollover objective:
- keep single-shard recovery <= RTO budget
- keep p95/p99 query latency within SLO
- keep merge/indexing pressure below operational thresholds
- preserve minimum free-disk headroom
- keep snapshot/restore drill within RPO/RTO plan
measure:
GET _cat/shards/atlasmart-orders-*?v
GET _cat/indices/atlasmart-orders-*?v
GET _nodes/stats/indices,fs,jvm
6. Coordinate rollover with retention and restore
Rollover creates generations; it does not by itself delete anything or prove that deletion is compliant. Define when an old generation becomes read-only, when it is snapshotted, how restore is tested, when it can be removed, and which legal/tenant requirements override ordinary retention. A backup is independent recovery state; replicas and older rollover generations are not backups.
Check your understanding
- What changes during alias rollover?
- Why is max_size not a universal shard recommendation?
- What happens to a data stream on rollover?
- Are ILM and OpenSearch ISM interchangeable?
- What must accompany retention deletion?
Review the answers
1. A new concrete index becomes the alias write index and the previous generation stops receiving writes.
2. Storage size alone does not encode recovery throughput, query cost, merge pressure, hardware or RTO.
3. A new backing index becomes the write index and the generation increases.
4. No. They solve related lifecycle problems with product-specific policy APIs and capabilities.
5. Recovery/restore evidence, compliance requirements and explicit ownership of the deletion policy.
Production judgment
Choose rollover from evidence: ingest rate, primary-shard growth, segment/merge behavior, search fan-out, recovery speed, snapshot duration and headroom. Use application-independent logical names so rollover does not require client deployment. Test policy behavior on staging with realistic time and volume, and alert on rollover failures or a write index that grows beyond the expected envelope.
Summary and next step
Rollover manages future generations. The next lesson handles a different problem: copying existing documents into a newly prepared index while controlling migration pressure and validating what actually moved.
Authoritative references
- Elastic aliases — Alias filters, routing, write-index behavior and alias management.
- Elastic update aliases API — Multi-action alias changes through POST /_aliases.
- Elastic rollover API — Manual rollover for data streams and index aliases; conditions and write-index semantics.
- Elastic reindex API — Reindex requirements, throttling, slicing, remote sources and destination preparation.
- Elastic update by query — Snapshot semantics, conflicts, slicing, throttling and task monitoring.
- Elastic task API — Task progress and status for long-running operations.
- OpenSearch aliases API — Atomic alias action sets, filters, routing and write-index behavior.
- OpenSearch rollover API — Rollover for aliases/data streams with age, docs and size conditions.
- OpenSearch reindex API — Local/remote reindex, slicing, throttling, background tasks and validation fields.
- OpenSearch update by query — Snapshot-based update-by-query, conflicts and partial completion semantics.
- OpenSearch delete by query — Snapshot-based deletion, conflicts and non-rollback behavior.
- OpenSearch cancel tasks — Cancellation behavior and cancellable task checks.