Chapter 12 · Aliases, Rollover, Reindex, Update/Delete by Query, and Zero-Downtime Index Changes
Update-by-Query/Delete-by-Query, Task APIs, Cancellation, Version Conflicts, and Cluster Impact
Run mass mutation and deletion as observable background tasks: understand snapshots, conflicts, partial progress, throttling, cancellation limits, and why successful operations are not automatically rolled back.
Learning outcomes
AtlasMart discovers that a subset of products needs a
schema-compatible backfill and another subset must be removed
for a retention or legal requirement.
_update_by_query and
_delete_by_query look like single REST calls, but
internally they search a snapshot and execute batches of writes.
They can conflict with concurrent changes, consume cluster
resources, finish partially, and continue as background tasks.
Explain snapshot and batch semantics for update-by-query and delete-by-query.
Use version-conflict policy deliberately and understand that successful earlier batches are not rolled back.
Run by-query work asynchronously, monitor task progress and cancel only with realistic expectations.
Use throttling and slicing to bound impact on search, indexing, merges and recovery.
Prefer safer alternatives when a destructive bulk mutation would make rollback or compliance evidence weak.
Examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 using the copied CA certificate, OpenSearch on https://localhost:9201 using the disposable demo certificate only for local learning, pinned server versions, and no moving latest tags. Both products expose broadly similar by-query APIs, but response details, retry behavior and task persistence/cleanup can differ. Always inspect the target version’s task and by-query documentation.
The generation environment does not run both search servers. Requests below are deterministic lab specifications reviewed against current product documentation. Expected output describes invariants, not fabricated captured results. Execute against disposable AtlasMart resources and record your own task IDs, counts, conflicts, throttling time, p95/p99 latency, CPU, disk growth and rollback evidence.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. A by-query operation is not a transaction
PUT atlasmart-products-v1
{
"settings":{"number_of_shards":1,"number_of_replicas":0},
"mappings":{"dynamic":"strict","properties":{
"sku":{"type":"keyword"},
"name":{"type":"text","fields":{"keyword":{"type":"keyword"}}},
"category":{"type":"keyword"},
"price":{"type":"scaled_float","scaling_factor":100},
"available":{"type":"boolean"},
"updated_at":{"type":"date"}
}}
}
POST _bulk?refresh=true
{"index":{"_index":"atlasmart-products-v1","_id":"P-1001"}}
{"sku":"P-1001","name":"Wireless Noise Cancelling Headphones","category":"audio","price":199.99,"available":true,"updated_at":"2026-09-11T07:00:00Z"}
{"index":{"_index":"atlasmart-products-v1","_id":"P-1002"}}
{"sku":"P-1002","name":"USB-C Travel Charger","category":"power","price":49.95,"available":true,"updated_at":"2026-09-11T07:01:00Z"}
{"index":{"_index":"atlasmart-products-v1","_id":"P-1003"}}
{"sku":"P-1003","name":"Ergonomic Mechanical Keyboard","category":"input","price":129.00,"available":false,"updated_at":"2026-09-11T07:02:00Z"}
POST atlasmart-products-v1/_update_by_query?wait_for_completion=false&requests_per_second=10&conflicts=proceed
{
"query":{"term":{"category":"audio"}},
"script":{
"lang":"painless",
"source":"ctx._source.updated_at = params.ts",
"params":{"ts":"2026-09-11T08:30:00Z"}
}
}
The operation takes a point-in-time/snapshot view according to the product implementation, then writes matching documents in batches. If a matching document changes before its batch update, version conflict handling determines whether the operation aborts or records the conflict and proceeds. Already successful updates remain successful even if a later batch fails.
2. Monitor progress as a task
GET _tasks?actions=*byquery&detailed=true
GET _tasks/<node_id:task_id>
POST _tasks/<node_id:task_id>/_cancel
Cancellation is cooperative and not every task is cancelable. Even after a cancel request, some in-flight batches may complete. Therefore, “cancel” is a resource-control action, not an undo button. If you require business rollback, design a reversible data change or migrate to a new index instead of mutating the only copy.
3. Delete-by-query needs an evidence and recovery plan
POST atlasmart-products-v1/_delete_by_query?wait_for_completion=false&requests_per_second=5&conflicts=proceed
{
"query":{"term":{"available":false}}
}
A successful delete removes matching documents from the index’s live view; storage reclamation occurs through normal segment merging later. A replica is not an undo source because it receives the same deletion. If deletion must be recoverable, prove snapshot/restore or preserve a source-of-truth copy before starting. For legal deletion, retention of backups can itself be policy-sensitive; follow the applicable compliance process rather than making a search-engine-only decision.
4. Throttling and slicing trade time for pressure
| Control | Benefit | Risk |
|---|---|---|
| requests_per_second | Bounds write sub-request rate and creates breathing room | Migration lasts longer; too high still overloads cluster. |
| slices | Parallelizes work | Raises concurrent search/write/merge pressure and complicates observation. |
| conflicts=abort | Stops on a version conflict | May stop after earlier successful batches; not full rollback. |
| conflicts=proceed | Completes non-conflicting work | Requires explicit reconciliation of skipped/conflicting docs. |
| cancel task | Stops cancellable work cooperatively | Does not reverse completed batches. |
Choose values from cluster headroom and the urgency of the operation. Measure p95/p99 search latency, indexing throughput, rejection counters, merge/disk pressure and task progress before increasing concurrency.
5. Deliberately wrong approach: update every document in place before migration
If the update is part of an incompatible schema migration, mutating v1 can destroy your rollback point. A safer pattern is to preserve v1, create v2 with the new contract, transform during reindex, validate, then cut over. Use update-by-query mainly for compatible backfills or repairs where the rollback strategy is explicit.
scope query reviewed with sample hits: PASS
source snapshot / authoritative replay source: PASS
estimated matched documents: RECORDED
requests_per_second limit: SET
slices/concurrency: JUSTIFIED
version-conflict policy: DOCUMENTED
p95/p99 + CPU/disk guardrails: SET
cancellation procedure: TESTED
post-operation validation query: READY
rollback/recovery path: TESTED
security approval / tenant scope: PASS
6. Security and tenant boundaries
By-query endpoints can alter or delete many documents with one request. Use least-privilege service identities, explicit index names where possible, query review, audit logging and tenant scoping. Do not expose raw user-provided Query DSL to these endpoints. For cross-tenant indices, a mistaken filter can become a mass data incident rather than a search bug.
Check your understanding
- Are successful batches rolled back if update-by-query later fails?
- Does cancelling a task undo completed work?
- Why can conflicts=proceed be dangerous?
- When is reindex-to-new-index safer than update-by-query?
- Why are replicas not a recovery mechanism for delete-by-query?
Review the answers
1. No. Earlier successful updates persist.
2. No. Cancellation is cooperative resource control, not transaction rollback.
3. It can leave a mixed state unless skipped/conflicting documents are reconciled explicitly.
4. When the change is schema-incompatible, rollback-sensitive, or should preserve the old copy unchanged.
5. They replicate the deletion; independent snapshots or source-of-truth replay are needed for recovery.
Production judgment
Treat mass mutation as a change-management event. Estimate scope before launch, run asynchronously, set explicit rate limits, watch cluster/resource SLOs, and persist the final response including version conflicts and failures. After completion, re-run the exact selection query and business validation checks. For irreversible deletions, require recovery/compliance sign-off before starting.
Summary and next step
You now have the operational controls for long-running data changes. The final lesson combines them into a complete zero-downtime v1→v2 migration with explicit write coordination, cutover evidence and rollback.
Authoritative references
- Elastic aliases — Alias filters, routing, write-index behavior and alias management.
- Elastic update aliases API — Multi-action alias changes through POST /_aliases.
- Elastic rollover API — Manual rollover for data streams and index aliases; conditions and write-index semantics.
- Elastic reindex API — Reindex requirements, throttling, slicing, remote sources and destination preparation.
- Elastic update by query — Snapshot semantics, conflicts, slicing, throttling and task monitoring.
- Elastic task API — Task progress and status for long-running operations.
- OpenSearch aliases API — Atomic alias action sets, filters, routing and write-index behavior.
- OpenSearch rollover API — Rollover for aliases/data streams with age, docs and size conditions.
- OpenSearch reindex API — Local/remote reindex, slicing, throttling, background tasks and validation fields.
- OpenSearch update by query — Snapshot-based update-by-query, conflicts and partial completion semantics.
- OpenSearch delete by query — Snapshot-based deletion, conflicts and non-rollback behavior.
- OpenSearch cancel tasks — Cancellation behavior and cancellable task checks.