Chapter 12 · Aliases, Rollover, Reindex, Update/Delete by Query, and Zero-Downtime Index Changes

Update-by-Query/Delete-by-Query, Task APIs, Cancellation, Version Conflicts, and Cluster Impact

Run mass mutation and deletion as observable background tasks: understand snapshots, conflicts, partial progress, throttling, cancellation limits, and why successful operations are not automatically rolled back.

Intermediate110–135 minutesZero-downtime migration labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart discovers that a subset of products needs a schema-compatible backfill and another subset must be removed for a retention or legal requirement. _update_by_query and _delete_by_query look like single REST calls, but internally they search a snapshot and execute batches of writes. They can conflict with concurrent changes, consume cluster resources, finish partially, and continue as background tasks.

01

Explain snapshot and batch semantics for update-by-query and delete-by-query.

02

Use version-conflict policy deliberately and understand that successful earlier batches are not rolled back.

03

Run by-query work asynchronously, monitor task progress and cancel only with realistic expectations.

04

Use throttling and slicing to bound impact on search, indexing, merges and recovery.

05

Prefer safer alternatives when a destructive bulk mutation would make rollback or compliance evidence weak.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 with the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 using the copied CA certificate, OpenSearch on https://localhost:9201 using the disposable demo certificate only for local learning, pinned server versions, and no moving latest tags. Both products expose broadly similar by-query APIs, but response details, retry behavior and task persistence/cleanup can differ. Always inspect the target version’s task and by-query documentation.

Execution and measurement note

The generation environment does not run both search servers. Requests below are deterministic lab specifications reviewed against current product documentation. Expected output describes invariants, not fabricated captured results. Execute against disposable AtlasMart resources and record your own task IDs, counts, conflicts, throttling time, p95/p99 latency, CPU, disk growth and rollback evidence.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. A by-query operation is not a transaction

Create a small disposable fixture
PUT atlasmart-products-v1
{
  "settings":{"number_of_shards":1,"number_of_replicas":0},
  "mappings":{"dynamic":"strict","properties":{
    "sku":{"type":"keyword"},
    "name":{"type":"text","fields":{"keyword":{"type":"keyword"}}},
    "category":{"type":"keyword"},
    "price":{"type":"scaled_float","scaling_factor":100},
    "available":{"type":"boolean"},
    "updated_at":{"type":"date"}
  }}
}

POST _bulk?refresh=true
{"index":{"_index":"atlasmart-products-v1","_id":"P-1001"}}
{"sku":"P-1001","name":"Wireless Noise Cancelling Headphones","category":"audio","price":199.99,"available":true,"updated_at":"2026-09-11T07:00:00Z"}
{"index":{"_index":"atlasmart-products-v1","_id":"P-1002"}}
{"sku":"P-1002","name":"USB-C Travel Charger","category":"power","price":49.95,"available":true,"updated_at":"2026-09-11T07:01:00Z"}
{"index":{"_index":"atlasmart-products-v1","_id":"P-1003"}}
{"sku":"P-1003","name":"Ergonomic Mechanical Keyboard","category":"input","price":129.00,"available":false,"updated_at":"2026-09-11T07:02:00Z"}
Backfill a schema-compatible field
POST atlasmart-products-v1/_update_by_query?wait_for_completion=false&requests_per_second=10&conflicts=proceed
{
  "query":{"term":{"category":"audio"}},
  "script":{
    "lang":"painless",
    "source":"ctx._source.updated_at = params.ts",
    "params":{"ts":"2026-09-11T08:30:00Z"}
  }
}

The operation takes a point-in-time/snapshot view according to the product implementation, then writes matching documents in batches. If a matching document changes before its batch update, version conflict handling determines whether the operation aborts or records the conflict and proceeds. Already successful updates remain successful even if a later batch fails.

2. Monitor progress as a task

Task inspection and cancellation
GET _tasks?actions=*byquery&detailed=true
GET _tasks/<node_id:task_id>
POST _tasks/<node_id:task_id>/_cancel

Cancellation is cooperative and not every task is cancelable. Even after a cancel request, some in-flight batches may complete. Therefore, “cancel” is a resource-control action, not an undo button. If you require business rollback, design a reversible data change or migrate to a new index instead of mutating the only copy.

3. Delete-by-query needs an evidence and recovery plan

Disposable delete example
POST atlasmart-products-v1/_delete_by_query?wait_for_completion=false&requests_per_second=5&conflicts=proceed
{
  "query":{"term":{"available":false}}
}

A successful delete removes matching documents from the index’s live view; storage reclamation occurs through normal segment merging later. A replica is not an undo source because it receives the same deletion. If deletion must be recoverable, prove snapshot/restore or preserve a source-of-truth copy before starting. For legal deletion, retention of backups can itself be policy-sensitive; follow the applicable compliance process rather than making a search-engine-only decision.

4. Throttling and slicing trade time for pressure

Control Benefit Risk
requests_per_second Bounds write sub-request rate and creates breathing room Migration lasts longer; too high still overloads cluster.
slices Parallelizes work Raises concurrent search/write/merge pressure and complicates observation.
conflicts=abort Stops on a version conflict May stop after earlier successful batches; not full rollback.
conflicts=proceed Completes non-conflicting work Requires explicit reconciliation of skipped/conflicting docs.
cancel task Stops cancellable work cooperatively Does not reverse completed batches.

Choose values from cluster headroom and the urgency of the operation. Measure p95/p99 search latency, indexing throughput, rejection counters, merge/disk pressure and task progress before increasing concurrency.

5. Deliberately wrong approach: update every document in place before migration

If the update is part of an incompatible schema migration, mutating v1 can destroy your rollback point. A safer pattern is to preserve v1, create v2 with the new contract, transform during reindex, validate, then cut over. Use update-by-query mainly for compatible backfills or repairs where the rollback strategy is explicit.

Pre-flight gate for destructive/by-query work
scope query reviewed with sample hits: PASS
source snapshot / authoritative replay source: PASS
estimated matched documents: RECORDED
requests_per_second limit: SET
slices/concurrency: JUSTIFIED
version-conflict policy: DOCUMENTED
p95/p99 + CPU/disk guardrails: SET
cancellation procedure: TESTED
post-operation validation query: READY
rollback/recovery path: TESTED
security approval / tenant scope: PASS

6. Security and tenant boundaries

By-query endpoints can alter or delete many documents with one request. Use least-privilege service identities, explicit index names where possible, query review, audit logging and tenant scoping. Do not expose raw user-provided Query DSL to these endpoints. For cross-tenant indices, a mistaken filter can become a mass data incident rather than a search bug.

Check your understanding

  1. Are successful batches rolled back if update-by-query later fails?
  2. Does cancelling a task undo completed work?
  3. Why can conflicts=proceed be dangerous?
  4. When is reindex-to-new-index safer than update-by-query?
  5. Why are replicas not a recovery mechanism for delete-by-query?
Review the answers

1. No. Earlier successful updates persist.

2. No. Cancellation is cooperative resource control, not transaction rollback.

3. It can leave a mixed state unless skipped/conflicting documents are reconciled explicitly.

4. When the change is schema-incompatible, rollback-sensitive, or should preserve the old copy unchanged.

5. They replicate the deletion; independent snapshots or source-of-truth replay are needed for recovery.

Production judgment

Treat mass mutation as a change-management event. Estimate scope before launch, run asynchronously, set explicit rate limits, watch cluster/resource SLOs, and persist the final response including version conflicts and failures. After completion, re-run the exact selection query and business validation checks. For irreversible deletions, require recovery/compliance sign-off before starting.

Summary and next step

You now have the operational controls for long-running data changes. The final lesson combines them into a complete zero-downtime v1→v2 migration with explicit write coordination, cutover evidence and rollback.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.