Chapter 16 · Search Execution, Profiling, Slow Logs, Async/Search-After/PIT, and Pagination
Async Search, Cancellation, Timeouts, Partial Results, and Long-Running Analytical Requests
Design long-running search flows around product-specific async APIs, polling, cancellation, timeout semantics, partial results, retention and bounded resource ownership.
Learning outcomes
An AtlasMart analyst launches a cross-month aggregation that takes tens of seconds. The web tier keeps a synchronous connection open, times out at 15 seconds, retries automatically, and accidentally doubles cluster work. The query may be legitimate; the request lifecycle is not. Long-running search needs ownership, polling, cancellation and retention semantics.
Distinguish a synchronous search timeout from true task cancellation and from asynchronous search lifecycle.
Use the product-specific Elasticsearch and OpenSearch asynchronous search APIs without assuming endpoint compatibility.
Interpret partial-results/completion state and define whether partial data is acceptable for each application surface.
Bound result retention, response size, concurrency and tenant ownership for long-running analytics.
Test client disconnect/retry behavior so timeouts do not silently multiply expensive server work.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. AtlasMart keeps https://localhost:9200 for
Elasticsearch with CA verification and
https://localhost:9201 for the disposable
OpenSearch demo certificate using
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The
established containers are atlasmart-es and
atlasmart-os. This chapter deliberately creates a
disposable index with three primary shards and zero replicas
to make shard fan-out observable on one local node; that is a
teaching topology, not a production sizing recommendation.
OpenSearch demo -k remains local-only; production
must validate certificates. No moving latest tags
are used.
The generation environment does not run the AtlasMart containers, so latency, profile nanoseconds, slow-log lines, task IDs and PIT IDs are not fabricated. Expected outputs describe invariant fields and directions of change. Run the bounded lab locally and record your own p50/p95/p99, shard counts, profile trees and resource statistics before accepting a performance conclusion.
1. Timeout, cancellation and async are different controls
| Control | What it means | What it does not guarantee |
|---|---|---|
Search timeout |
Ask search execution to stop after a duration and return timeout metadata/available work according to semantics. | Instant preemption of every underlying operation at the exact deadline. |
| Task cancellation | Explicitly request cancellation of a cancellable server task. | Rollback of work already performed or immediate cleanup if code is between cancellation checks. |
| Client socket timeout | Client stops waiting. | Server task has stopped. |
| Async search | Server owns a long-running search with an ID, status/results retention and polling. | Unlimited concurrency or free storage for saved responses. |
The dangerous pattern is a client timeout followed by an automatic retry while the original expensive server-side operation continues. A user sees one failure; the cluster sees two costly searches. Use request IDs/tracing plus Tasks/async IDs to verify server lifecycle.
2. Elasticsearch async search
Elasticsearch exposes POST /{index}/_async_search.
The submit call can wait briefly for completion and otherwise
returns an ID that can be polled through
GET /_async_search/{id}. Results can be retained
for a bounded keep_alive; access is tied to the
submitting user/API key when security is enabled. Current
Elasticsearch also caps the stored async response size by
search.max_async_search_response_size.
POST /atlasmart-search-exec-v1/_async_search?wait_for_completion_timeout=100ms&keep_alive=5m
{
"size":0,
"query":{"match_all":{}},
"aggs":{
"by_category":{
"terms":{"field":"category","size":10},
"aggs":{"avg_price":{"avg":{"field":"price"}}}
}
}
}
GET /_async_search/ASYNC_ID
DELETE /_async_search/ASYNC_ID
The delete API releases stored state and cancels an ongoing async search as applicable. Do not keep completed responses “just in case” beyond the application retention requirement.
3. OpenSearch asynchronous search is a plugin namespace
OpenSearch uses the Asynchronous Search plugin and the endpoint
namespace /_plugins/_asynchronous_search. Its
settings include per-coordinating-node concurrency, maximum
running time, maximum keep-alive and maximum wait-for-completion
timeout. This is not wire-compatible with Elasticsearch
async-search endpoints even though the workflow—submit, poll,
optionally receive partial results, delete—is conceptually
similar.
POST /_plugins/_asynchronous_search?wait_for_completion_timeout=100ms&keep_alive=5m
{
"index":"atlasmart-search-exec-v1",
"size":0,
"query":{"match_all":{}},
"aggs":{
"by_category":{
"terms":{"field":"category","size":10},
"aggs":{"avg_price":{"avg":{"field":"price"}}}
}
}
}
GET /_plugins/_asynchronous_search/ASYNC_ID
DELETE /_plugins/_asynchronous_search/ASYNC_ID
OpenSearch asynchronous search is plugin functionality. Managed distributions can expose different plugin availability, permissions or limits. Confirm the installed plugin and API schema before production automation.
4. Partial results are a product decision, not only a flag
A response can be incomplete because shards failed, the request timed out, an async search is still running, or cross-cluster components are unavailable. The UI must decide whether partial data is acceptable. A product-search page might reject incomplete price/filter results; an exploratory dashboard might display “partial as of …” with explicit shard/completion metadata.
Elasticsearch has allow_partial_search_results for
shard failures and async responses that expose completion state.
OpenSearch async search can return partial results while
running. In both cases, application code must not silently label
partial aggregates as final business totals.
GET atlasmart-search-exec-v1/_search
{
"timeout":"1ms",
"size":0,
"query":{"match_all":{}},
"aggs":{"by_category":{"terms":{"field":"category","size":10}}}
}
# Inspect timed_out, _shards.failed, failures and aggregation/hit state.
# Do not assume this tiny fixture will always time out.
5. Cancellation and Tasks API
List active search tasks during a deliberately long bounded lab, capture the relevant task ID, then use the task cancellation API only against that disposable task. Cancellation is cooperative; verify task disappearance/completion state instead of assuming the HTTP acknowledgement means all CPU stopped at that instant.
GET /_tasks?detailed=true&actions=*search*
# Cancel only the confirmed disposable task ID:
POST /_tasks/NODE_ID:TASK_ID/_cancel
GET /_tasks?detailed=true&actions=*search*
Task IDs are ephemeral. Match action, description, user/tenant context and timestamp before cancellation. A broad task cancel can terminate unrelated searches.
6. Long-running request contract
| Contract item | Recommended question |
|---|---|
| Ownership | Which authenticated user/tenant owns the async ID and who may poll/delete it? |
| Idempotency/retries | Can the client safely resubmit, or must it recover an existing ID after a network failure? |
| Completion | What fields prove final versus partial/in-progress state? |
| Retention | How long are completed results useful, and what is the cleanup path? |
| Admission | How many concurrent expensive searches per tenant/coordinator are allowed? |
| Size | What is the maximum result payload and should the query return aggregates instead of hits? |
| Cancellation | What happens on user navigation, logout, deadline expiry or upstream cancellation? |
7. Production judgment
Async search is a request-lifecycle tool, not a performance optimization. The same expensive query still consumes CPU, heap, I/O and shard capacity. Its value is that the application can represent progress, avoid fragile long HTTP waits, control retries, and explicitly cancel/expire work.
For recurring heavy analytics, precomputation, transforms/rollups, materialized summaries or a separate analytics system may be better than repeatedly launching the same distributed search. Measure total cluster cost, not only user-perceived responsiveness.
Check your understanding
- Does a client HTTP timeout prove the server search stopped?
- What is the Elasticsearch async endpoint family?
- Why is OpenSearch async search not a drop-in endpoint replacement?
- When are partial results unsafe?
- What does async search improve?
Review the answers
1. No. The client stopped waiting; server cancellation must be verified separately.
2. POST /{index}/_async_search plus GET/DELETE /_async_search/{id}.
3. It is exposed under the _plugins/_asynchronous_search namespace with plugin-specific settings and behavior.
4. When the application would present them as complete/correct business results without shard/completion disclosure.
5. Request lifecycle, polling, cancellation and retention; it does not make the underlying query free.
Summary and next step
You can now operate long-running search intentionally. Lesson 5 combines the chapter into a tuning exercise: form a hypothesis, change query/mapping/routing/shards/aggregation/pagination, and accept the change only after correctness and tail-latency evidence.
Authoritative references
- Elastic search API — Current search request options, search type, pre-filter shard behavior and partial-results controls.
- Elastic search profiling — Profile operator trees, collectors and DFS profiling; also documents what profile does not measure.
- Elastic pagination — from/size result window, search_after, PIT consistency and PIT cleanup guidance.
- Elastic point in time API — Opening PITs, keep_alive, changing PIT IDs and retained-segment resource implications.
- Elastic async search submit — Long-running asynchronous search submission and response-size constraints.
- Elastic async search results — Polling, ownership/security and keep_alive behavior for async search.
- Elastic slow logs — Shard-level query/fetch slow logs and current query-logging guidance.
- OpenSearch pagination — from/size, search_after and PIT-backed pagination tradeoffs.
- OpenSearch PIT — Create/list/delete PIT APIs, security permissions and resource lifetime.
- OpenSearch profile API — Search component timings and explicit omissions such as network/queue/coordinator idle time.
- OpenSearch asynchronous search — Plugin endpoint, partial results and long-running search model.
- OpenSearch async settings — Maximum running time, concurrency, retention and wait timeout settings.
- OpenSearch logs — Request-level and shard-level search slow logs plus task-resource logging.
- OpenSearch tasks API — Task inspection and cancellation mechanisms.