Chapter 16 · Search Execution, Profiling, Slow Logs, Async/Search-After/PIT, and Pagination
from/size Limits, search_after, Point-in-Time Readers, Stable Sort Keys, and Pagination Correctness
Replace deep from/size paging with deterministic search_after and point-in-time readers, then test pagination under concurrent writes and clean up retained search contexts.
Learning outcomes
AtlasMart adds a “page 500” button to product search. The API
uses from=4990&size=10. Shards repeatedly
collect thousands of candidates only to discard almost all of
them, and shoppers sometimes see duplicates or missing items
when products are updated between clicks. Pagination is now both
a cost problem and a correctness problem.
Explain why deep from/size multiplies per-shard candidate work and why result-window limits are safeguards.
Build deterministic search_after pagination with a total ordering and a stable tie breaker.
Use product-specific PIT APIs to freeze the searchable view across pages.
Test concurrent writes to prove PIT consistency and understand what new data is intentionally hidden.
Close PITs promptly and account for retained segments, file handles and heap/resource lifetime.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. AtlasMart keeps https://localhost:9200 for
Elasticsearch with CA verification and
https://localhost:9201 for the disposable
OpenSearch demo certificate using
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The
established containers are atlasmart-es and
atlasmart-os. This chapter deliberately creates a
disposable index with three primary shards and zero replicas
to make shard fan-out observable on one local node; that is a
teaching topology, not a production sizing recommendation.
OpenSearch demo -k remains local-only; production
must validate certificates. No moving latest tags
are used.
The generation environment does not run the AtlasMart containers, so latency, profile nanoseconds, slow-log lines, task IDs and PIT IDs are not fabricated. Expected outputs describe invariant fields and directions of change. Run the bounded lab locally and record your own p50/p95/p99, shard counts, profile trees and resource statistics before accepting a performance conclusion.
1. Why from/size gets expensive
For a sorted distributed query, each shard must produce enough
local candidates for the coordinator to construct the requested
global page. A request for from=90,size=10 can
require each shard to consider its local top 100 candidates even
though only ten hits reach the user. Increasing the offset
therefore grows heap/CPU work and coordinator merge pressure
with no user-visible benefit.
Elasticsearch documents a default result-window safeguard of
10,000 hits via index.max_result_window. OpenSearch
likewise defaults to a 10,000 result window. This lab
intentionally lowers the index setting to 100 so the failure is
easy and safe to reproduce.
GET atlasmart-search-exec-v1/_search
{
"from":95,
"size":10,
"query":{"match_all":{}},
"sort":[{"updated_at":"asc"},{"sku":"asc"}]
}
Raising index.max_result_window until page 500
works removes a safety boundary but does not remove the
per-shard deep-page work. Use cursor-style pagination or a
different export/analytics workflow.
2. search_after: a cursor over sort values
search_after asks for results strictly after the
sort values of the last hit from the previous page. Every page
must preserve the same query and sort semantics. A stable, total
ordering is essential. If several documents share the same
primary sort value and there is no deterministic tie breaker,
replicas can order equal values differently and pages can
duplicate or omit hits.
Use a unique sortable field such as sku as the
final application-level tie breaker in this lab. Do not assume
_id has doc values; if you need ID-based sorting,
map a keyword copy explicitly.
GET atlasmart-search-exec-v1/_search
{
"size":10,
"query":{"term":{"available":true}},
"sort":[{"updated_at":"asc"},{"sku":"asc"}]
}
# Copy the last hit's sort array exactly, for example:
GET atlasmart-search-exec-v1/_search
{
"size":10,
"query":{"term":{"available":true}},
"sort":[{"updated_at":"asc"},{"sku":"asc"}],
"search_after":["2026-09-01T00:12:00.000Z","P-1012"]
}
3. search_after alone is live, PIT freezes the view
Without PIT, each page is a fresh search against the then-current visible index state. A refresh between pages can insert, update or delete documents relative to the sort cursor and produce inconsistent traversal. A point in time keeps a logical view of the index stable across page requests.
The APIs are not identical. Elasticsearch opens a PIT with
POST /index/_pit?keep_alive=... and returns
id. OpenSearch opens one with
POST /index/_search/point_in_time?keep_alive=...
and returns pit_id. Search requests then use a
pit object. Elasticsearch documentation also
requires using the most recently returned PIT ID because it can
change between searches.
POST /atlasmart-search-exec-v1/_pit?keep_alive=2m
# capture response.id as PIT_ID
GET /_search
{
"size":10,
"pit":{"id":"PIT_ID","keep_alive":"2m"},
"query":{"term":{"available":true}},
"sort":[{"updated_at":"asc"},{"sku":"asc"}]
}
# Use the latest PIT id returned by the search plus the last hit's sort array.
DELETE /_pit
{"id":"LATEST_PIT_ID"}
POST /atlasmart-search-exec-v1/_search/point_in_time?keep_alive=2m
# capture response.pit_id as PIT_ID
GET /_search
{
"size":10,
"pit":{"id":"PIT_ID","keep_alive":"2m"},
"query":{"term":{"available":true}},
"sort":[{"updated_at":"asc"},{"sku":"asc"}]
}
DELETE /_search/point_in_time
{"pit_id":["PIT_ID"]}
4. Concurrent-write correctness test
- Open a PIT and fetch page 1. Store its PIT ID and the last hit sort values.
-
Index a new document whose
updated_atwould sort inside the already-open traversal; refresh it. - Fetch page 2 through the same PIT. The new document should not suddenly appear because the PIT preserves the earlier view.
- Run a new live search or open a new PIT; the new document should now be visible.
- Delete the PIT and the synthetic document during cleanup.
PUT atlasmart-search-exec-v1/_doc/P-9999?refresh=true
{
"sku":"P-9999",
"name":"AtlasMart wireless headset late insert",
"category":"audio",
"price":49.95,
"available":true,
"updated_at":"2026-09-01T00:05:30Z",
"popularity":50
}
5. PIT is consistency with a resource bill
A PIT keeps older Lucene segments available while the view is alive. That can retain disk space and file handles; updates/deletes can also increase bookkeeping. OpenSearch exposes PIT count/segment APIs and caps open contexts by setting. Elasticsearch exposes open search-context statistics and warns that PITs can retain old segments and require extra heap/file handles.
Therefore keep_alive is not “session duration.” It
only needs to be long enough to reach the next request. Refresh
it on page access and close the PIT when the user/export is
finished. Put server-side limits around abandoned cursors.
Store the exact query/filter/sort version with the cursor. If the user changes filters or sort order, start a new PIT/cursor. A cursor from one query contract is not valid for another.
6. Lab acceptance criteria
PUT atlasmart-search-exec-v1
{
"settings": {
"number_of_shards": 3,
"number_of_replicas": 0,
"index.max_result_window": 100
},
"mappings": {
"properties": {
"sku": {"type":"keyword"},
"name": {"type":"text","fields":{"raw":{"type":"keyword"}}},
"category": {"type":"keyword"},
"price": {"type":"double"},
"available": {"type":"boolean"},
"updated_at": {"type":"date"},
"popularity": {"type":"integer"}
}
}
}
# Save as chapter16_fixture.py and redirect stdout to fixture.ndjson.
import json
from datetime import datetime, timedelta, timezone
base = datetime(2026, 9, 1, tzinfo=timezone.utc)
for i in range(240):
sku = f"P-{1000+i:04d}"
meta = {"index":{"_index":"atlasmart-search-exec-v1","_id":sku}}
doc = {
"sku": sku,
"name": f"AtlasMart {'wireless' if i%3==0 else 'wired'} headset model {i:03d}",
"category": ["audio","mobile","office"][i%3],
"price": round(20 + (i%80)*1.25, 2),
"available": i%5 != 0,
"updated_at": (base + timedelta(minutes=i)).isoformat().replace('+00:00','Z'),
"popularity": (i*17)%101
}
print(json.dumps(meta,separators=(',',':')))
print(json.dumps(doc,separators=(',',':')))
# Then POST fixture.ndjson to /_bulk?refresh=true with Content-Type application/x-ndjson.
-
The deliberate
from=95,size=10request fails because the lab result window is 100. - A search_after traversal returns each fixture SKU at most once for the fixed query/sort.
- PIT traversal remains stable when P-9999 is inserted and refreshed after PIT creation.
- A newly opened PIT can see P-9999.
- PIT resources are explicitly deleted; no lab depends on waiting for expiration.
- Client-side p95/p99 for late pages is compared between from/size and search_after under the same workload and warmness.
7. Production judgment
User-facing pagination and bulk export are different workloads.
For interactive deep traversal, PIT plus stable
search_after gives consistent cursor semantics. For
full exports, consider sliced PIT/search_after, reindex/export
tooling, snapshots or an analytics store depending on the
objective. Scroll remains useful for some batch workflows but is
not the preferred real-time deep-pagination mechanism in current
Elastic guidance.
Do not keep PITs open indefinitely. Resource lifetime is part of API design: cursor TTL, maximum open cursors per tenant, cancellation on client abandonment and cleanup telemetry belong in the service contract.
Check your understanding
- Why does deep from/size cost more?
- What must search_after preserve?
- Why is a tie breaker necessary?
- What problem does PIT solve?
- What is the resource risk of long PITs?
Review the answers
1. Each shard must retain/consider candidates up to the requested offset plus page size so the coordinator can build the global page.
2. The same query/filter and sort contract, using the previous page’s final sort values.
3. Equal primary sort values otherwise have no deterministic total order across shard copies/pages.
4. It keeps the searchable dataset view stable across page requests despite concurrent refreshes/writes.
5. They can retain old segments, disk space, file handles and related search-context bookkeeping.
Summary and next step
You now have correct deep pagination under concurrent writes. Lesson 4 addresses a different long-running pattern: analytical searches that should be submitted, polled, cancelled and expired rather than tied to one synchronous HTTP request.
Authoritative references
- Elastic search API — Current search request options, search type, pre-filter shard behavior and partial-results controls.
- Elastic search profiling — Profile operator trees, collectors and DFS profiling; also documents what profile does not measure.
- Elastic pagination — from/size result window, search_after, PIT consistency and PIT cleanup guidance.
- Elastic point in time API — Opening PITs, keep_alive, changing PIT IDs and retained-segment resource implications.
- Elastic async search submit — Long-running asynchronous search submission and response-size constraints.
- Elastic async search results — Polling, ownership/security and keep_alive behavior for async search.
- Elastic slow logs — Shard-level query/fetch slow logs and current query-logging guidance.
- OpenSearch pagination — from/size, search_after and PIT-backed pagination tradeoffs.
- OpenSearch PIT — Create/list/delete PIT APIs, security permissions and resource lifetime.
- OpenSearch profile API — Search component timings and explicit omissions such as network/queue/coordinator idle time.
- OpenSearch asynchronous search — Plugin endpoint, partial results and long-running search model.
- OpenSearch async settings — Maximum running time, concurrency, retention and wait timeout settings.
- OpenSearch logs — Request-level and shard-level search slow logs plus task-resource logging.
- OpenSearch tasks API — Task inspection and cancellation mechanisms.