Chapter 16 · Search Execution, Profiling, Slow Logs, Async/Search-After/PIT, and Pagination

from/size Limits, search_after, Point-in-Time Readers, Stable Sort Keys, and Pagination Correctness

Replace deep from/size paging with deterministic search_after and point-in-time readers, then test pagination under concurrent writes and clean up retained search contexts.

Intermediate → Advanced115–150 minutesDistributed search & pagination labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart adds a “page 500” button to product search. The API uses from=4990&size=10. Shards repeatedly collect thousands of candidates only to discard almost all of them, and shoppers sometimes see duplicates or missing items when products are updated between clicks. Pagination is now both a cost problem and a correctness problem.

01

Explain why deep from/size multiplies per-shard candidate work and why result-window limits are safeguards.

02

Build deterministic search_after pagination with a total ordering and a stable tie breaker.

03

Use product-specific PIT APIs to freeze the searchable view across pages.

04

Test concurrent writes to prove PIT consistency and understand what new data is intentionally hidden.

05

Close PITs promptly and account for retained segments, file handles and heap/resource lifetime.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. AtlasMart keeps https://localhost:9200 for Elasticsearch with CA verification and https://localhost:9201 for the disposable OpenSearch demo certificate using OPENSEARCH_INITIAL_ADMIN_PASSWORD. The established containers are atlasmart-es and atlasmart-os. This chapter deliberately creates a disposable index with three primary shards and zero replicas to make shard fan-out observable on one local node; that is a teaching topology, not a production sizing recommendation. OpenSearch demo -k remains local-only; production must validate certificates. No moving latest tags are used.

Execution note

The generation environment does not run the AtlasMart containers, so latency, profile nanoseconds, slow-log lines, task IDs and PIT IDs are not fabricated. Expected outputs describe invariant fields and directions of change. Run the bounded lab locally and record your own p50/p95/p99, shard counts, profile trees and resource statistics before accepting a performance conclusion.

1. Why from/size gets expensive

For a sorted distributed query, each shard must produce enough local candidates for the coordinator to construct the requested global page. A request for from=90,size=10 can require each shard to consider its local top 100 candidates even though only ten hits reach the user. Increasing the offset therefore grows heap/CPU work and coordinator merge pressure with no user-visible benefit.

Elasticsearch documents a default result-window safeguard of 10,000 hits via index.max_result_window. OpenSearch likewise defaults to a 10,000 result window. This lab intentionally lowers the index setting to 100 so the failure is easy and safe to reproduce.

Safe deep-window failure
GET atlasmart-search-exec-v1/_search
{
  "from":95,
  "size":10,
  "query":{"match_all":{}},
  "sort":[{"updated_at":"asc"},{"sku":"asc"}]
}
Wrong repair

Raising index.max_result_window until page 500 works removes a safety boundary but does not remove the per-shard deep-page work. Use cursor-style pagination or a different export/analytics workflow.

2. search_after: a cursor over sort values

search_after asks for results strictly after the sort values of the last hit from the previous page. Every page must preserve the same query and sort semantics. A stable, total ordering is essential. If several documents share the same primary sort value and there is no deterministic tie breaker, replicas can order equal values differently and pages can duplicate or omit hits.

Use a unique sortable field such as sku as the final application-level tie breaker in this lab. Do not assume _id has doc values; if you need ID-based sorting, map a keyword copy explicitly.

Page 1 and page 2 with search_after
GET atlasmart-search-exec-v1/_search
{
  "size":10,
  "query":{"term":{"available":true}},
  "sort":[{"updated_at":"asc"},{"sku":"asc"}]
}

# Copy the last hit's sort array exactly, for example:
GET atlasmart-search-exec-v1/_search
{
  "size":10,
  "query":{"term":{"available":true}},
  "sort":[{"updated_at":"asc"},{"sku":"asc"}],
  "search_after":["2026-09-01T00:12:00.000Z","P-1012"]
}

3. search_after alone is live, PIT freezes the view

Without PIT, each page is a fresh search against the then-current visible index state. A refresh between pages can insert, update or delete documents relative to the sort cursor and produce inconsistent traversal. A point in time keeps a logical view of the index stable across page requests.

The APIs are not identical. Elasticsearch opens a PIT with POST /index/_pit?keep_alive=... and returns id. OpenSearch opens one with POST /index/_search/point_in_time?keep_alive=... and returns pit_id. Search requests then use a pit object. Elasticsearch documentation also requires using the most recently returned PIT ID because it can change between searches.

Elasticsearch 9.5.3 PIT flow
POST /atlasmart-search-exec-v1/_pit?keep_alive=2m
# capture response.id as PIT_ID

GET /_search
{
  "size":10,
  "pit":{"id":"PIT_ID","keep_alive":"2m"},
  "query":{"term":{"available":true}},
  "sort":[{"updated_at":"asc"},{"sku":"asc"}]
}

# Use the latest PIT id returned by the search plus the last hit's sort array.
DELETE /_pit
{"id":"LATEST_PIT_ID"}
OpenSearch 3.8.0 PIT flow
POST /atlasmart-search-exec-v1/_search/point_in_time?keep_alive=2m
# capture response.pit_id as PIT_ID

GET /_search
{
  "size":10,
  "pit":{"id":"PIT_ID","keep_alive":"2m"},
  "query":{"term":{"available":true}},
  "sort":[{"updated_at":"asc"},{"sku":"asc"}]
}

DELETE /_search/point_in_time
{"pit_id":["PIT_ID"]}

4. Concurrent-write correctness test

  1. Open a PIT and fetch page 1. Store its PIT ID and the last hit sort values.
  2. Index a new document whose updated_at would sort inside the already-open traversal; refresh it.
  3. Fetch page 2 through the same PIT. The new document should not suddenly appear because the PIT preserves the earlier view.
  4. Run a new live search or open a new PIT; the new document should now be visible.
  5. Delete the PIT and the synthetic document during cleanup.
Concurrent document
PUT atlasmart-search-exec-v1/_doc/P-9999?refresh=true
{
  "sku":"P-9999",
  "name":"AtlasMart wireless headset late insert",
  "category":"audio",
  "price":49.95,
  "available":true,
  "updated_at":"2026-09-01T00:05:30Z",
  "popularity":50
}

5. PIT is consistency with a resource bill

A PIT keeps older Lucene segments available while the view is alive. That can retain disk space and file handles; updates/deletes can also increase bookkeeping. OpenSearch exposes PIT count/segment APIs and caps open contexts by setting. Elasticsearch exposes open search-context statistics and warns that PITs can retain old segments and require extra heap/file handles.

Therefore keep_alive is not “session duration.” It only needs to be long enough to reach the next request. Refresh it on page access and close the PIT when the user/export is finished. Put server-side limits around abandoned cursors.

Pagination contract

Store the exact query/filter/sort version with the cursor. If the user changes filters or sort order, start a new PIT/cursor. A cursor from one query contract is not valid for another.

6. Lab acceptance criteria

Disposable Chapter 16 index — run separately against each product
PUT atlasmart-search-exec-v1
{
  "settings": {
    "number_of_shards": 3,
    "number_of_replicas": 0,
    "index.max_result_window": 100
  },
  "mappings": {
    "properties": {
      "sku":        {"type":"keyword"},
      "name":       {"type":"text","fields":{"raw":{"type":"keyword"}}},
      "category":   {"type":"keyword"},
      "price":      {"type":"double"},
      "available":  {"type":"boolean"},
      "updated_at": {"type":"date"},
      "popularity": {"type":"integer"}
    }
  }
}
Deterministic fixture generator (Python 3, writes NDJSON)
# Save as chapter16_fixture.py and redirect stdout to fixture.ndjson.
import json
from datetime import datetime, timedelta, timezone
base = datetime(2026, 9, 1, tzinfo=timezone.utc)
for i in range(240):
    sku = f"P-{1000+i:04d}"
    meta = {"index":{"_index":"atlasmart-search-exec-v1","_id":sku}}
    doc = {
      "sku": sku,
      "name": f"AtlasMart {'wireless' if i%3==0 else 'wired'} headset model {i:03d}",
      "category": ["audio","mobile","office"][i%3],
      "price": round(20 + (i%80)*1.25, 2),
      "available": i%5 != 0,
      "updated_at": (base + timedelta(minutes=i)).isoformat().replace('+00:00','Z'),
      "popularity": (i*17)%101
    }
    print(json.dumps(meta,separators=(',',':')))
    print(json.dumps(doc,separators=(',',':')))
# Then POST fixture.ndjson to /_bulk?refresh=true with Content-Type application/x-ndjson.
  • The deliberate from=95,size=10 request fails because the lab result window is 100.
  • A search_after traversal returns each fixture SKU at most once for the fixed query/sort.
  • PIT traversal remains stable when P-9999 is inserted and refreshed after PIT creation.
  • A newly opened PIT can see P-9999.
  • PIT resources are explicitly deleted; no lab depends on waiting for expiration.
  • Client-side p95/p99 for late pages is compared between from/size and search_after under the same workload and warmness.

7. Production judgment

User-facing pagination and bulk export are different workloads. For interactive deep traversal, PIT plus stable search_after gives consistent cursor semantics. For full exports, consider sliced PIT/search_after, reindex/export tooling, snapshots or an analytics store depending on the objective. Scroll remains useful for some batch workflows but is not the preferred real-time deep-pagination mechanism in current Elastic guidance.

Do not keep PITs open indefinitely. Resource lifetime is part of API design: cursor TTL, maximum open cursors per tenant, cancellation on client abandonment and cleanup telemetry belong in the service contract.

Check your understanding

  1. Why does deep from/size cost more?
  2. What must search_after preserve?
  3. Why is a tie breaker necessary?
  4. What problem does PIT solve?
  5. What is the resource risk of long PITs?
Review the answers

1. Each shard must retain/consider candidates up to the requested offset plus page size so the coordinator can build the global page.

2. The same query/filter and sort contract, using the previous page’s final sort values.

3. Equal primary sort values otherwise have no deterministic total order across shard copies/pages.

4. It keeps the searchable dataset view stable across page requests despite concurrent refreshes/writes.

5. They can retain old segments, disk space, file handles and related search-context bookkeeping.

Summary and next step

You now have correct deep pagination under concurrent writes. Lesson 4 addresses a different long-running pattern: analytical searches that should be submitted, polled, cancelled and expired rather than tied to one synchronous HTTP request.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.