Chapter 16 · Search Execution, Profiling, Slow Logs, Async/Search-After/PIT, and Pagination
Tune a Slow Search by Changing Query, Mapping, Routing, Shards, Aggregation Strategy, or Pagination—Then Measure
Use a hypothesis-driven AtlasMart tuning workflow that changes one causal variable at a time and validates correctness, p95/p99 latency, shard fan-out and resource cost before accepting an optimization.
Learning outcomes
AtlasMart’s search SLO is violated, but the team has six
plausible fixes: simplify the query, add a keyword field, route
by tenant, change shard count, replace a large terms tree with
composite paging, or replace deep from with PIT +
search_after. Applying all six at once might
improve latency—and destroy the ability to know why. This lesson
turns optimization into a falsifiable experiment.
Choose the correct tuning lever from observed query, fetch, shard, aggregation or pagination evidence.
Define a baseline that includes correctness, p50/p95/p99, throughput, failures, shard fan-out and resource metrics.
Change one causal variable while keeping dataset, query set, warmness and concurrency controlled.
Reject optimizations that improve averages but regress relevance, freshness, tenant isolation or tail latency.
Produce a rollback-ready tuning report that distinguishes product/version-specific behavior from portable design principles.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. AtlasMart keeps https://localhost:9200 for
Elasticsearch with CA verification and
https://localhost:9201 for the disposable
OpenSearch demo certificate using
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The
established containers are atlasmart-es and
atlasmart-os. This chapter deliberately creates a
disposable index with three primary shards and zero replicas
to make shard fan-out observable on one local node; that is a
teaching topology, not a production sizing recommendation.
OpenSearch demo -k remains local-only; production
must validate certificates. No moving latest tags
are used.
The generation environment does not run the AtlasMart containers, so latency, profile nanoseconds, slow-log lines, task IDs and PIT IDs are not fabricated. Expected outputs describe invariant fields and directions of change. Run the bounded lab locally and record your own p50/p95/p99, shard counts, profile trees and resource statistics before accepting a performance conclusion.
1. Start with a symptom and a hypothesis
| Symptom/evidence | Candidate hypothesis | High-value experiment |
|---|---|---|
| Profile shows wildcard/script dominates query phase. | Query shape is expensive. | Replace with indexed normalized field/prefix strategy; same dataset/query intent. |
| Fetch slow logs/bytes dominate. | Payload retrieval is expensive. | Return only required fields or redesign source payload; preserve semantics. |
| Search fans across many shards but requests are tenant-scoped. | Fan-out is unnecessary. | Test correct custom routing on a disposable version; measure skew too. |
| Huge terms aggregation causes memory/latency spikes. | Bucket strategy is unbounded. | Bound buckets or use composite for complete key enumeration. |
| Late pages degrade strongly and hit result window. | Deep from/size is the cause. | PIT + stable search_after under concurrent writes. |
| One shard is CPU hot while peers idle. | Routing/data skew is the cause. | Compare key distribution and rebalance routing model; do not merely add replicas. |
Changing mappings, shards, query and client concurrency together can produce a faster benchmark with no causal explanation. Keep a written hypothesis and expected evidence before changing the system.
2. Baseline contract
- Pinned Elasticsearch 9.5.3 or OpenSearch 3.8.0; record distribution, plugins and JVM from node APIs.
- Same fixture or production-like snapshot, same shard topology unless shard count is the variable.
- Same query corpus with expected hits/aggregations/relevance judgments.
- Same client concurrency, request rate, timeout/retry policy and warm-up procedure.
-
Record p50/p95/p99 and throughput—not only
took. -
Record
_shards, failures/skips, CPU, GC, thread-pool rejection, breaker/backpressure and storage state. - Run enough repetitions to see variability; report cold/warm state explicitly.
3. Lab A: pagination hypothesis
PUT atlasmart-search-exec-v1
{
"settings": {
"number_of_shards": 3,
"number_of_replicas": 0,
"index.max_result_window": 100
},
"mappings": {
"properties": {
"sku": {"type":"keyword"},
"name": {"type":"text","fields":{"raw":{"type":"keyword"}}},
"category": {"type":"keyword"},
"price": {"type":"double"},
"available": {"type":"boolean"},
"updated_at": {"type":"date"},
"popularity": {"type":"integer"}
}
}
}
# Save as chapter16_fixture.py and redirect stdout to fixture.ndjson.
import json
from datetime import datetime, timedelta, timezone
base = datetime(2026, 9, 1, tzinfo=timezone.utc)
for i in range(240):
sku = f"P-{1000+i:04d}"
meta = {"index":{"_index":"atlasmart-search-exec-v1","_id":sku}}
doc = {
"sku": sku,
"name": f"AtlasMart {'wireless' if i%3==0 else 'wired'} headset model {i:03d}",
"category": ["audio","mobile","office"][i%3],
"price": round(20 + (i%80)*1.25, 2),
"available": i%5 != 0,
"updated_at": (base + timedelta(minutes=i)).isoformat().replace('+00:00','Z'),
"popularity": (i*17)%101
}
print(json.dumps(meta,separators=(',',':')))
print(json.dumps(doc,separators=(',',':')))
# Then POST fixture.ndjson to /_bulk?refresh=true with Content-Type application/x-ndjson.
GET atlasmart-search-exec-v1/_search
{
"from":80,
"size":10,
"query":{"term":{"available":true}},
"sort":[{"updated_at":"asc"},{"sku":"asc"}],
"profile":true
}
Capture profile only for structural evidence. Then measure unprofiled repeated requests. Build the equivalent traversal with PIT + search_after. Validate the exact same SKU order for the fixed PIT and compare late-page client p95/p99. The expected mechanism is reduced deep-offset candidate work, not a guaranteed numeric speedup on a 240-document laptop fixture.
4. Lab B: aggregation and fetch hypotheses
GET atlasmart-search-exec-v1/_search
{
"size":10,
"_source":["sku","name","price"],
"query":{"match":{"name":"headset"}},
"aggs":{
"category":{"terms":{"field":"category","size":10}},
"price_stats":{"stats":{"field":"price"}}
}
}
If an application needs every distinct high-cardinality key, a
huge terms.size is not pagination. Use
composite with after_key and accept
its ordering/feature tradeoffs. If fetch dominates, trim
response fields only when the API contract allows it. Every
performance change must re-run correctness assertions.
5. Routing/shard/mapping changes require migration evidence
A routing or shard-count change is not a runtime toggle. It changes document placement and usually requires a new index plus reindex/cutover as taught in Chapters 09 and 12. A mapping change such as adding a keyword subfield can be additive for future documents in some cases, but retroactive values may require reindexing depending on the change.
Measure recovery and failure-domain consequences before accepting a topology optimization. Reducing shard fan-out can improve search while making one shard too large for recovery objectives. There is no universal “one shard per X GB” answer.
Never trade Chapter 13 recovery objectives or Chapter 07 relevance quality for a synthetic search-latency win without documenting the compromise.
6. Before/after evidence table
| Measure | Baseline | Candidate | Acceptance rule |
|---|---|---|---|
| Correct result set / aggregate | Record fixture assertions. | Must match unless semantics intentionally change. | No silent correctness regression. |
| Relevance metrics | Judged baseline where ranking changes. | Re-run NDCG/MRR/Recall@k as applicable. | Stay within declared regression budget. |
| Client p95/p99 | Measured under fixed concurrency. | Measured under identical load. | Meet SLO with confidence/repetition, not one sample. |
| Shard fan-out | _shards.total/skipped. |
Same or intentionally changed. | Explain why change is safe. |
| CPU/GC/rejections | Time-aligned node evidence. | Same observation window. | No hidden saturation transfer. |
| Freshness/consistency | Live or PIT semantics documented. | Re-test concurrent changes. | No accidental stale/missing-page behavior. |
7. Failure injection and rollback
Safe performance work includes failure behavior. During the disposable lab, cancel one long-running task, expire/delete a PIT, and issue one request that exceeds the reduced result window. Confirm the client handles each as an explicit error/state rather than infinite retry.
For production rollout, keep the previous query template/index alias/API version available. Roll back if correctness, relevance, error rate or tail latency breaches the acceptance criteria even if average latency improves.
8. Cleanup
DELETE atlasmart-search-exec-v1
# Also delete any open PIT/async IDs created during the lab using the product-specific APIs.
# Restore any temporary slow-log thresholds to -1.
Do not delete shared AtlasMart indices from prior chapters. The Chapter 16 fixture is intentionally isolated so failure experiments and changed result-window settings remain reversible.
9. Production judgment and bridge to Chapter 17
A search tuning change is complete only when it has a mechanism, repeatable evidence, correctness/relevance checks, operational side-effect analysis and rollback. “Profile got smaller” is not enough. The production decision surface includes p95/p99, throughput, shard/recovery cost, freshness, failure behavior, tenant security and version/API compatibility.
Chapter 17 moves from per-request/runtime behavior to index lifecycle. The same discipline applies: rollover, tiering, shrink/force-merge-like actions and retention must be justified by recovery, cost and compliance evidence—and Elastic ILM and OpenSearch ISM must not be treated as the same control plane.
Check your understanding
- Why should a tuning experiment change one variable?
- Why is p99 paired with throughput?
- When can routing be a good search optimization?
- Why can a mapping optimization require reindexing?
- What must accompany a production search optimization?
Review the answers
1. So observed differences can support or falsify a causal hypothesis instead of becoming an uninterpretable bundle of changes.
2. A system can increase throughput by queueing more work while making tail latency or errors unacceptable.
3. When the application has a correct stable routing key that reduces fan-out without creating dangerous skew/hotspots.
4. Existing indexed structures are not retroactively rebuilt for every mapping change; the new representation may require a new index/reindex path.
5. Correctness/relevance checks, operational impact, version-specific assumptions, measurable acceptance thresholds and rollback.
Summary
Chapter 16 made distributed search cost visible from shard selection through query/fetch, established a diagnostic toolkit, replaced deep paging with stable PIT-backed cursors, defined long-running async/cancellation semantics, and ended with a falsifiable tuning workflow. You are ready to apply the same operational discipline to lifecycle automation.
Authoritative references
- Elastic search API — Current search request options, search type, pre-filter shard behavior and partial-results controls.
- Elastic search profiling — Profile operator trees, collectors and DFS profiling; also documents what profile does not measure.
- Elastic pagination — from/size result window, search_after, PIT consistency and PIT cleanup guidance.
- Elastic point in time API — Opening PITs, keep_alive, changing PIT IDs and retained-segment resource implications.
- Elastic async search submit — Long-running asynchronous search submission and response-size constraints.
- Elastic async search results — Polling, ownership/security and keep_alive behavior for async search.
- Elastic slow logs — Shard-level query/fetch slow logs and current query-logging guidance.
- OpenSearch pagination — from/size, search_after and PIT-backed pagination tradeoffs.
- OpenSearch PIT — Create/list/delete PIT APIs, security permissions and resource lifetime.
- OpenSearch profile API — Search component timings and explicit omissions such as network/queue/coordinator idle time.
- OpenSearch asynchronous search — Plugin endpoint, partial results and long-running search model.
- OpenSearch async settings — Maximum running time, concurrency, retention and wait timeout settings.
- OpenSearch logs — Request-level and shard-level search slow logs plus task-resource logging.
- OpenSearch tasks API — Task inspection and cancellation mechanisms.