Chapter 16 · Search Execution, Profiling, Slow Logs, Async/Search-After/PIT, and Pagination

Tune a Slow Search by Changing Query, Mapping, Routing, Shards, Aggregation Strategy, or Pagination—Then Measure

Use a hypothesis-driven AtlasMart tuning workflow that changes one causal variable at a time and validates correctness, p95/p99 latency, shard fan-out and resource cost before accepting an optimization.

Intermediate → Advanced115–150 minutesDistributed search & pagination labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart’s search SLO is violated, but the team has six plausible fixes: simplify the query, add a keyword field, route by tenant, change shard count, replace a large terms tree with composite paging, or replace deep from with PIT + search_after. Applying all six at once might improve latency—and destroy the ability to know why. This lesson turns optimization into a falsifiable experiment.

01

Choose the correct tuning lever from observed query, fetch, shard, aggregation or pagination evidence.

02

Define a baseline that includes correctness, p50/p95/p99, throughput, failures, shard fan-out and resource metrics.

03

Change one causal variable while keeping dataset, query set, warmness and concurrency controlled.

04

Reject optimizations that improve averages but regress relevance, freshness, tenant isolation or tail latency.

05

Produce a rollback-ready tuning report that distinguishes product/version-specific behavior from portable design principles.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. AtlasMart keeps https://localhost:9200 for Elasticsearch with CA verification and https://localhost:9201 for the disposable OpenSearch demo certificate using OPENSEARCH_INITIAL_ADMIN_PASSWORD. The established containers are atlasmart-es and atlasmart-os. This chapter deliberately creates a disposable index with three primary shards and zero replicas to make shard fan-out observable on one local node; that is a teaching topology, not a production sizing recommendation. OpenSearch demo -k remains local-only; production must validate certificates. No moving latest tags are used.

Execution note

The generation environment does not run the AtlasMart containers, so latency, profile nanoseconds, slow-log lines, task IDs and PIT IDs are not fabricated. Expected outputs describe invariant fields and directions of change. Run the bounded lab locally and record your own p50/p95/p99, shard counts, profile trees and resource statistics before accepting a performance conclusion.

1. Start with a symptom and a hypothesis

Symptom/evidence Candidate hypothesis High-value experiment
Profile shows wildcard/script dominates query phase. Query shape is expensive. Replace with indexed normalized field/prefix strategy; same dataset/query intent.
Fetch slow logs/bytes dominate. Payload retrieval is expensive. Return only required fields or redesign source payload; preserve semantics.
Search fans across many shards but requests are tenant-scoped. Fan-out is unnecessary. Test correct custom routing on a disposable version; measure skew too.
Huge terms aggregation causes memory/latency spikes. Bucket strategy is unbounded. Bound buckets or use composite for complete key enumeration.
Late pages degrade strongly and hit result window. Deep from/size is the cause. PIT + stable search_after under concurrent writes.
One shard is CPU hot while peers idle. Routing/data skew is the cause. Compare key distribution and rebalance routing model; do not merely add replicas.
One change at a time

Changing mappings, shards, query and client concurrency together can produce a faster benchmark with no causal explanation. Keep a written hypothesis and expected evidence before changing the system.

2. Baseline contract

  • Pinned Elasticsearch 9.5.3 or OpenSearch 3.8.0; record distribution, plugins and JVM from node APIs.
  • Same fixture or production-like snapshot, same shard topology unless shard count is the variable.
  • Same query corpus with expected hits/aggregations/relevance judgments.
  • Same client concurrency, request rate, timeout/retry policy and warm-up procedure.
  • Record p50/p95/p99 and throughput—not only took.
  • Record _shards, failures/skips, CPU, GC, thread-pool rejection, breaker/backpressure and storage state.
  • Run enough repetitions to see variability; report cold/warm state explicitly.

3. Lab A: pagination hypothesis

Disposable Chapter 16 index — run separately against each product
PUT atlasmart-search-exec-v1
{
  "settings": {
    "number_of_shards": 3,
    "number_of_replicas": 0,
    "index.max_result_window": 100
  },
  "mappings": {
    "properties": {
      "sku":        {"type":"keyword"},
      "name":       {"type":"text","fields":{"raw":{"type":"keyword"}}},
      "category":   {"type":"keyword"},
      "price":      {"type":"double"},
      "available":  {"type":"boolean"},
      "updated_at": {"type":"date"},
      "popularity": {"type":"integer"}
    }
  }
}
Deterministic fixture generator (Python 3, writes NDJSON)
# Save as chapter16_fixture.py and redirect stdout to fixture.ndjson.
import json
from datetime import datetime, timedelta, timezone
base = datetime(2026, 9, 1, tzinfo=timezone.utc)
for i in range(240):
    sku = f"P-{1000+i:04d}"
    meta = {"index":{"_index":"atlasmart-search-exec-v1","_id":sku}}
    doc = {
      "sku": sku,
      "name": f"AtlasMart {'wireless' if i%3==0 else 'wired'} headset model {i:03d}",
      "category": ["audio","mobile","office"][i%3],
      "price": round(20 + (i%80)*1.25, 2),
      "available": i%5 != 0,
      "updated_at": (base + timedelta(minutes=i)).isoformat().replace('+00:00','Z'),
      "popularity": (i*17)%101
    }
    print(json.dumps(meta,separators=(',',':')))
    print(json.dumps(doc,separators=(',',':')))
# Then POST fixture.ndjson to /_bulk?refresh=true with Content-Type application/x-ndjson.
Baseline late page
GET atlasmart-search-exec-v1/_search
{
  "from":80,
  "size":10,
  "query":{"term":{"available":true}},
  "sort":[{"updated_at":"asc"},{"sku":"asc"}],
  "profile":true
}

Capture profile only for structural evidence. Then measure unprofiled repeated requests. Build the equivalent traversal with PIT + search_after. Validate the exact same SKU order for the fixed PIT and compare late-page client p95/p99. The expected mechanism is reduced deep-offset candidate work, not a guaranteed numeric speedup on a 240-document laptop fixture.

4. Lab B: aggregation and fetch hypotheses

Bounded facet query
GET atlasmart-search-exec-v1/_search
{
  "size":10,
  "_source":["sku","name","price"],
  "query":{"match":{"name":"headset"}},
  "aggs":{
    "category":{"terms":{"field":"category","size":10}},
    "price_stats":{"stats":{"field":"price"}}
  }
}

If an application needs every distinct high-cardinality key, a huge terms.size is not pagination. Use composite with after_key and accept its ordering/feature tradeoffs. If fetch dominates, trim response fields only when the API contract allows it. Every performance change must re-run correctness assertions.

5. Routing/shard/mapping changes require migration evidence

A routing or shard-count change is not a runtime toggle. It changes document placement and usually requires a new index plus reindex/cutover as taught in Chapters 09 and 12. A mapping change such as adding a keyword subfield can be additive for future documents in some cases, but retroactive values may require reindexing depending on the change.

Measure recovery and failure-domain consequences before accepting a topology optimization. Reducing shard fan-out can improve search while making one shard too large for recovery objectives. There is no universal “one shard per X GB” answer.

Cross-chapter guardrail

Never trade Chapter 13 recovery objectives or Chapter 07 relevance quality for a synthetic search-latency win without documenting the compromise.

6. Before/after evidence table

Measure Baseline Candidate Acceptance rule
Correct result set / aggregate Record fixture assertions. Must match unless semantics intentionally change. No silent correctness regression.
Relevance metrics Judged baseline where ranking changes. Re-run NDCG/MRR/Recall@k as applicable. Stay within declared regression budget.
Client p95/p99 Measured under fixed concurrency. Measured under identical load. Meet SLO with confidence/repetition, not one sample.
Shard fan-out _shards.total/skipped. Same or intentionally changed. Explain why change is safe.
CPU/GC/rejections Time-aligned node evidence. Same observation window. No hidden saturation transfer.
Freshness/consistency Live or PIT semantics documented. Re-test concurrent changes. No accidental stale/missing-page behavior.

7. Failure injection and rollback

Safe performance work includes failure behavior. During the disposable lab, cancel one long-running task, expire/delete a PIT, and issue one request that exceeds the reduced result window. Confirm the client handles each as an explicit error/state rather than infinite retry.

For production rollout, keep the previous query template/index alias/API version available. Roll back if correctness, relevance, error rate or tail latency breaches the acceptance criteria even if average latency improves.

8. Cleanup

Remove disposable Chapter 16 resources
DELETE atlasmart-search-exec-v1
# Also delete any open PIT/async IDs created during the lab using the product-specific APIs.
# Restore any temporary slow-log thresholds to -1.

Do not delete shared AtlasMart indices from prior chapters. The Chapter 16 fixture is intentionally isolated so failure experiments and changed result-window settings remain reversible.

9. Production judgment and bridge to Chapter 17

A search tuning change is complete only when it has a mechanism, repeatable evidence, correctness/relevance checks, operational side-effect analysis and rollback. “Profile got smaller” is not enough. The production decision surface includes p95/p99, throughput, shard/recovery cost, freshness, failure behavior, tenant security and version/API compatibility.

Chapter 17 moves from per-request/runtime behavior to index lifecycle. The same discipline applies: rollover, tiering, shrink/force-merge-like actions and retention must be justified by recovery, cost and compliance evidence—and Elastic ILM and OpenSearch ISM must not be treated as the same control plane.

Check your understanding

  1. Why should a tuning experiment change one variable?
  2. Why is p99 paired with throughput?
  3. When can routing be a good search optimization?
  4. Why can a mapping optimization require reindexing?
  5. What must accompany a production search optimization?
Review the answers

1. So observed differences can support or falsify a causal hypothesis instead of becoming an uninterpretable bundle of changes.

2. A system can increase throughput by queueing more work while making tail latency or errors unacceptable.

3. When the application has a correct stable routing key that reduces fan-out without creating dangerous skew/hotspots.

4. Existing indexed structures are not retroactively rebuilt for every mapping change; the new representation may require a new index/reindex path.

5. Correctness/relevance checks, operational impact, version-specific assumptions, measurable acceptance thresholds and rollback.

Summary

Chapter 16 made distributed search cost visible from shard selection through query/fetch, established a diagnostic toolkit, replaced deep paging with stable PIT-backed cursors, defined long-running async/cancellation semantics, and ended with a falsifiable tuning workflow. You are ready to apply the same operational discipline to lifecycle automation.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.