Locate a safe operating region with mixed load, tails, quality, recovery, and comparable cost evidence.
Run Production-Like Benchmarks with Tail Latency, Relevance, Recovery, Indexing Lag, Saturation, and Cost per Workload
Build capacity and performance plans from measured workload dimensions and tail behavior rather than generic shard, heap, bulk, or hardware rules.
Learning outcomes
Build a reproducible mixed-workload benchmark with declared dataset, warmup, offered load, concurrency, and quality checks.
Locate the saturation knee from achieved throughput, p95/p99, queueing, rejection, lag, and resource evidence.
Measure retrieval quality and freshness alongside latency so performance changes cannot silently degrade correctness.
Run a safe recovery test and separate normal-state capacity from failure-state headroom.
Convert benchmark evidence into a capacity/cost recommendation and explicit triggers for rebenchmarking.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: a benchmark that cannot answer “safe at peak?”
The final performance test must answer more than “what is the maximum QPS?” AtlasMart needs a safe operating region: load levels where search tails, indexing lag, relevance, error/rejection rate, resource pressure, and recovery objectives remain within their contracts. The saturation knee is the region where additional offered load produces diminishing throughput while queueing, tails, lag, or rejections rise disproportionately.
2. Benchmark protocol: compare controlled states
| Phase | Purpose | Evidence |
|---|---|---|
| prepare | create identical mapping/data and verify counts/judgments | versions, topology, docs/bytes, settings, checksums |
| warmup | reach declared cache/JIT/connection state | warmup duration, cache state, no scored metrics yet |
| load A | expected steady/peak-like load | p50/p95/p99, throughput, lag, quality, resources |
| load B | higher offered load to expose headroom/knee | same metrics + queues/rejections |
| one change | topology/tuning hypothesis | exact setting/code diff |
| repeat A/B | causal comparison | same driver/data/window and confidence/repeat count |
| recovery | bounded spike or optional node loss in resilient lab | time to SLO recovery and shard/recovery evidence |
| cleanup | return lab to baseline | settings restored, test indices removed |
3. Driver hygiene prevents benchmark-driver bottlenecks
Run the driver separately from the server when practical. Record its CPU/network use so client saturation is not misdiagnosed as cluster saturation. Use a time-based steady window, bounded ramp-up, persistent connections, representative request bodies, and realistic think-time/arrival patterns. Rally and OpenSearch Benchmark can automate this, but the mandatory lab can use a transparent Python driver so no paid service or specialized benchmark package is required.
from math import ceil
def percentile_ms(samples, p):
xs = sorted(samples)
if not xs:
return None
i = max(0, min(len(xs)-1, ceil((p/100) * len(xs)) - 1))
return xs[i]
# record each completed request in milliseconds
# never invent missing values; report None if the run did not produce evidence
result = {
"p50": percentile_ms(latencies_ms, 50),
"p95": percentile_ms(latencies_ms, 95),
"p99": percentile_ms(latencies_ms, 99),
"errors": error_count,
"completed": len(latencies_ms)
}
For serious capacity decisions, repeat runs and use a tool/harness that models open/closed-loop arrival behavior correctly. The small helper teaches evidence handling; it is not a substitute for a validated benchmark framework.
4. Mixed workload: write, lexical, facet, and vector together
Run the classes concurrently in the proportions declared in Lesson 1. A cluster can pass each workload alone and fail when merges, vector search, aggregations, and user search compete simultaneously. Keep query IDs and relevance judgments stable so a tuning change cannot “improve” latency by returning fewer or worse results.
duration_s: 300
warmup_s: 60
load_A:
offered_search_qps: <target-A>
offered_write_ops_s: <target-A>
clients: <A>
load_B:
offered_search_qps: <target-B>
offered_write_ops_s: <target-B>
clients: <B>
mix:
lexical_filter: 0.55
facet_aggregation: 0.20
vector_knn: 0.10
writes_updates: 0.15
quality:
golden_queries: atlasmart-judgments-v1
require_ndcg_at_10: <floor>
require_vector_recall_at_10: <floor>
freshness:
require_write_visibility_p99_ms: <target>
The percentages are merely a reproducible teaching fixture; replace them with production telemetry for real planning.
5. Saturation is a multi-signal diagnosis
| Signal | Healthy interpretation | Knee/overload evidence |
|---|---|---|
| achieved throughput | tracks offered load | plateaus or falls while offered load rises |
| p95/p99 | within SLO and stable | sharp nonlinear rise |
| queues/rejections | bounded/rare per budget | sustained growth or 429/rejections |
| indexing lag | within freshness target | backlog grows after load stops |
| CPU/heap/GC | stable with recovery margin | sustained saturation, long GC, breaker pressure |
| disk/merge | steady-state debt bounded | merge/recovery cannot catch up |
| quality | NDCG/Recall floors hold | candidate/timeout shortcuts reduce quality |
Do not define the knee from CPU alone. Sometimes the first limiting resource is storage, memory/GC, network, shard coordination, an expensive query class, or the client itself.
6. Recovery test: prove the system returns to its SLO
The mandatory single-workstation path uses a safe finite load spike: raise offered load to B for a bounded window, return to A, and measure how long p99, indexing lag, queues, merge activity, and rejection rate take to return to the established healthy band. That tests backlog recovery without pretending a one-node lab has high availability.
For an optional multi-node disposable lab, stop exactly one eligible data node only when the cluster has redundant shard copies and a replay/snapshot path. Keep client load at the declared failure-load target, record shard recovery and user p99, then restore the node. Stop immediately if the failure threatens data or host stability.
7. One justified change, then repeat
Select the change from evidence. Examples: a refresh interval that still meets freshness; a different bulk size/concurrency; a more efficient query/aggregation; a revised shard layout; more CPU; faster storage; additional data nodes; or a tier/workload separation. Change one major variable at a time. If you change hardware, shard count, query shape, and refresh together, you cannot attribute the result.
8. Cost per workload, not “cheapest node”
Normalize cost over the same measurement window. For self-managed systems include host/storage/network/backup/operations assumptions; for managed services include instance/storage/transfer and feature/subscription constraints relevant to the chosen deployment. Useful derived units include cost per 1,000 searches at the target SLO, cost per million indexed documents, or monthly cost for the declared retention and peak profile. Do not compare costs for configurations with different durability or availability.
{
"run_id": "atlasmart-perf-YYYYMMDD-NN",
"platform": {"product": null, "version": null, "topology": null},
"dataset": {"documents": null, "source_bytes": null, "indexed_bytes": null, "warm_state": null},
"offered_load": {"search_qps": null, "write_ops_s": null, "clients": null},
"results": {
"achieved_search_qps": null,
"achieved_write_ops_s": null,
"latency_ms": {"p50": null, "p95": null, "p99": null},
"indexing_lag_ms": {"p95": null, "p99": null},
"errors": null,
"rejections": null,
"relevance": {"ndcg_at_10": null, "vector_recall_at_10": null},
"recovery_seconds": null
},
"resources": {"cpu": null, "heap": null, "gc": null, "disk_io": null, "network": null},
"cost": {"currency": null, "window_cost": null, "cost_per_1k_queries": null},
"notes": {"warmup": null, "cache_state": null, "one_change_from_baseline": null}
}
9. Acceptance gate
decision: <accept / reject / retest>
platform: <Elasticsearch 9.5.3 | OpenSearch 3.8.0>
configuration_id: <immutable-id>
safe_capacity:
search_qps: <measured>
write_ops_s: <measured>
concurrent_clients: <measured>
headroom:
normal_peak_percent: <measured>
failure_state_percent: <measured>
quality_and_correctness:
ndcg_at_10: <measured>
vector_recall_at_10: <measured>
indexing_visibility_p99_ms: <measured>
recovery:
bounded_spike_recovery_s: <measured>
node_failure_recovery_s: <measured-or-not-run>
cost:
measurement_window_cost: <measured/estimated-with-source>
rollback:
change_to_revert: <exact-setting/topology>
rebenchmark_triggers:
- dataset growth threshold
- query/write mix change
- mapping/analyzer/vector model change
- shard/replica/topology/storage change
- server/client/plugin upgrade
- SLO/RPO/RTO change
10. Bridge to Chapter 31: capacity becomes a capstone requirement
Chapter 31 begins with Define Search/Analytics Requirements, Relevance Judgments, Freshness, Retention, Security, SLOs, RPO/RTO, and Cost Constraints. Carry the Chapter 30 benchmark contract, safe-capacity envelope, recovery evidence, and cost assumptions into the capstone. The capstone should not choose Elasticsearch or OpenSearch from a feature checklist; it should use measured relevance, operability, security, recovery, performance, licensing, and migration evidence.
Check your understanding
- What defines the saturation knee?
- Why is a load spike recovery test useful in the one-node mandatory lab?
- Why run mixed workloads concurrently?
- When is a tuning change accepted?
- What evidence moves into Chapter 31?
Review the answers
1. The region where additional offered load yields diminishing achieved throughput while tail latency, lag, queueing, rejections, or resource pressure rise disproportionately.
2. It measures backlog and SLO recovery without falsely claiming node-failure high availability.
3. Indexing, merges, aggregations, vector search, and interactive search compete for shared CPU, memory, storage, network, and queues.
4. Only when repeated comparable runs improve the target while meeting freshness, quality, correctness, durability/availability, recovery, and cost constraints.
5. The workload contract, safe operating envelope, tail/relevance/freshness measurements, failure/recovery evidence, topology assumptions, cost model, and rebenchmark triggers.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Elasticsearch performance optimizations
- Elasticsearch tune for indexing speed
- Elasticsearch tune for search speed
- Elasticsearch size your shards
- Elasticsearch JVM settings
- Elasticsearch resilience guidance
- Elasticsearch bulk API
- Elasticsearch refresh parameter
- Rally documentation
- OpenSearch 3.8 version history
- OpenSearch indexing performance tuning
- OpenSearch Bulk API
- OpenSearch Refresh Index API
- OpenSearch search shard routing
- OpenSearch index settings
- OpenSearch shard indexing backpressure
- OpenSearch Benchmark quickstart
- OpenSearch Performance Analyzer metrics
- OpenSearch vector search performance tuning