Locate a safe operating region with mixed load, tails, quality, recovery, and comparable cost evidence.

Run Production-Like Benchmarks with Tail Latency, Relevance, Recovery, Indexing Lag, Saturation, and Cost per Workload

Build capacity and performance plans from measured workload dimensions and tail behavior rather than generic shard, heap, bulk, or hardware rules.

Intermediate → Advanced190–250 minutesMixed workload saturation/recovery benchmark · Chapter 30 · Lesson 05Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · free/local benchmark pathLast reviewed: September 2026

Learning outcomes

01

Build a reproducible mixed-workload benchmark with declared dataset, warmup, offered load, concurrency, and quality checks.

02

Locate the saturation knee from achieved throughput, p95/p99, queueing, rejection, lag, and resource evidence.

03

Measure retrieval quality and freshness alongside latency so performance changes cannot silently degrade correctness.

04

Run a safe recovery test and separate normal-state capacity from failure-state headroom.

05

Convert benchmark evidence into a capacity/cost recommendation and explicit triggers for rebenchmarking.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned performance baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The established local endpoints remain Elasticsearch at https://localhost:9200 and OpenSearch at https://localhost:9201 on the atlasmart-search Docker network. The generation environment did not execute live clusters, so this chapter never invents throughput, p95/p99, GC, disk, vector-recall, or cost results: numeric fields shown in report templates are deliberately blank/null until measured.

1. AtlasMart problem: a benchmark that cannot answer “safe at peak?”

The final performance test must answer more than “what is the maximum QPS?” AtlasMart needs a safe operating region: load levels where search tails, indexing lag, relevance, error/rejection rate, resource pressure, and recovery objectives remain within their contracts. The saturation knee is the region where additional offered load produces diminishing throughput while queueing, tails, lag, or rejections rise disproportionately.

2. Benchmark protocol: compare controlled states

Phase Purpose Evidence
prepare create identical mapping/data and verify counts/judgments versions, topology, docs/bytes, settings, checksums
warmup reach declared cache/JIT/connection state warmup duration, cache state, no scored metrics yet
load A expected steady/peak-like load p50/p95/p99, throughput, lag, quality, resources
load B higher offered load to expose headroom/knee same metrics + queues/rejections
one change topology/tuning hypothesis exact setting/code diff
repeat A/B causal comparison same driver/data/window and confidence/repeat count
recovery bounded spike or optional node loss in resilient lab time to SLO recovery and shard/recovery evidence
cleanup return lab to baseline settings restored, test indices removed

3. Driver hygiene prevents benchmark-driver bottlenecks

Run the driver separately from the server when practical. Record its CPU/network use so client saturation is not misdiagnosed as cluster saturation. Use a time-based steady window, bounded ramp-up, persistent connections, representative request bodies, and realistic think-time/arrival patterns. Rally and OpenSearch Benchmark can automate this, but the mandatory lab can use a transparent Python driver so no paid service or specialized benchmark package is required.

Minimal percentile helper for a transparent local driver
from math import ceil

def percentile_ms(samples, p):
    xs = sorted(samples)
    if not xs:
        return None
    i = max(0, min(len(xs)-1, ceil((p/100) * len(xs)) - 1))
    return xs[i]

# record each completed request in milliseconds
# never invent missing values; report None if the run did not produce evidence
result = {
    "p50": percentile_ms(latencies_ms, 50),
    "p95": percentile_ms(latencies_ms, 95),
    "p99": percentile_ms(latencies_ms, 99),
    "errors": error_count,
    "completed": len(latencies_ms)
}

For serious capacity decisions, repeat runs and use a tool/harness that models open/closed-loop arrival behavior correctly. The small helper teaches evidence handling; it is not a substitute for a validated benchmark framework.

4. Mixed workload: write, lexical, facet, and vector together

Run the classes concurrently in the proportions declared in Lesson 1. A cluster can pass each workload alone and fail when merges, vector search, aggregations, and user search compete simultaneously. Keep query IDs and relevance judgments stable so a tuning change cannot “improve” latency by returning fewer or worse results.

Illustrative fixed workload mix
duration_s: 300
warmup_s: 60
load_A:
  offered_search_qps: <target-A>
  offered_write_ops_s: <target-A>
  clients: <A>
load_B:
  offered_search_qps: <target-B>
  offered_write_ops_s: <target-B>
  clients: <B>
mix:
  lexical_filter: 0.55
  facet_aggregation: 0.20
  vector_knn: 0.10
  writes_updates: 0.15
quality:
  golden_queries: atlasmart-judgments-v1
  require_ndcg_at_10: <floor>
  require_vector_recall_at_10: <floor>
freshness:
  require_write_visibility_p99_ms: <target>

The percentages are merely a reproducible teaching fixture; replace them with production telemetry for real planning.

5. Saturation is a multi-signal diagnosis

Signal Healthy interpretation Knee/overload evidence
achieved throughput tracks offered load plateaus or falls while offered load rises
p95/p99 within SLO and stable sharp nonlinear rise
queues/rejections bounded/rare per budget sustained growth or 429/rejections
indexing lag within freshness target backlog grows after load stops
CPU/heap/GC stable with recovery margin sustained saturation, long GC, breaker pressure
disk/merge steady-state debt bounded merge/recovery cannot catch up
quality NDCG/Recall floors hold candidate/timeout shortcuts reduce quality

Do not define the knee from CPU alone. Sometimes the first limiting resource is storage, memory/GC, network, shard coordination, an expensive query class, or the client itself.

6. Recovery test: prove the system returns to its SLO

The mandatory single-workstation path uses a safe finite load spike: raise offered load to B for a bounded window, return to A, and measure how long p99, indexing lag, queues, merge activity, and rejection rate take to return to the established healthy band. That tests backlog recovery without pretending a one-node lab has high availability.

For an optional multi-node disposable lab, stop exactly one eligible data node only when the cluster has redundant shard copies and a replay/snapshot path. Keep client load at the declared failure-load target, record shard recovery and user p99, then restore the node. Stop immediately if the failure threatens data or host stability.

Wrong approach: run until the cluster crashes to discover “maximum throughput.” Repair: define stop conditions before the run—swap/OOM risk, sustained rejection, disk watermark risk, correctness loss, unrelated SLO breach—and stop at the safe saturation boundary.

7. One justified change, then repeat

Select the change from evidence. Examples: a refresh interval that still meets freshness; a different bulk size/concurrency; a more efficient query/aggregation; a revised shard layout; more CPU; faster storage; additional data nodes; or a tier/workload separation. Change one major variable at a time. If you change hardware, shard count, query shape, and refresh together, you cannot attribute the result.

8. Cost per workload, not “cheapest node”

Normalize cost over the same measurement window. For self-managed systems include host/storage/network/backup/operations assumptions; for managed services include instance/storage/transfer and feature/subscription constraints relevant to the chosen deployment. Useful derived units include cost per 1,000 searches at the target SLO, cost per million indexed documents, or monthly cost for the declared retention and peak profile. Do not compare costs for configurations with different durability or availability.

Benchmark report: fill only measured values
{
  "run_id": "atlasmart-perf-YYYYMMDD-NN",
  "platform": {"product": null, "version": null, "topology": null},
  "dataset": {"documents": null, "source_bytes": null, "indexed_bytes": null, "warm_state": null},
  "offered_load": {"search_qps": null, "write_ops_s": null, "clients": null},
  "results": {
    "achieved_search_qps": null,
    "achieved_write_ops_s": null,
    "latency_ms": {"p50": null, "p95": null, "p99": null},
    "indexing_lag_ms": {"p95": null, "p99": null},
    "errors": null,
    "rejections": null,
    "relevance": {"ndcg_at_10": null, "vector_recall_at_10": null},
    "recovery_seconds": null
  },
  "resources": {"cpu": null, "heap": null, "gc": null, "disk_io": null, "network": null},
  "cost": {"currency": null, "window_cost": null, "cost_per_1k_queries": null},
  "notes": {"warmup": null, "cache_state": null, "one_change_from_baseline": null}
}

9. Acceptance gate

Capacity decision record
decision: <accept / reject / retest>
platform: <Elasticsearch 9.5.3 | OpenSearch 3.8.0>
configuration_id: <immutable-id>
safe_capacity:
  search_qps: <measured>
  write_ops_s: <measured>
  concurrent_clients: <measured>
headroom:
  normal_peak_percent: <measured>
  failure_state_percent: <measured>
quality_and_correctness:
  ndcg_at_10: <measured>
  vector_recall_at_10: <measured>
  indexing_visibility_p99_ms: <measured>
recovery:
  bounded_spike_recovery_s: <measured>
  node_failure_recovery_s: <measured-or-not-run>
cost:
  measurement_window_cost: <measured/estimated-with-source>
rollback:
  change_to_revert: <exact-setting/topology>
rebenchmark_triggers:
  - dataset growth threshold
  - query/write mix change
  - mapping/analyzer/vector model change
  - shard/replica/topology/storage change
  - server/client/plugin upgrade
  - SLO/RPO/RTO change

10. Bridge to Chapter 31: capacity becomes a capstone requirement

Chapter 31 begins with Define Search/Analytics Requirements, Relevance Judgments, Freshness, Retention, Security, SLOs, RPO/RTO, and Cost Constraints. Carry the Chapter 30 benchmark contract, safe-capacity envelope, recovery evidence, and cost assumptions into the capstone. The capstone should not choose Elasticsearch or OpenSearch from a feature checklist; it should use measured relevance, operability, security, recovery, performance, licensing, and migration evidence.

Check your understanding

  1. What defines the saturation knee?
  2. Why is a load spike recovery test useful in the one-node mandatory lab?
  3. Why run mixed workloads concurrently?
  4. When is a tuning change accepted?
  5. What evidence moves into Chapter 31?
Review the answers

1. The region where additional offered load yields diminishing achieved throughput while tail latency, lag, queueing, rejections, or resource pressure rise disproportionately.

2. It measures backlog and SLO recovery without falsely claiming node-failure high availability.

3. Indexing, merges, aggregations, vector search, and interactive search compete for shared CPU, memory, storage, network, and queues.

4. Only when repeated comparable runs improve the target while meeting freshness, quality, correctness, durability/availability, recovery, and cost constraints.

5. The workload contract, safe operating envelope, tail/relevance/freshness measurements, failure/recovery evidence, topology assumptions, cost model, and rebenchmark triggers.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.