Capacity starts with a reproducible workload contract, not a node-count guess.

Workload Characterization: QPS, Write Rate, Document Size, Fields, Retention, Aggregations, Vector Dimensions, and Concurrency

Build capacity and performance plans from measured workload dimensions and tail behavior rather than generic shard, heap, bulk, or hardware rules.

Intermediate → Advanced150–200 minutesWorkload characterization & baseline · Chapter 30 · Lesson 01Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · free/local benchmark pathLast reviewed: September 2026

Learning outcomes

01

Turn business traffic into a measurable workload contract covering QPS, writes, concurrency, freshness, retention, and failure headroom.

02

Explain why document size, field count/cardinality, aggregation shape, and vector dimensions change CPU, heap, disk, network, and index-build cost.

03

Distinguish offered load from achieved throughput and median latency from tail latency and queueing behavior.

04

Design a benchmark dataset and warmup policy that cannot hide saturation behind tiny fully cached data.

05

Produce the AtlasMart baseline record that Lessons 2–5 will change one controlled variable at a time.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned performance baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The established local endpoints remain Elasticsearch at https://localhost:9200 and OpenSearch at https://localhost:9201 on the atlasmart-search Docker network. The generation environment did not execute live clusters, so this chapter never invents throughput, p95/p99, GC, disk, vector-recall, or cost results: numeric fields shown in report templates are deliberately blank/null until measured.

1. AtlasMart problem: “how many nodes do we need?” is not yet a capacity question

AtlasMart expects product-search traffic to double during a campaign while catalog writes, price updates, faceted analytics, and eight-dimensional teaching vectors share the same cluster. Asking for a node count before describing the workload is equivalent to sizing a road without knowing vehicles, arrival rate, or acceptable queueing. Capacity planning begins by turning application behavior into a workload contract that can be replayed.

QPS is completed query requests per second. Write rate is indexing/update/delete operations per second. Concurrency is how many operations are outstanding at once. Those values are related but not interchangeable: at the same QPS, longer service times require more concurrent requests, so queues and connection pools can become the bottleneck even when aggregate throughput looks unchanged.

2. Characterize the whole request, not just its count

Dimension Record before the run Why it changes capacity conclusions
Traffic QPS/write ops/s, burst ratio, concurrency, think time same throughput with different concurrency can produce very different queueing and p99
Documents source bytes, indexed bytes, field count, cardinality, update ratio affects network, parsing, heap, segment structure, disk, merge and fetch cost
Search mix lexical/filter/facet/aggregation/vector percentages CPU, cacheability, memory and coordination profiles differ
Freshness write→search visibility target and observed indexing lag refresh frequency trades freshness for segment/merge work
Retention hot/warm/cold duration and daily growth drives disk, shard count, recovery time and tiering
Quality golden-query relevance and vector recall@k a faster system is not an improvement if result quality regresses
Failure model one node/zone loss target and recovery objective normal-state capacity is not sufficient headroom
Cost hardware/service cost over the same measurement window allows cost per indexed GB, query, or business transaction to be compared

Record both offered load and achieved throughput. If the driver attempts 500 searches/s but the server only completes 360/s while queues grow, “500 QPS” is not capacity; it is an overload input. Record p50, p95, and p99 because averages can remain attractive while a minority of users wait behind saturated CPU, disk, thread pools, or coordinating-node work.

3. Document size, fields, and cardinality are work multipliers

A 2 KB product and a 200 KB product are not equivalent “one document” writes. Larger sources consume more network bandwidth, parsing time, disk, filesystem-cache space, and fetch I/O. Field count changes mapping/segment metadata; high-cardinality keyword fields change terms dictionaries, global ordinals, and aggregation state. Dynamic field explosions also increase cluster-state/mapping overhead. Measure source bytes and indexed store bytes rather than extrapolating from document count alone.

Retention converts a rate into a standing dataset. If AtlasMart produces R indexed bytes/day and retains D days, raw retained bytes are roughly R × D before replicas, merge/translog overhead, snapshots, tier copies, and operational free space. Treat that formula as a planning identity, not a final disk multiplier: the measured compression ratio and shard layout belong in the benchmark report.

4. Aggregations and vectors add different resource dimensions

A filtered top-10 product query, a high-cardinality terms aggregation, and an approximate kNN request can all be “one search” yet stress different subsystems. Aggregations can hold bucket state in memory and coordinate reductions; vector search introduces dimension count, vector encoding, ANN graph/index size, candidate count, filter selectivity, and recall targets. A capacity model must therefore preserve the production mix instead of benchmarking one cheap query repeatedly.

Work item Primary evidence Quality/correctness guardrail
lexical retrieval p95/p99, CPU, searched shards golden top-k / NDCG
facets/aggregations latency, heap/breakers, bucket counts exact/known fixture totals
bulk indexing docs/s, MB/s, indexing lag, merge/disk item failures and source count
vector ANN latency, build time, memory/disk recall@k against exact ground truth
mixed workload all of the above plus rejections no hidden degradation in any workload class

5. Warm, cold, and “too small to fail” benchmarks

If the entire AtlasMart test index fits in page cache, repeated searches may mostly measure memory access rather than the storage behavior expected in production. Conversely, a deliberately cold first request may overstate normal latency. Declare the state: dataset size relative to RAM, whether the run follows a fixed warmup, whether caches were intentionally cold, and whether a restart occurred. Compare like with like.

Wrong approach: index 5,000 tiny products, run the same query 1,000 times, report mean latency, then multiply the result to predict a 500 GB production cluster. Repair: scale data and field/cardinality shape enough to exercise the intended storage/cache regime, replay the real query mix and concurrency, report tails/rejections/freshness, and test recovery headroom.

6. AtlasMart baseline fixture

All Chapter 30 mandatory exercises preserve the course's established free/local security boundary: Elasticsearch 9.5.3 at https://localhost:9200 authenticated with ELASTIC_PASSWORD and the copied CA atlasmart-es-http-ca; OpenSearch 3.8.0 at https://localhost:9201 authenticated with OPENSEARCH_INITIAL_ADMIN_PASSWORD. The shared Docker network remains atlasmart-search. OpenSearch's -k examples are for the disposable demo certificate only and are not production TLS guidance. The default local fixture uses one primary and zero replicas because it is a workstation lab; any availability or node-failure conclusion must be tested in a disposable multi-node variant rather than inferred from this single-node setup.

AtlasMart performance fixture mapping
PUT atlasmart-perf-v1
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0,
    "refresh_interval": "1s"
  },
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "sku":        {"type":"keyword"},
      "tenant_id":  {"type":"keyword"},
      "name":       {"type":"text", "fields":{"raw":{"type":"keyword"}}},
      "category":   {"type":"keyword"},
      "brand":      {"type":"keyword"},
      "price":      {"type":"scaled_float", "scaling_factor":100},
      "available":  {"type":"boolean"},
      "event_time": {"type":"date"},
      "popularity": {"type":"float"},
      "embedding":  {"type":"dense_vector", "dims":8, "index":true, "similarity":"cosine"}
    }
  }
}
Cross-product mapping note. The Elasticsearch dense_vector example above is intentionally product-specific. For OpenSearch, create an equivalent knn_vector mapping using the Chapter 25 pinned engine/method settings rather than copying this JSON. The benchmark contract is equivalent dimensions, source vectors, filters, candidate count, and judgments—not identical mapping syntax.
Record the workload contract before tuning
platforms:
  elasticsearch: 9.5.3
  opensearch: 3.8.0
workload:
  search_qps: <measure/target>
  write_ops_s: <measure/target>
  concurrent_clients: <measure/target>
  query_mix:
    lexical_filter: <percent>
    facet_aggregation: <percent>
    vector_knn: <percent>
  source_document_bytes: <distribution>
  indexed_field_count: <count>
  high_cardinality_fields: [sku, tenant_id]
  vector_dimensions: 8
  refresh_visibility_slo_ms: <target>
  retention_days: <target>
  failure_target: <node/zone loss>
acceptance:
  p95_ms: <target>
  p99_ms: <target>
  max_rejection_rate: <target>
  relevance_floor: <metric>
  recovery_rto_s: <target>

7. Measure the platform, not only the client

Capture server-side evidence before and after each run
GET _cluster/health
GET _cat/nodes?v&h=name,cpu,heap.percent,ram.percent,disk.used_percent,node.role,master
GET _cat/thread_pool/search,write?v&h=node_name,name,active,queue,rejected,completed
GET _nodes/stats/jvm,process,os,fs,indices,thread_pool,indexing_pressure
GET atlasmart-perf-v1/_stats?level=shards
GET _cat/recovery/atlasmart-perf-v1?v
GET atlasmart-perf-v1/_segments

Client latency tells you the user symptom; server metrics explain the mechanism. Align timestamps so a p99 spike can be compared with CPU saturation, GC, disk throughput, merge activity, search/write queueing, rejections, and shard recovery. A green cluster only means shard allocation is healthy; it does not prove the search SLO is healthy.

8. Production judgment

The output of Lesson 1 is not a tuning recommendation. It is a versioned benchmark contract: representative data, traffic mix, concurrency, freshness, quality, failure target, and cost unit. Rebenchmark after mapping/analyzer changes, vector-model or dimension changes, shard/replica changes, hardware/storage changes, significant data growth, new query classes, server upgrades, or changes in managed-service topology.

Check your understanding

  1. Why is offered QPS different from achieved throughput?
  2. Why is document count an insufficient sizing metric?
  3. Why must p99 be reported with throughput?
  4. What must accompany ANN latency?
  5. What changes in Lesson 2?
Review the answers

1. Offered QPS is what the load generator attempts; achieved throughput is what the system actually completes. Under saturation they diverge as queues, timeouts, or rejections grow.

2. Document byte size, field count/cardinality, update shape, mappings, vectors, compression, and retention determine resource use.

3. A system can maintain aggregate throughput while a tail of requests waits behind saturated resources; p99 reveals that user-facing queueing.

4. A quality metric such as recall@k against exact ground truth, plus filter selectivity and vector/index settings.

5. The workload contract stays fixed while indexing-side variables such as bulk size/concurrency, refresh, replicas, pipelines, merge pressure, durability, and backpressure are tested.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.