Capacity starts with a reproducible workload contract, not a node-count guess.
Workload Characterization: QPS, Write Rate, Document Size, Fields, Retention, Aggregations, Vector Dimensions, and Concurrency
Build capacity and performance plans from measured workload dimensions and tail behavior rather than generic shard, heap, bulk, or hardware rules.
Learning outcomes
Turn business traffic into a measurable workload contract covering QPS, writes, concurrency, freshness, retention, and failure headroom.
Explain why document size, field count/cardinality, aggregation shape, and vector dimensions change CPU, heap, disk, network, and index-build cost.
Distinguish offered load from achieved throughput and median latency from tail latency and queueing behavior.
Design a benchmark dataset and warmup policy that cannot hide saturation behind tiny fully cached data.
Produce the AtlasMart baseline record that Lessons 2–5 will change one controlled variable at a time.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: “how many nodes do we need?” is not yet a capacity question
AtlasMart expects product-search traffic to double during a campaign while catalog writes, price updates, faceted analytics, and eight-dimensional teaching vectors share the same cluster. Asking for a node count before describing the workload is equivalent to sizing a road without knowing vehicles, arrival rate, or acceptable queueing. Capacity planning begins by turning application behavior into a workload contract that can be replayed.
QPS is completed query requests per second. Write rate is indexing/update/delete operations per second. Concurrency is how many operations are outstanding at once. Those values are related but not interchangeable: at the same QPS, longer service times require more concurrent requests, so queues and connection pools can become the bottleneck even when aggregate throughput looks unchanged.
2. Characterize the whole request, not just its count
| Dimension | Record before the run | Why it changes capacity conclusions |
|---|---|---|
| Traffic | QPS/write ops/s, burst ratio, concurrency, think time | same throughput with different concurrency can produce very different queueing and p99 |
| Documents | source bytes, indexed bytes, field count, cardinality, update ratio | affects network, parsing, heap, segment structure, disk, merge and fetch cost |
| Search mix | lexical/filter/facet/aggregation/vector percentages | CPU, cacheability, memory and coordination profiles differ |
| Freshness | write→search visibility target and observed indexing lag | refresh frequency trades freshness for segment/merge work |
| Retention | hot/warm/cold duration and daily growth | drives disk, shard count, recovery time and tiering |
| Quality | golden-query relevance and vector recall@k | a faster system is not an improvement if result quality regresses |
| Failure model | one node/zone loss target and recovery objective | normal-state capacity is not sufficient headroom |
| Cost | hardware/service cost over the same measurement window | allows cost per indexed GB, query, or business transaction to be compared |
Record both offered load and achieved throughput. If the driver attempts 500 searches/s but the server only completes 360/s while queues grow, “500 QPS” is not capacity; it is an overload input. Record p50, p95, and p99 because averages can remain attractive while a minority of users wait behind saturated CPU, disk, thread pools, or coordinating-node work.
3. Document size, fields, and cardinality are work multipliers
A 2 KB product and a 200 KB product are not equivalent “one document” writes. Larger sources consume more network bandwidth, parsing time, disk, filesystem-cache space, and fetch I/O. Field count changes mapping/segment metadata; high-cardinality keyword fields change terms dictionaries, global ordinals, and aggregation state. Dynamic field explosions also increase cluster-state/mapping overhead. Measure source bytes and indexed store bytes rather than extrapolating from document count alone.
Retention converts a rate into a standing dataset. If AtlasMart
produces R indexed bytes/day and retains
D days, raw retained bytes are roughly
R × D before replicas, merge/translog overhead,
snapshots, tier copies, and operational free space. Treat that
formula as a planning identity, not a final disk multiplier: the
measured compression ratio and shard layout belong in the
benchmark report.
4. Aggregations and vectors add different resource dimensions
A filtered top-10 product query, a high-cardinality
terms aggregation, and an approximate kNN request
can all be “one search” yet stress different subsystems.
Aggregations can hold bucket state in memory and coordinate
reductions; vector search introduces dimension count, vector
encoding, ANN graph/index size, candidate count, filter
selectivity, and recall targets. A capacity model must therefore
preserve the production mix instead of benchmarking one cheap
query repeatedly.
| Work item | Primary evidence | Quality/correctness guardrail |
|---|---|---|
| lexical retrieval | p95/p99, CPU, searched shards | golden top-k / NDCG |
| facets/aggregations | latency, heap/breakers, bucket counts | exact/known fixture totals |
| bulk indexing | docs/s, MB/s, indexing lag, merge/disk | item failures and source count |
| vector ANN | latency, build time, memory/disk | recall@k against exact ground truth |
| mixed workload | all of the above plus rejections | no hidden degradation in any workload class |
5. Warm, cold, and “too small to fail” benchmarks
If the entire AtlasMart test index fits in page cache, repeated searches may mostly measure memory access rather than the storage behavior expected in production. Conversely, a deliberately cold first request may overstate normal latency. Declare the state: dataset size relative to RAM, whether the run follows a fixed warmup, whether caches were intentionally cold, and whether a restart occurred. Compare like with like.
6. AtlasMart baseline fixture
All Chapter 30 mandatory exercises preserve the course's
established free/local security boundary: Elasticsearch 9.5.3 at
https://localhost:9200 authenticated with
ELASTIC_PASSWORD and the copied CA
atlasmart-es-http-ca; OpenSearch 3.8.0 at
https://localhost:9201 authenticated with
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The shared
Docker network remains atlasmart-search.
OpenSearch's -k examples are for the disposable
demo certificate only and are not production TLS guidance. The
default local fixture uses one primary and zero replicas because
it is a workstation lab; any availability or node-failure
conclusion must be tested in a disposable multi-node variant
rather than inferred from this single-node setup.
PUT atlasmart-perf-v1
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0,
"refresh_interval": "1s"
},
"mappings": {
"dynamic": "strict",
"properties": {
"sku": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"name": {"type":"text", "fields":{"raw":{"type":"keyword"}}},
"category": {"type":"keyword"},
"brand": {"type":"keyword"},
"price": {"type":"scaled_float", "scaling_factor":100},
"available": {"type":"boolean"},
"event_time": {"type":"date"},
"popularity": {"type":"float"},
"embedding": {"type":"dense_vector", "dims":8, "index":true, "similarity":"cosine"}
}
}
}
dense_vector example above is intentionally
product-specific. For OpenSearch, create an equivalent
knn_vector mapping using the Chapter 25 pinned
engine/method settings rather than copying this JSON. The
benchmark contract is equivalent dimensions, source vectors,
filters, candidate count, and judgments—not identical mapping
syntax.
platforms:
elasticsearch: 9.5.3
opensearch: 3.8.0
workload:
search_qps: <measure/target>
write_ops_s: <measure/target>
concurrent_clients: <measure/target>
query_mix:
lexical_filter: <percent>
facet_aggregation: <percent>
vector_knn: <percent>
source_document_bytes: <distribution>
indexed_field_count: <count>
high_cardinality_fields: [sku, tenant_id]
vector_dimensions: 8
refresh_visibility_slo_ms: <target>
retention_days: <target>
failure_target: <node/zone loss>
acceptance:
p95_ms: <target>
p99_ms: <target>
max_rejection_rate: <target>
relevance_floor: <metric>
recovery_rto_s: <target>
7. Measure the platform, not only the client
GET _cluster/health
GET _cat/nodes?v&h=name,cpu,heap.percent,ram.percent,disk.used_percent,node.role,master
GET _cat/thread_pool/search,write?v&h=node_name,name,active,queue,rejected,completed
GET _nodes/stats/jvm,process,os,fs,indices,thread_pool,indexing_pressure
GET atlasmart-perf-v1/_stats?level=shards
GET _cat/recovery/atlasmart-perf-v1?v
GET atlasmart-perf-v1/_segments
Client latency tells you the user symptom; server metrics explain the mechanism. Align timestamps so a p99 spike can be compared with CPU saturation, GC, disk throughput, merge activity, search/write queueing, rejections, and shard recovery. A green cluster only means shard allocation is healthy; it does not prove the search SLO is healthy.
8. Production judgment
The output of Lesson 1 is not a tuning recommendation. It is a versioned benchmark contract: representative data, traffic mix, concurrency, freshness, quality, failure target, and cost unit. Rebenchmark after mapping/analyzer changes, vector-model or dimension changes, shard/replica changes, hardware/storage changes, significant data growth, new query classes, server upgrades, or changes in managed-service topology.
Check your understanding
- Why is offered QPS different from achieved throughput?
- Why is document count an insufficient sizing metric?
- Why must p99 be reported with throughput?
- What must accompany ANN latency?
- What changes in Lesson 2?
Review the answers
1. Offered QPS is what the load generator attempts; achieved throughput is what the system actually completes. Under saturation they diverge as queues, timeouts, or rejections grow.
2. Document byte size, field count/cardinality, update shape, mappings, vectors, compression, and retention determine resource use.
3. A system can maintain aggregate throughput while a tail of requests waits behind saturated resources; p99 reveals that user-facing queueing.
4. A quality metric such as recall@k against exact ground truth, plus filter selectivity and vector/index settings.
5. The workload contract stays fixed while indexing-side variables such as bulk size/concurrency, refresh, replicas, pipelines, merge pressure, durability, and backpressure are tested.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Elasticsearch performance optimizations
- Elasticsearch tune for indexing speed
- Elasticsearch tune for search speed
- Elasticsearch size your shards
- Elasticsearch JVM settings
- Elasticsearch resilience guidance
- Elasticsearch bulk API
- Elasticsearch refresh parameter
- Rally documentation
- OpenSearch 3.8 version history
- OpenSearch indexing performance tuning
- OpenSearch Bulk API
- OpenSearch Refresh Index API
- OpenSearch search shard routing
- OpenSearch index settings
- OpenSearch shard indexing backpressure
- OpenSearch Benchmark quickstart
- OpenSearch Performance Analyzer metrics
- OpenSearch vector search performance tuning