Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health
Thread Pools, Queues, Rejections, Indexing Pressure, Search Backpressure, and Client-Side Concurrency
Trace work from client concurrency into server thread pools, queues and memory-pressure controls, then interpret rejections as backpressure signals and apply bounded admission control rather than hiding overload with larger queues.
Learning outcomes
AtlasMart launches a sale. Client concurrency jumps from 8 to 80. Search throughput rises briefly, then p99 latency explodes and 429 responses appear. An operator proposes increasing every server queue so “requests stop failing.” The queue only stores more waiting work; it does not create CPU, heap, disk bandwidth or shard capacity. This is a backpressure problem: the system must limit admitted work before waiting/retry amplification collapses useful throughput.
Explain fixed/scaling thread pools, queueing and rejection as execution-capacity controls.
Distinguish Elasticsearch indexing pressure from OpenSearch shard indexing and search backpressure.
Interpret 429/rejections as overload evidence instead of hiding them with unbounded queues.
Design client concurrency, bulk size and retry/backoff limits that cooperate with server admission controls.
Run a bounded concurrency experiment and identify the knee where tail latency/rejections worsen without useful throughput gain.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual JVM/runtime with
GET _nodes/jvm and
GET _nodes/stats/jvm,process,os rather than
hard-coding a JDK patch. Labs use one primary and zero
replicas unless a step explicitly changes topology. OpenSearch
demo -k TLS bypass remains disposable-lab-only;
production must validate certificates. No moving
latest tags are used.
This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.
1. Thread pools are lanes; queues are waiting rooms
A node has separate execution pools for important categories such as search and write work. A fixed thread pool has a bounded worker count and queue; when all workers and queue slots are occupied, new work is rejected. Scaling pools can vary workers within product-defined bounds for their task type.
Thread-pool sizing depends on allocated processors and version-specific defaults. Elasticsearch 9.x changed some queue defaults relative to older releases—for example, current search queue sizing scales with the search pool size unless explicitly overridden. Therefore screenshots/blog posts from older versions are not a reliable tuning contract.
GET _cat/thread_pool/search,write?v&h=node_name,name,active,queue,rejected,completed,size,queue_size
GET _nodes/stats/thread_pool?human
GET _nodes/hot_threads?threads=5&ignore_idle_threads=true
GET _nodes/stats/jvm,process,os?human
2. Why “make the queue bigger” is usually not the first repair
A larger queue can reduce immediate rejections while increasing waiting time, retained request state and timeout/retry overlap. If the service already cannot drain work as fast as clients send it, a bigger queue often moves the symptom from fast rejection to very slow success/failure.
The useful questions are: what resource limits drain rate—CPU, storage, shard fan-out, heap/GC, network, expensive query shape, hot routing? Is the client sending more concurrent work than the node can process? Can work be isolated or scaled? Queue changes come after that evidence.
Setting an unbounded queue (where configurable) converts explicit backpressure into potentially unbounded waiting/memory. Treat queues as finite shock absorbers, not extra capacity.
3. Elasticsearch indexing pressure
Elasticsearch accounts outstanding indexing bytes through
coordinating, primary and replica stages. The documented
indexing_pressure.memory.limit defaults to a
fraction of heap and causes new indexing work to be rejected
when outstanding bytes exceed limits. The accounting
intentionally persists across downstream stages so a
coordinating request is not “forgotten” while primaries/replicas
still owe work.
Do not immediately raise the limit on 429s. Inspect bulk size, concurrency, shard skew, replica/storage health and write-thread-pool stats. A smaller bulk with bounded concurrency often protects tail latency and search availability better than one giant outstanding pipeline.
GET _nodes/stats/indexing_pressure?human
GET _nodes/stats/thread_pool,jvm,process?human
GET _cat/thread_pool/write?v&h=node_name,active,queue,rejected,completed,size,queue_size
4. OpenSearch adds explicit shard/search backpressure surfaces
OpenSearch 3.8 exposes ordinary node indexing pressure plus shard indexing backpressure, which can track per-shard strain and—when enabled/enforced—reject indexing work based on memory plus secondary degradation signals. The feature can run in shadow mode to collect evidence before enforcement.
OpenSearch also exposes search backpressure: it tracks CPU, heap and elapsed-time signals for search tasks and can cancel resource-intensive tasks when the node is under duress and configured thresholds are breached. These names and settings are OpenSearch-specific; do not copy them into Elasticsearch configuration as if the systems were drop-in equivalents.
GET _nodes/stats/indexing_pressure,shard_indexing_pressure,search_backpressure?human
GET _cluster/settings?include_defaults=true&flat_settings=true
GET _cat/thread_pool/search,write?v&h=node_name,name,active,queue,rejected,completed,size,queue_size
Do not enable aggressive rejection/cancellation settings in a shared production cluster from a tutorial. Learn the metrics and shadow/default behavior locally; production changes require workload-specific SLOs, rollback and canary evidence.
5. Client concurrency is part of the cluster control loop
The client decides how much work exists simultaneously. If every application instance doubles concurrency after a timeout, it can create a retry storm exactly when the cluster has the least spare capacity. Admission control belongs at both ends.
Use bounded worker pools/semaphores, bounded bulk/request sizes and retries with exponential backoff plus jitter. Retry only operations whose semantics make retry safe; preserve idempotency keys/IDs and inspect bulk item failures. Respect explicit server rejection instead of immediately resubmitting at the same rate.
| Signal | Interpretation | Safer first action |
|---|---|---|
| Search queue rising, few rejections | Arrival rate is approaching/exceeding drain rate. | Reduce/constrain client concurrency; inspect query/shard/CPU cost. |
| 429 + write rejections | Write path is saturated/protected. | Reduce bulk concurrency/size; inspect indexing pressure, hot shards and storage. |
| High CPU + hot search threads | Compute-heavy search or excessive fan-out. | Repair query/routing/aggregation or scale/isolate; do not only enlarge queue. |
| Low CPU + high wait/storage latency | I/O or blocked dependency may dominate. | Investigate storage/network/recovery before adding compute threads. |
| Retries increase as success rate falls | Positive feedback/retry storm. | Backoff, cap retry budget and shed/defer non-critical work. |
6. AtlasMart bounded-concurrency experiment
The lab intentionally increases concurrency only on a disposable local node and uses a hard stop. The objective is to find a knee: a point where additional concurrency yields little/no throughput improvement while p95/p99, queues or rejections worsen.
Use the same fixed request set from Lessons 1–3. Test a short ladder such as concurrency 1 → 2 → 4 → 8 only if the previous step remains healthy; do not assume 8 is safe on every machine. Each step has a fixed request count and cooldown. Stop at the first meaningful rejection/host-pressure signal.
GET _nodes/stats/jvm,process,os,breaker,thread_pool,indexing_pressure?human
GET _nodes/stats/indices/search,indexing,query_cache,request_cache?human
GET _nodes/hot_threads?threads=3&ignore_idle_threads=true
# OpenSearch additionally:
GET _nodes/stats/shard_indexing_pressure,search_backpressure?human
- Warm with concurrency 1 and a fixed request set.
- Run a fixed count; record throughput and p50/p95/p99 from the client.
- Increase one concurrency step only if no stop condition triggered.
- Capture queue/rejection/CPU/heap/GC deltas, not only latency.
- Stop on 429 burst, sustained host swapping, container memory kill risk, unrelated workload impact or sharply worsening tail latency.
- Choose an operating concurrency below the knee with recovery/headroom, not the maximum observed throughput point.
7. Production judgment: tune, scale or shed?
If the bottleneck is an expensive query, adding nodes may postpone but not fix abuse. If correctly modeled queries consume all CPU at legitimate demand, scaling or workload isolation may be justified. If one tenant can saturate the cluster, admission control and tenant budgets are security/reliability requirements, not optional performance polish.
Retries and timeouts change load. A client timeout does not guarantee the server instantly stops all work; Chapter 16 will examine cancellation, tasks and long-running search semantics. For now, treat timeouts as part of the feedback loop and keep retry budgets bounded.
Check your understanding
- Why can a larger queue increase p99 latency?
- What does an indexing-pressure rejection mean?
- What is OpenSearch search backpressure?
- Why must the client limit concurrency too?
- Where should you operate relative to the measured concurrency knee?
Review the answers
1. It lets more requests wait behind saturated workers, increasing queueing time and retained state without increasing drain capacity.
2. The node is protecting itself because outstanding indexing work exceeded a configured/accounted limit; it is evidence to reduce/reshape load or add capacity.
3. An OpenSearch-specific mechanism that tracks search task resource use and can cancel expensive search work when the node is under duress.
4. Unbounded clients and retries can overwhelm server queues/memory faster than the server can drain work, creating a positive feedback loop.
5. Below it, leaving headroom for bursts, recovery, merges, GC and other workloads rather than at the maximum throughput edge.
Summary and next step
Thread pools and queues make overload visible, indexing/search protections reject or cancel work, and the client controls how much work exists at once. Lesson 5 combines those signals with GC, CPU, disk and latency into a causal incident method instead of tuning one metric at a time.
Authoritative references
- Elastic JVM settings — Automatic heap sizing, filesystem-cache headroom, container memory, and compressed ordinary object pointer guidance.
- Elastic node query cache settings — Filter-context query-cache eligibility, segment-level caching and invalidation behavior.
- Elastic shard request cache — Shard-level request-result caching and request/index controls.
- Elastic field data cache settings — On-heap fielddata/global-ordinal cache behavior and breaker interaction.
- Elastic circuit breaker settings — Parent, fielddata, request and in-flight breaker semantics.
- Elastic thread pool settings — Current thread-pool types, sizing and queue behavior.
- Elastic indexing pressure — Outstanding indexing-byte accounting and rejection behavior.
- Elastic Nodes Stats API — JVM, process, caches, breakers, thread pools and indexing-pressure observations.
- OpenSearch circuit breaker settings — Parent/child breaker settings and failure-protection intent.
- OpenSearch caching overview — On-heap request/query/fielddata caching concepts.
- OpenSearch index request cache — Shard-level request-result caching, invalidation and metrics.
- OpenSearch field data cache — Fielddata/global ordinals and the keyword-subfield alternative.
- OpenSearch thread pool settings — Thread pools, queue sizing and monitoring guidance.
- OpenSearch search backpressure — OpenSearch-specific search task resource tracking and cancellation.
- OpenSearch shard indexing backpressure — OpenSearch-specific per-shard indexing pressure and rejection mechanism.
- OpenSearch Nodes Stats API — JVM, process, caches, indexing pressure and backpressure statistics.
- OpenSearch Nodes Hot Threads API — CPU/wait/block thread samples for diagnosis.