Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health
Build a Saturation Diagnosis from GC, CPU, Heap, Rejections, Latency, Disk, and Workload Evidence
Build an evidence chain for AtlasMart runtime saturation, run a bounded mixed-workload experiment, change one variable, and prove or reject the causal hypothesis without confusing storage, JVM, queue or client-side pressure.
Learning outcomes
AtlasMart now has all the pieces to misdiagnose a slowdown: high heap could mean healthy cache use or GC pressure; high CPU could mean useful throughput or one pathological query; a low cache hit rate might be irrelevant; a queue might be the first saturation signal or merely a brief burst. The goal is therefore not a dashboard screenshot—it is a causal narrative supported by synchronized workload, latency and server evidence.
Build a time-aligned saturation evidence bundle spanning client latency/throughput/errors, CPU, heap/GC, queues/rejections, breakers and storage.
Form a falsifiable hypothesis that names the resource, mechanism and predicted observations.
Run a resource-bounded mixed-workload test with explicit stop conditions.
Change one variable and distinguish causal improvement from warm-cache/noise effects.
Decide whether the durable action is query/mapping repair, admission control, workload isolation, capacity, storage repair or a JVM/config change.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual JVM/runtime with
GET _nodes/jvm and
GET _nodes/stats/jvm,process,os rather than
hard-coding a JDK patch. Labs use one primary and zero
replicas unless a step explicitly changes topology. OpenSearch
demo -k TLS bypass remains disposable-lab-only;
production must validate certificates. No moving
latest tags are used.
This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.
1. Start with the service symptom, not the JVM knob
Write the incident in client terms first: “From 14:03–14:11, product-search p99 rose from its normal band, 429s appeared, and indexing freshness exceeded the agreed window while request rate increased.” Do not begin with “heap is 78%” because that already presumes the cause.
Then align server evidence to the same window. A runtime diagnosis is strongest when independent signals agree: queue depth/rejections, CPU/hot threads, heap/GC, breaker/indexing-pressure counters, disk latency, shard/indexing/search rates and client concurrency.
| Evidence family | What to capture | What it can falsify |
|---|---|---|
| Client | Request rate, concurrency, p50/p95/p99, status/errors, retries, payload/query mix. | Server-only theories when load/retry behavior changed first. |
| JVM | Heap used/max, GC collections/time, JVM identity. | “GC caused it” if GC remains stable across the incident. |
| Execution | Thread-pool active/queue/rejected/completed; tasks/hot threads. | “No saturation” when queues/rejections rise. |
| Protection | Breaker trips/estimates; indexing pressure; OpenSearch backpressure stats. | “The cluster had plenty of admission headroom.” |
| Storage/index | Merge/refresh/search/indexing stats; disk/host telemetry. | Heap-only theory when I/O wait/storage tail rises. |
| Topology | Shard counts/routing/skew/node health/recovery. | Uniform-load theory when one hot shard/node dominates. |
2. One API evidence bundle
GET _cluster/health?level=shards
GET _cat/nodes?v&h=name,roles,cpu,heap.percent,ram.percent,load_1m,node.role
GET _nodes/stats/jvm,process,os,breaker,thread_pool,indexing_pressure?human
GET _nodes/stats/indices/search,indexing,merge,refresh,query_cache,request_cache,fielddata?human
GET _nodes/hot_threads?threads=5&ignore_idle_threads=true
GET _tasks?detailed=true&actions=*search*,*write*,*bulk*
GET _nodes/stats/shard_indexing_pressure,search_backpressure?human
GET _cluster/settings?include_defaults=true&flat_settings=true
Not every API field is identical across products/releases. Preserve raw JSON alongside your normalized incident summary so future reviewers can see what the server actually returned. Scrub credentials and sensitive query/user data before sharing evidence.
3. Turn observations into a falsifiable hypothesis
Bad hypothesis: “Elasticsearch needs more heap.” Good hypothesis: “The sale traffic increased client search concurrency beyond the search pool’s sustainable drain rate; queued requests raised p99 and retries amplified arrival rate, while GC and disk remained stable. Capping client concurrency should reduce queue/rejection growth and p99 without reducing steady-state useful throughput.”
The good hypothesis predicts both confirming and disconfirming evidence. If queue depth and p99 remain high after concurrency is capped while disk wait spikes, the hypothesis is incomplete or wrong.
4. A decision matrix for common saturation shapes
| Observed shape | Likely mechanism to test | Do not jump straight to |
|---|---|---|
| Heap climbs + GC time/pauses + breaker pressure | Heap-heavy query/aggregation, fielddata, oversized caches, concurrency, cluster state. | Bigger heap / disabled breaker. |
| CPU high + search hot threads + queue/rejections | Expensive query, shard fan-out, script/aggregation cost, excessive concurrency. | Bigger queue. |
| CPU moderate + disk wait/merge high + latency high | Storage/merge/recovery or cold filesystem cache. | More search threads. |
| Write rejections + indexing pressure + hot shard | Bulk size/concurrency, routing skew, replica/storage drain rate. | Raise indexing-pressure limit. |
| OpenSearch search-backpressure cancellations | Resource-intensive search under node duress. | Disable cancellation protection. |
| Low cache hit rate but SLO healthy | Dynamic workload where caching may not matter. | Increase cache sizes simply to improve hit rate. |
5. Build the resource-bounded AtlasMart mixed workload
The mandatory lab uses atlasmart-runtime-lab-v1 and
a finite query set: one full-text product search, one filtered
size-zero facet aggregation, one exact lookup and a small
deterministic bulk batch. Use a single-node local container
only. Do not target a shared cluster.
The workload runner can be any local client capable of recording per-request latency and status. Keep the server API payloads exactly fixed across candidates. If you implement a custom runner, cap workers and request count in constants—do not generate unbounded load.
GET atlasmart-runtime-lab-v1/_search
{
"size": 10,
"query": {
"bool": {
"must": [{"match":{"name":"wireless keyboard"}}],
"filter": [
{"term":{"available":true}},
{"range":{"price":{"lte":150}}}
]
}
},
"sort": [{"_score":"desc"},{"sku":"asc"}]
}
GET atlasmart-runtime-lab-v1/_search?request_cache=true
{
"size": 0,
"query":{"term":{"available":true}},
"aggs":{
"category":{"terms":{"field":"category","size":10}},
"price":{"percentiles":{"field":"price","percents":[50,95]}}
}
}
- Baseline: concurrency 1, fixed request count, warm-up then measurement.
- Candidate pressure step: increase to the next bounded concurrency only if stop conditions remain clear.
- Safe saturation signal may be rising queue/p99 or a controlled 429/rejection; you do not need an OOM or breaker trip.
- Stop before host swapping, container kill risk, prolonged GC, unrelated impact or runaway retries.
- After cooldown, apply one repair—e.g., cap client concurrency—and rerun the identical mix.
6. Calculate latency percentiles without pretending they are universal targets
Sort the measured client durations. The 95th percentile approximates the latency below which 95% of requests completed; p99 focuses further into the tail. Use the same percentile method and sample size for baseline/candidate. Do not publish a universal “Elasticsearch p99 should be X ms”: acceptable latency depends on query, data, topology, hardware and product requirements.
Throughput alone is insufficient. A candidate that increases operations/second by 5% while doubling p99 and introducing rejections may be worse for an interactive service. Conversely, a batch pipeline may accept higher latency if useful throughput and recovery behavior improve within its SLA.
7. Change one variable, then try to falsify yourself
Suppose concurrency 8 creates search queue growth and 429s while CPU is high, but concurrency 4 has similar useful throughput and much lower p99. Set the application limit to 4 for the lab and rerun. If queues/rejections disappear and throughput stays similar, the evidence supports client admission control as part of the repair.
Now try to falsify it: repeat after a restart/cold cache, after enough time for merge/background work to settle, and with the same data/query mix. If the result depends entirely on warmed filesystem cache, the conclusion must say so. If a hot shard dominates, revisit Chapter 13 routing rather than attributing everything to concurrency.
8. What to scale and what to tune
| Finding | Primary durable actions |
|---|---|
| Bad query/mapping shape | Fix query, mapping/analyzer, aggregation/pagination strategy; add regression tests. |
| Client-driven overload | Bound concurrency/bulk size; backoff+jitter; queue/shed/defer noncritical work. |
| Hot tenant/shard | Routing/model fix, tenant budgets/isolation, rebalance/reindex if justified. |
| Legitimate CPU-bound capacity shortage | Scale/resize after query efficiency and admission controls are validated. |
| Storage/merge bottleneck | Storage/headroom/lifecycle/refresh/merge investigation from Chapter 14; do not hide with heap. |
| Heap-heavy legitimate working set | Mapping/cache/query repair first; then supported heap/capacity evaluation with filesystem-cache headroom preserved. |
9. Operational and security boundaries
Runtime evidence can contain query strings, tenant names, hostnames and potentially sensitive fields. Restrict node/task/hot-thread APIs to operators, store diagnostics with access control, and redact before external sharing. On managed Elastic/Amazon/OpenSearch offerings, some host/JVM controls are intentionally abstracted; use provider metrics and supported knobs rather than trying to reproduce self-managed internals exactly.
A node replica is not a backup. Capacity tuning must preserve snapshot/restore and recovery objectives. A “fast” cluster that cannot recover within its failure budget or has no disk/memory headroom is not production-ready.
Production judgment and bridge to Chapter 16
A credible saturation diagnosis says what changed, which resource became constrained, how the mechanism produced the client symptom, what evidence supports it, what evidence would disprove it, and how the repair was verified. It does not end with a generic tuning checklist.
Chapter 16 moves from node-level saturation into search
execution itself: distributed query/fetch phases, profiling,
slow logs, tasks, deep pagination, search_after,
point-in-time readers and cancellation. The runtime evidence
learned here becomes the context for interpreting those
search-specific tools.
Check your understanding
- Why start with the client-visible symptom?
- What makes a hypothesis falsifiable?
- What is a safe saturation signal for the local lab?
- Why compare p99 and throughput together?
- When is scaling preferable to tuning?
Review the answers
1. It anchors diagnosis in the service objective and prevents one internal metric from being assumed causal before time-aligned evidence exists.
2. It names a mechanism and predicts observations that should change—or fail to change—when one controlled variable is altered.
3. Rising queue/tail latency or a bounded rejection under capped load; an OOM or disabled safety control is not required.
4. More throughput can be purchased by harmful queueing/rejection tail latency; service quality requires both plus correctness/errors.
5. After query/model/admission problems are addressed and legitimate demand still consumes the measured capacity needed to meet SLOs with recovery headroom.
Summary and next step
You can now separate heap, filesystem cache, logical caches, breaker pressure, queues, indexing/search backpressure and client concurrency, then connect them to synchronized workload and tail-latency evidence. That is the runtime foundation for Chapter 16’s query/fetch profiling and pagination work.
Authoritative references
- Elastic JVM settings — Automatic heap sizing, filesystem-cache headroom, container memory, and compressed ordinary object pointer guidance.
- Elastic node query cache settings — Filter-context query-cache eligibility, segment-level caching and invalidation behavior.
- Elastic shard request cache — Shard-level request-result caching and request/index controls.
- Elastic field data cache settings — On-heap fielddata/global-ordinal cache behavior and breaker interaction.
- Elastic circuit breaker settings — Parent, fielddata, request and in-flight breaker semantics.
- Elastic thread pool settings — Current thread-pool types, sizing and queue behavior.
- Elastic indexing pressure — Outstanding indexing-byte accounting and rejection behavior.
- Elastic Nodes Stats API — JVM, process, caches, breakers, thread pools and indexing-pressure observations.
- OpenSearch circuit breaker settings — Parent/child breaker settings and failure-protection intent.
- OpenSearch caching overview — On-heap request/query/fielddata caching concepts.
- OpenSearch index request cache — Shard-level request-result caching, invalidation and metrics.
- OpenSearch field data cache — Fielddata/global ordinals and the keyword-subfield alternative.
- OpenSearch thread pool settings — Thread pools, queue sizing and monitoring guidance.
- OpenSearch search backpressure — OpenSearch-specific search task resource tracking and cancellation.
- OpenSearch shard indexing backpressure — OpenSearch-specific per-shard indexing pressure and rejection mechanism.
- OpenSearch Nodes Stats API — JVM, process, caches, indexing pressure and backpressure statistics.
- OpenSearch Nodes Hot Threads API — CPU/wait/block thread samples for diagnosis.