Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health

Build a Saturation Diagnosis from GC, CPU, Heap, Rejections, Latency, Disk, and Workload Evidence

Build an evidence chain for AtlasMart runtime saturation, run a bounded mixed-workload experiment, change one variable, and prove or reject the causal hypothesis without confusing storage, JVM, queue or client-side pressure.

Intermediate115–150 minutesJVM/runtime health & bounded saturation labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart now has all the pieces to misdiagnose a slowdown: high heap could mean healthy cache use or GC pressure; high CPU could mean useful throughput or one pathological query; a low cache hit rate might be irrelevant; a queue might be the first saturation signal or merely a brief burst. The goal is therefore not a dashboard screenshot—it is a causal narrative supported by synchronized workload, latency and server evidence.

01

Build a time-aligned saturation evidence bundle spanning client latency/throughput/errors, CPU, heap/GC, queues/rejections, breakers and storage.

02

Form a falsifiable hypothesis that names the resource, mechanism and predicted observations.

03

Run a resource-bounded mixed-workload test with explicit stop conditions.

04

Change one variable and distinguish causal improvement from warm-cache/noise effects.

05

Decide whether the durable action is query/mapping repair, admission control, workload isolation, capacity, storage repair or a JVM/config change.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual JVM/runtime with GET _nodes/jvm and GET _nodes/stats/jvm,process,os rather than hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly changes topology. OpenSearch demo -k TLS bypass remains disposable-lab-only; production must validate certificates. No moving latest tags are used.

Execution note

This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.

1. Start with the service symptom, not the JVM knob

Write the incident in client terms first: “From 14:03–14:11, product-search p99 rose from its normal band, 429s appeared, and indexing freshness exceeded the agreed window while request rate increased.” Do not begin with “heap is 78%” because that already presumes the cause.

Then align server evidence to the same window. A runtime diagnosis is strongest when independent signals agree: queue depth/rejections, CPU/hot threads, heap/GC, breaker/indexing-pressure counters, disk latency, shard/indexing/search rates and client concurrency.

Evidence family What to capture What it can falsify
Client Request rate, concurrency, p50/p95/p99, status/errors, retries, payload/query mix. Server-only theories when load/retry behavior changed first.
JVM Heap used/max, GC collections/time, JVM identity. “GC caused it” if GC remains stable across the incident.
Execution Thread-pool active/queue/rejected/completed; tasks/hot threads. “No saturation” when queues/rejections rise.
Protection Breaker trips/estimates; indexing pressure; OpenSearch backpressure stats. “The cluster had plenty of admission headroom.”
Storage/index Merge/refresh/search/indexing stats; disk/host telemetry. Heap-only theory when I/O wait/storage tail rises.
Topology Shard counts/routing/skew/node health/recovery. Uniform-load theory when one hot shard/node dominates.

2. One API evidence bundle

Elasticsearch/OpenSearch common evidence
GET _cluster/health?level=shards
GET _cat/nodes?v&h=name,roles,cpu,heap.percent,ram.percent,load_1m,node.role
GET _nodes/stats/jvm,process,os,breaker,thread_pool,indexing_pressure?human
GET _nodes/stats/indices/search,indexing,merge,refresh,query_cache,request_cache,fielddata?human
GET _nodes/hot_threads?threads=5&ignore_idle_threads=true
GET _tasks?detailed=true&actions=*search*,*write*,*bulk*
OpenSearch-specific additions
GET _nodes/stats/shard_indexing_pressure,search_backpressure?human
GET _cluster/settings?include_defaults=true&flat_settings=true

Not every API field is identical across products/releases. Preserve raw JSON alongside your normalized incident summary so future reviewers can see what the server actually returned. Scrub credentials and sensitive query/user data before sharing evidence.

3. Turn observations into a falsifiable hypothesis

Bad hypothesis: “Elasticsearch needs more heap.” Good hypothesis: “The sale traffic increased client search concurrency beyond the search pool’s sustainable drain rate; queued requests raised p99 and retries amplified arrival rate, while GC and disk remained stable. Capping client concurrency should reduce queue/rejection growth and p99 without reducing steady-state useful throughput.”

The good hypothesis predicts both confirming and disconfirming evidence. If queue depth and p99 remain high after concurrency is capped while disk wait spikes, the hypothesis is incomplete or wrong.

4. A decision matrix for common saturation shapes

Observed shape Likely mechanism to test Do not jump straight to
Heap climbs + GC time/pauses + breaker pressure Heap-heavy query/aggregation, fielddata, oversized caches, concurrency, cluster state. Bigger heap / disabled breaker.
CPU high + search hot threads + queue/rejections Expensive query, shard fan-out, script/aggregation cost, excessive concurrency. Bigger queue.
CPU moderate + disk wait/merge high + latency high Storage/merge/recovery or cold filesystem cache. More search threads.
Write rejections + indexing pressure + hot shard Bulk size/concurrency, routing skew, replica/storage drain rate. Raise indexing-pressure limit.
OpenSearch search-backpressure cancellations Resource-intensive search under node duress. Disable cancellation protection.
Low cache hit rate but SLO healthy Dynamic workload where caching may not matter. Increase cache sizes simply to improve hit rate.

5. Build the resource-bounded AtlasMart mixed workload

The mandatory lab uses atlasmart-runtime-lab-v1 and a finite query set: one full-text product search, one filtered size-zero facet aggregation, one exact lookup and a small deterministic bulk batch. Use a single-node local container only. Do not target a shared cluster.

The workload runner can be any local client capable of recording per-request latency and status. Keep the server API payloads exactly fixed across candidates. If you implement a custom runner, cap workers and request count in constants—do not generate unbounded load.

Search request A · lexical + filters
GET atlasmart-runtime-lab-v1/_search
{
  "size": 10,
  "query": {
    "bool": {
      "must": [{"match":{"name":"wireless keyboard"}}],
      "filter": [
        {"term":{"available":true}},
        {"range":{"price":{"lte":150}}}
      ]
    }
  },
  "sort": [{"_score":"desc"},{"sku":"asc"}]
}
Search request B · bounded analytical facet
GET atlasmart-runtime-lab-v1/_search?request_cache=true
{
  "size": 0,
  "query":{"term":{"available":true}},
  "aggs":{
    "category":{"terms":{"field":"category","size":10}},
    "price":{"percentiles":{"field":"price","percents":[50,95]}}
  }
}
  • Baseline: concurrency 1, fixed request count, warm-up then measurement.
  • Candidate pressure step: increase to the next bounded concurrency only if stop conditions remain clear.
  • Safe saturation signal may be rising queue/p99 or a controlled 429/rejection; you do not need an OOM or breaker trip.
  • Stop before host swapping, container kill risk, prolonged GC, unrelated impact or runaway retries.
  • After cooldown, apply one repair—e.g., cap client concurrency—and rerun the identical mix.

6. Calculate latency percentiles without pretending they are universal targets

Sort the measured client durations. The 95th percentile approximates the latency below which 95% of requests completed; p99 focuses further into the tail. Use the same percentile method and sample size for baseline/candidate. Do not publish a universal “Elasticsearch p99 should be X ms”: acceptable latency depends on query, data, topology, hardware and product requirements.

Throughput alone is insufficient. A candidate that increases operations/second by 5% while doubling p99 and introducing rejections may be worse for an interactive service. Conversely, a batch pipeline may accept higher latency if useful throughput and recovery behavior improve within its SLA.

7. Change one variable, then try to falsify yourself

Suppose concurrency 8 creates search queue growth and 429s while CPU is high, but concurrency 4 has similar useful throughput and much lower p99. Set the application limit to 4 for the lab and rerun. If queues/rejections disappear and throughput stays similar, the evidence supports client admission control as part of the repair.

Now try to falsify it: repeat after a restart/cold cache, after enough time for merge/background work to settle, and with the same data/query mix. If the result depends entirely on warmed filesystem cache, the conclusion must say so. If a hot shard dominates, revisit Chapter 13 routing rather than attributing everything to concurrency.

8. What to scale and what to tune

Finding Primary durable actions
Bad query/mapping shape Fix query, mapping/analyzer, aggregation/pagination strategy; add regression tests.
Client-driven overload Bound concurrency/bulk size; backoff+jitter; queue/shed/defer noncritical work.
Hot tenant/shard Routing/model fix, tenant budgets/isolation, rebalance/reindex if justified.
Legitimate CPU-bound capacity shortage Scale/resize after query efficiency and admission controls are validated.
Storage/merge bottleneck Storage/headroom/lifecycle/refresh/merge investigation from Chapter 14; do not hide with heap.
Heap-heavy legitimate working set Mapping/cache/query repair first; then supported heap/capacity evaluation with filesystem-cache headroom preserved.

9. Operational and security boundaries

Runtime evidence can contain query strings, tenant names, hostnames and potentially sensitive fields. Restrict node/task/hot-thread APIs to operators, store diagnostics with access control, and redact before external sharing. On managed Elastic/Amazon/OpenSearch offerings, some host/JVM controls are intentionally abstracted; use provider metrics and supported knobs rather than trying to reproduce self-managed internals exactly.

A node replica is not a backup. Capacity tuning must preserve snapshot/restore and recovery objectives. A “fast” cluster that cannot recover within its failure budget or has no disk/memory headroom is not production-ready.

Production judgment and bridge to Chapter 16

A credible saturation diagnosis says what changed, which resource became constrained, how the mechanism produced the client symptom, what evidence supports it, what evidence would disprove it, and how the repair was verified. It does not end with a generic tuning checklist.

Chapter 16 moves from node-level saturation into search execution itself: distributed query/fetch phases, profiling, slow logs, tasks, deep pagination, search_after, point-in-time readers and cancellation. The runtime evidence learned here becomes the context for interpreting those search-specific tools.

Check your understanding

  1. Why start with the client-visible symptom?
  2. What makes a hypothesis falsifiable?
  3. What is a safe saturation signal for the local lab?
  4. Why compare p99 and throughput together?
  5. When is scaling preferable to tuning?
Review the answers

1. It anchors diagnosis in the service objective and prevents one internal metric from being assumed causal before time-aligned evidence exists.

2. It names a mechanism and predicts observations that should change—or fail to change—when one controlled variable is altered.

3. Rising queue/tail latency or a bounded rejection under capped load; an OOM or disabled safety control is not required.

4. More throughput can be purchased by harmful queueing/rejection tail latency; service quality requires both plus correctness/errors.

5. After query/model/admission problems are addressed and legitimate demand still consumes the measured capacity needed to meet SLOs with recovery headroom.

Summary and next step

You can now separate heap, filesystem cache, logical caches, breaker pressure, queues, indexing/search backpressure and client concurrency, then connect them to synchronized workload and tail-latency evidence. That is the runtime foundation for Chapter 16’s query/fetch profiling and pagination work.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.