Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health
Heap vs Filesystem Cache, Object Overhead, Compressed OOPs Concepts, and Why More Heap Is Not Always Better
Separate JVM heap, native/off-heap memory and the operating-system filesystem cache, then reason from GC, object representation and Lucene I/O evidence instead of “more heap is always faster.”
Learning outcomes
AtlasMart has a 16 GiB development host. A well-intentioned operator raises the search-node heap from 4 GiB toward nearly all available RAM because “the index is large.” GC pauses increase, the operating-system page cache shrinks, and search tail latency becomes less predictable. Nothing about the product catalog mapping changed. The failure is a memory-partitioning mistake: JVM heap and the filesystem cache solve different problems.
This lesson builds a runtime mental model before touching any tuning knob. JVM heap stores Java objects such as request state, aggregation structures, caches and cluster metadata. Native/off-heap memory includes JVM/native runtime needs and some buffers. The filesystem cache is RAM managed by the operating system for file pages—critically, Lucene segment data. Giving the heap more RAM can take RAM away from Lucene file access.
Separate JVM heap, native/off-heap memory and the operating-system filesystem cache.
Explain why Java object overhead and compressed ordinary object pointers affect heap efficiency without turning a JVM implementation detail into a magic sizing rule.
Read node JVM/process/OS evidence and distinguish pressure from healthy utilization.
Relate Chapter 14 Lucene file access to filesystem-cache headroom and GC behavior.
Design a bounded heap experiment that changes one variable and evaluates tail latency, GC and cache effects together.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual JVM/runtime with
GET _nodes/jvm and
GET _nodes/stats/jvm,process,os rather than
hard-coding a JDK patch. Labs use one primary and zero
replicas unless a step explicitly changes topology. OpenSearch
demo -k TLS bypass remains disposable-lab-only;
production must validate certificates. No moving
latest tags are used.
This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.
1. Three memory domains, three different jobs
| Memory domain | What it serves | Typical evidence | Wrong conclusion |
|---|---|---|---|
| JVM heap | Java objects: query/aggregation state, mappings/cluster state, fielddata/global ordinals, caches, task/accounting structures. |
_nodes/stats/jvm, GC collections/time, heap
used/max.
|
“All free RAM should become heap.” |
| Native/off-heap/process memory | JVM runtime, direct/network buffers, thread stacks, memory mappings and other process allocations. |
_nodes/stats/process,os, cgroup/container
metrics, OS process metrics.
|
“RSS above Xmx means a leak.” |
| OS filesystem cache | Recently/frequently accessed Lucene files and other filesystem pages. | Host page-cache/storage metrics plus cold/warm search behavior. | “The filesystem cache is redundant because Elasticsearch/OpenSearch has query caches.” |
Lucene is designed to benefit from the operating system cache. That cache is not a duplicate of the Query DSL result caches: it caches file pages, not logical query answers. A node with a giant heap can therefore have more Java space yet slower disk-backed search because fewer useful Lucene pages remain resident.
Elastic 9.5 recommends automatic heap sizing for most deployments and documents 50% of available memory as an upper bound, not a target; smaller heaps can perform better when they leave more room for filesystem cache. OpenSearch documentation commonly uses half of system memory as a starting point for a dedicated node, but also states that workload requirements matter. Neither product offers a universal percentage that substitutes for measurement.
2. Observe before changing heap
GET _nodes/jvm
GET _nodes/stats/jvm,process,os?human
GET _nodes/stats/indices,breaker,thread_pool?human
GET _nodes/hot_threads?threads=3&ignore_idle_threads=true
Record the node name, process/JVM identity,
heap_used_in_bytes, heap_max_in_bytes,
GC counters, process CPU, breaker estimates/trips and the
workload window that produced them. A single high heap sample is
not automatically a problem; a causal diagnosis needs a time
relationship with GC, latency, rejection or allocation symptoms.
GC is garbage collection: the JVM reclaims objects that are no longer reachable. More heap can reduce collection frequency in some workloads but can also increase the amount of live data and the cost/variability of collections. Do not infer collector health from one counter; use trends and the bundled runtime’s logs/metrics.
3. Object overhead and compressed ordinary object pointers
Java objects consume more than their payload: object headers, references, alignment and container structures add overhead. Search workloads can create many small objects for coordination, scripts, aggregation state and response construction. That is why “we only cache 3 GiB of values” does not imply 3 GiB of heap cost.
Compressed ordinary object pointers (compressed oops) are a HotSpot JVM technique that can encode object references more compactly when the heap/layout allows it. Elastic documents a conservative threshold—26 GiB is safe on most systems and the actual threshold can be higher—and tells operators to verify the runtime rather than assume a fixed boundary. The correct lesson is not “always set 26 GiB”; it is “check whether compressed oops are active and evaluate the whole node.”
GET _nodes/_all/jvm
# Look for the per-node JVM field:
# jvm.using_compressed_ordinary_object_pointers : true|false
# Also inspect Elasticsearch startup logs for the compressed-oops status.
Compressed references are a JVM implementation concern, not an Elasticsearch-only search feature. OpenSearch 3.8 ships with a supported JDK in its distribution; inspect the actual runtime and JVM flags/logs. Do not copy an Elasticsearch-specific heap threshold into OpenSearch as a guaranteed product contract.
4. Wrong approach: maximize heap because RAM is available
Suppose AtlasMart gives an 8 GiB container 7 GiB of heap. The
operator expects fewer cache misses. Instead, the process and OS
have little room for native memory and filesystem cache. Lucene
file reads become more dependent on storage, and container/host
memory pressure may terminate the process even though the JVM
itself has not exceeded Xmx.
The repair is evidence-driven: restore a bounded heap, preserve host/container headroom, run the same query/indexing mix, and compare p95/p99 latency, throughput, GC time/frequency, process memory and storage behavior. The useful outcome is not “smaller heap always wins”; it is a measured operating region for this workload and machine.
A warmed filesystem cache can hide storage cost, while a cold run can exaggerate startup/recovery behavior. Report whether the test is cold, warming or steady-state and preserve the same state when comparing candidates.
5. AtlasMart bounded memory experiment
Use the disposable local runtime lab only. Keep the existing production-like AtlasMart indices untouched. If your container runtime supports memory limits, give each test container the same total memory limit and change only the heap between candidate runs. Restarting changes filesystem/JVM warmness, so include a repeatable warm-up phase and then a fixed measurement phase.
For Elasticsearch use the documented container/JVM configuration for the pinned 9.5.3 image and its generated CA. For OpenSearch 3.8.0 use the pinned image and a local-only demo certificate. Do not run both candidate configurations simultaneously if that changes host memory pressure.
PUT atlasmart-runtime-lab-v1
{
"settings": {"number_of_shards":1,"number_of_replicas":0},
"mappings": {"properties": {
"sku":{"type":"keyword"},
"category":{"type":"keyword"},
"name":{"type":"text","fields":{"raw":{"type":"keyword"}}},
"price":{"type":"scaled_float","scaling_factor":100},
"available":{"type":"boolean"},
"event_time":{"type":"date"}
}}
}
# Reuse deterministic AtlasMart documents from earlier chapters or load your
# local synthetic fixture. Keep document IDs and query set identical between runs.
GET _nodes/stats/jvm,process,os,breaker,thread_pool?human
GET _nodes/stats/indices/query_cache,request_cache,fielddata?human
GET atlasmart-runtime-lab-v1/_stats/search,indexing,request_cache,query_cache,fielddata?human
GET _nodes/hot_threads?threads=3&ignore_idle_threads=true
- Warm using the same finite query set; record when measurement begins.
- Run the same finite mixed search/indexing request count and the same client concurrency.
- Record client p50/p95/p99, throughput and HTTP/transport errors or 429 responses.
- Capture JVM/GC, process/OS, breaker and thread-pool deltas over the same window.
- Change only heap allocation, repeat after equivalent warm-up, and compare directions—not fabricated target numbers.
- Stop immediately if the host swaps heavily, the container is OOM-killed, breaker trips cascade, or unrelated workloads are affected.
6. Production judgment
Heap sizing is a resource-allocation decision among competing consumers, not a cache-maximization contest. More heap may help a heap-bound workload; it may also reduce filesystem-cache effectiveness, lengthen GC effects, or crowd native/process memory. The safest default is the product’s supported automatic/default sizing unless measurements show a reason to change it.
Tail latency matters because user experience and timeout/retry behavior are driven by slow requests, not averages. Measure p95/p99 together with throughput and errors. If a smaller heap improves steady-state search but causes breaker trips during a known aggregation peak, the answer may be workload isolation, query repair or additional capacity—not simply reversing to an oversized heap.
Check your understanding
- Why can increasing heap make Lucene search slower?
- What does compressed-oops status tell you?
- Why is resident process memory allowed to exceed Xmx?
- What evidence makes a heap change credible?
- When should you keep default/automatic heap sizing?
Review the answers
1. It can reduce RAM available to the operating-system filesystem cache, increasing storage reads for Lucene segment pages even though Java has more memory.
2. It tells you how the current HotSpot runtime represents object references; it is one heap-efficiency clue, not a universal heap target.
3. The process also uses native/off-heap memory, thread stacks, memory mappings and other JVM/runtime allocations outside the Java heap.
4. The same workload and total-memory envelope, comparable warmness, client latency/throughput/errors, JVM/GC and process/OS evidence, with one changed variable.
5. When there is no measured workload-specific problem that a supported change demonstrably improves without harming other objectives.
Summary and next step
You now have the memory partition: heap for Java objects, native/off-heap for process/runtime needs, and filesystem cache for Lucene file pages. Lesson 2 separates the logical caches that live on or interact with heap so “cache” stops being one ambiguous tuning word.
Authoritative references
- Elastic JVM settings — Automatic heap sizing, filesystem-cache headroom, container memory, and compressed ordinary object pointer guidance.
- Elastic node query cache settings — Filter-context query-cache eligibility, segment-level caching and invalidation behavior.
- Elastic shard request cache — Shard-level request-result caching and request/index controls.
- Elastic field data cache settings — On-heap fielddata/global-ordinal cache behavior and breaker interaction.
- Elastic circuit breaker settings — Parent, fielddata, request and in-flight breaker semantics.
- Elastic thread pool settings — Current thread-pool types, sizing and queue behavior.
- Elastic indexing pressure — Outstanding indexing-byte accounting and rejection behavior.
- Elastic Nodes Stats API — JVM, process, caches, breakers, thread pools and indexing-pressure observations.
- OpenSearch circuit breaker settings — Parent/child breaker settings and failure-protection intent.
- OpenSearch caching overview — On-heap request/query/fielddata caching concepts.
- OpenSearch index request cache — Shard-level request-result caching, invalidation and metrics.
- OpenSearch field data cache — Fielddata/global ordinals and the keyword-subfield alternative.
- OpenSearch thread pool settings — Thread pools, queue sizing and monitoring guidance.
- OpenSearch search backpressure — OpenSearch-specific search task resource tracking and cancellation.
- OpenSearch shard indexing backpressure — OpenSearch-specific per-shard indexing pressure and rejection mechanism.
- OpenSearch Nodes Stats API — JVM, process, caches, indexing pressure and backpressure statistics.
- OpenSearch Nodes Hot Threads API — CPU/wait/block thread samples for diagnosis.