Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health

Heap vs Filesystem Cache, Object Overhead, Compressed OOPs Concepts, and Why More Heap Is Not Always Better

Separate JVM heap, native/off-heap memory and the operating-system filesystem cache, then reason from GC, object representation and Lucene I/O evidence instead of “more heap is always faster.”

Intermediate115–150 minutesJVM/runtime health & bounded saturation labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart has a 16 GiB development host. A well-intentioned operator raises the search-node heap from 4 GiB toward nearly all available RAM because “the index is large.” GC pauses increase, the operating-system page cache shrinks, and search tail latency becomes less predictable. Nothing about the product catalog mapping changed. The failure is a memory-partitioning mistake: JVM heap and the filesystem cache solve different problems.

This lesson builds a runtime mental model before touching any tuning knob. JVM heap stores Java objects such as request state, aggregation structures, caches and cluster metadata. Native/off-heap memory includes JVM/native runtime needs and some buffers. The filesystem cache is RAM managed by the operating system for file pages—critically, Lucene segment data. Giving the heap more RAM can take RAM away from Lucene file access.

01

Separate JVM heap, native/off-heap memory and the operating-system filesystem cache.

02

Explain why Java object overhead and compressed ordinary object pointers affect heap efficiency without turning a JVM implementation detail into a magic sizing rule.

03

Read node JVM/process/OS evidence and distinguish pressure from healthy utilization.

04

Relate Chapter 14 Lucene file access to filesystem-cache headroom and GC behavior.

05

Design a bounded heap experiment that changes one variable and evaluates tail latency, GC and cache effects together.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual JVM/runtime with GET _nodes/jvm and GET _nodes/stats/jvm,process,os rather than hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly changes topology. OpenSearch demo -k TLS bypass remains disposable-lab-only; production must validate certificates. No moving latest tags are used.

Execution note

This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.

1. Three memory domains, three different jobs

Memory domain What it serves Typical evidence Wrong conclusion
JVM heap Java objects: query/aggregation state, mappings/cluster state, fielddata/global ordinals, caches, task/accounting structures. _nodes/stats/jvm, GC collections/time, heap used/max. “All free RAM should become heap.”
Native/off-heap/process memory JVM runtime, direct/network buffers, thread stacks, memory mappings and other process allocations. _nodes/stats/process,os, cgroup/container metrics, OS process metrics. “RSS above Xmx means a leak.”
OS filesystem cache Recently/frequently accessed Lucene files and other filesystem pages. Host page-cache/storage metrics plus cold/warm search behavior. “The filesystem cache is redundant because Elasticsearch/OpenSearch has query caches.”

Lucene is designed to benefit from the operating system cache. That cache is not a duplicate of the Query DSL result caches: it caches file pages, not logical query answers. A node with a giant heap can therefore have more Java space yet slower disk-backed search because fewer useful Lucene pages remain resident.

Elastic 9.5 recommends automatic heap sizing for most deployments and documents 50% of available memory as an upper bound, not a target; smaller heaps can perform better when they leave more room for filesystem cache. OpenSearch documentation commonly uses half of system memory as a starting point for a dedicated node, but also states that workload requirements matter. Neither product offers a universal percentage that substitutes for measurement.

2. Observe before changing heap

Dev Tools · runtime identity and memory evidence
GET _nodes/jvm
GET _nodes/stats/jvm,process,os?human
GET _nodes/stats/indices,breaker,thread_pool?human
GET _nodes/hot_threads?threads=3&ignore_idle_threads=true

Record the node name, process/JVM identity, heap_used_in_bytes, heap_max_in_bytes, GC counters, process CPU, breaker estimates/trips and the workload window that produced them. A single high heap sample is not automatically a problem; a causal diagnosis needs a time relationship with GC, latency, rejection or allocation symptoms.

GC is garbage collection: the JVM reclaims objects that are no longer reachable. More heap can reduce collection frequency in some workloads but can also increase the amount of live data and the cost/variability of collections. Do not infer collector health from one counter; use trends and the bundled runtime’s logs/metrics.

3. Object overhead and compressed ordinary object pointers

Java objects consume more than their payload: object headers, references, alignment and container structures add overhead. Search workloads can create many small objects for coordination, scripts, aggregation state and response construction. That is why “we only cache 3 GiB of values” does not imply 3 GiB of heap cost.

Compressed ordinary object pointers (compressed oops) are a HotSpot JVM technique that can encode object references more compactly when the heap/layout allows it. Elastic documents a conservative threshold—26 GiB is safe on most systems and the actual threshold can be higher—and tells operators to verify the runtime rather than assume a fixed boundary. The correct lesson is not “always set 26 GiB”; it is “check whether compressed oops are active and evaluate the whole node.”

Elasticsearch · verify rather than guess
GET _nodes/_all/jvm

# Look for the per-node JVM field:
# jvm.using_compressed_ordinary_object_pointers : true|false
# Also inspect Elasticsearch startup logs for the compressed-oops status.
OpenSearch/JVM boundary

Compressed references are a JVM implementation concern, not an Elasticsearch-only search feature. OpenSearch 3.8 ships with a supported JDK in its distribution; inspect the actual runtime and JVM flags/logs. Do not copy an Elasticsearch-specific heap threshold into OpenSearch as a guaranteed product contract.

4. Wrong approach: maximize heap because RAM is available

Suppose AtlasMart gives an 8 GiB container 7 GiB of heap. The operator expects fewer cache misses. Instead, the process and OS have little room for native memory and filesystem cache. Lucene file reads become more dependent on storage, and container/host memory pressure may terminate the process even though the JVM itself has not exceeded Xmx.

The repair is evidence-driven: restore a bounded heap, preserve host/container headroom, run the same query/indexing mix, and compare p95/p99 latency, throughput, GC time/frequency, process memory and storage behavior. The useful outcome is not “smaller heap always wins”; it is a measured operating region for this workload and machine.

Do not benchmark one warm run

A warmed filesystem cache can hide storage cost, while a cold run can exaggerate startup/recovery behavior. Report whether the test is cold, warming or steady-state and preserve the same state when comparing candidates.

5. AtlasMart bounded memory experiment

Use the disposable local runtime lab only. Keep the existing production-like AtlasMart indices untouched. If your container runtime supports memory limits, give each test container the same total memory limit and change only the heap between candidate runs. Restarting changes filesystem/JVM warmness, so include a repeatable warm-up phase and then a fixed measurement phase.

For Elasticsearch use the documented container/JVM configuration for the pinned 9.5.3 image and its generated CA. For OpenSearch 3.8.0 use the pinned image and a local-only demo certificate. Do not run both candidate configurations simultaneously if that changes host memory pressure.

Create the small reusable runtime fixture
PUT atlasmart-runtime-lab-v1
{
  "settings": {"number_of_shards":1,"number_of_replicas":0},
  "mappings": {"properties": {
    "sku":{"type":"keyword"},
    "category":{"type":"keyword"},
    "name":{"type":"text","fields":{"raw":{"type":"keyword"}}},
    "price":{"type":"scaled_float","scaling_factor":100},
    "available":{"type":"boolean"},
    "event_time":{"type":"date"}
  }}
}

# Reuse deterministic AtlasMart documents from earlier chapters or load your
# local synthetic fixture. Keep document IDs and query set identical between runs.
Evidence snapshot before/after each run
GET _nodes/stats/jvm,process,os,breaker,thread_pool?human
GET _nodes/stats/indices/query_cache,request_cache,fielddata?human
GET atlasmart-runtime-lab-v1/_stats/search,indexing,request_cache,query_cache,fielddata?human
GET _nodes/hot_threads?threads=3&ignore_idle_threads=true
  • Warm using the same finite query set; record when measurement begins.
  • Run the same finite mixed search/indexing request count and the same client concurrency.
  • Record client p50/p95/p99, throughput and HTTP/transport errors or 429 responses.
  • Capture JVM/GC, process/OS, breaker and thread-pool deltas over the same window.
  • Change only heap allocation, repeat after equivalent warm-up, and compare directions—not fabricated target numbers.
  • Stop immediately if the host swaps heavily, the container is OOM-killed, breaker trips cascade, or unrelated workloads are affected.

6. Production judgment

Heap sizing is a resource-allocation decision among competing consumers, not a cache-maximization contest. More heap may help a heap-bound workload; it may also reduce filesystem-cache effectiveness, lengthen GC effects, or crowd native/process memory. The safest default is the product’s supported automatic/default sizing unless measurements show a reason to change it.

Tail latency matters because user experience and timeout/retry behavior are driven by slow requests, not averages. Measure p95/p99 together with throughput and errors. If a smaller heap improves steady-state search but causes breaker trips during a known aggregation peak, the answer may be workload isolation, query repair or additional capacity—not simply reversing to an oversized heap.

Check your understanding

  1. Why can increasing heap make Lucene search slower?
  2. What does compressed-oops status tell you?
  3. Why is resident process memory allowed to exceed Xmx?
  4. What evidence makes a heap change credible?
  5. When should you keep default/automatic heap sizing?
Review the answers

1. It can reduce RAM available to the operating-system filesystem cache, increasing storage reads for Lucene segment pages even though Java has more memory.

2. It tells you how the current HotSpot runtime represents object references; it is one heap-efficiency clue, not a universal heap target.

3. The process also uses native/off-heap memory, thread stacks, memory mappings and other JVM/runtime allocations outside the Java heap.

4. The same workload and total-memory envelope, comparable warmness, client latency/throughput/errors, JVM/GC and process/OS evidence, with one changed variable.

5. When there is no measured workload-specific problem that a supported change demonstrably improves without harming other objectives.

Summary and next step

You now have the memory partition: heap for Java objects, native/off-heap for process/runtime needs, and filesystem cache for Lucene file pages. Lesson 2 separates the logical caches that live on or interact with heap so “cache” stops being one ambiguous tuning word.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.