Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health
Query/Request/Fielddata/Shard Request Caches: Eligibility, Invalidation, and Memory Tradeoffs
Distinguish filter/query caching, shard-level request caching, fielddata/global ordinals and the filesystem cache, and measure whether a cache helps the AtlasMart workload rather than optimizing hit rate in isolation.
Learning outcomes
AtlasMart dashboards repeat a size-zero sales aggregation and
become very fast after the first run. A product-search page
repeats the same lexical filter but still performs work after
every refresh. Another engineer enables fielddata on
name because a terms aggregation failed, and heap
consumption spikes. All three observations mention “cache,” but
they are different mechanisms with different eligibility and
invalidation rules.
Distinguish filesystem cache, query/filter cache, shard/index request cache and fielddata/global ordinals.
Explain cache eligibility, invalidation and segment/refresh relationships for current Elasticsearch/OpenSearch behavior.
Use cache hit/miss/eviction statistics as evidence without treating hit rate as the business objective.
Avoid fielddata on high-cardinality analyzed text by modeling sortable/aggregatable values explicitly.
Design repeatable cache experiments that preserve query body, refresh state, warmness and response semantics.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual JVM/runtime with
GET _nodes/jvm and
GET _nodes/stats/jvm,process,os rather than
hard-coding a JDK patch. Labs use one primary and zero
replicas unless a step explicitly changes topology. OpenSearch
demo -k TLS bypass remains disposable-lab-only;
production must validate certificates. No moving
latest tags are used.
This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.
1. A cache taxonomy that prevents category errors
| Mechanism | What is reused | Scope / storage | Typical invalidation or turnover |
|---|---|---|---|
| OS filesystem cache | File pages such as Lucene segment data. | Operating-system RAM; outside JVM heap. | Memory pressure, file lifecycle, host/container behavior—not Query DSL invalidation. |
| Query/filter cache | Reusable filter-context query results/bit sets. | Node/shard/segment-oriented internal cache; on heap. | Segment lifecycle/merges and cache policy/eviction. |
| Shard/index request cache | A shard-level result for a complete search request, especially aggregation-style requests. | Per shard/index; heap-backed unless a product-specific tiered cache is configured. | Refresh/data changes, settings changes, LRU/size and request eligibility. |
| Fielddata/global ordinals | Per-document field access/global term ordinals used for sorting/aggregations in relevant mappings. | Node heap. | Field changes/refresh and cache eviction/rebuild depending on structure; breaker protects growth. |
These mechanisms can all improve latency, but they cache different things. Clearing a request cache does not evict the OS page cache. Warming Lucene file pages does not make an ineligible request-cache query eligible. A fielddata problem cannot be repaired by increasing request-cache size.
2. Query cache: filter context, repeated structure, segment lifecycle
In Elasticsearch, the node query cache caches eligible queries used in filter context. Queries used for scoring are not automatically interchangeable with cached filters. Eligibility also depends on repetition and segment characteristics. Because results are tied to segments, merges can invalidate entries.
OpenSearch also exposes query-cache statistics/settings and uses cacheable query/filter execution as an optimization. Treat the exact cache policy as product/version behavior, not an API guarantee that every repeated filter will be cached.
GET _nodes/stats/indices/query_cache?human
GET atlasmart-runtime-lab-v1/_stats/query_cache?human
GET atlasmart-runtime-lab-v1/_search
{
"size": 0,
"query": {"bool":{"filter":[
{"term":{"available":true}},
{"term":{"category":"electronics"}}
]}},
"aggs":{"price":{"stats":{"field":"price"}}}
}
Run the same request repeatedly and inspect cache statistics. The engine may decide a query is not worth caching, or the fixture may be too small to cross internal eligibility thresholds. “Miss” is an observation, not a lab failure.
3. Shard/index request cache: whole-result reuse
Elasticsearch calls this the shard request cache. OpenSearch documentation calls the corresponding feature the index request cache, although the cached result is still shard-local. The cache is especially valuable for repeatable aggregation/dashboard requests because a shard can reuse a prior local result and the coordinating node can reduce those cached shard results.
By default, size-zero aggregation-like requests are the common
cacheable case. Product/version rules differ for non-zero hit
sizes and non-deterministic constructs. Relative time such as
now, profiling, scroll/DFS modes or script
nondeterminism can remove or reduce cacheability. Preserve the
exact request body when measuring because the body participates
in cache identity.
GET atlasmart-runtime-lab-v1/_search?request_cache=true
{
"size": 0,
"query":{"term":{"available":true}},
"aggs":{"by_category":{"terms":{"field":"category","size":10}}}
}
GET atlasmart-runtime-lab-v1/_stats/request_cache?human
GET _nodes/stats/indices/request_cache?human
Repeat without indexing between runs, then index one document and wait for/issue the refresh required by your experiment. Observe that data-changing refresh semantics can invalidate request-cache entries. Do not claim that a cache hit returns stale search data; the cache lifecycle is integrated with index refresh/change visibility.
4. Fielddata and global ordinals: not a generic “make text aggregatable” switch
Analyzed text is optimized for full-text matching
through an inverted index. Sorting or terms aggregation needs
efficient per-document values. Enabling
fielddata on high-cardinality text asks the engine
to build an on-heap structure by uninverting text terms, which
can consume substantial heap. Both products therefore default
fielddata off for text and recommend modeling an exact-value
keyword subfield for sorting/aggregation.
Global ordinals map per-segment term ordinals into a shard-wide view for terms-like operations. They can be built lazily after refresh or eagerly when the mapping requests it. Eager construction can move latency from the first aggregation into refresh/indexing work; it is a latency-placement decision, not free speed.
"name": {
"type": "text",
"fields": {
"raw": {"type": "keyword", "ignore_above": 256}
}
}
# Aggregate/sort on name.raw, search full text on name.
# Do NOT set fielddata:true on name merely to make a demo aggregation pass.
GET _cat/fielddata?v&bytes=mb
GET _nodes/stats/indices/fielddata?human
GET atlasmart-runtime-lab-v1/_stats/fielddata?human
5. Wrong approach: optimize cache hit rate as the goal
A team disables refreshes and canonicalizes every dashboard request until request-cache hit rate is excellent. Unfortunately, users now see stale data beyond the product requirement. Another team increases cache sizes until heap pressure causes GC and breaker risk. Both optimized an internal metric while degrading the service.
Cache metrics are diagnostic inputs. The service objective is correct results within freshness, latency, throughput, cost and failure-isolation constraints. A lower hit rate may be acceptable for highly dynamic queries; a high hit rate is useless if the cached work is cheap or if it displaces more valuable memory.
Repeatedly clearing caches can create rebuild storms and hide the actual working-set problem. Use it only in a controlled experiment where cold/warm state is the variable under test.
6. AtlasMart cache experiment
-
Capture
_stats/query_cache,request_cache,fielddataand node JVM stats. - Run a fixed filter+aggregation request N times without writes; record client latency distribution and cache deltas.
- Run a semantically similar request with an intentionally non-cacheable/changed element and compare.
- Index one known document and refresh; rerun the first request and inspect request-cache behavior.
-
Compare aggregation on
keywordto the prohibited idea of fielddata on text; do not enable fielddata in the shared fixture. - Interpret whether any latency difference is due to logical result caching, filesystem warmness, both, or neither.
You should be able to point to cache hit/miss/eviction counters and a controlled request sequence. Exact counts and latency depend on product version, fixture size, refreshes, segment topology and previous activity.
7. Elastic vs OpenSearch details that matter
| Area | Elasticsearch 9.5.3 | OpenSearch 3.8.0 |
|---|---|---|
| Query cache | Node query cache documented for filter-context results with per-segment eligibility/history. | Query cache exposed through index/node settings/stats; optimize by measurement rather than assuming identical policy thresholds. |
| Request cache | Shard request cache; default size-zero requests are the common auto-cache case, with per-request override. |
Index request cache; 2.19+ adds
indices.requests.cache.maximum_cacheable_size, defaulting to zero-size automatic caching.
|
| Fielddata | On-heap fielddata/global ordinals; fielddata breaker protects growth. | Same fundamental hazard; docs explicitly recommend keyword subfields instead of text fielddata. |
| Tiered cache | Do not assume OpenSearch cache-tier settings exist. | OpenSearch can expose tiered-request-cache statistics/features in supported configurations; treat as product-specific. |
Production judgment
If a cache is the only thing keeping a query within the latency objective, verify what happens after refresh, relocation, node restart and eviction. That is the real miss path users will eventually experience. Capacity planning must include both warm and recovery/cold behavior.
A cache that avoids expensive repeated work can be valuable, but the next lesson covers the safety layer underneath: circuit breakers exist because some requests can allocate more heap than caches and ordinary GC can safely absorb.
Check your understanding
- What does the filesystem cache store that the request cache does not?
- Why can a merge affect query-cache entries?
- Why is fielddata on text dangerous?
- Why can eager global ordinals increase indexing/refresh work?
- What is a better success metric than cache hit rate?
Review the answers
1. Filesystem pages such as Lucene segment data; the request cache stores a shard-level logical search result.
2. Query-cache entries are associated with segment execution; merging changes the segment set.
3. It can build large on-heap per-document structures from analyzed terms, especially with high cardinality, increasing GC and breaker risk.
4. They move ordinal construction earlier into refresh rather than lazily paying it on the first aggregation.
5. End-to-end correctness/freshness plus latency, throughput, error rate and resource cost for the actual workload.
Summary and next step
You can now say exactly which cache you mean, what it reuses, when it invalidates and which resource it consumes. Lesson 3 uses that vocabulary to explain circuit breakers, aggregation memory, in-flight requests and why safety limits should not be disabled to make one query succeed.
Authoritative references
- Elastic JVM settings — Automatic heap sizing, filesystem-cache headroom, container memory, and compressed ordinary object pointer guidance.
- Elastic node query cache settings — Filter-context query-cache eligibility, segment-level caching and invalidation behavior.
- Elastic shard request cache — Shard-level request-result caching and request/index controls.
- Elastic field data cache settings — On-heap fielddata/global-ordinal cache behavior and breaker interaction.
- Elastic circuit breaker settings — Parent, fielddata, request and in-flight breaker semantics.
- Elastic thread pool settings — Current thread-pool types, sizing and queue behavior.
- Elastic indexing pressure — Outstanding indexing-byte accounting and rejection behavior.
- Elastic Nodes Stats API — JVM, process, caches, breakers, thread pools and indexing-pressure observations.
- OpenSearch circuit breaker settings — Parent/child breaker settings and failure-protection intent.
- OpenSearch caching overview — On-heap request/query/fielddata caching concepts.
- OpenSearch index request cache — Shard-level request-result caching, invalidation and metrics.
- OpenSearch field data cache — Fielddata/global ordinals and the keyword-subfield alternative.
- OpenSearch thread pool settings — Thread pools, queue sizing and monitoring guidance.
- OpenSearch search backpressure — OpenSearch-specific search task resource tracking and cancellation.
- OpenSearch shard indexing backpressure — OpenSearch-specific per-shard indexing pressure and rejection mechanism.
- OpenSearch Nodes Stats API — JVM, process, caches, indexing pressure and backpressure statistics.
- OpenSearch Nodes Hot Threads API — CPU/wait/block thread samples for diagnosis.