Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health

Query/Request/Fielddata/Shard Request Caches: Eligibility, Invalidation, and Memory Tradeoffs

Distinguish filter/query caching, shard-level request caching, fielddata/global ordinals and the filesystem cache, and measure whether a cache helps the AtlasMart workload rather than optimizing hit rate in isolation.

Intermediate115–150 minutesJVM/runtime health & bounded saturation labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart dashboards repeat a size-zero sales aggregation and become very fast after the first run. A product-search page repeats the same lexical filter but still performs work after every refresh. Another engineer enables fielddata on name because a terms aggregation failed, and heap consumption spikes. All three observations mention “cache,” but they are different mechanisms with different eligibility and invalidation rules.

01

Distinguish filesystem cache, query/filter cache, shard/index request cache and fielddata/global ordinals.

02

Explain cache eligibility, invalidation and segment/refresh relationships for current Elasticsearch/OpenSearch behavior.

03

Use cache hit/miss/eviction statistics as evidence without treating hit rate as the business objective.

04

Avoid fielddata on high-cardinality analyzed text by modeling sortable/aggregatable values explicitly.

05

Design repeatable cache experiments that preserve query body, refresh state, warmness and response semantics.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual JVM/runtime with GET _nodes/jvm and GET _nodes/stats/jvm,process,os rather than hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly changes topology. OpenSearch demo -k TLS bypass remains disposable-lab-only; production must validate certificates. No moving latest tags are used.

Execution note

This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.

1. A cache taxonomy that prevents category errors

Mechanism What is reused Scope / storage Typical invalidation or turnover
OS filesystem cache File pages such as Lucene segment data. Operating-system RAM; outside JVM heap. Memory pressure, file lifecycle, host/container behavior—not Query DSL invalidation.
Query/filter cache Reusable filter-context query results/bit sets. Node/shard/segment-oriented internal cache; on heap. Segment lifecycle/merges and cache policy/eviction.
Shard/index request cache A shard-level result for a complete search request, especially aggregation-style requests. Per shard/index; heap-backed unless a product-specific tiered cache is configured. Refresh/data changes, settings changes, LRU/size and request eligibility.
Fielddata/global ordinals Per-document field access/global term ordinals used for sorting/aggregations in relevant mappings. Node heap. Field changes/refresh and cache eviction/rebuild depending on structure; breaker protects growth.

These mechanisms can all improve latency, but they cache different things. Clearing a request cache does not evict the OS page cache. Warming Lucene file pages does not make an ineligible request-cache query eligible. A fielddata problem cannot be repaired by increasing request-cache size.

2. Query cache: filter context, repeated structure, segment lifecycle

In Elasticsearch, the node query cache caches eligible queries used in filter context. Queries used for scoring are not automatically interchangeable with cached filters. Eligibility also depends on repetition and segment characteristics. Because results are tied to segments, merges can invalidate entries.

OpenSearch also exposes query-cache statistics/settings and uses cacheable query/filter execution as an optimization. Treat the exact cache policy as product/version behavior, not an API guarantee that every repeated filter will be cached.

Observe query cache, not just wall-clock time
GET _nodes/stats/indices/query_cache?human
GET atlasmart-runtime-lab-v1/_stats/query_cache?human

GET atlasmart-runtime-lab-v1/_search
{
  "size": 0,
  "query": {"bool":{"filter":[
    {"term":{"available":true}},
    {"term":{"category":"electronics"}}
  ]}},
  "aggs":{"price":{"stats":{"field":"price"}}}
}
Do not promise a hit

Run the same request repeatedly and inspect cache statistics. The engine may decide a query is not worth caching, or the fixture may be too small to cross internal eligibility thresholds. “Miss” is an observation, not a lab failure.

3. Shard/index request cache: whole-result reuse

Elasticsearch calls this the shard request cache. OpenSearch documentation calls the corresponding feature the index request cache, although the cached result is still shard-local. The cache is especially valuable for repeatable aggregation/dashboard requests because a shard can reuse a prior local result and the coordinating node can reduce those cached shard results.

By default, size-zero aggregation-like requests are the common cacheable case. Product/version rules differ for non-zero hit sizes and non-deterministic constructs. Relative time such as now, profiling, scroll/DFS modes or script nondeterminism can remove or reduce cacheability. Preserve the exact request body when measuring because the body participates in cache identity.

Request-cache evidence
GET atlasmart-runtime-lab-v1/_search?request_cache=true
{
  "size": 0,
  "query":{"term":{"available":true}},
  "aggs":{"by_category":{"terms":{"field":"category","size":10}}}
}

GET atlasmart-runtime-lab-v1/_stats/request_cache?human
GET _nodes/stats/indices/request_cache?human

Repeat without indexing between runs, then index one document and wait for/issue the refresh required by your experiment. Observe that data-changing refresh semantics can invalidate request-cache entries. Do not claim that a cache hit returns stale search data; the cache lifecycle is integrated with index refresh/change visibility.

4. Fielddata and global ordinals: not a generic “make text aggregatable” switch

Analyzed text is optimized for full-text matching through an inverted index. Sorting or terms aggregation needs efficient per-document values. Enabling fielddata on high-cardinality text asks the engine to build an on-heap structure by uninverting text terms, which can consume substantial heap. Both products therefore default fielddata off for text and recommend modeling an exact-value keyword subfield for sorting/aggregation.

Global ordinals map per-segment term ordinals into a shard-wide view for terms-like operations. They can be built lazily after refresh or eagerly when the mapping requests it. Eager construction can move latency from the first aggregation into refresh/indexing work; it is a latency-placement decision, not free speed.

Safe mapping pattern
"name": {
  "type": "text",
  "fields": {
    "raw": {"type": "keyword", "ignore_above": 256}
  }
}

# Aggregate/sort on name.raw, search full text on name.
# Do NOT set fielddata:true on name merely to make a demo aggregation pass.
Observe fielddata/global-ordinal memory
GET _cat/fielddata?v&bytes=mb
GET _nodes/stats/indices/fielddata?human
GET atlasmart-runtime-lab-v1/_stats/fielddata?human

5. Wrong approach: optimize cache hit rate as the goal

A team disables refreshes and canonicalizes every dashboard request until request-cache hit rate is excellent. Unfortunately, users now see stale data beyond the product requirement. Another team increases cache sizes until heap pressure causes GC and breaker risk. Both optimized an internal metric while degrading the service.

Cache metrics are diagnostic inputs. The service objective is correct results within freshness, latency, throughput, cost and failure-isolation constraints. A lower hit rate may be acceptable for highly dynamic queries; a high hit rate is useless if the cached work is cheap or if it displaces more valuable memory.

Cache clearing is a diagnostic tool, not routine maintenance

Repeatedly clearing caches can create rebuild storms and hide the actual working-set problem. Use it only in a controlled experiment where cold/warm state is the variable under test.

6. AtlasMart cache experiment

  1. Capture _stats/query_cache,request_cache,fielddata and node JVM stats.
  2. Run a fixed filter+aggregation request N times without writes; record client latency distribution and cache deltas.
  3. Run a semantically similar request with an intentionally non-cacheable/changed element and compare.
  4. Index one known document and refresh; rerun the first request and inspect request-cache behavior.
  5. Compare aggregation on keyword to the prohibited idea of fielddata on text; do not enable fielddata in the shared fixture.
  6. Interpret whether any latency difference is due to logical result caching, filesystem warmness, both, or neither.
Expected evidence, not fixed numbers

You should be able to point to cache hit/miss/eviction counters and a controlled request sequence. Exact counts and latency depend on product version, fixture size, refreshes, segment topology and previous activity.

7. Elastic vs OpenSearch details that matter

Area Elasticsearch 9.5.3 OpenSearch 3.8.0
Query cache Node query cache documented for filter-context results with per-segment eligibility/history. Query cache exposed through index/node settings/stats; optimize by measurement rather than assuming identical policy thresholds.
Request cache Shard request cache; default size-zero requests are the common auto-cache case, with per-request override. Index request cache; 2.19+ adds indices.requests.cache.maximum_cacheable_size, defaulting to zero-size automatic caching.
Fielddata On-heap fielddata/global ordinals; fielddata breaker protects growth. Same fundamental hazard; docs explicitly recommend keyword subfields instead of text fielddata.
Tiered cache Do not assume OpenSearch cache-tier settings exist. OpenSearch can expose tiered-request-cache statistics/features in supported configurations; treat as product-specific.

Production judgment

If a cache is the only thing keeping a query within the latency objective, verify what happens after refresh, relocation, node restart and eviction. That is the real miss path users will eventually experience. Capacity planning must include both warm and recovery/cold behavior.

A cache that avoids expensive repeated work can be valuable, but the next lesson covers the safety layer underneath: circuit breakers exist because some requests can allocate more heap than caches and ordinary GC can safely absorb.

Check your understanding

  1. What does the filesystem cache store that the request cache does not?
  2. Why can a merge affect query-cache entries?
  3. Why is fielddata on text dangerous?
  4. Why can eager global ordinals increase indexing/refresh work?
  5. What is a better success metric than cache hit rate?
Review the answers

1. Filesystem pages such as Lucene segment data; the request cache stores a shard-level logical search result.

2. Query-cache entries are associated with segment execution; merging changes the segment set.

3. It can build large on-heap per-document structures from analyzed terms, especially with high cardinality, increasing GC and breaker risk.

4. They move ordinal construction earlier into refresh rather than lazily paying it on the first aggregation.

5. End-to-end correctness/freshness plus latency, throughput, error rate and resource cost for the actual workload.

Summary and next step

You can now say exactly which cache you mean, what it reuses, when it invalidates and which resource it consumes. Lesson 3 uses that vocabulary to explain circuit breakers, aggregation memory, in-flight requests and why safety limits should not be disabled to make one query succeed.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.