Chapter 15 · Heap, Caches, Circuit Breakers, Thread Pools, Backpressure, and JVM/Runtime Health
Circuit Breakers, In-Flight Requests, Aggregation Memory, Fielddata Hazards, and Failure Protection
Treat circuit breakers and request accounting as protective admission controls, diagnose breaker evidence, and repair mappings/queries instead of disabling safeguards or moving failures into JVM out-of-memory territory.
Learning outcomes
AtlasMart adds a user-configurable aggregation endpoint. A request with many high-cardinality buckets begins consuming large amounts of heap. One engineer proposes raising or disabling the breaker because the request is “legitimate.” That changes the failure mode from a controlled rejection into potential JVM/process instability. This lesson treats a breaker trip as protective evidence, not an inconvenience to erase.
Explain parent and child circuit breakers as estimated admission controls rather than hard memory isolation.
Distinguish fielddata, request/aggregation and in-flight request pressure.
Diagnose a circuit-breaking exception using breaker stats, query shape, mapping and concurrent workload evidence.
Replace fielddata/high-cardinality bucket hazards with safer mapping, query and pagination strategies.
Test a repair while preserving correctness and service-level latency/error objectives.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual JVM/runtime with
GET _nodes/jvm and
GET _nodes/stats/jvm,process,os rather than
hard-coding a JDK patch. Labs use one primary and zero
replicas unless a step explicitly changes topology. OpenSearch
demo -k TLS bypass remains disposable-lab-only;
production must validate certificates. No moving
latest tags are used.
This chapter was authored against current official product documentation, but the generation environment does not run the AtlasMart Elasticsearch/OpenSearch containers. Therefore exact latency, GC duration, queue depth, cache-hit count and breaker values are not fabricated. Expected outputs below describe invariant fields, direction of change, status classes and acceptance checks. Run the bounded lab on your own pinned local containers to obtain measurements for your machine.
1. What a circuit breaker does—and does not do
A circuit breaker estimates or measures a category of memory pressure and rejects additional work before that work is likely to endanger the node. The parent breaker provides a total safety boundary; child breakers account for categories such as fielddata, request structures and in-flight network requests. Both products expose closely related breaker families, but defaults and newer breaker categories must be checked per release.
A breaker is not a perfect per-request memory sandbox. Estimates have overhead factors, some memory is outside breaker accounting, native/off-heap allocations still exist, and concurrent components share the process. A breaker therefore reduces risk; it does not make out-of-memory impossible.
GET _nodes/stats/breaker?human
GET _nodes/stats/jvm,process?human
GET _cluster/settings?include_defaults=true&flat_settings=true
# During an incident, preserve the exact circuit_breaking_exception response.
# Record breaker name, bytes_wanted, bytes_limit, durability/status and request shape
# when those fields are returned by the current product/version.
2. Parent breaker and real-memory accounting
With real-memory accounting enabled—the default documented by both current products—the parent breaker considers current JVM memory use rather than merely summing reservations from child breakers. That makes it a broad admission signal when heap is already crowded by multiple components.
Do not read a documented default percentage as “safe to allocate this much to my aggregation.” It is a node protection threshold. The headroom below it is shared with live objects, caches, cluster state, query coordination and GC behavior.
This can convert a deterministic 429/error into long GC pauses or process failure. First reduce request cardinality/memory, repair field mappings, cap concurrency, isolate workloads or add capacity. Change breaker settings only when measured workload evidence and the supported product guidance justify it.
3. Request/aggregation memory: why high-cardinality buckets matter
A distributed terms/composite/pipeline aggregation can create
per-shard bucket/state structures and then reduce them on a
coordinating node. Memory grows with the query shape,
cardinality, shard fan-out and concurrent requests—not only with
response JSON size. Chapter 08 showed that a huge
terms.size is not pagination; this chapter adds the
runtime consequence: it can consume heap and trigger
request/parent protection.
Safer strategies include bounding bucket counts, using
composite pagination for exhaustive bucket
traversal where semantics fit, filtering early, precomputing
known analytics, routing/isolation, and protecting analytical
workloads from interactive search. Do not simply raise
search.max_buckets or breaker limits without
evidence.
GET atlasmart-runtime-lab-v1/_search
{
"size": 0,
"query": {"term":{"available":true}},
"aggs": {
"categories": {
"terms": {"field":"category", "size":10},
"aggs": {"price":{"stats":{"field":"price"}}}
}
}
}
4. Fielddata breaker: the mapping mistake that looks like a tuning problem
If a terms aggregation targets an analyzed
text field and you enable fielddata,
the engine can build on-heap fielddata from the inverted index.
High-cardinality text can consume enough heap to trip the
fielddata breaker. The correct default repair is to model an
exact keyword field/subfield and reindex if
necessary.
Global ordinals also use heap and can be expensive to rebuild. If the first query after refresh is slow, decide whether eager ordinals are appropriate based on refresh and query objectives; do not mistake the symptom for a reason to remove breaker protection.
GET _cat/fielddata?v&bytes=mb
GET _nodes/stats/indices/fielddata,breaker?human
# Full-text matching
GET atlasmart-runtime-lab-v1/_search
{"query":{"match":{"name":"wireless keyboard"}}}
# Exact aggregation/sort
GET atlasmart-runtime-lab-v1/_search
{"size":0,"aggs":{"names":{"terms":{"field":"name.raw","size":10}}}}
5. In-flight requests: network bytes are also memory pressure
The in-flight request breaker accounts for requests currently moving through the transport/HTTP handling path. Large bulk bodies, huge search responses and high client concurrency can create simultaneous pressure even if each individual request is acceptable. That is one reason “increase concurrency until throughput stops rising” is unsafe as a production tuning strategy.
Prefer bounded bulk/request sizes, client concurrency limits and retry/backoff policies. For bulk indexing, remember Chapter 04: HTTP success can still contain item-level failures; admission control must not erase item inspection or idempotency logic.
6. Controlled breaker/failure lab without an OOM experiment
Do not intentionally disable breakers or allocate until the process crashes. A deterministic learning path can use current breaker statistics and a bounded aggregation fixture. If your local node naturally returns a breaker error under a capped experiment, preserve it. If it does not, study the request shape and breaker estimates without forcing a trip.
A safe lab success criterion is: you can show the query/mapping feature that increases memory demand, the relevant breaker/cache counters, and a repaired query/mapping that returns the same required business answer with lower bounded state.
GET _nodes/stats/breaker,jvm,thread_pool?human
GET _nodes/stats/indices/fielddata,request_cache,query_cache?human
GET atlasmart-runtime-lab-v1/_stats/fielddata,search?human
GET _nodes/hot_threads?threads=3&ignore_idle_threads=true
- Keep the dataset and answer semantics fixed.
- Candidate A: intentionally wide but still resource-bounded aggregation; record server/client evidence.
- Candidate B: bounded terms or composite-style strategy appropriate to the business requirement.
- Do not change breaker limits between A and B.
- Accept B only if result semantics are correct and memory/latency/rejection evidence improves or stays safely bounded.
7. Elasticsearch vs OpenSearch protection surfaces
| Protection surface | Elasticsearch 9.5.3 | OpenSearch 3.8.0 |
|---|---|---|
| Parent/fielddata/request/in-flight breakers | Current breaker families documented in cluster settings and node stats. | Closely corresponding breaker families documented; verify defaults in 3.8 docs/settings. |
| Indexing pressure | Built-in outstanding-byte accounting rejects coordinating/primary/replica indexing work at configured limits. | Has node indexing-pressure controls and an additional shard-indexing-backpressure subsystem that can run shadow/enforced. |
| Search overload | Thread-pool rejection, breakers, task/cancellation/resource controls; do not invent an OpenSearch-named feature. | Explicit search-backpressure subsystem tracks CPU/heap/elapsed time and can cancel resource-intensive search tasks under duress. |
Production judgment
A breaker trip is a symptom with a protective action attached. Diagnose the request, mapping, concurrency and node state. If the same legitimate workload repeatedly hits a correctly configured breaker after query/mapping repairs, capacity/workload isolation may be the right answer.
The next lesson follows work into thread pools and queues. That matters because overload can surface as rising queue time and 429 rejections before heap breakers ever trip.
Check your understanding
- Why is disabling a circuit breaker unsafe?
- Why can a small JSON response still be memory-expensive?
- What is the preferred repair for text fielddata used only to aggregate exact values?
- Does a breaker guarantee no OOM?
- What should remain unchanged in a breaker repair experiment?
Review the answers
1. It removes a controlled rejection boundary and can let memory growth progress toward long GC pauses or JVM/process out-of-memory failure.
2. Distributed aggregation/search execution may build large per-shard/coordinator structures that are not represented by final response size.
3. Model/reindex an exact keyword field or subfield and aggregate that instead.
4. No. It estimates/accounts selected memory categories; native/off-heap and untracked/concurrent allocations still exist.
5. Business result semantics, dataset, topology and breaker settings—change the query/mapping/concurrency variable you are testing.
Summary and next step
Circuit breakers are failure protection, not throughput knobs. You can now read a trip as evidence about a memory-demanding request and repair the request/mapping/concurrency path. Lesson 4 adds execution queues and product-specific backpressure so you can recognize overload before retry storms amplify it.
Authoritative references
- Elastic JVM settings — Automatic heap sizing, filesystem-cache headroom, container memory, and compressed ordinary object pointer guidance.
- Elastic node query cache settings — Filter-context query-cache eligibility, segment-level caching and invalidation behavior.
- Elastic shard request cache — Shard-level request-result caching and request/index controls.
- Elastic field data cache settings — On-heap fielddata/global-ordinal cache behavior and breaker interaction.
- Elastic circuit breaker settings — Parent, fielddata, request and in-flight breaker semantics.
- Elastic thread pool settings — Current thread-pool types, sizing and queue behavior.
- Elastic indexing pressure — Outstanding indexing-byte accounting and rejection behavior.
- Elastic Nodes Stats API — JVM, process, caches, breakers, thread pools and indexing-pressure observations.
- OpenSearch circuit breaker settings — Parent/child breaker settings and failure-protection intent.
- OpenSearch caching overview — On-heap request/query/fielddata caching concepts.
- OpenSearch index request cache — Shard-level request-result caching, invalidation and metrics.
- OpenSearch field data cache — Fielddata/global ordinals and the keyword-subfield alternative.
- OpenSearch thread pool settings — Thread pools, queue sizing and monitoring guidance.
- OpenSearch search backpressure — OpenSearch-specific search task resource tracking and cancellation.
- OpenSearch shard indexing backpressure — OpenSearch-specific per-shard indexing pressure and rejection mechanism.
- OpenSearch Nodes Stats API — JVM, process, caches, indexing pressure and backpressure statistics.
- OpenSearch Nodes Hot Threads API — CPU/wait/block thread samples for diagnosis.