Correlate AtlasMart user symptoms with JVM, CPU, disk, shard, indexing, search, cache, merge, refresh, and recovery evidence without mistaking raw metrics for an SLO.

Cluster / Node / Index Metrics: JVM, CPU, Disk, Shards, Indexing, Search, Cache, Merge, Refresh, and Recovery

Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.

Intermediate → Advanced120–160 minutesMetrics & freshness-evidence lab · Chapter 24 · Lesson 01Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Build an evidence hierarchy that starts with user symptoms and correlates them with cluster, node, index, and query metrics.

02

Distinguish JVM heap pressure, CPU saturation, filesystem/disk pressure, shard topology, indexing/search demand, cache behavior, merge debt, refresh work, and recovery work.

03

Use cluster/node/index stats as evidence without treating any single counter or green cluster health as proof of a healthy user experience.

04

Collect a resource-bounded AtlasMart metric snapshot and preserve before/after evidence around one controlled incident.

05

Choose telemetry that supports an SLO or diagnosis while limiting monitoring cardinality, permissions, and failure-headroom consumption.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned observability baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0, using the course's established local TLS/auth conventions and bundled JVMs. Elastic metrics examples use supported cluster, node, index, task, slow-log, Kibana rule, saved-object/data-view, and SLO surfaces. OpenSearch examples use node/index stats, Alerting monitors, Dashboards, and Query Insights where installed/enabled. OpenSearch's native SLO feature is currently documented as experimental; Elastic SLO management has license and node-role prerequisites, so the mandatory lab also includes a product-neutral SLI/error-budget worksheet. No live cluster is available in this generation environment: commands are reproducible, expected invariants are stated, and latency/throughput values are labeled MEASURED instead of fabricated.

1. AtlasMart incident: green cluster, slow customers

At 14:02, AtlasMart shoppers report that catalog search sometimes takes several seconds. The cluster health endpoint is green. That observation proves that all expected primary and replica shards are assigned; it does not prove that request latency, freshness, relevance, CPU headroom, disk latency, queues, or caches are healthy. An operations model therefore begins with a symptom and asks which evidence can falsify candidate causes.

Wrong approach: “green means healthy.”

Green health is an allocation state. Keep it on the dashboard, but pair it with user-visible availability/latency/freshness and workload evidence. A fully allocated cluster can still be overloaded, serving stale data, rejecting requests, or returning poor results.

2. Evidence layers and what each can prove

Layer Useful evidence What it can establish What it cannot establish alone
User/API success rate, p50/p95/p99 latency, timeouts, result count, relevance probes customer-visible impact root cause
Query slow log, Profile on a replica lab query, OpenSearch Query Insights, task state expensive query shapes and phases whole-node resource causality
Index/shard search/indexing rates, merges, refresh, segments, caches, recoveries workload and shard-local pressure OS-wide contention
Node/JVM/OS heap/GC, CPU, fs, thread-pool queues/rejections, breakers runtime/resource saturation signals business impact without correlation
Cluster health, allocation, pending tasks, recovery topology/control-plane state tail latency or relevance quality

3. Collect the AtlasMart evidence bundle

Use a monitoring principal with monitor-like cluster privileges and read-only access to the lab indices. Do not give a dashboard superuser credentials. Capture timestamps so application and cluster evidence can be aligned.

Elastic: bounded node/index evidence
GET /_cluster/health
GET /_cluster/stats
GET /_nodes/stats/jvm,os,process,fs,thread_pool,breaker,indexing_pressure,indices/search,indexing,merge,refresh,query_cache,request_cache,recovery
GET /atlasmart-telemetry-v24/_stats/search,indexing,merge,refresh,query_cache,request_cache,segments,store
GET /_cat/thread_pool/search,write?v=true&h=node_name,name,active,queue,rejected,completed
OpenSearch: equivalent evidence surfaces
GET /_cluster/health
GET /_cluster/stats
GET /_nodes/stats/jvm,os,process,fs,thread_pool,breaker,indices
GET /atlasmart-telemetry-v24/_stats/search,indexing,merge,refresh,query_cache,request_cache,segments,store
GET /_cat/thread_pool/search,write?v=true&h=node_name,name,active,queue,rejected,completed
# If Query Insights is installed/enabled:
GET /_insights/top_queries?type=latency

Field names can differ across product versions. Save the raw JSON with the version response and a timestamp rather than building a parser around undocumented fields.

4. Interpret the high-value metric families

Signal Mechanism Useful correlation Common false conclusion
JVM heap + GC managed object pressure and collection heap sawtooth, GC time, breaker trips, tail latency “high heap always means add heap”
CPU query, indexing, merge, GC, scripting, coordination work hot threads + request mix + p99 “CPU < 100% means no saturation”
Disk/fs segment writes/reads, merges, recovery, snapshots merge/recovery rate + disk latency + queueing “free capacity equals fast storage”
Shards fan-out, routing, recovery units active shards + per-shard work + topology “more shards always increase throughput”
Caches reused query/request structures and OS page cache hit/miss/eviction + workload repetition “maximize cache hit rate”
Merge/refresh immutable-segment lifecycle merge time/current + refresh count/time + writes “refresh and flush are the same”
Recovery shard copy/reconstruction bytes/time + network/disk + foreground p99 “faster recovery is always safer”

5. Controlled incident: create one safe freshness signal

The mandatory local exercise avoids dangerous saturation. Create a disposable one-primary/zero-replica index with a long refresh interval, index a marker without refresh=true, and show that the document can be acknowledged by the write path before ordinary search sees it. Then issue a manual refresh and verify visibility. This demonstrates freshness lag without driving CPU or heap to failure.

Disposable freshness incident
PUT /atlasmart-observe-v24
{
  "settings": {"number_of_shards": 1, "number_of_replicas": 0, "refresh_interval": "30s"},
  "mappings": {"properties": {"@timestamp": {"type": "date"}, "marker": {"type": "keyword"}}}
}

POST /atlasmart-observe-v24/_doc/freshness-1
{"@timestamp":"2026-09-11T12:00:00Z","marker":"incident"}

GET /atlasmart-observe-v24/_search?q=marker:incident
# Expected invariant before refresh: the hit MAY be absent because refresh has not occurred.

POST /atlasmart-observe-v24/_refresh
GET /atlasmart-observe-v24/_search?q=marker:incident
# Expected invariant after successful refresh: freshness-1 is searchable.
Evidence rule

Record the write acknowledgment time, first successful search time, index refresh counters, and client-observed latency. The difference is measured freshness for this fixture, not a universal refresh SLA.

6. Telemetry cost and independent monitoring

Monitoring can become part of the outage if it runs expensive broad queries at high frequency on a saturated cluster. Prefer bounded field sets, interval-appropriate sampling, dedicated monitoring credentials, and external synthetic probes. Preserve some monitoring path outside the search cluster so a cluster outage does not erase the only evidence that it is down. High-cardinality dimensions such as raw query text, user IDs, or request IDs need explicit retention and privacy controls.

7. Production judgment

A useful dashboard is a causal map, not a wall of gauges. Begin with availability, p95/p99, freshness, rejection/error rate, and a relevance indicator; then add the minimal cluster/node/index/query panels needed to diagnose those symptoms. Capture baselines by workload and time-of-day, use rate-of-change for counters, and keep incident evidence long enough to compare pre-incident, incident, and recovery windows.

Check your understanding

  1. Why can a green cluster still violate the search SLO?
  2. What is the first observability layer during an incident?
  3. Why capture raw stats with timestamps and versions?
  4. What does the refresh lab prove?
  5. Why limit monitoring cost?
Review the answers

1. Green describes shard allocation, not user latency, freshness, error rate, or relevance.

2. A user-visible symptom or SLI, because resource metrics need impact context.

3. Field semantics and topology are version-dependent; timestamps let evidence be correlated.

4. It proves search visibility can lag acknowledged indexing until refresh; it does not establish a universal lag value.

5. Telemetry competes for the same cluster resources and can consume failure headroom during incidents.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.