Correlate AtlasMart user symptoms with JVM, CPU, disk, shard, indexing, search, cache, merge, refresh, and recovery evidence without mistaking raw metrics for an SLO.
Cluster / Node / Index Metrics: JVM, CPU, Disk, Shards, Indexing, Search, Cache, Merge, Refresh, and Recovery
Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.
Learning outcomes
Build an evidence hierarchy that starts with user symptoms and correlates them with cluster, node, index, and query metrics.
Distinguish JVM heap pressure, CPU saturation, filesystem/disk pressure, shard topology, indexing/search demand, cache behavior, merge debt, refresh work, and recovery work.
Use cluster/node/index stats as evidence without treating any single counter or green cluster health as proof of a healthy user experience.
Collect a resource-bounded AtlasMart metric snapshot and preserve before/after evidence around one controlled incident.
Choose telemetry that supports an SLO or diagnosis while limiting monitoring cardinality, permissions, and failure-headroom consumption.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0, using
the course's established local TLS/auth conventions and bundled
JVMs. Elastic metrics examples use supported cluster, node,
index, task, slow-log, Kibana rule, saved-object/data-view, and
SLO surfaces. OpenSearch examples use node/index stats, Alerting
monitors, Dashboards, and Query Insights where
installed/enabled. OpenSearch's native SLO feature is currently
documented as experimental; Elastic SLO
management has license and node-role prerequisites, so the
mandatory lab also includes a product-neutral SLI/error-budget
worksheet. No live cluster is available in this generation
environment: commands are reproducible, expected invariants are
stated, and latency/throughput values are labeled
MEASURED instead of fabricated.
1. AtlasMart incident: green cluster, slow customers
At 14:02, AtlasMart shoppers report that catalog search sometimes takes several seconds. The cluster health endpoint is green. That observation proves that all expected primary and replica shards are assigned; it does not prove that request latency, freshness, relevance, CPU headroom, disk latency, queues, or caches are healthy. An operations model therefore begins with a symptom and asks which evidence can falsify candidate causes.
Green health is an allocation state. Keep it on the dashboard, but pair it with user-visible availability/latency/freshness and workload evidence. A fully allocated cluster can still be overloaded, serving stale data, rejecting requests, or returning poor results.
2. Evidence layers and what each can prove
| Layer | Useful evidence | What it can establish | What it cannot establish alone |
|---|---|---|---|
| User/API | success rate, p50/p95/p99 latency, timeouts, result count, relevance probes | customer-visible impact | root cause |
| Query | slow log, Profile on a replica lab query, OpenSearch Query Insights, task state | expensive query shapes and phases | whole-node resource causality |
| Index/shard | search/indexing rates, merges, refresh, segments, caches, recoveries | workload and shard-local pressure | OS-wide contention |
| Node/JVM/OS | heap/GC, CPU, fs, thread-pool queues/rejections, breakers | runtime/resource saturation signals | business impact without correlation |
| Cluster | health, allocation, pending tasks, recovery | topology/control-plane state | tail latency or relevance quality |
3. Collect the AtlasMart evidence bundle
Use a monitoring principal with monitor-like
cluster privileges and read-only access to the lab indices. Do
not give a dashboard superuser credentials. Capture timestamps
so application and cluster evidence can be aligned.
GET /_cluster/health
GET /_cluster/stats
GET /_nodes/stats/jvm,os,process,fs,thread_pool,breaker,indexing_pressure,indices/search,indexing,merge,refresh,query_cache,request_cache,recovery
GET /atlasmart-telemetry-v24/_stats/search,indexing,merge,refresh,query_cache,request_cache,segments,store
GET /_cat/thread_pool/search,write?v=true&h=node_name,name,active,queue,rejected,completed
GET /_cluster/health
GET /_cluster/stats
GET /_nodes/stats/jvm,os,process,fs,thread_pool,breaker,indices
GET /atlasmart-telemetry-v24/_stats/search,indexing,merge,refresh,query_cache,request_cache,segments,store
GET /_cat/thread_pool/search,write?v=true&h=node_name,name,active,queue,rejected,completed
# If Query Insights is installed/enabled:
GET /_insights/top_queries?type=latency
Field names can differ across product versions. Save the raw JSON with the version response and a timestamp rather than building a parser around undocumented fields.
4. Interpret the high-value metric families
| Signal | Mechanism | Useful correlation | Common false conclusion |
|---|---|---|---|
| JVM heap + GC | managed object pressure and collection | heap sawtooth, GC time, breaker trips, tail latency | “high heap always means add heap” |
| CPU | query, indexing, merge, GC, scripting, coordination work | hot threads + request mix + p99 | “CPU < 100% means no saturation” |
| Disk/fs | segment writes/reads, merges, recovery, snapshots | merge/recovery rate + disk latency + queueing | “free capacity equals fast storage” |
| Shards | fan-out, routing, recovery units | active shards + per-shard work + topology | “more shards always increase throughput” |
| Caches | reused query/request structures and OS page cache | hit/miss/eviction + workload repetition | “maximize cache hit rate” |
| Merge/refresh | immutable-segment lifecycle | merge time/current + refresh count/time + writes | “refresh and flush are the same” |
| Recovery | shard copy/reconstruction | bytes/time + network/disk + foreground p99 | “faster recovery is always safer” |
5. Controlled incident: create one safe freshness signal
The mandatory local exercise avoids dangerous saturation. Create
a disposable one-primary/zero-replica index with a long refresh
interval, index a marker without refresh=true, and
show that the document can be acknowledged by the write path
before ordinary search sees it. Then issue a manual refresh and
verify visibility. This demonstrates
freshness lag without driving CPU or heap to
failure.
PUT /atlasmart-observe-v24
{
"settings": {"number_of_shards": 1, "number_of_replicas": 0, "refresh_interval": "30s"},
"mappings": {"properties": {"@timestamp": {"type": "date"}, "marker": {"type": "keyword"}}}
}
POST /atlasmart-observe-v24/_doc/freshness-1
{"@timestamp":"2026-09-11T12:00:00Z","marker":"incident"}
GET /atlasmart-observe-v24/_search?q=marker:incident
# Expected invariant before refresh: the hit MAY be absent because refresh has not occurred.
POST /atlasmart-observe-v24/_refresh
GET /atlasmart-observe-v24/_search?q=marker:incident
# Expected invariant after successful refresh: freshness-1 is searchable.
Record the write acknowledgment time, first successful search time, index refresh counters, and client-observed latency. The difference is measured freshness for this fixture, not a universal refresh SLA.
6. Telemetry cost and independent monitoring
Monitoring can become part of the outage if it runs expensive broad queries at high frequency on a saturated cluster. Prefer bounded field sets, interval-appropriate sampling, dedicated monitoring credentials, and external synthetic probes. Preserve some monitoring path outside the search cluster so a cluster outage does not erase the only evidence that it is down. High-cardinality dimensions such as raw query text, user IDs, or request IDs need explicit retention and privacy controls.
7. Production judgment
A useful dashboard is a causal map, not a wall of gauges. Begin with availability, p95/p99, freshness, rejection/error rate, and a relevance indicator; then add the minimal cluster/node/index/query panels needed to diagnose those symptoms. Capture baselines by workload and time-of-day, use rate-of-change for counters, and keep incident evidence long enough to compare pre-incident, incident, and recovery windows.
Check your understanding
- Why can a green cluster still violate the search SLO?
- What is the first observability layer during an incident?
- Why capture raw stats with timestamps and versions?
- What does the refresh lab prove?
- Why limit monitoring cost?
Review the answers
1. Green describes shard allocation, not user latency, freshness, error rate, or relevance.
2. A user-visible symptom or SLI, because resource metrics need impact context.
3. Field semantics and topology are version-dependent; timestamps let evidence be correlated.
4. It proves search visibility can lag acknowledged indexing until refresh; it does not establish a universal lag value.
5. Telemetry competes for the same cluster resources and can consume failure headroom during incidents.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic cluster stats API — Cluster-wide node, shard, store, JVM, CPU, and plugin context.
- Elastic node stats API — JVM, filesystem, thread pools, breakers, indexing pressure, search/indexing, caches, merge, refresh, and recovery metrics.
- Kibana alerting — Rules, schedules, alerts, and connector actions.
- Elastic SLO access — Current license, transform/ingest-role, space, and index-security requirements.
- Kibana saved objects — Dashboards, visualizations, data views, import/export, and permissions.
- OpenSearch Nodes Stats API — JVM, filesystem, process, thread-pool, and index metric evidence.
- OpenSearch Alerting — Monitor, trigger, alert, and action model.
- OpenSearch Query Insights: top N queries — Latency, CPU, and memory query evidence.
- OpenSearch Query Insights Dashboards — Live/top-N/configuration UI and query details.
- OpenSearch Observability — Dashboards, alerting/detection, traces, and experimental SLO scope.
- Elastic Stack 9.5.3 release and OpenSearch 3.8.0 artifacts — pinned September/August 2026 server baselines.