Run an AtlasMart observability game day from user symptom through dashboard evidence, alert, diagnosis, corrective action, SLO recovery, and a reusable incident bundle.
Build an Operations Dashboard and Incident Runbook Linking User Symptoms to Cluster and Query Evidence
Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.
Learning outcomes
Assemble an operations dashboard whose panels map directly from user symptoms to query, runtime, storage, and topology evidence.
Write an incident runbook with explicit gates, safe diagnostic commands, stop conditions, ownership, and recovery verification.
Use OpenSearch Query Insights/slow logs and Elastic slow logs/Profile/tasks only as diagnostic evidence, not as fabricated benchmark proof.
Run a controlled AtlasMart freshness incident from detection through alert, diagnosis, remediation, SLO recovery, and post-incident evidence.
Produce a versioned operational evidence bundle suitable for regression testing and future chapters on vector/hybrid search.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0, using
the course's established local TLS/auth conventions and bundled
JVMs. Elastic metrics examples use supported cluster, node,
index, task, slow-log, Kibana rule, saved-object/data-view, and
SLO surfaces. OpenSearch examples use node/index stats, Alerting
monitors, Dashboards, and Query Insights where
installed/enabled. OpenSearch's native SLO feature is currently
documented as experimental; Elastic SLO
management has license and node-role prerequisites, so the
mandatory lab also includes a product-neutral SLI/error-budget
worksheet. No live cluster is available in this generation
environment: commands are reproducible, expected invariants are
stated, and latency/throughput values are labeled
MEASURED instead of fabricated.
1. AtlasMart game day: customer symptom first
The final lab starts with a synthetic catalog probe reporting stale results. The dashboard should immediately show the freshness SLI in violation while availability remains healthy. The operator follows the runbook; they do not begin by changing heap, shard count, queues, or refresh settings at random.
2. Dashboard-to-runbook map
| Symptom panel | First drill-down | Second drill-down | Likely owner |
|---|---|---|---|
| availability | HTTP/error/rejection classes | cluster health/thread pools/breakers | search platform |
| p99 latency | top/slow queries | CPU/GC/disk/merge/cache | search + application |
| freshness | ingest success + refresh/index lag | pipeline/backpressure/storage | ingest + search |
| relevance | golden query diff | mapping/analyzer/model/version changes | search relevance |
3. Runbook skeleton with gates
# AtlasMart search freshness incident
Owner: search-platform
Severity: based on SLO impact, not raw refresh counters
1. Confirm user symptom
- Run synthetic marker probe.
- Record environment, cluster, index/data stream, UTC window.
2. Bound blast radius
- Which services/tenants/indexes are affected?
- Availability and p99 still within objective?
3. Collect evidence (read-only first)
- cluster health/stats
- node JVM/CPU/fs/thread pools
- index indexing/refresh/merge/recovery stats
- slow query / Query Insights evidence if relevant
4. Form one hypothesis
- Example: refresh interval was changed on disposable lab index.
5. Apply the smallest reversible correction
6. Verify SLO recovery with the same synthetic probe
7. Restore temporary settings and collect final evidence
8. Open post-incident action if configuration drift caused the event
STOP if:
- evidence indicates possible data loss,
- required recovery/snapshot state is unknown,
- production headroom is falling,
- the next action is irreversible or outside operator authority.
4. Diagnostic evidence by product
For Elasticsearch, correlate node/index stats, search slow logs, tasks/hot threads, and targeted Profile API output from a representative non-production or safely bounded request. Profile adds overhead and is not a production benchmark. For OpenSearch, add Query Insights when the plugin is installed/enabled; current top-N query monitoring can rank latency, CPU, or memory consumers. Query Insights itself consumes resources, so configure retention/window/N deliberately.
GET /_insights/top_queries?type=latency
GET /_insights/top_queries?type=cpu
GET /_insights/top_queries?type=memory
# Treat records as diagnostic evidence; correlate with client p99 and node stats.
5. Execute the safe freshness game day
-
Create
atlasmart-observe-v24with one primary, zero replicas, and a deliberately long refresh interval. - Start the synthetic marker probe and metric capture.
- Index a marker without forcing refresh. Observe whether the freshness SLI crosses the lab threshold.
- Let the alert/rule/monitor (or deterministic local evaluator) enter firing state.
- Follow the runbook and identify the refresh setting as the causal change.
- Perform a manual refresh, restore the normal interval, and rerun the same probe.
- Record alert recovery plus before/after stats. Delete the disposable index.
PUT /atlasmart-observe-v24
{
"settings": {"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"30s"},
"mappings": {"properties":{"@timestamp":{"type":"date"},"marker":{"type":"keyword"}}}
}
# ... run marker probe and evidence collection ...
POST /atlasmart-observe-v24/_refresh
PUT /atlasmart-observe-v24/_settings
{"index":{"refresh_interval":"1s"}}
# Verify recovery, then clean up
DELETE /atlasmart-observe-v24
The mandatory failure injection changes only search freshness on a disposable local index. CPU/heap/rejection saturation may be discussed or simulated with deterministic traces, but the chapter does not instruct learners to exhaust a workstation or disable circuit breakers.
6. Evidence bundle
Store enough evidence to reproduce the diagnosis without capturing secrets. Include version response, topology summary, SLI time series, raw bounded stats, alert state transitions, runbook steps, configuration diff, and cleanup verification. Redact API keys, passwords, authorization headers, user PII, and raw query payloads that contain sensitive data.
atlasmart-v24-incident/
versions.json
topology.txt
sli-before.json
node-stats-before.json
index-stats-before.json
alert-fired.json
hypothesis.md
config-diff.txt
sli-after.json
alert-recovered.json
cleanup.txt
README.md # timeline, owner, findings, limitations
7. What proves recovery?
Recovery is not “the cluster returned green.” Use the same user-visible probe that detected the incident, verify the relevant SLI is back within objective for the recovery window, confirm error/rejection rates did not worsen, and verify no temporary debugging or refresh settings remain. If relevance was affected, rerun the golden query suite.
8. Operational maturity and failure headroom
Keep at least one externally observable probe path, version dashboard/runbook artifacts, rehearse incidents periodically, and make ownership visible. Observability must survive the failure it is meant to explain: do not place every dashboard, notification path, and evidence store behind the same cluster and network dependency.
9. Bridge to Chapter 25
Chapter 25 introduces vector search, where observability must expand beyond ordinary search latency. The same evidence model will track recall/quality, vector dimensions/model versions, ANN candidate settings, memory, index-build cost, filter selectivity, and p99. The discipline remains unchanged: user-visible quality first, mechanism evidence second, measured tradeoffs instead of folklore.
Check your understanding
- What starts the final incident runbook?
- Why is Profile not a benchmark?
- What is the mandatory failure injection?
- What proves recovery?
- What new evidence will vector search add next chapter?
Review the answers
1. Confirmation of the user-visible symptom with the same synthetic or client SLI used by the SLO.
2. It adds instrumentation overhead and exposes operator timings for diagnosis, not representative production performance.
3. A reversible long refresh interval on a disposable local index that creates freshness lag.
4. The original SLI returns within objective and temporary settings are restored; green health alone is insufficient.
5. Quality/recall, model/vector contracts, ANN settings, memory/build cost, filter selectivity, and tail latency.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic cluster stats API — Cluster-wide node, shard, store, JVM, CPU, and plugin context.
- Elastic node stats API — JVM, filesystem, thread pools, breakers, indexing pressure, search/indexing, caches, merge, refresh, and recovery metrics.
- Kibana alerting — Rules, schedules, alerts, and connector actions.
- Elastic SLO access — Current license, transform/ingest-role, space, and index-security requirements.
- Kibana saved objects — Dashboards, visualizations, data views, import/export, and permissions.
- OpenSearch Nodes Stats API — JVM, filesystem, process, thread-pool, and index metric evidence.
- OpenSearch Alerting — Monitor, trigger, alert, and action model.
- OpenSearch Query Insights: top N queries — Latency, CPU, and memory query evidence.
- OpenSearch Query Insights Dashboards — Live/top-N/configuration UI and query details.
- OpenSearch Observability — Dashboards, alerting/detection, traces, and experimental SLO scope.
- Elastic Stack 9.5.3 release and OpenSearch 3.8.0 artifacts — pinned September/August 2026 server baselines.