Run an AtlasMart observability game day from user symptom through dashboard evidence, alert, diagnosis, corrective action, SLO recovery, and a reusable incident bundle.

Build an Operations Dashboard and Incident Runbook Linking User Symptoms to Cluster and Query Evidence

Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.

Intermediate → Advanced135–180 minutesOperations dashboard & incident-runbook game day · Chapter 24 · Lesson 05Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Assemble an operations dashboard whose panels map directly from user symptoms to query, runtime, storage, and topology evidence.

02

Write an incident runbook with explicit gates, safe diagnostic commands, stop conditions, ownership, and recovery verification.

03

Use OpenSearch Query Insights/slow logs and Elastic slow logs/Profile/tasks only as diagnostic evidence, not as fabricated benchmark proof.

04

Run a controlled AtlasMart freshness incident from detection through alert, diagnosis, remediation, SLO recovery, and post-incident evidence.

05

Produce a versioned operational evidence bundle suitable for regression testing and future chapters on vector/hybrid search.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned observability baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0, using the course's established local TLS/auth conventions and bundled JVMs. Elastic metrics examples use supported cluster, node, index, task, slow-log, Kibana rule, saved-object/data-view, and SLO surfaces. OpenSearch examples use node/index stats, Alerting monitors, Dashboards, and Query Insights where installed/enabled. OpenSearch's native SLO feature is currently documented as experimental; Elastic SLO management has license and node-role prerequisites, so the mandatory lab also includes a product-neutral SLI/error-budget worksheet. No live cluster is available in this generation environment: commands are reproducible, expected invariants are stated, and latency/throughput values are labeled MEASURED instead of fabricated.

1. AtlasMart game day: customer symptom first

The final lab starts with a synthetic catalog probe reporting stale results. The dashboard should immediately show the freshness SLI in violation while availability remains healthy. The operator follows the runbook; they do not begin by changing heap, shard count, queues, or refresh settings at random.

2. Dashboard-to-runbook map

Symptom panel First drill-down Second drill-down Likely owner
availability HTTP/error/rejection classes cluster health/thread pools/breakers search platform
p99 latency top/slow queries CPU/GC/disk/merge/cache search + application
freshness ingest success + refresh/index lag pipeline/backpressure/storage ingest + search
relevance golden query diff mapping/analyzer/model/version changes search relevance

3. Runbook skeleton with gates

runbooks/search-freshness.md
# AtlasMart search freshness incident
Owner: search-platform
Severity: based on SLO impact, not raw refresh counters

1. Confirm user symptom
   - Run synthetic marker probe.
   - Record environment, cluster, index/data stream, UTC window.
2. Bound blast radius
   - Which services/tenants/indexes are affected?
   - Availability and p99 still within objective?
3. Collect evidence (read-only first)
   - cluster health/stats
   - node JVM/CPU/fs/thread pools
   - index indexing/refresh/merge/recovery stats
   - slow query / Query Insights evidence if relevant
4. Form one hypothesis
   - Example: refresh interval was changed on disposable lab index.
5. Apply the smallest reversible correction
6. Verify SLO recovery with the same synthetic probe
7. Restore temporary settings and collect final evidence
8. Open post-incident action if configuration drift caused the event

STOP if:
- evidence indicates possible data loss,
- required recovery/snapshot state is unknown,
- production headroom is falling,
- the next action is irreversible or outside operator authority.

4. Diagnostic evidence by product

For Elasticsearch, correlate node/index stats, search slow logs, tasks/hot threads, and targeted Profile API output from a representative non-production or safely bounded request. Profile adds overhead and is not a production benchmark. For OpenSearch, add Query Insights when the plugin is installed/enabled; current top-N query monitoring can rank latency, CPU, or memory consumers. Query Insights itself consumes resources, so configure retention/window/N deliberately.

OpenSearch Query Insights read-only check
GET /_insights/top_queries?type=latency
GET /_insights/top_queries?type=cpu
GET /_insights/top_queries?type=memory
# Treat records as diagnostic evidence; correlate with client p99 and node stats.

5. Execute the safe freshness game day

  1. Create atlasmart-observe-v24 with one primary, zero replicas, and a deliberately long refresh interval.
  2. Start the synthetic marker probe and metric capture.
  3. Index a marker without forcing refresh. Observe whether the freshness SLI crosses the lab threshold.
  4. Let the alert/rule/monitor (or deterministic local evaluator) enter firing state.
  5. Follow the runbook and identify the refresh setting as the causal change.
  6. Perform a manual refresh, restore the normal interval, and rerun the same probe.
  7. Record alert recovery plus before/after stats. Delete the disposable index.
Lab setup and cleanup
PUT /atlasmart-observe-v24
{
  "settings": {"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"30s"},
  "mappings": {"properties":{"@timestamp":{"type":"date"},"marker":{"type":"keyword"}}}
}

# ... run marker probe and evidence collection ...
POST /atlasmart-observe-v24/_refresh
PUT /atlasmart-observe-v24/_settings
{"index":{"refresh_interval":"1s"}}

# Verify recovery, then clean up
DELETE /atlasmart-observe-v24
Do not manufacture the incident with dangerous load.

The mandatory failure injection changes only search freshness on a disposable local index. CPU/heap/rejection saturation may be discussed or simulated with deterministic traces, but the chapter does not instruct learners to exhaust a workstation or disable circuit breakers.

6. Evidence bundle

Store enough evidence to reproduce the diagnosis without capturing secrets. Include version response, topology summary, SLI time series, raw bounded stats, alert state transitions, runbook steps, configuration diff, and cleanup verification. Redact API keys, passwords, authorization headers, user PII, and raw query payloads that contain sensitive data.

Evidence manifest
atlasmart-v24-incident/
  versions.json
  topology.txt
  sli-before.json
  node-stats-before.json
  index-stats-before.json
  alert-fired.json
  hypothesis.md
  config-diff.txt
  sli-after.json
  alert-recovered.json
  cleanup.txt
  README.md  # timeline, owner, findings, limitations

7. What proves recovery?

Recovery is not “the cluster returned green.” Use the same user-visible probe that detected the incident, verify the relevant SLI is back within objective for the recovery window, confirm error/rejection rates did not worsen, and verify no temporary debugging or refresh settings remain. If relevance was affected, rerun the golden query suite.

8. Operational maturity and failure headroom

Keep at least one externally observable probe path, version dashboard/runbook artifacts, rehearse incidents periodically, and make ownership visible. Observability must survive the failure it is meant to explain: do not place every dashboard, notification path, and evidence store behind the same cluster and network dependency.

9. Bridge to Chapter 25

Chapter 25 introduces vector search, where observability must expand beyond ordinary search latency. The same evidence model will track recall/quality, vector dimensions/model versions, ANN candidate settings, memory, index-build cost, filter selectivity, and p99. The discipline remains unchanged: user-visible quality first, mechanism evidence second, measured tradeoffs instead of folklore.

Check your understanding

  1. What starts the final incident runbook?
  2. Why is Profile not a benchmark?
  3. What is the mandatory failure injection?
  4. What proves recovery?
  5. What new evidence will vector search add next chapter?
Review the answers

1. Confirmation of the user-visible symptom with the same synthetic or client SLI used by the SLO.

2. It adds instrumentation overhead and exposes operator timings for diagnosis, not representative production performance.

3. A reversible long refresh interval on a disposable local index that creates freshness lag.

4. The original SLI returns within objective and temporary settings are restored; green health alone is insufficient.

5. Quality/recall, model/vector contracts, ANN settings, memory/build cost, filter selectivity, and tail latency.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.