Create actionable AtlasMart alerts with thresholds, rates, persistence, noise controls, recovery evidence, and product-correct Kibana rule/OpenSearch monitor semantics.
Alerting/Rules/Monitors, Thresholds vs Anomaly/Rate Conditions, Noise Control, and Incident Routing
Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.
Learning outcomes
Design alerts from actionable service symptoms and diagnostic evidence rather than mirroring every dashboard metric.
Distinguish static thresholds, rates, burn-rate style conditions, and anomaly detection, including their data and maturity requirements.
Compare Kibana rules/connectors with OpenSearch monitors/triggers/actions without assuming API or execution equivalence.
Control alert noise through windows, persistence, recovery notifications, grouping, deduplication, routing, and ownership.
Inject one safe AtlasMart incident, produce firing and recovery evidence, and verify that the runbook leads to the causal metric.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0, using
the course's established local TLS/auth conventions and bundled
JVMs. Elastic metrics examples use supported cluster, node,
index, task, slow-log, Kibana rule, saved-object/data-view, and
SLO surfaces. OpenSearch examples use node/index stats, Alerting
monitors, Dashboards, and Query Insights where
installed/enabled. OpenSearch's native SLO feature is currently
documented as experimental; Elastic SLO
management has license and node-role prerequisites, so the
mandatory lab also includes a product-neutral SLI/error-budget
worksheet. No live cluster is available in this generation
environment: commands are reproducible, expected invariants are
stated, and latency/throughput values are labeled
MEASURED instead of fabricated.
1. AtlasMart problem: 400 alerts, no incident owner
A brief traffic spike pushes CPU, search queue depth, cache evictions, and p95 upward. Four hundred alerts fire, each to a different channel, while customer availability remains healthy. This is an alert-design failure: the monitoring system converted correlated symptoms into independent pages.
Most metrics are diagnostic context. Page on symptoms or rapidly consumed error budget; route lower-confidence resource anomalies as investigation signals unless they directly require operator action.
2. Condition types and when they fit
| Condition | Good use | Risk | AtlasMart example |
|---|---|---|---|
| Threshold | clear hard boundary | flaps near boundary | error rate > agreed limit for 10m |
| Rate/delta | counters and sudden changes | counter resets/window mistakes | rejections/minute rising |
| Multi-window SLO burn | fast/slow error-budget consumption | needs correct SLI math | availability budget burn |
| Anomaly | seasonal/unknown baselines | false positives, model warm-up, maturity | unexpected query-volume pattern |
An anomaly is not automatically an incident. It is evidence that behavior differs from a learned baseline and still needs impact/context.
3. Elastic: rules, alerts, actions, connectors
Kibana alerting is the generally available rules system. A rule evaluates conditions on a schedule; matching conditions produce alert instances; actions use connectors for notification or automation. Connector availability and individual solution rule types can vary by deployment/license. Elastic Stack 9.5 also documents a separate experimental alerting system; this chapter does not silently substitute that for GA Kibana rules.
Rule: atlasmart-search-latency
Schedule: every 1m
Lookback: 5m
Condition: p99(client_search_latency_ms) > SLO threshold AND request_count >= minimum_sample
Group by: environment, service
Recovery action: enabled
Owner: search-platform
Runbook: runbooks/search-tail-latency.md
Connector: local test webhook or no-op in mandatory free/local lab
4. OpenSearch: monitors, triggers, alerts, actions
OpenSearch Alerting uses monitors that run on a schedule. Triggers evaluate query results and actions send notifications. Current documentation includes per-query, per-bucket, PPL, per-document, and composite patterns, with security and cross-cluster limitations varying by monitor type. Dashboard alerting visualizations and Notifications are separate plugin/UI capabilities.
POST /_plugins/_alerting/monitors
{
"type": "monitor",
"name": "atlasmart-latency-v24",
"monitor_type": "query_level_monitor",
"enabled": true,
"schedule": {"period": {"interval": 1, "unit": "MINUTES"}},
"inputs": [{"search": {
"indices": ["atlasmart-ops-v24"],
"query": {"size": 0, "query": {"range": {"@timestamp": {"gte": "now-5m"}}}}
}}],
"triggers": []
}
# Add a trigger only after validating the exact aggregation path and sample-size guard.
Do not manually edit OpenSearch alerting system indexes. Use the plugin APIs and snapshot/export supported configuration where needed.
5. Noise-control mechanics
| Technique | Purpose |
|---|---|
| Persistence window | ignore one-sample spikes |
| Minimum sample | avoid unstable percentiles/rates |
| Grouping | one incident per service/environment rather than per host |
| Deduplication key | keep retries from creating new incidents |
| Action throttle | limit repeated notifications while condition persists |
| Recovery notification | prove the condition cleared and close the incident loop |
| Severity routing | page only when urgency and actionability justify interruption |
6. Controlled incident and causal verification
Reuse the disposable 30-second-refresh fixture from Lesson 1. The alert condition is not “refresh count changed”; it is “freshness lag exceeded the lab objective for a sustained window.” The runbook then checks refresh settings/counters and search visibility. After manual refresh or restoration of the normal interval, the recovery condition must be observed.
incident_id=atlasmart-v24-freshness-001
start_utc=MEASURED
symptom=freshness_lag_exceeded
alert_fired_utc=MEASURED
first_evidence=index_refresh_interval_30s
corrective_action=manual_refresh_then_restore_normal_interval
alert_recovered_utc=MEASURED
customer_impact=lab_only
notes=No production saturation; disposable index only.
7. Incident routing contract
Every paging alert needs an owner, urgency, runbook, deduplication key, and expected operator action. If there is no meaningful action, it is probably a dashboard event or ticket rather than a page. Route security incidents, relevance regressions, ingestion lag, and capacity warnings to different owners even when they share the same search cluster.
8. Production judgment
Start with a small alert set linked to SLOs and add resource alerts only when they predict or prevent user impact. Review false positives/negatives after incidents. Alert evaluation itself consumes query and scheduling resources; monitor its task health and avoid running expensive analytical queries every minute on the same cluster whose headroom you are protecting.
Check your understanding
- Why not page on every dashboard metric?
- What protects percentile alerts from tiny samples?
- What is the OpenSearch alerting execution model?
- Why send recovery notifications?
- Why monitor the monitoring system itself?
Review the answers
1. Most metrics are diagnostic context; paging should be actionable and tied to impact or imminent risk.
2. A minimum request-count guard plus a suitable evaluation window.
3. A scheduled monitor provides input, triggers evaluate conditions, and actions notify or automate.
4. They prove the alert condition cleared and close the operational feedback loop.
5. Rules/monitors consume resources and can fail or worsen an overloaded cluster.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic cluster stats API — Cluster-wide node, shard, store, JVM, CPU, and plugin context.
- Elastic node stats API — JVM, filesystem, thread pools, breakers, indexing pressure, search/indexing, caches, merge, refresh, and recovery metrics.
- Kibana alerting — Rules, schedules, alerts, and connector actions.
- Elastic SLO access — Current license, transform/ingest-role, space, and index-security requirements.
- Kibana saved objects — Dashboards, visualizations, data views, import/export, and permissions.
- OpenSearch Nodes Stats API — JVM, filesystem, process, thread-pool, and index metric evidence.
- OpenSearch Alerting — Monitor, trigger, alert, and action model.
- OpenSearch Query Insights: top N queries — Latency, CPU, and memory query evidence.
- OpenSearch Query Insights Dashboards — Live/top-N/configuration UI and query details.
- OpenSearch Observability — Dashboards, alerting/detection, traces, and experimental SLO scope.
- Elastic Stack 9.5.3 release and OpenSearch 3.8.0 artifacts — pinned September/August 2026 server baselines.