Create actionable AtlasMart alerts with thresholds, rates, persistence, noise controls, recovery evidence, and product-correct Kibana rule/OpenSearch monitor semantics.

Alerting/Rules/Monitors, Thresholds vs Anomaly/Rate Conditions, Noise Control, and Incident Routing

Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.

Intermediate → Advanced120–160 minutesAlert-noise & recovery lab · Chapter 24 · Lesson 03Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Design alerts from actionable service symptoms and diagnostic evidence rather than mirroring every dashboard metric.

02

Distinguish static thresholds, rates, burn-rate style conditions, and anomaly detection, including their data and maturity requirements.

03

Compare Kibana rules/connectors with OpenSearch monitors/triggers/actions without assuming API or execution equivalence.

04

Control alert noise through windows, persistence, recovery notifications, grouping, deduplication, routing, and ownership.

05

Inject one safe AtlasMart incident, produce firing and recovery evidence, and verify that the runbook leads to the causal metric.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned observability baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0, using the course's established local TLS/auth conventions and bundled JVMs. Elastic metrics examples use supported cluster, node, index, task, slow-log, Kibana rule, saved-object/data-view, and SLO surfaces. OpenSearch examples use node/index stats, Alerting monitors, Dashboards, and Query Insights where installed/enabled. OpenSearch's native SLO feature is currently documented as experimental; Elastic SLO management has license and node-role prerequisites, so the mandatory lab also includes a product-neutral SLI/error-budget worksheet. No live cluster is available in this generation environment: commands are reproducible, expected invariants are stated, and latency/throughput values are labeled MEASURED instead of fabricated.

1. AtlasMart problem: 400 alerts, no incident owner

A brief traffic spike pushes CPU, search queue depth, cache evictions, and p95 upward. Four hundred alerts fire, each to a different channel, while customer availability remains healthy. This is an alert-design failure: the monitoring system converted correlated symptoms into independent pages.

Wrong approach: alert on every metric threshold.

Most metrics are diagnostic context. Page on symptoms or rapidly consumed error budget; route lower-confidence resource anomalies as investigation signals unless they directly require operator action.

2. Condition types and when they fit

Condition Good use Risk AtlasMart example
Threshold clear hard boundary flaps near boundary error rate > agreed limit for 10m
Rate/delta counters and sudden changes counter resets/window mistakes rejections/minute rising
Multi-window SLO burn fast/slow error-budget consumption needs correct SLI math availability budget burn
Anomaly seasonal/unknown baselines false positives, model warm-up, maturity unexpected query-volume pattern

An anomaly is not automatically an incident. It is evidence that behavior differs from a learned baseline and still needs impact/context.

3. Elastic: rules, alerts, actions, connectors

Kibana alerting is the generally available rules system. A rule evaluates conditions on a schedule; matching conditions produce alert instances; actions use connectors for notification or automation. Connector availability and individual solution rule types can vary by deployment/license. Elastic Stack 9.5 also documents a separate experimental alerting system; this chapter does not silently substitute that for GA Kibana rules.

Illustrative Kibana rule contract (conceptual, not copied into a hidden index)
Rule: atlasmart-search-latency
Schedule: every 1m
Lookback: 5m
Condition: p99(client_search_latency_ms) > SLO threshold AND request_count >= minimum_sample
Group by: environment, service
Recovery action: enabled
Owner: search-platform
Runbook: runbooks/search-tail-latency.md
Connector: local test webhook or no-op in mandatory free/local lab

4. OpenSearch: monitors, triggers, alerts, actions

OpenSearch Alerting uses monitors that run on a schedule. Triggers evaluate query results and actions send notifications. Current documentation includes per-query, per-bucket, PPL, per-document, and composite patterns, with security and cross-cluster limitations varying by monitor type. Dashboard alerting visualizations and Notifications are separate plugin/UI capabilities.

OpenSearch: minimal per-query monitor skeleton
POST /_plugins/_alerting/monitors
{
  "type": "monitor",
  "name": "atlasmart-latency-v24",
  "monitor_type": "query_level_monitor",
  "enabled": true,
  "schedule": {"period": {"interval": 1, "unit": "MINUTES"}},
  "inputs": [{"search": {
    "indices": ["atlasmart-ops-v24"],
    "query": {"size": 0, "query": {"range": {"@timestamp": {"gte": "now-5m"}}}}
  }}],
  "triggers": []
}
# Add a trigger only after validating the exact aggregation path and sample-size guard.

Do not manually edit OpenSearch alerting system indexes. Use the plugin APIs and snapshot/export supported configuration where needed.

5. Noise-control mechanics

Technique Purpose
Persistence window ignore one-sample spikes
Minimum sample avoid unstable percentiles/rates
Grouping one incident per service/environment rather than per host
Deduplication key keep retries from creating new incidents
Action throttle limit repeated notifications while condition persists
Recovery notification prove the condition cleared and close the incident loop
Severity routing page only when urgency and actionability justify interruption

6. Controlled incident and causal verification

Reuse the disposable 30-second-refresh fixture from Lesson 1. The alert condition is not “refresh count changed”; it is “freshness lag exceeded the lab objective for a sustained window.” The runbook then checks refresh settings/counters and search visibility. After manual refresh or restoration of the normal interval, the recovery condition must be observed.

Incident evidence record
incident_id=atlasmart-v24-freshness-001
start_utc=MEASURED
symptom=freshness_lag_exceeded
alert_fired_utc=MEASURED
first_evidence=index_refresh_interval_30s
corrective_action=manual_refresh_then_restore_normal_interval
alert_recovered_utc=MEASURED
customer_impact=lab_only
notes=No production saturation; disposable index only.

7. Incident routing contract

Every paging alert needs an owner, urgency, runbook, deduplication key, and expected operator action. If there is no meaningful action, it is probably a dashboard event or ticket rather than a page. Route security incidents, relevance regressions, ingestion lag, and capacity warnings to different owners even when they share the same search cluster.

8. Production judgment

Start with a small alert set linked to SLOs and add resource alerts only when they predict or prevent user impact. Review false positives/negatives after incidents. Alert evaluation itself consumes query and scheduling resources; monitor its task health and avoid running expensive analytical queries every minute on the same cluster whose headroom you are protecting.

Check your understanding

  1. Why not page on every dashboard metric?
  2. What protects percentile alerts from tiny samples?
  3. What is the OpenSearch alerting execution model?
  4. Why send recovery notifications?
  5. Why monitor the monitoring system itself?
Review the answers

1. Most metrics are diagnostic context; paging should be actionable and tied to impact or imminent risk.

2. A minimum request-count guard plus a suitable evaluation window.

3. A scheduled monitor provides input, triggers evaluate conditions, and actions notify or automate.

4. They prove the alert condition cleared and close the operational feedback loop.

5. Rules/monitors consume resources and can fail or worsen an overloaded cluster.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.