Define AtlasMart search reliability through availability, tail latency, freshness, indexing lag, errors/rejections, and relevance quality—with explicit populations and error budgets.
Search SLOs: Availability, p95/p99 Latency, Freshness, Indexing Lag, Error/Rejection Rate, and Relevance Quality
Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.
Learning outcomes
Define availability, p95/p99 latency, freshness, indexing lag, error/rejection rate, and relevance quality as separate SLIs with explicit populations and windows.
Convert an SLI target into an error budget without confusing averages, percentiles, rates, or internal cluster metrics.
Use client/synthetic evidence as the primary service view and correlate it with Elasticsearch/OpenSearch diagnostic telemetry.
Handle Elastic SLO feature prerequisites and OpenSearch experimental SLO status without making paid/preview features mandatory.
Create a deterministic AtlasMart SLO worksheet and regression gate that remains portable across both products.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0, using
the course's established local TLS/auth conventions and bundled
JVMs. Elastic metrics examples use supported cluster, node,
index, task, slow-log, Kibana rule, saved-object/data-view, and
SLO surfaces. OpenSearch examples use node/index stats, Alerting
monitors, Dashboards, and Query Insights where
installed/enabled. OpenSearch's native SLO feature is currently
documented as experimental; Elastic SLO
management has license and node-role prerequisites, so the
mandatory lab also includes a product-neutral SLI/error-budget
worksheet. No live cluster is available in this generation
environment: commands are reproducible, expected invariants are
stated, and latency/throughput values are labeled
MEASURED instead of fabricated.
1. From “fast search” to an operational contract
“Search should be fast” cannot be alerted on or reviewed. An SLI is a measured indicator, an SLO is a target for that indicator over a defined population/window, and an error budget is the amount of allowed bad service implied by the target. These are service contracts, not Elasticsearch settings.
2. AtlasMart SLI catalog
| SLI | Good-event definition | Evidence source | Boundary to state |
|---|---|---|---|
| Availability | eligible search request returns accepted success semantics | client/synthetic probe | which status/partial responses count |
| Latency | eligible request completes within objective | client timing | endpoint, cache state, geography, percentile window |
| Freshness | known write becomes searchable within objective | write/search marker probe | refresh semantics and expected indexing path |
| Indexing lag | event-time to searchable-time lag within objective | pipeline timestamps + probe | clock sync and late-event policy |
| Error/rejection | request not rejected/failed under defined semantics | client plus server counters | retry policy and partial results |
| Relevance | golden query meets judged metric/constraint | offline evaluation + production guard | query set, judgments, model/index version |
3. Tail latency: p95 and p99 are not averages
Percentiles answer distribution questions. p99 is the value at or below which 99% of eligible observations fall for the chosen window/population. Do not derive p99 by averaging per-node p99 values, and do not compare percentiles from different request populations as if they are the same SLI. Record both tail latency and throughput because a system can meet latency only by rejecting work.
Averages hide tails. AtlasMart's primary interactive-search SLI uses client p95/p99 plus success/rejection counts.
4. Error-budget math
For an availability target of T over
N eligible requests, the allowed bad-event budget
is N × (1 − T). For a time-based SLO, multiply the
window duration by (1 − T). This arithmetic is
deterministic; the actual good/bad counts are measured.
slo_name: atlasmart_interactive_search
window: 30d
population: prod catalog search requests excluding synthetic/admin traffic
availability_target: 99.90%
latency_objective: p99 <= AGREED_MS
freshness_objective: marker_searchable_within <= AGREED_SECONDS
rejection_objective: rejection_rate <= AGREED_RATE
relevance_gate: golden_query_suite >= AGREED_SCORE
budget_math:
availability_bad_fraction = 1 - 0.999 = 0.001
allowed_bad_requests = eligible_requests * 0.001
all_runtime_values: MEASURED
5. Elastic and OpenSearch SLO product features
Elastic's Observability SLO feature can manage SLO definitions and error-budget summaries, but current documentation requires an appropriate license and Elasticsearch nodes with transform and ingest roles. Kibana spaces scope the SLO UI but do not replace index authorization. OpenSearch documentation currently labels its SLO feature experimental and ties it to Prometheus-compatible ruler concepts. Therefore the course does not make either UI a prerequisite: the canonical lab artifact is the product-neutral SLI/SLO worksheet plus raw evidence.
6. Relevance belongs in reliability
A search endpoint can return 200 in 40 ms and still fail users because ranking changed. Maintain a versioned golden query set with judgments and a chosen metric (for example recall@k, NDCG@k, or a simpler must-contain constraint appropriate to the product). Run it before analyzer, synonym, vector-model, or ranking changes and after incident recovery. Do not turn one offline metric into a universal business KPI; state the query set and judgment process.
query_id,q,required_top_k_product
q001,"waterproof hiking backpack",P-1005
q002,"usb c charger",P-1002
q003,"noise cancelling headphones",P-1004
# Gate example:
# PASS when every required product appears within the agreed top-k for the pinned fixture.
# Production quality still needs broader judgments and business validation.
7. Freshness probe design
Use a uniquely identified marker document. Record the write acknowledgment, then poll a normal search path at a bounded interval until the marker appears or the objective expires. Clean up the marker. This measures end-to-end search freshness for that path; it is stronger than reading only a refresh counter.
t_write_ack = now()
index(unique_marker)
while now() - t_write_ack < objective:
if ordinary_search(unique_marker).found:
record_freshness_lag(now() - t_write_ack)
break
sleep(bounded_poll_interval)
else:
record_bad_event("freshness")
cleanup(unique_marker)
8. Production judgment
Keep SLOs few, stable, and owned. Exclude traffic only with a documented reason, version the definition, and preserve historical changes. Use internal metrics as explanatory evidence, not as substitutes for the SLI. Pair fast-burn alerts with longer-window budget views so operators distinguish acute incidents from slow degradation.
Check your understanding
- What is the difference between an SLI and an SLO?
- Why not average node-level p99 values?
- Why include throughput with latency?
- Why is relevance an SLO candidate?
- Why is the product-neutral worksheet mandatory?
Review the answers
1. The SLI is the measured indicator; the SLO is its target over a defined population/window.
2. Percentiles are distribution statistics and are not composable by simple averaging.
3. A system can appear fast by rejecting or shedding work.
4. Fast successful responses can still fail user intent if ranking quality regresses.
5. Elastic SLO prerequisites and OpenSearch experimental status differ, so the learning path must remain free/local and portable.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and version checks
- Elastic cluster stats API — Cluster-wide node, shard, store, JVM, CPU, and plugin context.
- Elastic node stats API — JVM, filesystem, thread pools, breakers, indexing pressure, search/indexing, caches, merge, refresh, and recovery metrics.
- Kibana alerting — Rules, schedules, alerts, and connector actions.
- Elastic SLO access — Current license, transform/ingest-role, space, and index-security requirements.
- Kibana saved objects — Dashboards, visualizations, data views, import/export, and permissions.
- OpenSearch Nodes Stats API — JVM, filesystem, process, thread-pool, and index metric evidence.
- OpenSearch Alerting — Monitor, trigger, alert, and action model.
- OpenSearch Query Insights: top N queries — Latency, CPU, and memory query evidence.
- OpenSearch Query Insights Dashboards — Live/top-N/configuration UI and query details.
- OpenSearch Observability — Dashboards, alerting/detection, traces, and experimental SLO scope.
- Elastic Stack 9.5.3 release and OpenSearch 3.8.0 artifacts — pinned September/August 2026 server baselines.