Define AtlasMart search reliability through availability, tail latency, freshness, indexing lag, errors/rejections, and relevance quality—with explicit populations and error budgets.

Search SLOs: Availability, p95/p99 Latency, Freshness, Indexing Lag, Error/Rejection Rate, and Relevance Quality

Create an operational evidence model that links user-visible search symptoms to cluster, node, index, query, ingest, and relevance signals and turns them into actionable SLOs.

Intermediate → Advanced120–165 minutesSLO/error-budget & relevance-gate lab · Chapter 24 · Lesson 04Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Define availability, p95/p99 latency, freshness, indexing lag, error/rejection rate, and relevance quality as separate SLIs with explicit populations and windows.

02

Convert an SLI target into an error budget without confusing averages, percentiles, rates, or internal cluster metrics.

03

Use client/synthetic evidence as the primary service view and correlate it with Elasticsearch/OpenSearch diagnostic telemetry.

04

Handle Elastic SLO feature prerequisites and OpenSearch experimental SLO status without making paid/preview features mandatory.

05

Create a deterministic AtlasMart SLO worksheet and regression gate that remains portable across both products.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned observability baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0, using the course's established local TLS/auth conventions and bundled JVMs. Elastic metrics examples use supported cluster, node, index, task, slow-log, Kibana rule, saved-object/data-view, and SLO surfaces. OpenSearch examples use node/index stats, Alerting monitors, Dashboards, and Query Insights where installed/enabled. OpenSearch's native SLO feature is currently documented as experimental; Elastic SLO management has license and node-role prerequisites, so the mandatory lab also includes a product-neutral SLI/error-budget worksheet. No live cluster is available in this generation environment: commands are reproducible, expected invariants are stated, and latency/throughput values are labeled MEASURED instead of fabricated.

1. From “fast search” to an operational contract

“Search should be fast” cannot be alerted on or reviewed. An SLI is a measured indicator, an SLO is a target for that indicator over a defined population/window, and an error budget is the amount of allowed bad service implied by the target. These are service contracts, not Elasticsearch settings.

2. AtlasMart SLI catalog

SLI Good-event definition Evidence source Boundary to state
Availability eligible search request returns accepted success semantics client/synthetic probe which status/partial responses count
Latency eligible request completes within objective client timing endpoint, cache state, geography, percentile window
Freshness known write becomes searchable within objective write/search marker probe refresh semantics and expected indexing path
Indexing lag event-time to searchable-time lag within objective pipeline timestamps + probe clock sync and late-event policy
Error/rejection request not rejected/failed under defined semantics client plus server counters retry policy and partial results
Relevance golden query meets judged metric/constraint offline evaluation + production guard query set, judgments, model/index version

3. Tail latency: p95 and p99 are not averages

Percentiles answer distribution questions. p99 is the value at or below which 99% of eligible observations fall for the chosen window/population. Do not derive p99 by averaging per-node p99 values, and do not compare percentiles from different request populations as if they are the same SLI. Record both tail latency and throughput because a system can meet latency only by rejecting work.

Wrong approach: dashboard the mean and call it latency health.

Averages hide tails. AtlasMart's primary interactive-search SLI uses client p95/p99 plus success/rejection counts.

4. Error-budget math

For an availability target of T over N eligible requests, the allowed bad-event budget is N × (1 − T). For a time-based SLO, multiply the window duration by (1 − T). This arithmetic is deterministic; the actual good/bad counts are measured.

Portable SLO worksheet
slo_name: atlasmart_interactive_search
window: 30d
population: prod catalog search requests excluding synthetic/admin traffic
availability_target: 99.90%
latency_objective: p99 <= AGREED_MS
freshness_objective: marker_searchable_within <= AGREED_SECONDS
rejection_objective: rejection_rate <= AGREED_RATE
relevance_gate: golden_query_suite >= AGREED_SCORE
budget_math:
  availability_bad_fraction = 1 - 0.999 = 0.001
  allowed_bad_requests = eligible_requests * 0.001
all_runtime_values: MEASURED

5. Elastic and OpenSearch SLO product features

Elastic's Observability SLO feature can manage SLO definitions and error-budget summaries, but current documentation requires an appropriate license and Elasticsearch nodes with transform and ingest roles. Kibana spaces scope the SLO UI but do not replace index authorization. OpenSearch documentation currently labels its SLO feature experimental and ties it to Prometheus-compatible ruler concepts. Therefore the course does not make either UI a prerequisite: the canonical lab artifact is the product-neutral SLI/SLO worksheet plus raw evidence.

6. Relevance belongs in reliability

A search endpoint can return 200 in 40 ms and still fail users because ranking changed. Maintain a versioned golden query set with judgments and a chosen metric (for example recall@k, NDCG@k, or a simpler must-contain constraint appropriate to the product). Run it before analyzer, synonym, vector-model, or ranking changes and after incident recovery. Do not turn one offline metric into a universal business KPI; state the query set and judgment process.

Deterministic relevance gate fixture
query_id,q,required_top_k_product
q001,"waterproof hiking backpack",P-1005
q002,"usb c charger",P-1002
q003,"noise cancelling headphones",P-1004

# Gate example:
# PASS when every required product appears within the agreed top-k for the pinned fixture.
# Production quality still needs broader judgments and business validation.

7. Freshness probe design

Use a uniquely identified marker document. Record the write acknowledgment, then poll a normal search path at a bounded interval until the marker appears or the objective expires. Clean up the marker. This measures end-to-end search freshness for that path; it is stronger than reading only a refresh counter.

Freshness-probe pseudocode
t_write_ack = now()
index(unique_marker)
while now() - t_write_ack < objective:
    if ordinary_search(unique_marker).found:
        record_freshness_lag(now() - t_write_ack)
        break
    sleep(bounded_poll_interval)
else:
    record_bad_event("freshness")
cleanup(unique_marker)

8. Production judgment

Keep SLOs few, stable, and owned. Exclude traffic only with a documented reason, version the definition, and preserve historical changes. Use internal metrics as explanatory evidence, not as substitutes for the SLI. Pair fast-burn alerts with longer-window budget views so operators distinguish acute incidents from slow degradation.

Check your understanding

  1. What is the difference between an SLI and an SLO?
  2. Why not average node-level p99 values?
  3. Why include throughput with latency?
  4. Why is relevance an SLO candidate?
  5. Why is the product-neutral worksheet mandatory?
Review the answers

1. The SLI is the measured indicator; the SLO is its target over a defined population/window.

2. Percentiles are distribution statistics and are not composable by simple averaging.

3. A system can appear fast by rejecting or shedding work.

4. Fast successful responses can still fail user intent if ranking quality regresses.

5. Elastic SLO prerequisites and OpenSearch experimental status differ, so the learning path must remain free/local and portable.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.