Prove operations under load, security boundaries, recovery, maintenance, and incident conditions.
Load-Test, Profile, Secure, Snapshot/Restore, Fail Over, Upgrade, Monitor, and Execute Operational Incident Drills
Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.
Learning outcomes
Run a bounded mixed-load test with explicit tail-latency, freshness, relevance, rejection, and recovery gates.
Use Profile/Explain/slow logs/node statistics as diagnostic evidence without confusing them with production benchmarks.
Prove least privilege with negative tests and credential-rotation evidence on the chosen product path.
Execute an isolated snapshot/restore drill and record RPO/RTO rather than assuming replicas are backups.
Run a controlled incident/upgrade game day with stop conditions, monitoring, rollback checkpoints, and a post-incident evidence trail.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: a platform is not finished when the feature tests pass
Lesson 4 asks whether AtlasMart survives the operational conditions that invalidate happy-path demos: load spikes, expensive queries, credential misuse, deleted indices, node/region assumptions, version changes, and operator mistakes. Every drill has a precondition, a bounded action, a stop condition, an observation set, and a rollback.
2. Mixed-load acceptance matrix
dataset: atlasmart-capstone-v1
warmup_s: <fixed>
measurement_s: <fixed>
load_levels:
normal_peak:
search_qps: <target>
write_ops_s: <target>
stress_gate:
search_qps: <bounded-higher-target>
write_ops_s: <bounded-higher-target>
workload_mix:
lexical_filter: <percent>
facet_aggregation: <percent>
vector_hybrid: <percent>
writes_updates: <percent>
acceptance:
p95_ms: <target>
p99_ms: <target>
indexing_visibility_p99_ms: <target>
rejection_rate: <target>
ndcg_at_10_floor: <target>
vector_recall_at_10_floor: <target>
recovery_after_spike_s: <target>
stop_conditions:
- swap_or_oom_risk
- disk_watermark_risk
- sustained_rejections_beyond_budget
- correctness_or_security_failure
- unrelated_host_impact
Record offered load and achieved throughput separately. Report p50/p95/p99, not averages alone. A tuning change is rejected if it wins throughput by violating freshness, relevance, durability, recovery, or security.
3. Profile and diagnostics answer “why,” not “how fast”
Use query Profile output, Explain, shard slow logs, hot threads, task APIs, node/index stats, GC, filesystem, thread pools, indexing pressure, OpenSearch search/shard indexing backpressure where relevant, and client traces to identify the mechanism behind a symptom. Their instrumentation overhead and scope mean they are not substitutes for an external latency benchmark. Align timestamps so client p99 can be correlated with server-side evidence.
GET _cluster/health
GET _cat/nodes?v
GET _cat/shards?v
GET _nodes/stats/jvm,process,os,fs,indices,thread_pool,indexing_pressure
GET _tasks?detailed=true&actions=*search*
GET _nodes/hot_threads
GET atlasmart-products-v1/_stats
GET _cat/recovery?v
4. Security drill: allowed success plus forbidden failure
Create separate read/search and ingest machine principals. Do not use administrator credentials in the application. Prove the reader can search only the intended indices/tenant boundary and cannot write/delete/manage security. Prove the ingester can write required targets but cannot read unrelated protected data or change cluster settings. Rotate a short-lived/scoped credential and prove the old credential no longer works according to the product's revocation/expiration semantics.
| Test | Expected result | Evidence |
|---|---|---|
| reader searches authorized tenant | 200 + only authorized IDs | request + identity + result IDs |
| reader writes/deletes | 403/denied | audit/security event where available |
| ingester writes target | success | bulk per-item success |
| ingester cluster-admin action | 403/denied | request + role/permission inspection |
| cross-tenant semantic hit | absent from candidates/context | retrieval trace |
| old rotated credential | fails after invalidation/expiry contract | timestamped request |
5. Snapshot/restore drill: recovery, not replication
Take a snapshot to the disposable filesystem-compatible repository configured in Chapter 18, record the last included write timestamp, remove or rename only a disposable capstone index, restore under an isolated target name, then validate mappings/settings, document count, sample IDs/checksum-like fixture hashes, aliases/templates where in scope, and representative queries. Do not blindly restore global/security state. The observed distance from incident time to last recoverable write is RPO evidence; time from declared recovery start to application-ready validation is RTO evidence.
{
"snapshot": "atlasmart-capstone-<timestamp>",
"source_index": "atlasmart-products-v1",
"restore_index": "atlasmart-products-restore-drill-v1",
"last_recoverable_write_time": "<measured>",
"recovery_start_time": "<measured>",
"application_ready_time": "<measured>",
"rpo_seconds": null,
"rto_seconds": null,
"document_count_match": null,
"mapping_contract_match": null,
"sample_query_match": null,
"security_state_restored": false
}
6. Failover and upgrade game day
The mandatory one-node lab simulates the decision gates without claiming quorum/failover. The optional multi-node disposable lab may restart exactly one data node or perform a version-safe rolling maintenance sequence only after checking compatibility, plugin/client versions, snapshot state, shard redundancy, and recovery headroom. Stop when health/recovery/SLO criteria fail. Never assume downgrade is supported; rollback may require restoring a pre-change cluster/data path.
PRECHECK
[ ] current snapshot verified
[ ] exact server/UI/client/plugin versions recorded
[ ] compatibility/breaking changes reviewed
[ ] shard redundancy/recovery headroom adequate for the drill
[ ] alerting and client smoke tests active
[ ] stop/rollback owner assigned
AFTER EACH NODE/CHANGE
[ ] expected version/roles/plugins
[ ] quorum/cluster-manager or master safety intact
[ ] shard recovery within gate
[ ] application read/write smoke tests pass
[ ] p95/p99/freshness/rejections within gate
[ ] no security regression
STOP on any failed gate; do not “continue and see.”
7. Monitoring and incident timeline
Build one operations view that starts with user-facing SLIs, then links to cluster/node/index/query/ingest/relevance evidence. Alert on actionable symptoms and rates, not every metric. Keep monitoring cost/cardinality bounded so the observability workload does not consume the cluster's remaining failure headroom.
| User symptom | First SLI | Next evidence | Likely runbook branch |
|---|---|---|---|
| slow search | p95/p99 by query class | queues, CPU, disk, hot threads, profile/slow log | query shape, saturation, storage, shard fan-out |
| stale catalog | visibility/indexing lag | write failures, refresh, merge/backpressure | ingest backlog/freshness |
| missing tenant data | security/retrieval trace | identity/role mapping, filters, index targets | authorization or query contract |
| bad ranking | NDCG/MRR regression | component result sets, analyzer/model version | relevance rollback |
| restore too slow | RTO timer | repository/network/recovery stats | capacity/repository plan |
8. Wrong approach: “push through” failed gates
9. Production judgment
The operational capstone is accepted only when recovery, security, performance, monitoring, and maintenance are repeatable procedures rather than tribal knowledge. Lesson 5 uses that evidence to present a platform decision and, just as importantly, an exit strategy.
Check your understanding
- Why are Profile results not benchmark numbers?
- Why must security tests include expected failures?
- Why are replicas not a backup?
- What should happen when an upgrade gate fails?
- What makes monitoring actionable?
Review the answers
1. Profiling adds instrumentation and describes execution work; it is diagnostic evidence, not an external user-latency measurement.
2. Least privilege is proven by both required actions succeeding and forbidden actions being denied.
3. Replicas mirror logical changes such as deletions/corruption; snapshots provide an independent recovery point with different failure coverage.
4. Stop the sequence, preserve evidence, and use the preplanned rollback/recovery decision path rather than continuing.
5. It begins with user SLIs, links to diagnostic evidence, has ownership and runbook actions, and controls telemetry/noise cost.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Download Elasticsearch 9.5.3
- Elasticsearch mappings
- Elasticsearch text analysis
- Elasticsearch index templates
- Elasticsearch ingest pipelines
- Elasticsearch data streams
- Elasticsearch ILM
- Elasticsearch Query DSL
- Elasticsearch aggregations
- Elasticsearch vector search
- Elasticsearch hybrid search
- Elasticsearch security
- Elasticsearch snapshot and restore
- Elasticsearch performance guidance
- Elasticsearch subscription feature matrix
- OpenSearch 3.8 version history
- OpenSearch downloads and Apache 2.0 licensing
- OpenSearch mappings and field types
- OpenSearch index templates
- OpenSearch data streams
- OpenSearch ingest pipelines
- OpenSearch Index State Management
- OpenSearch query DSL
- OpenSearch aggregations
- OpenSearch vector search
- OpenSearch hybrid search
- OpenSearch Security plugin
- OpenSearch snapshot and restore
- OpenSearch performance tuning
- OpenSearch Benchmark