Production begins by turning business promises into measurable and rejectable contracts.
Define Search/Analytics Requirements, Relevance Judgments, Freshness, Retention, Security, SLOs, RPO/RTO, and Cost Constraints
Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.
Learning outcomes
Translate AtlasMart business/search goals into explicit correctness, relevance, freshness, security, SLO, RPO/RTO, retention, and cost requirements.
Separate an SLO from an SLI, an RPO from an RTO, and relevance quality from infrastructure availability.
Build a version/license/feature matrix before architecture decisions so Elastic and OpenSearch differences are visible rather than assumed away.
Define acceptance gates and negative tests that can fail the capstone before implementation proceeds.
Produce the ADR and evidence contract that Lessons 2–5 will satisfy.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: “production-ready” must become a falsifiable claim
AtlasMart now has catalog search, facets, telemetry, vector retrieval, RAG support, multiple tenants, and an operations team that must recover from failures. A demo that returns ten relevant products is not a production architecture. The capstone starts by converting each business promise into a measurable contract: what must be correct, how quickly new data must become searchable, which users may see which documents, how ranking quality is judged, what outage/data-loss envelope is acceptable, and what the platform is allowed to cost.
An SLI is the measured indicator, such as successful search requests with p99 under the target. An SLO is the agreed objective for that indicator. RPO bounds tolerable data loss measured backward from an incident; RTO bounds the time to restore an acceptable service. None of these is implied by a green cluster state.
2. Requirements must include search quality
Availability without useful ranking is not success. Keep at least three independent quality views: lexical judgments for common catalog queries, vector ground truth for semantic candidates, and end-to-end RAG checks that verify context provenance/faithfulness separately from retrieval recall. BM25 scores, cosine/dot-product scores, and fused ranks are not interchangeable units, so acceptance gates should use judged metrics and stable query sets rather than raw score thresholds copied across retrieval modes.
| Requirement class | Example capstone SLI | Typical failure hidden by infrastructure-only monitoring |
|---|---|---|
| Correctness | schema/analyzer/query tests pass | wrong analyzer silently changes matching |
| Relevance | NDCG@10 / MRR-like metric on judged queries | results remain fast but business ranking regresses |
| Freshness | write→search visibility p99 | indexing backlog grows while searches remain healthy |
| Availability | successful requests meeting contract | cluster green but clients receive timeouts/403s |
| Recovery | isolated restore RPO/RTO drill | replicas survive a node loss but deletion/corruption has no backup path |
| Security | tenant and forbidden-operation negative tests | UI tenant looks isolated while backing index is over-permissioned |
| Cost | normalized cost per workload at required SLO | cheapest node fails recovery headroom or quality targets |
3. Freeze versions, licensing, and responsibility boundaries
Record the exact server, UI, client, plugin, JVM, and deployment distribution. Elasticsearch 9.5.3 and OpenSearch 3.8.0 share ancestry but are not drop-in equivalents. Elastic ILM/data tiers, Elastic security/application privileges, retrievers/inference integrations, and subscription-sensitive features must be evaluated on Elastic's current matrix. OpenSearch ISM, Security plugin role mappings/tenants, ML Commons, k-NN/neural/hybrid features, and managed-service deviations have their own APIs and lifecycle. The capstone must never claim API parity merely because two requests look similar.
elastic:
server: 9.5.3
ui: Kibana 9.5.3
distribution_license: Elastic License
feature_matrix_reviewed_at: 2026-09-12
paid_or_managed_features_used_by_mandatory_lab: false
opensearch:
server: 3.8.0
ui: OpenSearch Dashboards 3.8.0
distribution_license: Apache-2.0
feature_matrix_reviewed_at: 2026-09-12
paid_or_managed_features_used_by_mandatory_lab: false
shared_claims:
rest_api_identity: false
snapshot_cross_product_portability: false
lifecycle_policy_identity: false
security_model_identity: false
4. Wrong capstone: a happy-path feature tour
| Evidence artifact | Minimum contents | Decision it supports |
|---|---|---|
| Architecture decision record (ADR) | requirements, constraints, alternatives, chosen design, rejected design, assumptions | why this platform/topology is appropriate |
| Version/license matrix | server, UI, client, plugin, license/subscription, managed-service deviations | what behavior can actually be relied on |
| Schema/search tests | mapping, analyzer tokens, queries, aggs, pagination, edge cases | correctness and compatibility |
| Relevance evaluation | judged queries, lexical/vector/hybrid baselines, Recall/NDCG/MRR-like metrics | whether ranking improved rather than merely changed |
| Security-negative tests | forbidden index/cluster actions, tenant leakage tests, credential rotation | blast radius and isolation |
| Performance/recovery report | p50/p95/p99, throughput, lag, rejections, resource evidence, recovery | capacity and failure headroom |
| Snapshot/restore drill | repository verification, snapshot ID, isolated restore, counts/schema/query checks, RPO/RTO | recoverability |
| SLO dashboard/runbook | user symptom, SLI, alert, evidence path, owner, action | operability |
| Migration/exit runbook | portable source-of-truth data, nonportable features, ETL/reindex path, cutover/rollback | long-term exit cost |
5. Reproducible local capstone boundary
All mandatory capstone work keeps the established free/local
security boundary: Elasticsearch 9.5.3 at
https://localhost:9200 authenticated with
ELASTIC_PASSWORD and the copied CA
atlasmart-es-http-ca; OpenSearch 3.8.0 at
https://localhost:9201 authenticated with
OPENSEARCH_INITIAL_ADMIN_PASSWORD. Both remain on
the disposable atlasmart-search Docker network. Any
OpenSearch -k example is limited to the upstream
demo self-signed certificate in the disposable lab and is never
production TLS guidance. The mandatory one-node path cannot
prove node-failure high availability; multi-node failover is an
optional disposable extension or a deterministic runbook
simulation.
capstone_id: atlasmart-search-platform-v1
pinned_versions:
elasticsearch: 9.5.3
kibana: 9.5.3
opensearch: 3.8.0
opensearch_dashboards: 3.8.0
business:
workloads: [catalog_search, faceted_analytics, telemetry_search, rag_support]
tenant_boundary: tenant_id
quality:
lexical_judgments: atlasmart-judgments-v1
vector_ground_truth: atlasmart-vector-ground-truth-v1
ndcg_at_10_floor: <measured/approved>
recall_at_10_floor: <measured/approved>
freshness:
catalog_visibility_p99_ms: <target>
telemetry_visibility_p99_ms: <target>
reliability:
availability_slo: <target>
rpo_seconds: <target>
rto_seconds: <target>
security:
public_anonymous_access: false
application_admin_credentials: false
tenant_filter_required: true
credential_rotation_days: <policy>
capacity:
normal_peak_qps: <target>
write_ops_s: <target>
failure_headroom_percent: <target>
retention:
catalog: <business-policy>
telemetry: <policy>
cost:
monthly_budget: <currency/value>
measurement_method: <documented-source>
exit:
maximum_acceptable_migration_window: <hours>
source_data_export_format: ndjson/jsonl
nonportable_features: <inventory>
The placeholders are intentional. Populate them with organization-approved targets before measuring. A capstone that invents targets after seeing results is optimizing the acceptance criteria rather than the system.
6. Acceptance gates before Lesson 2
GO only if all are explicit:
[ ] user-facing search/analytics requirements
[ ] judged relevance set and vector ground truth
[ ] freshness targets
[ ] retention and legal/compliance constraints
[ ] tenant/security model and forbidden operations
[ ] SLO + error-budget definition
[ ] RPO/RTO and backup scope
[ ] peak/failure workload contract
[ ] cost-accounting method
[ ] pinned version/license matrix
[ ] rollback/exit requirements
STOP if a mandatory requirement depends on an unapproved paid/managed feature
without a documented free/local learning equivalent.
7. Production judgment
Requirements are architecture inputs, not documentation written after implementation. If AtlasMart's RTO requires a multi-node standby but the available budget only funds a one-node workstation, the correct result is a visible constraint conflict—not a misleading “high availability” claim. Lesson 2 now turns the approved contracts into mappings, analyzers, templates, ingest, lifecycle, and core search behavior.
Check your understanding
- Why is cluster health not an availability SLO?
- Why must relevance be in the capstone contract?
- What is the difference between RPO and RTO?
- Why pin versions and licensing before implementation?
- What makes the ADR falsifiable?
Review the answers
1. Cluster health primarily describes shard allocation state; users can still see timeouts, authorization failures, stale data, or bad relevance while health is green.
2. A search platform can become faster or cheaper by returning worse results. Judged metrics prevent performance tuning from silently sacrificing business search quality.
3. RPO bounds tolerable data loss; RTO bounds time to restore an acceptable service.
4. Behavior, feature availability, security boundaries, clients, plugins, and managed-service responsibilities are version/distribution dependent.
5. It includes acceptance gates and evidence that can reject the chosen architecture, not only reasons supporting it.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Download Elasticsearch 9.5.3
- Elasticsearch mappings
- Elasticsearch text analysis
- Elasticsearch index templates
- Elasticsearch ingest pipelines
- Elasticsearch data streams
- Elasticsearch ILM
- Elasticsearch Query DSL
- Elasticsearch aggregations
- Elasticsearch vector search
- Elasticsearch hybrid search
- Elasticsearch security
- Elasticsearch snapshot and restore
- Elasticsearch performance guidance
- Elasticsearch subscription feature matrix
- OpenSearch 3.8 version history
- OpenSearch downloads and Apache 2.0 licensing
- OpenSearch mappings and field types
- OpenSearch index templates
- OpenSearch data streams
- OpenSearch ingest pipelines
- OpenSearch Index State Management
- OpenSearch query DSL
- OpenSearch aggregations
- OpenSearch vector search
- OpenSearch hybrid search
- OpenSearch Security plugin
- OpenSearch snapshot and restore
- OpenSearch performance tuning
- OpenSearch Benchmark