Production begins by turning business promises into measurable and rejectable contracts.

Define Search/Analytics Requirements, Relevance Judgments, Freshness, Retention, Security, SLOs, RPO/RTO, and Cost Constraints

Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.

Intermediate → Advanced160–220 minutesRequirements, ADR & acceptance gates · Chapter 31 · Lesson 01Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · free/local mandatory pathLast reviewed: September 2026

Learning outcomes

01

Translate AtlasMart business/search goals into explicit correctness, relevance, freshness, security, SLO, RPO/RTO, retention, and cost requirements.

02

Separate an SLO from an SLI, an RPO from an RTO, and relevance quality from infrastructure availability.

03

Build a version/license/feature matrix before architecture decisions so Elastic and OpenSearch differences are visible rather than assumed away.

04

Define acceptance gates and negative tests that can fail the capstone before implementation proceeds.

05

Produce the ADR and evidence contract that Lessons 2–5 will satisfy.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Capstone baseline. Examples are frozen to Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), with their bundled JVMs. The established disposable local endpoints remain Elasticsearch at https://localhost:9200 and OpenSearch at https://localhost:9201 on the atlasmart-search Docker network. The generation environment did not execute live clusters, so no latency, throughput, relevance, restore-time, or cost number is presented as measured unless the learner records it.

1. AtlasMart problem: “production-ready” must become a falsifiable claim

AtlasMart now has catalog search, facets, telemetry, vector retrieval, RAG support, multiple tenants, and an operations team that must recover from failures. A demo that returns ten relevant products is not a production architecture. The capstone starts by converting each business promise into a measurable contract: what must be correct, how quickly new data must become searchable, which users may see which documents, how ranking quality is judged, what outage/data-loss envelope is acceptable, and what the platform is allowed to cost.

An SLI is the measured indicator, such as successful search requests with p99 under the target. An SLO is the agreed objective for that indicator. RPO bounds tolerable data loss measured backward from an incident; RTO bounds the time to restore an acceptable service. None of these is implied by a green cluster state.

2. Requirements must include search quality

Availability without useful ranking is not success. Keep at least three independent quality views: lexical judgments for common catalog queries, vector ground truth for semantic candidates, and end-to-end RAG checks that verify context provenance/faithfulness separately from retrieval recall. BM25 scores, cosine/dot-product scores, and fused ranks are not interchangeable units, so acceptance gates should use judged metrics and stable query sets rather than raw score thresholds copied across retrieval modes.

Requirement class Example capstone SLI Typical failure hidden by infrastructure-only monitoring
Correctness schema/analyzer/query tests pass wrong analyzer silently changes matching
Relevance NDCG@10 / MRR-like metric on judged queries results remain fast but business ranking regresses
Freshness write→search visibility p99 indexing backlog grows while searches remain healthy
Availability successful requests meeting contract cluster green but clients receive timeouts/403s
Recovery isolated restore RPO/RTO drill replicas survive a node loss but deletion/corruption has no backup path
Security tenant and forbidden-operation negative tests UI tenant looks isolated while backing index is over-permissioned
Cost normalized cost per workload at required SLO cheapest node fails recovery headroom or quality targets

3. Freeze versions, licensing, and responsibility boundaries

Record the exact server, UI, client, plugin, JVM, and deployment distribution. Elasticsearch 9.5.3 and OpenSearch 3.8.0 share ancestry but are not drop-in equivalents. Elastic ILM/data tiers, Elastic security/application privileges, retrievers/inference integrations, and subscription-sensitive features must be evaluated on Elastic's current matrix. OpenSearch ISM, Security plugin role mappings/tenants, ML Commons, k-NN/neural/hybrid features, and managed-service deviations have their own APIs and lifecycle. The capstone must never claim API parity merely because two requests look similar.

Version/license decision record
elastic:
  server: 9.5.3
  ui: Kibana 9.5.3
  distribution_license: Elastic License
  feature_matrix_reviewed_at: 2026-09-12
  paid_or_managed_features_used_by_mandatory_lab: false
opensearch:
  server: 3.8.0
  ui: OpenSearch Dashboards 3.8.0
  distribution_license: Apache-2.0
  feature_matrix_reviewed_at: 2026-09-12
  paid_or_managed_features_used_by_mandatory_lab: false
shared_claims:
  rest_api_identity: false
  snapshot_cross_product_portability: false
  lifecycle_policy_identity: false
  security_model_identity: false

4. Wrong capstone: a happy-path feature tour

Wrong approach: choose Elasticsearch or OpenSearch from a feature checklist, ingest sample data, run one keyword query and one vector query, then call the platform production-ready. This proves only that selected APIs answered. Repair: require an evidence bundle that can reject the design: judged quality, security-negative tests, mixed-load tails, snapshot restore, controlled incident recovery, upgrade/version gates, and a migration/exit plan.
Evidence artifact Minimum contents Decision it supports
Architecture decision record (ADR) requirements, constraints, alternatives, chosen design, rejected design, assumptions why this platform/topology is appropriate
Version/license matrix server, UI, client, plugin, license/subscription, managed-service deviations what behavior can actually be relied on
Schema/search tests mapping, analyzer tokens, queries, aggs, pagination, edge cases correctness and compatibility
Relevance evaluation judged queries, lexical/vector/hybrid baselines, Recall/NDCG/MRR-like metrics whether ranking improved rather than merely changed
Security-negative tests forbidden index/cluster actions, tenant leakage tests, credential rotation blast radius and isolation
Performance/recovery report p50/p95/p99, throughput, lag, rejections, resource evidence, recovery capacity and failure headroom
Snapshot/restore drill repository verification, snapshot ID, isolated restore, counts/schema/query checks, RPO/RTO recoverability
SLO dashboard/runbook user symptom, SLI, alert, evidence path, owner, action operability
Migration/exit runbook portable source-of-truth data, nonportable features, ETL/reindex path, cutover/rollback long-term exit cost

5. Reproducible local capstone boundary

All mandatory capstone work keeps the established free/local security boundary: Elasticsearch 9.5.3 at https://localhost:9200 authenticated with ELASTIC_PASSWORD and the copied CA atlasmart-es-http-ca; OpenSearch 3.8.0 at https://localhost:9201 authenticated with OPENSEARCH_INITIAL_ADMIN_PASSWORD. Both remain on the disposable atlasmart-search Docker network. Any OpenSearch -k example is limited to the upstream demo self-signed certificate in the disposable lab and is never production TLS guidance. The mandatory one-node path cannot prove node-failure high availability; multi-node failover is an optional disposable extension or a deterministic runbook simulation.

AtlasMart capstone requirements contract
capstone_id: atlasmart-search-platform-v1
pinned_versions:
  elasticsearch: 9.5.3
  kibana: 9.5.3
  opensearch: 3.8.0
  opensearch_dashboards: 3.8.0
business:
  workloads: [catalog_search, faceted_analytics, telemetry_search, rag_support]
  tenant_boundary: tenant_id
quality:
  lexical_judgments: atlasmart-judgments-v1
  vector_ground_truth: atlasmart-vector-ground-truth-v1
  ndcg_at_10_floor: <measured/approved>
  recall_at_10_floor: <measured/approved>
freshness:
  catalog_visibility_p99_ms: <target>
  telemetry_visibility_p99_ms: <target>
reliability:
  availability_slo: <target>
  rpo_seconds: <target>
  rto_seconds: <target>
security:
  public_anonymous_access: false
  application_admin_credentials: false
  tenant_filter_required: true
  credential_rotation_days: <policy>
capacity:
  normal_peak_qps: <target>
  write_ops_s: <target>
  failure_headroom_percent: <target>
retention:
  catalog: <business-policy>
  telemetry: <policy>
cost:
  monthly_budget: <currency/value>
  measurement_method: <documented-source>
exit:
  maximum_acceptable_migration_window: <hours>
  source_data_export_format: ndjson/jsonl
  nonportable_features: <inventory>

The placeholders are intentional. Populate them with organization-approved targets before measuring. A capstone that invents targets after seeing results is optimizing the acceptance criteria rather than the system.

6. Acceptance gates before Lesson 2

Gate the implementation
GO only if all are explicit:
[ ] user-facing search/analytics requirements
[ ] judged relevance set and vector ground truth
[ ] freshness targets
[ ] retention and legal/compliance constraints
[ ] tenant/security model and forbidden operations
[ ] SLO + error-budget definition
[ ] RPO/RTO and backup scope
[ ] peak/failure workload contract
[ ] cost-accounting method
[ ] pinned version/license matrix
[ ] rollback/exit requirements

STOP if a mandatory requirement depends on an unapproved paid/managed feature
without a documented free/local learning equivalent.

7. Production judgment

Requirements are architecture inputs, not documentation written after implementation. If AtlasMart's RTO requires a multi-node standby but the available budget only funds a one-node workstation, the correct result is a visible constraint conflict—not a misleading “high availability” claim. Lesson 2 now turns the approved contracts into mappings, analyzers, templates, ingest, lifecycle, and core search behavior.

Check your understanding

  1. Why is cluster health not an availability SLO?
  2. Why must relevance be in the capstone contract?
  3. What is the difference between RPO and RTO?
  4. Why pin versions and licensing before implementation?
  5. What makes the ADR falsifiable?
Review the answers

1. Cluster health primarily describes shard allocation state; users can still see timeouts, authorization failures, stale data, or bad relevance while health is green.

2. A search platform can become faster or cheaper by returning worse results. Judged metrics prevent performance tuning from silently sacrificing business search quality.

3. RPO bounds tolerable data loss; RTO bounds time to restore an acceptable service.

4. Behavior, feature availability, security boundaries, clients, plugins, and managed-service responsibilities are version/distribution dependent.

5. It includes acceptance gates and evidence that can reject the chosen architecture, not only reasons supporting it.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.