Choose the platform from evidence—and preserve a credible way to leave it.

Present Elastic vs OpenSearch Architecture Choices with Evidence, Feature/License/Operations Tradeoffs, and Exit/Migration Strategy

Integrate the whole course into a production search platform whose model, relevance, vector/RAG retrieval, security, scaling, recovery, upgrade, monitoring, and platform choice are defended by evidence.

Intermediate → Advanced180–240 minutesPlatform decision & exit strategy · Chapter 31 · Lesson 05Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · free/local mandatory pathLast reviewed: September 2026

Learning outcomes

01

Compare Elasticsearch and OpenSearch with a weighted evidence matrix instead of a generic feature checklist.

02

Include licensing/support, security, vector/RAG, lifecycle, observability, managed responsibility, staffing, and migration cost in the architecture decision.

03

Separate “feature exists” from “feature is available in our exact distribution/subscription and proven under our workload.”

04

Create a long-term exit strategy with portable data contracts, nonportable-feature inventory, reindex/ETL validation, and rollback windows.

05

Present the complete capstone evidence pack and define the triggers that force re-evaluation after the course ends.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Capstone baseline. Examples are frozen to Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), with their bundled JVMs. The established disposable local endpoints remain Elasticsearch at https://localhost:9200 and OpenSearch at https://localhost:9201 on the atlasmart-search Docker network. The generation environment did not execute live clusters, so no latency, throughput, relevance, restore-time, or cost number is presented as measured unless the learner records it.

1. AtlasMart final decision: choose the operating model, not the logo

The capstone ends with a decision that can be audited months later. Elasticsearch and OpenSearch both provide mature distributed search foundations, but current product surfaces, licensing, managed-service responsibilities, security models, lifecycle automation, vector/AI stacks, analytics languages, plugins, and migration paths differ. The winning platform is the one that satisfies AtlasMart's weighted requirements with acceptable operational and exit cost—not the one with the longest feature list.

2. Evidence-weighted platform matrix

Decision matrix template
criteria:
  search_quality:
    weight: <0-100>
    elastic_evidence: <judged metrics + notes>
    opensearch_evidence: <judged metrics + notes>
  performance_and_capacity:
    weight: <0-100>
    elastic_evidence: <p95/p99/throughput/recovery/cost>
    opensearch_evidence: <p95/p99/throughput/recovery/cost>
  security_and_tenancy:
    weight: <0-100>
    elastic_evidence: <RBAC/DLS-FLS/subscription/negative tests>
    opensearch_evidence: <Security plugin/DLS-FLS/tenants/negative tests>
  lifecycle_and_recovery:
    weight: <0-100>
    elastic_evidence: <ILM/data tiers/snapshot RPO-RTO>
    opensearch_evidence: <ISM/snapshot RPO-RTO>
  vector_hybrid_rag:
    weight: <0-100>
    elastic_evidence: <quality/latency/feature boundaries>
    opensearch_evidence: <quality/latency/ML Commons/plugin boundaries>
  operations_and_staffing:
    weight: <0-100>
    elastic_evidence: <runbooks/support/skills>
    opensearch_evidence: <runbooks/support/skills>
  licensing_and_support:
    weight: <0-100>
    elastic_evidence: <Elastic License/subscription/support assumptions>
    opensearch_evidence: <Apache-2.0/project/provider support assumptions>
  exit_and_migration:
    weight: <0-100>
    elastic_evidence: <nonportable features + ETL/reindex effort>
    opensearch_evidence: <nonportable features + ETL/reindex effort>
rule:
  normalize_scores_only_after_evidence_is_complete: true
  feature_presence_without_proof_score: 0

Weights are business choices. Keep raw evidence beside any numerical score so a weighted total cannot hide a hard constraint such as an unapproved license, unacceptable RPO, or missing tenant isolation.

3. Feature/license/operations tradeoffs are specific, not ideological

Decision area Elasticsearch 9.5.3 questions OpenSearch 3.8.0 questions
Distribution/license Does Elastic License/subscription fit redistribution, support, and required feature policy? Does Apache-2.0 plus chosen support/managed provider fit governance?
Security Which RBAC/API-key/DLS/FLS/audit features are in the approved subscription/deployment? How will Security-plugin users/roles/backends/tenants/audit be initialized, backed up, and upgraded?
Lifecycle Do ILM/data tiers match topology and operations? Do ISM states/transitions/actions match retention/tiering intent?
Vector/AI Which dense/sparse/hybrid/inference/RAG features are approved and measured? Which k-NN/neural/ML Commons/search-pipeline features are approved and measured?
Observability Which Kibana/Elastic observability/SLO features are used and licensed? Which Dashboards/Alerting/Query Insights/observability plugins are used?
Managed service Which responsibilities move to Elastic Cloud and which remain ours? Which responsibilities/deviations exist in Amazon OpenSearch Service or another provider?
Migration What data and feature state remains portable through neutral export/reindex? What data and feature state remains portable through neutral export/reindex?

4. Exit strategy: preserve a portable source of truth

Do not make the only recoverable copy of business data a product-specific snapshot or proprietary feature-state index. Keep replayable source data or neutral exports where feasible, version mapping/analyzer intent in source control, and inventory product-specific features explicitly. Current snapshots are governed by product/version compatibility; they are not a guaranteed Elasticsearch↔OpenSearch interchange format. Between current Elasticsearch 9.5.3 and OpenSearch 3.8.0, the safe capstone migration model is explicit reindex/ETL/export-import with semantic validation, not assumed snapshot restore.

Exit inventory
portable_contracts:
  - source documents / neutral export
  - canonical field semantics
  - analyzer token tests
  - judged relevance queries
  - tenant/security intent
  - SLO/RPO/RTO targets
  - application-level API contract tests
product_specific_inventory:
  elastic:
    - ILM/data-tier policy artifacts
    - Elastic security/application privilege artifacts
    - retriever/inference/semantic feature configuration
    - Kibana saved objects/feature state
  opensearch:
    - ISM policies
    - Security-plugin users/roles/role mappings/tenants
    - ML Commons/neural/search-pipeline configuration
    - OpenSearch Dashboards saved objects/plugin state
migration_validation:
  - mapping/analyzer tests
  - document counts and deterministic sample hashes
  - lexical/vector/hybrid relevance regression
  - security-negative tests
  - p95/p99/freshness/rejection baselines
  - application/client integration tests
rollback:
  keep_source_writable_until: <validated-cutover-plus-window>
  trigger_on: <quality/performance/security/RPO gates>

5. Capstone evidence pack

Evidence artifact Minimum contents Decision it supports
Architecture decision record (ADR) requirements, constraints, alternatives, chosen design, rejected design, assumptions why this platform/topology is appropriate
Version/license matrix server, UI, client, plugin, license/subscription, managed-service deviations what behavior can actually be relied on
Schema/search tests mapping, analyzer tokens, queries, aggs, pagination, edge cases correctness and compatibility
Relevance evaluation judged queries, lexical/vector/hybrid baselines, Recall/NDCG/MRR-like metrics whether ranking improved rather than merely changed
Security-negative tests forbidden index/cluster actions, tenant leakage tests, credential rotation blast radius and isolation
Performance/recovery report p50/p95/p99, throughput, lag, rejections, resource evidence, recovery capacity and failure headroom
Snapshot/restore drill repository verification, snapshot ID, isolated restore, counts/schema/query checks, RPO/RTO recoverability
SLO dashboard/runbook user symptom, SLI, alert, evidence path, owner, action operability
Migration/exit runbook portable source-of-truth data, nonportable features, ETL/reindex path, cutover/rollback long-term exit cost
Final evidence-pack manifest
capstone/
  00-requirements-and-adr/
  01-version-license-matrix/
  02-schema-analyzer-template-tests/
  03-ingest-lifecycle-artifacts/
  04-query-aggregation-tests/
  05-relevance-vector-hybrid-evaluation/
  06-security-negative-tests/
  07-load-profile-capacity-report/
  08-snapshot-restore-rpo-rto/
  09-incident-upgrade-game-day/
  10-slo-dashboard-runbook/
  11-platform-decision-matrix/
  12-migration-exit-runbook/

Every directory contains:
  assumptions.md
  commands-or-code/
  expected-contracts/
  measured-results-or-NOT-RUN.md
  decision.md

The NOT-RUN record is important. It is more trustworthy than invented output. When a paid, managed, hardware-specific, or multi-node condition cannot be reproduced locally, preserve the exact test plan and state what remains unproven.

6. Wrong approach: choose from feature presence alone

Wrong approach: score a platform higher because it advertises a feature, regardless of subscription, plugin maturity, operator skill, measured quality, failure behavior, or migration cost. Repair: a feature earns decision value only when it is available in the pinned distribution, allowed by governance, integrated into the security/operations model, and demonstrated against AtlasMart acceptance gates.

7. Re-evaluation triggers

Trigger Why decision may change
server/UI/client/plugin upgrade APIs, defaults, compatibility, and security behavior can change
new workload or 2× data/traffic shape change capacity and shard/tier assumptions may no longer hold
new embedding/reranker/model quality, dimensions, cost, privacy, and reindex requirements change
SLO/RPO/RTO or compliance change topology, snapshots, retention, and security boundaries change
license/subscription/provider change feature availability and total cost/responsibility change
major incident or failed restore runbook/capacity assumptions have been falsified
migration/exit cost exceeds threshold architecture lock-in risk is no longer acceptable

8. Course completion: what “production-ready” means here

The final result is not “Elasticsearch is better” or “OpenSearch is better.” It is a defensible AtlasMart architecture with explicit assumptions, measured quality/performance, least privilege, tested recovery, bounded incident procedures, observable SLOs, version/license awareness, and a realistic exit path. If one of those remains untested, mark it as an open risk rather than hiding it behind a successful demo.

Return to the course curriculum to review any evidence area that remains weak. The course is complete when the platform decision can be challenged and reproduced by another engineer using the artifacts above.

Check your understanding

  1. Why is a weighted matrix better than a generic feature checklist?
  2. Why must snapshots be excluded from the generic cross-product exit assumption?
  3. What should remain portable even if product features differ?
  4. When should the platform decision be revisited?
  5. What is the capstone definition of production-ready?
Review the answers

1. It ties platform choice to AtlasMart priorities and measured evidence while keeping hard constraints and operating cost visible.

2. Snapshot restore is governed by product/version compatibility and feature-state semantics; it is not a guaranteed Elasticsearch↔OpenSearch interchange format.

3. Source data/neutral exports, canonical field semantics, analyzer tests, judged queries, security intent, SLO/RPO/RTO targets, and application contract tests.

4. After material version, workload, model, SLO/compliance, licensing/provider, incident/recovery, or migration-cost changes.

5. A reproducible evidence-backed system with measured quality/performance, least privilege, tested recovery/operations, observability, explicit version/license assumptions, and an exit strategy.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.