Execute migration through measurable gates with real rollback triggers and retained source authority.

Build a Migration Runbook with Data Validation, Relevance Regression, Performance Baselines, RPO, and Rollback Triggers

Turn Elasticsearch↔OpenSearch migration into an evidence-based compatibility program covering APIs, mappings, queries, clients, plugins, snapshots, security, managed-service boundaries, data sync, relevance and rollback.

Intermediate → Advanced175–235 minutesEnd-to-end migration runbook · Chapter 29 · Lesson 05Elasticsearch/Kibana 9.5.3 · elasticsearch-py 9.5.1 · OpenSearch/Dashboards 3.8.0 · opensearch-py 3.2.0Last reviewed: September 2026

Learning outcomes

01

Convert the entire migration into gated phases with measurable entry/exit criteria.

02

Validate data, mappings, queries, relevance, performance, security, and sync lag before endpoint cutover.

03

Define explicit rollback triggers and the data path required to make rollback real.

04

Run a small AtlasMart migration rehearsal that produces evidence instead of relying on operator memory.

05

Carry the measured source/target workload into Chapter 30 capacity planning instead of transferring benchmark folklore.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned migration baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. Where application-client behavior matters, the current documented Python baselines are elasticsearch-py 9.5.1 and opensearch-py 3.2.0; the mandatory migration lab itself uses curl/HTTP plus Python 3 standard-library code so the data path is inspectable and client-neutral. The generation environment did not execute live clusters, so timing, throughput, sync-lag, and relevance values in this chapter are acceptance criteria to measure locally—not fabricated results.

1. AtlasMart migration runbook: approval gates, not a checklist you can ignore

A runbook is executable only if each step has an owner, command/test, expected evidence, stop condition, and rollback action. “Validate target” is not executable. “Run golden-query suite v3; require zero authorization escapes, NDCG@10 degradation ≤ agreed threshold, and p99 within the approved SLO at load level L2” is executable.

2. Phase 0 — freeze the baseline

Artifact Required evidence
Source inventory server/UI/client/plugin versions; mappings/templates/pipelines/lifecycle/security/repositories
Workload baseline query mix, write rate, p50/p95/p99, errors/rejections, freshness/indexing lag
Relevance baseline golden queries + judgments + metrics
Data baseline count, selected-field hash/checksum sample, max update sequence/time
Recovery baseline fresh snapshot/checkpoint and tested restore/rollback path

3. Phase 1 — create target contracts before copying data

Create destination mappings, analyzers, templates, lifecycle policy, security roles, and aliases independently. Do not let automatic index creation decide production mapping. Translate non-portable features explicitly in the compatibility ledger.

All Chapter 29 labs preserve the established local endpoints and security assumptions: Elasticsearch 9.5.3 at https://localhost:9200 with ELASTIC_PASSWORD and the copied CA file atlasmart-es-http-ca; OpenSearch 3.8.0 at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD. The shared Docker network remains atlasmart-search. OpenSearch's demo certificate trust bypass (-k) is acceptable only for this disposable local lab. The migration fixture uses one primary and zero replicas to fit a single workstation; production redundancy, recovery headroom, and managed-service networking must be designed separately.

Portable AtlasMart migration source mapping
PUT atlasmart-migrate-source-v1
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0
  },
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "sku":       {"type":"keyword"},
      "tenant_id": {"type":"keyword"},
      "name":      {"type":"text", "fields":{"raw":{"type":"keyword"}}},
      "category":  {"type":"keyword"},
      "price":     {"type":"scaled_float", "scaling_factor":100},
      "available": {"type":"boolean"},
      "updated_at": {"type":"date"}
    }
  }
}
Seed a deterministic five-document fixture
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1001"}}
{"sku":"P-1001","tenant_id":"tenant-a","name":"Waterproof Hiking Boot","category":"footwear","price":129.90,"available":true,"updated_at":"2026-09-12T10:00:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1002"}}
{"sku":"P-1002","tenant_id":"tenant-a","name":"Trail Running Shoe","category":"footwear","price":99.50,"available":true,"updated_at":"2026-09-12T10:01:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1003"}}
{"sku":"P-1003","tenant_id":"tenant-a","name":"Insulated Water Bottle","category":"outdoor","price":32.00,"available":true,"updated_at":"2026-09-12T10:02:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1004"}}
{"sku":"P-1004","tenant_id":"tenant-b","name":"Lightweight Hiking Pack","category":"outdoor","price":74.00,"available":false,"updated_at":"2026-09-12T10:03:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1005"}}
{"sku":"P-1005","tenant_id":"tenant-b","name":"Merino Hiking Sock","category":"apparel","price":18.50,"available":true,"updated_at":"2026-09-12T10:04:00Z"}

For the lab, source is atlasmart-migrate-source-v1; target is atlasmart-migrate-target-v1. The exact source/target product direction can be reversed because the neutral ETL harness uses the shared search/Bulk subset.

4. Phase 2 — backfill and reconcile

Backfill acceptance commands
# Replace auth/TLS flags with the source or target product's established lab pattern.
GET atlasmart-migrate-source-v1/_count
GET atlasmart-migrate-target-v1/_count

POST atlasmart-migrate-source-v1/_search
{"size":0,"aggs":{"max_update":{"max":{"field":"updated_at"}},"tenants":{"terms":{"field":"tenant_id"}}}}

POST atlasmart-migrate-target-v1/_search
{"size":0,"aggs":{"max_update":{"max":{"field":"updated_at"}},"tenants":{"terms":{"field":"tenant_id"}}}}

Counts and aggregate checkpoints must agree, then sample/canonical-hash selected fields by ID. If production supports writes during backfill, continue incremental sync until measured lag satisfies the cutover RPO.

5. Phase 3 — dual-read behavioral validation

Golden-query comparison record
{
  "query_id":"catalog-hiking-tenant-a",
  "source_top_ids":["P-1001","P-1002"],
  "target_top_ids":["P-1001","P-1002"],
  "source_metrics":{"ndcg_at_3":"MEASURE","p99_ms":"MEASURE"},
  "target_metrics":{"ndcg_at_3":"MEASURE","p99_ms":"MEASURE"},
  "auth_negative_tests":{"tenant_b_visible":false,"write_allowed":false},
  "decision":"PASS_OR_FAIL_BY_DECLARED_THRESHOLDS"
}

Do not compare raw relevance scores across products as if they were calibrated. Compare judged ranking outcomes and user/business invariants.

6. Phase 4 — performance and saturation gate

Use the same dataset, cache/warmup disclosure, request mix, concurrency, TLS/auth, and hardware class where possible. Capture p50/p95/p99, throughput, errors/rejections, CPU, heap/GC, disk/network, merge/recovery, and freshness. The target need not reproduce the source internal metrics; it must satisfy the approved user-visible and recovery objectives.

Gate Example decision rule — replace with your measured SLO
availability no unexpected request failures during steady-state validation
tail latency target p99 within the approved application SLO
freshness new write searchable within approved lag
rejections no unexplained sustained 429/rejection increase at accepted load
relevance no agreed critical-query regression; aggregate metric above threshold
recovery node/service disruption drill recovers inside RTO budget

7. Phase 5 — security gate

Minimum negative-test set
reader identity:
  ALLOW search tenant-a catalog
  DENY index/update/delete
  DENY tenant-b documents
  DENY cluster admin and snapshot operations

ingester identity:
  ALLOW required write/bulk target
  DENY broad read if application does not need it
  DENY security/plugin administration

migration identity:
  ALLOW only source-read + target-write operations required for migration
  EXPIRE/REVOKE after migration

Security parity means equivalent least-privilege intent, not identical role JSON.

8. Phase 6 — cutover gate

Checkpoint Pass evidence Rollback trigger
sync lag ≤ agreed RPO and final delta checkpoint recorded lag exceeds RPO or reconciliation fails
application smoke read/write/facet/pagination flows pass functional error or unexpected response drift
auth all positive/negative tests pass any unauthorized access or required permission failure
performance tail latency/errors inside SLO sustained p99/error/rejection breach
relevance golden suite inside threshold critical relevance regression
observability target logs/alerts/runbook usable blindness to target failure

9. Rollback triggers must be machine-readable where possible

Example decision record
cutover_id: atlasmart-search-2026-09
rollback_window: "4h"
triggers:
  - "authorization escape or cross-tenant visibility"
  - "sustained application error rate above approved SLO"
  - "p99 above approved SLO for two consecutive windows"
  - "sync/reconciliation gap above RPO"
  - "critical relevance regression in golden suite"
rollback_actions:
  - "stop/redirect target writes according to write-ownership plan"
  - "replay target-only writes to source if required"
  - "restore source endpoint/service discovery"
  - "verify source smoke/security/relevance checks"
  - "preserve target evidence for incident analysis"

10. Do not delete the source at cutover

The rollback window closes only after the target has accumulated enough evidence under real traffic and source recovery is no longer required. Source retirement includes credentials, snapshots, repositories, monitoring, DNS, client configuration, data-retention obligations, and cost controls—not merely deleting the cluster.

Wrong approach. Define rollback as “switch DNS back” while target-only writes have no replay path and source credentials/snapshots have already been removed. Repair: keep source recoverable, journal/replay target-only writes, time-box the window, and require formal rollback-window closure before destruction.

11. Chapter 29 lab rehearsal

  1. Create the portable source and target mappings on opposite products.
  2. Run the five-document neutral ETL backfill.
  3. Compare counts, per-tenant buckets, selected fields by ID, and the golden hiking query.
  4. Translate one lifecycle intent (ILM↔ISM) on paper or in the disposable local clusters; do not copy JSON blindly.
  5. Create least-privilege reader/ingester test identities using each product’s own security model.
  6. Point a tiny test client at source, then target. Record response and latency distributions; do not fabricate values.
  7. Inject one failure: remove target read permission or alter a query mapping. Confirm the rollback trigger fires and restore source endpoint/config.
  8. Only after repair, rerun the complete gate set.

12. Migration evidence package

File/record Purpose
compatibility-matrix.json every feature classified and tested
mapping-analysis-diff.md schema/analyzer translation evidence
data-validation.json counts, hashes/samples, sync checkpoint
relevance-report.json golden query metrics and critical-query outcomes
performance-report.json load shape, p50/p95/p99, throughput, errors, resource evidence
security-negative-tests.json least-privilege proof
cutover-log.md timestamps, endpoint changes, smoke tests, owners
rollback-decision.yaml triggers, data-replay path, window closure

13. Bridge to Chapter 30: migration numbers become capacity inputs

Chapter 30 begins with Workload Characterization: QPS, Write Rate, Document Size, Fields, Retention, Aggregations, Vector Dimensions, and Concurrency. The migration performance report is not merely a go/no-go artifact; it becomes a measured workload baseline for capacity planning. Rebenchmark when query mix, mapping, shard layout, vectors, retention, hardware, or product version changes.

Check your understanding

  1. What makes a migration runbook executable?
  2. Why compare relevance metrics instead of raw scores?
  3. What must exist before cutover if rollback is required?
  4. When can the source be retired?
  5. What carries into Chapter 30?
Review the answers

1. Every phase has an owner, command/test, expected evidence, stop condition, and rollback action.

2. Raw score scales are not portable contracts; judged ranking outcomes are.

3. A real data path for target-only writes back to the source or another authoritative recovery mechanism.

4. After the rollback window formally closes and operational/security/retention cleanup is complete.

5. Measured query/write mix, latency/throughput, resource use, freshness, relevance, recovery behavior, and scaling assumptions.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.