Execute migration through measurable gates with real rollback triggers and retained source authority.
Build a Migration Runbook with Data Validation, Relevance Regression, Performance Baselines, RPO, and Rollback Triggers
Turn Elasticsearch↔OpenSearch migration into an evidence-based compatibility program covering APIs, mappings, queries, clients, plugins, snapshots, security, managed-service boundaries, data sync, relevance and rollback.
Learning outcomes
Convert the entire migration into gated phases with measurable entry/exit criteria.
Validate data, mappings, queries, relevance, performance, security, and sync lag before endpoint cutover.
Define explicit rollback triggers and the data path required to make rollback real.
Run a small AtlasMart migration rehearsal that produces evidence instead of relying on operator memory.
Carry the measured source/target workload into Chapter 30 capacity planning instead of transferring benchmark folklore.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart migration runbook: approval gates, not a checklist you can ignore
A runbook is executable only if each step has an owner, command/test, expected evidence, stop condition, and rollback action. “Validate target” is not executable. “Run golden-query suite v3; require zero authorization escapes, NDCG@10 degradation ≤ agreed threshold, and p99 within the approved SLO at load level L2” is executable.
2. Phase 0 — freeze the baseline
| Artifact | Required evidence |
|---|---|
| Source inventory | server/UI/client/plugin versions; mappings/templates/pipelines/lifecycle/security/repositories |
| Workload baseline | query mix, write rate, p50/p95/p99, errors/rejections, freshness/indexing lag |
| Relevance baseline | golden queries + judgments + metrics |
| Data baseline | count, selected-field hash/checksum sample, max update sequence/time |
| Recovery baseline | fresh snapshot/checkpoint and tested restore/rollback path |
3. Phase 1 — create target contracts before copying data
Create destination mappings, analyzers, templates, lifecycle policy, security roles, and aliases independently. Do not let automatic index creation decide production mapping. Translate non-portable features explicitly in the compatibility ledger.
All Chapter 29 labs preserve the established local endpoints and
security assumptions: Elasticsearch 9.5.3 at
https://localhost:9200 with
ELASTIC_PASSWORD and the copied CA file
atlasmart-es-http-ca; OpenSearch 3.8.0 at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The shared
Docker network remains atlasmart-search.
OpenSearch's demo certificate trust bypass (-k) is
acceptable only for this disposable local lab. The migration
fixture uses one primary and zero replicas to fit a single
workstation; production redundancy, recovery headroom, and
managed-service networking must be designed separately.
PUT atlasmart-migrate-source-v1
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0
},
"mappings": {
"dynamic": "strict",
"properties": {
"sku": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"name": {"type":"text", "fields":{"raw":{"type":"keyword"}}},
"category": {"type":"keyword"},
"price": {"type":"scaled_float", "scaling_factor":100},
"available": {"type":"boolean"},
"updated_at": {"type":"date"}
}
}
}
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1001"}}
{"sku":"P-1001","tenant_id":"tenant-a","name":"Waterproof Hiking Boot","category":"footwear","price":129.90,"available":true,"updated_at":"2026-09-12T10:00:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1002"}}
{"sku":"P-1002","tenant_id":"tenant-a","name":"Trail Running Shoe","category":"footwear","price":99.50,"available":true,"updated_at":"2026-09-12T10:01:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1003"}}
{"sku":"P-1003","tenant_id":"tenant-a","name":"Insulated Water Bottle","category":"outdoor","price":32.00,"available":true,"updated_at":"2026-09-12T10:02:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1004"}}
{"sku":"P-1004","tenant_id":"tenant-b","name":"Lightweight Hiking Pack","category":"outdoor","price":74.00,"available":false,"updated_at":"2026-09-12T10:03:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1005"}}
{"sku":"P-1005","tenant_id":"tenant-b","name":"Merino Hiking Sock","category":"apparel","price":18.50,"available":true,"updated_at":"2026-09-12T10:04:00Z"}
For the lab, source is atlasmart-migrate-source-v1;
target is atlasmart-migrate-target-v1. The exact
source/target product direction can be reversed because the
neutral ETL harness uses the shared search/Bulk subset.
4. Phase 2 — backfill and reconcile
# Replace auth/TLS flags with the source or target product's established lab pattern.
GET atlasmart-migrate-source-v1/_count
GET atlasmart-migrate-target-v1/_count
POST atlasmart-migrate-source-v1/_search
{"size":0,"aggs":{"max_update":{"max":{"field":"updated_at"}},"tenants":{"terms":{"field":"tenant_id"}}}}
POST atlasmart-migrate-target-v1/_search
{"size":0,"aggs":{"max_update":{"max":{"field":"updated_at"}},"tenants":{"terms":{"field":"tenant_id"}}}}
Counts and aggregate checkpoints must agree, then sample/canonical-hash selected fields by ID. If production supports writes during backfill, continue incremental sync until measured lag satisfies the cutover RPO.
5. Phase 3 — dual-read behavioral validation
{
"query_id":"catalog-hiking-tenant-a",
"source_top_ids":["P-1001","P-1002"],
"target_top_ids":["P-1001","P-1002"],
"source_metrics":{"ndcg_at_3":"MEASURE","p99_ms":"MEASURE"},
"target_metrics":{"ndcg_at_3":"MEASURE","p99_ms":"MEASURE"},
"auth_negative_tests":{"tenant_b_visible":false,"write_allowed":false},
"decision":"PASS_OR_FAIL_BY_DECLARED_THRESHOLDS"
}
Do not compare raw relevance scores across products as if they were calibrated. Compare judged ranking outcomes and user/business invariants.
6. Phase 4 — performance and saturation gate
Use the same dataset, cache/warmup disclosure, request mix, concurrency, TLS/auth, and hardware class where possible. Capture p50/p95/p99, throughput, errors/rejections, CPU, heap/GC, disk/network, merge/recovery, and freshness. The target need not reproduce the source internal metrics; it must satisfy the approved user-visible and recovery objectives.
| Gate | Example decision rule — replace with your measured SLO |
|---|---|
| availability | no unexpected request failures during steady-state validation |
| tail latency | target p99 within the approved application SLO |
| freshness | new write searchable within approved lag |
| rejections | no unexplained sustained 429/rejection increase at accepted load |
| relevance | no agreed critical-query regression; aggregate metric above threshold |
| recovery | node/service disruption drill recovers inside RTO budget |
7. Phase 5 — security gate
reader identity:
ALLOW search tenant-a catalog
DENY index/update/delete
DENY tenant-b documents
DENY cluster admin and snapshot operations
ingester identity:
ALLOW required write/bulk target
DENY broad read if application does not need it
DENY security/plugin administration
migration identity:
ALLOW only source-read + target-write operations required for migration
EXPIRE/REVOKE after migration
Security parity means equivalent least-privilege intent, not identical role JSON.
8. Phase 6 — cutover gate
| Checkpoint | Pass evidence | Rollback trigger |
|---|---|---|
| sync lag | ≤ agreed RPO and final delta checkpoint recorded | lag exceeds RPO or reconciliation fails |
| application smoke | read/write/facet/pagination flows pass | functional error or unexpected response drift |
| auth | all positive/negative tests pass | any unauthorized access or required permission failure |
| performance | tail latency/errors inside SLO | sustained p99/error/rejection breach |
| relevance | golden suite inside threshold | critical relevance regression |
| observability | target logs/alerts/runbook usable | blindness to target failure |
9. Rollback triggers must be machine-readable where possible
cutover_id: atlasmart-search-2026-09
rollback_window: "4h"
triggers:
- "authorization escape or cross-tenant visibility"
- "sustained application error rate above approved SLO"
- "p99 above approved SLO for two consecutive windows"
- "sync/reconciliation gap above RPO"
- "critical relevance regression in golden suite"
rollback_actions:
- "stop/redirect target writes according to write-ownership plan"
- "replay target-only writes to source if required"
- "restore source endpoint/service discovery"
- "verify source smoke/security/relevance checks"
- "preserve target evidence for incident analysis"
10. Do not delete the source at cutover
The rollback window closes only after the target has accumulated enough evidence under real traffic and source recovery is no longer required. Source retirement includes credentials, snapshots, repositories, monitoring, DNS, client configuration, data-retention obligations, and cost controls—not merely deleting the cluster.
11. Chapter 29 lab rehearsal
- Create the portable source and target mappings on opposite products.
- Run the five-document neutral ETL backfill.
- Compare counts, per-tenant buckets, selected fields by ID, and the golden hiking query.
- Translate one lifecycle intent (ILM↔ISM) on paper or in the disposable local clusters; do not copy JSON blindly.
- Create least-privilege reader/ingester test identities using each product’s own security model.
- Point a tiny test client at source, then target. Record response and latency distributions; do not fabricate values.
- Inject one failure: remove target read permission or alter a query mapping. Confirm the rollback trigger fires and restore source endpoint/config.
- Only after repair, rerun the complete gate set.
12. Migration evidence package
| File/record | Purpose |
|---|---|
| compatibility-matrix.json | every feature classified and tested |
| mapping-analysis-diff.md | schema/analyzer translation evidence |
| data-validation.json | counts, hashes/samples, sync checkpoint |
| relevance-report.json | golden query metrics and critical-query outcomes |
| performance-report.json | load shape, p50/p95/p99, throughput, errors, resource evidence |
| security-negative-tests.json | least-privilege proof |
| cutover-log.md | timestamps, endpoint changes, smoke tests, owners |
| rollback-decision.yaml | triggers, data-replay path, window closure |
13. Bridge to Chapter 30: migration numbers become capacity inputs
Chapter 30 begins with Workload Characterization: QPS, Write Rate, Document Size, Fields, Retention, Aggregations, Vector Dimensions, and Concurrency. The migration performance report is not merely a go/no-go artifact; it becomes a measured workload baseline for capacity planning. Rebenchmark when query mix, mapping, shard layout, vectors, retention, hardware, or product version changes.
Check your understanding
- What makes a migration runbook executable?
- Why compare relevance metrics instead of raw scores?
- What must exist before cutover if rollback is required?
- When can the source be retired?
- What carries into Chapter 30?
Review the answers
1. Every phase has an owner, command/test, expected evidence, stop condition, and rollback action.
2. Raw score scales are not portable contracts; judged ranking outcomes are.
3. A real data path for target-only writes back to the source or another authoritative recovery mechanism.
4. After the rollback window formally closes and operational/security/retention cleanup is complete.
5. Measured query/write mix, latency/throughput, resource use, freshness, relevance, recovery behavior, and scaling assumptions.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Elasticsearch snapshot and restore compatibility
- Elasticsearch restore snapshot guidance
- Elasticsearch reindex and reindex-from-remote
- Elasticsearch reindex settings
- Elasticsearch Python client compatibility
- Elasticsearch Python client release notes
- OpenSearch 3.8 version history
- OpenSearch upgrade or migrate guidance
- OpenSearch Reindex Documents API
- OpenSearch reindex data guidance
- OpenSearch language clients and compatibility
- OpenSearch Migration Assistant
- Migration Assistant supported migration paths
- Amazon OpenSearch Service snapshot migration