Choose data movement from compatibility, downtime, RPO, transform needs, and rollback—not convenience.
Snapshot Compatibility Boundaries vs Remote Reindex/ETL, Full Copy, Dual Write, and Incremental Synchronization
Turn Elasticsearch↔OpenSearch migration into an evidence-based compatibility program covering APIs, mappings, queries, clients, plugins, snapshots, security, managed-service boundaries, data sync, relevance and rollback.
Learning outcomes
Distinguish snapshot restore, reindex-from-remote, neutral ETL, full copy, dual write, and incremental synchronization by their guarantees.
Choose a data movement path from downtime, RPO, source load, transform needs, and source/target compatibility.
Explain why destination mappings/templates/lifecycle/security must be created independently of document copy.
Measure backfill completeness and catch-up lag without assuming copy completion implies cutover readiness.
Run a small AtlasMart cross-product migration using an inspectable common-API ETL path and explicit rollback checkpoints.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: moving bytes is only one phase
AtlasMart can copy five product documents in seconds, but production migration must also preserve IDs, routing assumptions, analyzers, aliases, write ownership, authorization, lifecycle intent, and application behavior while new writes continue. The data path must therefore be chosen from the required RPO (maximum acceptable data loss), RTO (time to restore service), downtime window, and transformation needs.
2. Compare the movement mechanisms
| Mechanism | Best fit | Main limits / questions |
|---|---|---|
| Snapshot restore | documented compatible product/version path; fast bulk recovery | index-version/product compatibility; repository ownership; not arbitrary ES↔OS interchange |
| Remote reindex |
supported source protocol, _source enabled,
destination mappings precreated
|
source load, throughput, allowlist/TLS, remote API compatibility |
| Neutral ETL | cross-product translation and transform control | slower; app-owned retry/checkpoint/idempotency logic |
| Full copy + outage | small/medium data with bounded maintenance window | write freeze and RPO≈0 only if freeze is enforced |
| Dual write | low-downtime migration when app can write two targets | partial failure semantics, idempotency, ordering, reconciliation |
| CDC/capture-replay | high write rate + low downtime with an appropriate capture layer | operational complexity and product/service support |
3. Remote reindex is asymmetric across products
Elasticsearch documents reindex-from-remote from a remote
Elasticsearch cluster and requires the
destination coordinator to allow the source host. OpenSearch 3.8
documents cross-cluster reindexing from remote OpenSearch and
Elasticsearch clusters through
reindex.remote.allowlist. This does not mean every
Elasticsearch major is guaranteed to work as a remote OpenSearch
source. Treat the exact pair as a tested integration.
POST _reindex
{
"source": {
"remote": {
"host": "https://source.example:9200",
"username": "reindex-reader",
"password": "${SECRET}"
},
"index": "atlasmart-products-v1",
"size": 500
},
"dest": {
"index": "atlasmart-products-v2"
}
}
The remote credentials require only source read/monitor capabilities. Put secrets in a secure runtime mechanism, not a committed JSON file. TLS verification belongs in node/service configuration.
4. Mandatory local lab: neutral ETL over the common search/bulk subset
All Chapter 29 labs preserve the established local endpoints and
security assumptions: Elasticsearch 9.5.3 at
https://localhost:9200 with
ELASTIC_PASSWORD and the copied CA file
atlasmart-es-http-ca; OpenSearch 3.8.0 at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The shared
Docker network remains atlasmart-search.
OpenSearch's demo certificate trust bypass (-k) is
acceptable only for this disposable local lab. The migration
fixture uses one primary and zero replicas to fit a single
workstation; production redundancy, recovery headroom, and
managed-service networking must be designed separately.
Create the portable source fixture on one product and a
separately defined
atlasmart-migrate-target-v1 mapping on the other.
The script below intentionally avoids product-specific clients
and snapshot formats. It performs a small sorted search,
validates each document, and writes Bulk NDJSON to the
destination. It is deliberately bounded to the five-document
lab; production backfill needs pagination/checkpointing,
backpressure, retry budgets, and resumability.
PUT atlasmart-migrate-source-v1
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0
},
"mappings": {
"dynamic": "strict",
"properties": {
"sku": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"name": {"type":"text", "fields":{"raw":{"type":"keyword"}}},
"category": {"type":"keyword"},
"price": {"type":"scaled_float", "scaling_factor":100},
"available": {"type":"boolean"},
"updated_at": {"type":"date"}
}
}
}
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1001"}}
{"sku":"P-1001","tenant_id":"tenant-a","name":"Waterproof Hiking Boot","category":"footwear","price":129.90,"available":true,"updated_at":"2026-09-12T10:00:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1002"}}
{"sku":"P-1002","tenant_id":"tenant-a","name":"Trail Running Shoe","category":"footwear","price":99.50,"available":true,"updated_at":"2026-09-12T10:01:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1003"}}
{"sku":"P-1003","tenant_id":"tenant-a","name":"Insulated Water Bottle","category":"outdoor","price":32.00,"available":true,"updated_at":"2026-09-12T10:02:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1004"}}
{"sku":"P-1004","tenant_id":"tenant-b","name":"Lightweight Hiking Pack","category":"outdoor","price":74.00,"available":false,"updated_at":"2026-09-12T10:03:00Z"}
{"index":{"_index":"atlasmart-migrate-source-v1","_id":"P-1005"}}
{"sku":"P-1005","tenant_id":"tenant-b","name":"Merino Hiking Sock","category":"apparel","price":18.50,"available":true,"updated_at":"2026-09-12T10:04:00Z"}
PUT atlasmart-migrate-target-v1
{
"settings": {"number_of_shards":1,"number_of_replicas":0},
"mappings": {
"dynamic":"strict",
"properties": {
"sku":{"type":"keyword"},
"tenant_id":{"type":"keyword"},
"name":{"type":"text","fields":{"raw":{"type":"keyword"}}},
"category":{"type":"keyword"},
"price":{"type":"scaled_float","scaling_factor":100},
"available":{"type":"boolean"},
"updated_at":{"type":"date"}
}
}
}
# atlasmart_migrate_small.py -- deterministic lab only, not a production copier
import base64, json, ssl, sys, urllib.request
src, src_user, src_pass, src_ca, dst, dst_user, dst_pass, dst_ca = sys.argv[1:9]
def context(ca):
if ca == "INSECURE_LAB_ONLY":
return ssl._create_unverified_context()
return ssl.create_default_context(cafile=ca)
def call(base, user, password, ca, method, path, body=None, ctype="application/json"):
token=base64.b64encode(f"{user}:{password}".encode()).decode()
data=None if body is None else body.encode()
req=urllib.request.Request(base+path, data=data, method=method,
headers={"Authorization":"Basic "+token,"Content-Type":ctype})
with urllib.request.urlopen(req, context=context(ca), timeout=20) as r:
return r.status, r.read().decode()
query=json.dumps({"size":100,"query":{"match_all":{}},"sort":[{"sku":"asc"}]})
status, raw=call(src,src_user,src_pass,src_ca,"POST","/atlasmart-migrate-source-v1/_search",query)
hits=json.loads(raw)["hits"]["hits"]
assert len(hits)==5, f"expected 5 lab docs, got {len(hits)}"
lines=[]
for hit in hits:
doc=hit["_source"]
required={"sku","tenant_id","name","category","price","available","updated_at"}
assert required <= doc.keys()
lines.append(json.dumps({"index":{"_index":"atlasmart-migrate-target-v1","_id":hit["_id"]}}))
lines.append(json.dumps(doc,separators=(",",":")))
bulk="\n".join(lines)+"\n"
status, raw=call(dst,dst_user,dst_pass,dst_ca,"POST","/_bulk?refresh=wait_for",bulk,"application/x-ndjson")
result=json.loads(raw)
assert not result.get("errors"), raw
print(json.dumps({"copied":len(hits),"target":"atlasmart-migrate-target-v1"}))
Example direction: Elasticsearch source uses
atlasmart-es-http-ca; OpenSearch target in this
disposable lab may pass INSECURE_LAB_ONLY because
its demo certificate is self-signed. Reverse the endpoints to
test the opposite direction after creating the source fixture
there. Production must use verified TLS.
5. Translate one non-portable feature explicitly: lifecycle intent
Do not copy index.lifecycle.name from Elastic and
expect OpenSearch to execute it. Express the business
requirement first—for example, “roll append-only logs by
measured size/age and delete after the approved retention
period”—then implement Elastic ILM or OpenSearch ISM
independently. The policies have different JSON, state models,
APIs, and troubleshooting surfaces.
| Intent | Elastic implementation | OpenSearch implementation |
|---|---|---|
| roll over write index/data stream | ILM rollover action | ISM rollover action/state |
| move colder / reduce cost | data-tier/ILM actions where supported | allocation/ISM actions or managed-service equivalent |
| delete after retention | ILM delete phase | ISM delete state/action |
| observe stuck policy | ILM explain/status | ISM explain/status |
6. Full copy is not synchronization
After a baseline copy, source writes create
sync lag. Define a high-water mark using an
application-owned monotonic value such as
updated_at plus a deterministic ID, or capture
writes through the application/event stream. A timestamp alone
can miss equal-time updates and clock anomalies; pair it with an
ID or sequence contract.
POST atlasmart-migrate-source-v1/_search
{
"size": 100,
"query": {
"range": {"updated_at": {"gte":"2026-09-12T10:02:00Z"}}
},
"sort": [
{"updated_at":"asc"},
{"sku":"asc"}
]
}
For production, use stable application sequencing or a proven capture/replay mechanism when write order and exactly-once-like effects matter. Dual writes without reconciliation merely duplicate failure modes.
7. Validation after copy
| Check | Evidence | Why count-only is insufficient |
|---|---|---|
| Document count | source/target _count |
duplicates or mismatched fields can preserve counts |
| Field/mapping | _mapping diff + schema tests |
same JSON source can index differently |
| Content | canonical selected-field hash by ID | detects silent transform drift |
| Queries | golden top-k IDs + aggregations | proves behavior, not just storage |
| Security | positive + 403 negative tests | admin copy may hide application authorization failures |
| Performance | same harness/warmup/concurrency | functional parity can still violate SLO |
8. Rollback starts before cutover
Keep the source authoritative until target validation is complete. If dual writing begins, define how source remains recoverable: either source continues to receive writes, or a replayable event log can reconstruct them. “We can change DNS back” is not rollback if writes after cutover exist only on the target.
Check your understanding
- Why does remote reindex still require destination mappings first?
- What makes dual write risky?
- Why is a timestamp-only incremental cursor weak?
- What does a successful document count prove?
- What does Lesson 3 add?
Review the answers
1. Because reindex copies document source; destination mappings/settings define how those documents are indexed and searched.
2. One target can succeed while the other fails, creating ordering/idempotency/reconciliation problems.
3. Equal timestamps, retries, or clock behavior can make ordering ambiguous; pair it with a stable sequence/ID contract.
4. Only that counts match; it does not prove mapping, content, relevance, authorization, or performance parity.
5. Application-client, query-language, deprecated-API, and feature-replacement testing on top of the data movement path.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Elasticsearch snapshot and restore compatibility
- Elasticsearch restore snapshot guidance
- Elasticsearch reindex and reindex-from-remote
- Elasticsearch reindex settings
- Elasticsearch Python client compatibility
- Elasticsearch Python client release notes
- OpenSearch 3.8 version history
- OpenSearch upgrade or migrate guidance
- OpenSearch Reindex Documents API
- OpenSearch reindex data guidance
- OpenSearch language clients and compatibility
- OpenSearch Migration Assistant
- Migration Assistant supported migration paths
- Amazon OpenSearch Service snapshot migration