Add cloud networking, IAM, encryption, repositories, DNS, and service restrictions to the migration model.

Managed-Service Migration: Networking, IAM, Encryption, Snapshot Repositories, DNS/Endpoint Cutover, and Rollback

Turn Elasticsearch↔OpenSearch migration into an evidence-based compatibility program covering APIs, mappings, queries, clients, plugins, snapshots, security, managed-service boundaries, data sync, relevance and rollback.

Intermediate → Advanced145–195 minutesManaged-service cutover design · Chapter 29 · Lesson 04Elasticsearch/Kibana 9.5.3 · elasticsearch-py 9.5.1 · OpenSearch/Dashboards 3.8.0 · opensearch-py 3.2.0Last reviewed: September 2026

Learning outcomes

01

Model managed-service migration as a network/identity/encryption project in addition to a data migration.

02

Separate self-managed credentials and repository access from cloud IAM, service roles, KMS, private networking, and endpoint policies.

03

Plan endpoint/DNS cutover around connection reuse, TTL, TLS names, health gates, and rollback data ownership.

04

Identify snapshot/reindex restrictions that are specific to Elastic Cloud or Amazon OpenSearch Service.

05

Build a deterministic managed-service checklist without requiring a paid cloud account for the mandatory lab.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned migration baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. Where application-client behavior matters, the current documented Python baselines are elasticsearch-py 9.5.1 and opensearch-py 3.2.0; the mandatory migration lab itself uses curl/HTTP plus Python 3 standard-library code so the data path is inspectable and client-neutral. The generation environment did not execute live clusters, so timing, throughput, sync-lag, and relevance values in this chapter are acceptance criteria to measure locally—not fabricated results.

1. AtlasMart problem: the target works from a laptop but not from production

A managed target can be perfectly healthy while application traffic fails because the VPC route is missing, private DNS resolves differently, the TLS certificate name does not match the endpoint, an IAM role lacks permissions, or a service disallows the repository/remote host pattern used in self-managed testing. Migration readiness therefore needs a connectivity contract before data cutover.

2. Managed-service boundary matrix

Concern Self-managed pattern Managed-service question
Network host/port + firewall public/private endpoint, VPC/VNet/PrivateLink, peering, egress rules
Identity native users/API keys/certs service IAM + product security + workload identity
Encryption local TLS/keystore provider TLS policy, KMS/customer-managed key options
Snapshots filesystem/S3/etc plugin provider-managed snapshots vs manual repository registration
Plugins install exact binary supported plugin/extension catalog only
Endpoint stable host you control service-generated endpoint, custom endpoint/DNS options
Upgrade node-by-node runbook provider-supported upgrade path/window

3. Networking acceptance tests come before migration traffic

Connectivity checklist
From the exact application runtime network:
1. Resolve target DNS and record addresses/TTL.
2. Establish TCP/TLS connection to the intended endpoint.
3. Validate certificate chain + hostname/SNI.
4. Authenticate as the application identity (not an admin).
5. GET cluster/product version only if permitted by least privilege.
6. Execute one allowed read and one intentionally forbidden write.
7. Measure connection setup and request latency separately.
8. Repeat through failover/proxy/load-balancer path if one exists.

A workstation success test does not validate production routes, security groups, proxies, private DNS, or workload identity.

4. IAM and product authorization are distinct layers

Amazon OpenSearch Service can involve AWS IAM/SigV4, domain access policy, VPC controls, and OpenSearch fine-grained access control. Elastic Cloud uses its own deployment/project authentication and network/security controls. Treat provider identity and search-engine authorization separately: a role that can reach the endpoint should not automatically receive broad index privileges.

Security rule. Never reuse one migration/admin credential across source, target, backfill, application, dashboard, and rollback tooling. Scope identities by function and rotate/revoke temporary migration credentials after cutover.

5. Snapshot repositories differ by service

Managed services often create automatic snapshots with provider-specific restore scope, and manual repositories can require IAM/service roles or supported storage locations. Amazon OpenSearch Service documents version-dependent snapshot migration paths. Elastic Cloud documents its own cross-deployment restore and remote-reindex workflows. Do not extrapolate self-managed repository behavior into a managed service.

For cross-product migration, keep the central rule from Lesson 2: a repository is not a generic interchange format. Use a provider-supported migration mechanism or a document-level copy/ETL path when product/version compatibility is uncertain.

6. Endpoint cutover: DNS is only the pointer

Gate before switch Evidence
Target caught up sync lag at/below agreed RPO; final delta applied
Read parity golden queries + relevance/aggregation checks pass
Security parity least-privilege positive/negative tests pass
Capacity target p95/p99/throughput under representative load pass
Observability target dashboards/alerts/log correlation verified
Rollback source remains usable and post-cutover-write strategy is defined
Cutover timeline template
T-60m  freeze schema/config changes; verify source snapshot/checkpoint
T-30m  final target validation + client connection-pool drain plan
T-10m  reduce DNS TTL only if your DNS strategy requires it and caches permit
T-05m  enter controlled write mode / final sync gate
T+00m  change service discovery/DNS/config endpoint
T+05m  verify target auth, reads, writes, p95/p99, error/rejection rate
T+15m  compare source/target write counts and sync journal
T+30m  decide: continue or rollback based on explicit thresholds
T+N    raise TTL / retire temporary migration access only after rollback window closes

7. Connection pools can outlive DNS changes

Clients may reuse established TCP connections and cache DNS longer than the authoritative TTL. Test the real client/runtime. A DNS update can produce a long mixed period where some instances still talk to the source. If the application performs writes, that becomes a data-consistency problem, not just a routing inconvenience.

Safer options include explicit application configuration rollout, a proxy/service-discovery layer with controlled draining, or a dual-write/capture layer whose behavior during the mixed interval is defined and measured.

8. Rollback in a managed environment

Rollback must restore traffic and authoritative data state. If target writes occurred after cutover, switching DNS back to source without replaying those writes violates RPO. Record the last source sequence/checkpoint, target-only writes, and reconciliation path. The rollback decision should be time-bounded because divergence and operational cost grow with every write.

Wrong approach. Lower DNS TTL, switch the CNAME, and call that a reversible migration. Repair: test resolver/client behavior, drain connection pools, gate on data/security/performance parity, and maintain a post-cutover write reconciliation path until the rollback window is formally closed.

9. Free/local deterministic learning path

The mandatory lab does not require Elastic Cloud or AWS. Simulate managed-service controls locally by running source and target behind different hostnames/ports, using distinct credentials, verifying TLS separately, and switching an application config variable from source to target. Document how the same control maps to your chosen managed service before production.

10. Local lab baseline for this lesson

Use the established disposable local endpoints: Elasticsearch 9.5.3 at https://localhost:9200 with ELASTIC_PASSWORD and CA file atlasmart-es-http-ca; OpenSearch 3.8.0 at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD. Both remain on Docker network atlasmart-search. The OpenSearch demo certificate may be bypassed with -k only in this local lab. Production migration requires verified TLS and separate least-privilege source, target, application, and migration identities.

11. Free/local managed-service cutover simulation

The mandatory lesson remains free/local. Simulate managed-service boundaries by treating the two local endpoints as separate administrative domains rather than requiring Elastic Cloud or AWS.

  1. From the same runtime that runs the test application, verify both https://localhost:9200 and https://localhost:9201 using their distinct trust and authentication rules.
  2. Create or select a read-only source identity and a write/read target identity appropriate to the smoke test. Do not use the administrative identities for the cutover rehearsal.
  3. Start the test application against the source endpoint and capture a known query result.
  4. Switch only an environment/config endpoint to the target, restart or drain the client connection pool, and rerun the same query/security checks.
  5. Inject one failure by using an invalid target credential or blocking the target route. Confirm the runbook selects rollback rather than retrying indefinitely.
  6. Restore the valid target path, rerun the checks, and record cutover and rollback timestamps.

What this proves: endpoint/config rollout, connection lifetime, TLS/auth boundaries, and rollback choreography. What it does not prove: a specific cloud provider's IAM, PrivateLink/VPC, KMS, repository, or managed snapshot behavior; those must be verified against the chosen service.

Check your understanding

  1. Why is managed migration not just data copy?
  2. What can survive a DNS change?
  3. Why keep provider IAM and engine roles separate?
  4. What makes rollback harder after target writes begin?
  5. What does Lesson 5 do?
Review the answers

1. Because network reachability, provider IAM, engine authorization, encryption, repository rules, endpoint behavior, and service-supported features can independently block cutover.

2. Existing client TCP connections and DNS caches, creating a mixed-routing period.

3. Network/service authorization and index/action authorization are different security layers.

4. The source may no longer contain the newest authoritative writes, so traffic reversal alone can lose data.

5. Turns every compatibility/data/security/performance/cutover assumption into an executable migration runbook with measurable rollback triggers.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.