Design an AtlasMart multi-region search and disaster-recovery architecture with explicit RPO/RTO, write ownership, security, failover gates, snapshots, and cost evidence.
Design Multi-Region Search/DR with Explicit RPO/RTO, Write Ownership, Security, Failover, and Cost Tradeoffs
Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.
Learning outcomes
This capstone converts the chapter into an AtlasMart multi-region decision record. The target is not maximum geographic complexity; it is a design whose write ownership, read path, security trust, failure behavior, RPO/RTO, and cost can all be tested.
Define explicit write ownership and eliminate accidental active-active conflict domains.
Choose between CCS, CCR, snapshots, dual-ingest, or centralized reporting by workload.
Quantify RPO/RTO and failover gates from measured replication and recovery evidence.
Design per-region security and configuration promotion without assuming CCR copies control-plane state.
Produce a reusable failover/failback runbook with cost and degraded-mode decisions.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep
the course's existing local TLS/auth assumptions.
Cross-cluster behavior is distribution-, license-, network-,
and managed-service-sensitive, so every exercise starts by
recording GET /, license/plugin state,
remote-cluster settings, TLS trust, and the exact feature path
being tested. Elastic's advanced API-key remote-cluster model
and CCR have subscription boundaries; OpenSearch's replication
plugin is bundled in the standard distribution but managed
services can impose different connection/IAM constraints.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Start with business objectives, not topology
| Requirement | AtlasMart example | Architecture consequence |
|---|---|---|
| RPO | ≤ 60 s catalog write loss | Replication lag alarm + promotion gate |
| RTO | ≤ 10 min regional search recovery | Pre-provisioned follower/search capacity + rehearsed DNS/app cutover |
| Global search completeness | All regions for finance; optional for discovery UI | Different skip_unavailable policies |
| Write ownership | Catalog writes owned by Region A | Followers/read replicas are not independent writers |
| Security | Region-scoped machine credentials | No global supercredential |
| Recovery from bad writes | Point-in-time restore required | Snapshots remain mandatory even with CCR |
RPO measures tolerable data loss/freshness; RTO measures restoration time. CCR can improve both for regional outages but may instantly replicate corruption. Snapshots protect different failure modes. CCS improves access to distributed data but does not improve data durability.
2. Choose a pattern deliberately
| Pattern | Strength | Primary risk |
|---|---|---|
| CCS only | No duplicate storage; live regional data | Every query depends on WAN/remote availability |
| CCR active-passive | Local reads and fast regional DR | Asynchronous lag; promotion/fencing required |
| Central reporting via CCR | Offloads operational clusters | Reporting freshness depends on replication lag |
| Dual-ingest from durable queue | Product-neutral write fan-out | Exactly-once/idempotency and ordering become application responsibilities |
| Snapshots only | Cheap point-in-time recovery | RTO/RPO generally slower than hot follower |
| CCS + CCR + snapshots | Strong coverage for read locality + DR + point-in-time restore | Highest complexity/cost; demands disciplined runbooks |
Neither Elastic CCR nor OpenSearch CCR merges conflicting writes to one logical follower index. “Bi-directional replication” is safe only when each direction has a clear leader for different indices/data domains or when the application explicitly partitions ownership.
3. AtlasMart target architecture
For the course scenario, choose Region A as the write owner for
atlasmart-products-v1. Region B receives a
read-only follower for search/DR. Regional telemetry remains
locally owned and is queried by CCS for global operations
dashboards. Nightly snapshots are copied off the primary failure
domain. Security roles, templates, pipelines, lifecycle
policies, and application configuration are promoted
independently through infrastructure-as-code.
dataset: atlasmart-products-v1
write_owner: region_a
replication_target: region_b
follower_name: atlasmart-products-dr-v1
catalog_rpo: 60s
catalog_rto: 10m
finance_global_search: required_all_regions
customer_discovery_global_search: degraded_allowed
snapshot_policy: independent_off_region_copy
security_model: region_scoped_credentials
config_replication: infrastructure_as_code, not CCR
failover_precondition: old_writer_fenced && lag_within_rpo && target_smoke_passes
4. Failover runbook: evidence before DNS
G0 Detect: prove Region A write/search path is unavailable or unsafe.
G1 Fence: disable/withdraw Region A write endpoint; prove no accepted writes.
G2 Freshness: read CCR stats; replication lag <= declared RPO.
G3 Target: cluster health, shard allocation, TLS/auth, templates/pipelines, app smoke pass.
G4 Promote: supported unfollow/stop-replication procedure only.
G5 Cutover: switch app/DNS/traffic to Region B; measure RTO.
G6 Observe: error rate, p95/p99, indexing lag, authorization failures.
G7 Re-protect: create a new replication direction or rebuild the former region before closing incident.
Do not fail over because “Region B looks green.” Green health says its assigned shards are healthy, not that follower data is current, credentials are correct, or the old writer is fenced. Likewise, do not fail back automatically when Region A returns. Reconcile ownership, rebuild replication, then execute a separate controlled cutover.
5. Cost model
Continuous replication consumes duplicate storage, follower compute, WAN transfer, and operator attention. CCS avoids duplicate storage but pushes remote-query latency and availability into every request. Managed services add cross-region transfer and connection constraints. Model cost per dataset and query class rather than choosing one global architecture.
monthly_replication_gb = MEASURED
cross_region_transfer_cost = PROVIDER_RATE * monthly_replication_gb
follower_compute_cost = MEASURED
follower_storage_cost = MEASURED
ccs_query_transfer_gb = MEASURED
operator_game_day_hours = MEASURED
# Compare architectures against RPO/RTO and latency, not cost alone.
6. Game day: controlled failover decision
Use either two disposable local OpenSearch clusters or the deterministic trace. Inject one leader outage after establishing a healthy follower. Record the last acknowledged leader write timestamp, last replicated follower document/checkpoint, time to detect, time to fence, time to validate, and time to serve from the target. Then restore the leader but do not automatically route writes back.
incident_id=ATLASMART-MULTIREGION-001
leader_last_ack_write=MEASURED
follower_last_visible_write=MEASURED
observed_rpo_seconds=MEASURED
detection_seconds=MEASURED
fence_seconds=MEASURED
target_validation_seconds=MEASURED
cutover_seconds=MEASURED
observed_rto_seconds=MEASURED
post_cutover_p95_ms=MEASURED
post_cutover_p99_ms=MEASURED
security_smoke=PASS|FAIL
config_drift=MEASURED
snapshot_recovery_point=VERIFIED|UNVERIFIED
The drill passes only if RPO/RTO are within declared budgets, old write ownership is fenced, target security/configuration checks pass, and operators know how to rebuild protection after failover. A “successful” DNS change without these proofs is not a DR success.
Check your understanding
- Why keep snapshots when CCR exists?
- What is the first failover data-safety action?
- When is CCS preferable to CCR?
- What configuration must be promoted separately?
- What closes the chapter?
Review the answers
1. CCR can replicate corruption/deletes quickly; snapshots provide independent point-in-time recovery.
2. Fence the old write owner before allowing the new region to accept writes.
3. When live distributed data is needed and the application can tolerate WAN latency/remote availability without duplicating storage.
4. Security and many templates/lifecycle/snapshot/cluster settings because CCR is not full control-plane replication.
5. A measured architecture contract and game-day evidence that tie topology directly to RPO/RTO, security, latency, and cost.
7. Bridge to Chapter 23
Chapter 22 made multi-cluster topology explicit. Chapter 23 changes the query interface: ES|QL on Elastic and PPL/SQL on OpenSearch. Those analytical languages can also cross cluster boundaries, but their support, security, cost, and subscription constraints must be evaluated rather than inferred from Query DSL.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic remote clusters — Current connection, security-model, managed-environment, and remote-cluster guidance.
- Elastic remote-cluster connection modes — Sniff versus proxy semantics and network reachability.
-
Elastic remote-cluster settings
—
skip_unavailable, roles, seeds, proxy settings, and defaults. - Elastic cross-cluster search — Search syntax, optional clusters, response metadata, and supported APIs.
- Elastic cross-cluster replication — Active-passive mechanics, limitations, and DR patterns.
- Elastic auto-follow patterns — Rolling-index/data-stream follower automation.
- Elastic subscriptions — Current cross-cluster licensing boundaries; verify the deployed license before labs.
- OpenSearch cross-cluster search — Security flow, remote connections, permissions, and examples.
- OpenSearch cross-cluster replication — Replication plugin model and operational behavior.
- OpenSearch CCR getting started — Sniff/proxy connectivity, roles, follower startup, and pull replication.
- OpenSearch auto-follow — Pattern-based automatic replication.
- Amazon OpenSearch Service cross-cluster search — Managed-service connection and IAM/network restrictions.
- Amazon OpenSearch Service cross-cluster replication — Managed CCR constraints and domain connection workflow.