Design an AtlasMart multi-region search and disaster-recovery architecture with explicit RPO/RTO, write ownership, security, failover gates, snapshots, and cost evidence.

Design Multi-Region Search/DR with Explicit RPO/RTO, Write Ownership, Security, Failover, and Cost Tradeoffs

Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.

Intermediate → Advanced130–180 minutesMulti-region DR game day · Chapter 22 · Lesson 05Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

This capstone converts the chapter into an AtlasMart multi-region decision record. The target is not maximum geographic complexity; it is a design whose write ownership, read path, security trust, failure behavior, RPO/RTO, and cost can all be tested.

01

Define explicit write ownership and eliminate accidental active-active conflict domains.

02

Choose between CCS, CCR, snapshots, dual-ingest, or centralized reporting by workload.

03

Quantify RPO/RTO and failover gates from measured replication and recovery evidence.

04

Design per-region security and configuration promotion without assuming CCR copies control-plane state.

05

Produce a reusable failover/failback runbook with cost and degraded-mode decisions.

Pinned multi-cluster baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course's existing local TLS/auth assumptions. Cross-cluster behavior is distribution-, license-, network-, and managed-service-sensitive, so every exercise starts by recording GET /, license/plugin state, remote-cluster settings, TLS trust, and the exact feature path being tested. Elastic's advanced API-key remote-cluster model and CCR have subscription boundaries; OpenSearch's replication plugin is bundled in the standard distribution but managed services can impose different connection/IAM constraints.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Start with business objectives, not topology

Requirement AtlasMart example Architecture consequence
RPO ≤ 60 s catalog write loss Replication lag alarm + promotion gate
RTO ≤ 10 min regional search recovery Pre-provisioned follower/search capacity + rehearsed DNS/app cutover
Global search completeness All regions for finance; optional for discovery UI Different skip_unavailable policies
Write ownership Catalog writes owned by Region A Followers/read replicas are not independent writers
Security Region-scoped machine credentials No global supercredential
Recovery from bad writes Point-in-time restore required Snapshots remain mandatory even with CCR

RPO measures tolerable data loss/freshness; RTO measures restoration time. CCR can improve both for regional outages but may instantly replicate corruption. Snapshots protect different failure modes. CCS improves access to distributed data but does not improve data durability.

2. Choose a pattern deliberately

Pattern Strength Primary risk
CCS only No duplicate storage; live regional data Every query depends on WAN/remote availability
CCR active-passive Local reads and fast regional DR Asynchronous lag; promotion/fencing required
Central reporting via CCR Offloads operational clusters Reporting freshness depends on replication lag
Dual-ingest from durable queue Product-neutral write fan-out Exactly-once/idempotency and ordering become application responsibilities
Snapshots only Cheap point-in-time recovery RTO/RPO generally slower than hot follower
CCS + CCR + snapshots Strong coverage for read locality + DR + point-in-time restore Highest complexity/cost; demands disciplined runbooks
Avoid same-dataset active-active writes without a conflict model.

Neither Elastic CCR nor OpenSearch CCR merges conflicting writes to one logical follower index. “Bi-directional replication” is safe only when each direction has a clear leader for different indices/data domains or when the application explicitly partitions ownership.

3. AtlasMart target architecture

For the course scenario, choose Region A as the write owner for atlasmart-products-v1. Region B receives a read-only follower for search/DR. Regional telemetry remains locally owned and is queried by CCS for global operations dashboards. Nightly snapshots are copied off the primary failure domain. Security roles, templates, pipelines, lifecycle policies, and application configuration are promoted independently through infrastructure-as-code.

Architecture contract
dataset: atlasmart-products-v1
write_owner: region_a
replication_target: region_b
follower_name: atlasmart-products-dr-v1
catalog_rpo: 60s
catalog_rto: 10m
finance_global_search: required_all_regions
customer_discovery_global_search: degraded_allowed
snapshot_policy: independent_off_region_copy
security_model: region_scoped_credentials
config_replication: infrastructure_as_code, not CCR
failover_precondition: old_writer_fenced && lag_within_rpo && target_smoke_passes

4. Failover runbook: evidence before DNS

Regional failover gates
G0 Detect: prove Region A write/search path is unavailable or unsafe.
G1 Fence: disable/withdraw Region A write endpoint; prove no accepted writes.
G2 Freshness: read CCR stats; replication lag <= declared RPO.
G3 Target: cluster health, shard allocation, TLS/auth, templates/pipelines, app smoke pass.
G4 Promote: supported unfollow/stop-replication procedure only.
G5 Cutover: switch app/DNS/traffic to Region B; measure RTO.
G6 Observe: error rate, p95/p99, indexing lag, authorization failures.
G7 Re-protect: create a new replication direction or rebuild the former region before closing incident.

Do not fail over because “Region B looks green.” Green health says its assigned shards are healthy, not that follower data is current, credentials are correct, or the old writer is fenced. Likewise, do not fail back automatically when Region A returns. Reconcile ownership, rebuild replication, then execute a separate controlled cutover.

5. Cost model

Continuous replication consumes duplicate storage, follower compute, WAN transfer, and operator attention. CCS avoids duplicate storage but pushes remote-query latency and availability into every request. Managed services add cross-region transfer and connection constraints. Model cost per dataset and query class rather than choosing one global architecture.

Cost worksheet
monthly_replication_gb = MEASURED
cross_region_transfer_cost = PROVIDER_RATE * monthly_replication_gb
follower_compute_cost = MEASURED
follower_storage_cost = MEASURED
ccs_query_transfer_gb = MEASURED
operator_game_day_hours = MEASURED

# Compare architectures against RPO/RTO and latency, not cost alone.

6. Game day: controlled failover decision

Use either two disposable local OpenSearch clusters or the deterministic trace. Inject one leader outage after establishing a healthy follower. Record the last acknowledged leader write timestamp, last replicated follower document/checkpoint, time to detect, time to fence, time to validate, and time to serve from the target. Then restore the leader but do not automatically route writes back.

DR evidence record
incident_id=ATLASMART-MULTIREGION-001
leader_last_ack_write=MEASURED
follower_last_visible_write=MEASURED
observed_rpo_seconds=MEASURED
detection_seconds=MEASURED
fence_seconds=MEASURED
target_validation_seconds=MEASURED
cutover_seconds=MEASURED
observed_rto_seconds=MEASURED
post_cutover_p95_ms=MEASURED
post_cutover_p99_ms=MEASURED
security_smoke=PASS|FAIL
config_drift=MEASURED
snapshot_recovery_point=VERIFIED|UNVERIFIED

The drill passes only if RPO/RTO are within declared budgets, old write ownership is fenced, target security/configuration checks pass, and operators know how to rebuild protection after failover. A “successful” DNS change without these proofs is not a DR success.

Check your understanding

  1. Why keep snapshots when CCR exists?
  2. What is the first failover data-safety action?
  3. When is CCS preferable to CCR?
  4. What configuration must be promoted separately?
  5. What closes the chapter?
Review the answers

1. CCR can replicate corruption/deletes quickly; snapshots provide independent point-in-time recovery.

2. Fence the old write owner before allowing the new region to accept writes.

3. When live distributed data is needed and the application can tolerate WAN latency/remote availability without duplicating storage.

4. Security and many templates/lifecycle/snapshot/cluster settings because CCR is not full control-plane replication.

5. A measured architecture contract and game-day evidence that tie topology directly to RPO/RTO, security, latency, and cost.

7. Bridge to Chapter 23

Chapter 22 made multi-cluster topology explicit. Chapter 23 changes the query interface: ES|QL on Elastic and PPL/SQL on OpenSearch. Those analytical languages can also cross cluster boundaries, but their support, security, cost, and subscription constraints must be evaluated rather than inferred from Query DSL.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.