Use Elastic cross-cluster replication as active-passive leader/follower replication, measure follower lag, document non-replicated configuration, and rehearse safe promotion boundaries.
Elastic Cross-Cluster Replication: Leader/Follower, Auto-Follow, Disaster Recovery, and Non-Replicated Configuration
Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.
Learning outcomes
AtlasMart wants Region B to continue serving catalog searches when Region A is unavailable. Elastic CCR can maintain read-only follower indices, but it is an active-passive replication mechanism—not multi-master conflict resolution and not full-cluster configuration replication.
Explain leader/follower mechanics, remote recovery, retention leases, and read-only follower behavior.
Configure or reason about a follower and auto-follow pattern with explicit subscription/version prerequisites.
Measure follower checkpoints/lag and derive a real RPO from evidence.
Document configuration that CCR does not replicate and must be promoted separately.
Rehearse failover/promotion without allowing simultaneous writes to the same logical dataset.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep
the course's existing local TLS/auth assumptions.
Cross-cluster behavior is distribution-, license-, network-,
and managed-service-sensitive, so every exercise starts by
recording GET /, license/plugin state,
remote-cluster settings, TLS trust, and the exact feature path
being tested. Elastic's advanced API-key remote-cluster model
and CCR have subscription boundaries; OpenSearch's replication
plugin is bundled in the standard distribution but managed
services can impose different connection/IAM constraints.
Cross-cluster replication is not a Basic/free feature in the
current Elastic subscription matrix. The mandatory no-cost
path in this lesson is the deterministic follower-state
simulation plus OpenSearch's bundled CCR lab in Lesson 4. If
your Elastic license supports CCR, run the APIs against
disposable clusters and record GET /_license.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Active-passive means one write owner per replicated index
The leader index accepts application writes. A follower cluster creates a follower index that pulls operations from the leader. The follower is read-only while following. Initial creation uses remote recovery to copy Lucene segment files, then replication follows operation history retained through soft-delete retention leases.
# On follower cluster; remote alias already configured
PUT /atlasmart-products-dr-v1/_ccr/follow
{
"remote_cluster": "region_a",
"leader_index": "atlasmart-products-v1"
}
GET /atlasmart-products-dr-v1/_ccr/info
GET /atlasmart-products-dr-v1/_ccr/stats
# Direct application writes to the follower should fail while following.
POST /atlasmart-products-dr-v1/_doc/test
{"name":"should-not-write"}
A follower that falls behind beyond available operation history can require recreation. That is a recovery event, not a reason to expand retention blindly; size history and RPO together.
2. Auto-follow automates follower creation, not all DR configuration
Elastic auto-follow patterns create follower indices for newly created matching leader indices and can follow rolling data-stream backing indices. Existing followers remain when a pattern changes. Auto-follow does not turn templates, ILM/SLM policies, roles, cluster settings, repositories, or system indices into replicated state.
PUT /_ccr/auto_follow/atlasmart-catalog
{
"remote_cluster": "region_a",
"leader_index_patterns": ["atlasmart-products-*"],
"follow_index_pattern": "{{leader_index}}"
}
GET /_ccr/auto_follow/atlasmart-catalog
Promote the non-data configuration through infrastructure-as-code or an explicit configuration release. Treat that release as independently testable and versioned.
3. What CCR replicates—and what it does not
| Artifact | CCR behavior | DR action |
|---|---|---|
| User-generated index documents/operations | Replicated to follower | Monitor lag and follower health |
| Mappings and supported leader index changes | Propagate with follower mechanics | Test compatibility before leader change |
| Aliases | Leader alias changes replicate, but write-index semantics cannot make follower writable | Validate read aliases after failover |
| Security users/roles | Not replicated | Deploy independently |
| Index templates | Not replicated | Promote via config pipeline |
| ILM/SLM policies | Not replicated | Install/validate separately |
| Snapshot repository settings | Not replicated | Provision independently |
| Cluster settings/system indices/searchable snapshots | Not a complete CCR payload | Use documented product-specific recovery/config processes |
Because security is independent, DR cannot be declared ready just because follower checkpoints are current. Application principals, TLS trust, templates, lifecycle rules, ingest pipelines, and routing expectations must be validated on the target cluster.
4. Measure lag as RPO evidence
RPO is the maximum acceptable data loss expressed in time or operations. CCR stats expose leader/follower checkpoints and operational state. Define a business RPO, then alarm before lag consumes that budget. Do not equate “follower task running” with “RPO met.”
GET /atlasmart-products-dr-v1/_ccr/stats
record:
leader_global_checkpoint=MEASURED
follower_global_checkpoint=MEASURED
operations_behind=MEASURED
last_successful_read=MEASURED
remote_recovery_state=MEASURED
rpo_gate: operations_behind <= DECLARED_BUDGET and follower is healthy
5. Failover is an application and write-ownership procedure
If Region A fails, do not immediately point writers at the follower while the old leader might still accept writes. First fence the old write path, verify lag against RPO, stop/pause the follow relationship as required, convert/unfollow according to the supported workflow, validate the target, then switch writes. The exact promotion sequence depends on version and topology and must be rehearsed.
1. Detect leader-region failure.
2. Fence old write endpoint / prove it cannot accept writes.
3. Read follower CCR stats; compare lag to RPO.
4. Validate security/config/templates/application smoke tests.
5. Promote/unfollow using supported APIs only.
6. Switch writes and reads.
7. Record data-loss window and RTO.
8. Rebuild replication before allowing a second failure domain.
You can design two directions for different leader indices, but replicated followers remain read-only. If two regions independently write the same logical entity set without a conflict model, CCR does not merge those conflicts for you.
Check your understanding
- Why is a follower read-only?
- Does auto-follow replicate index templates?
- What proves the RPO?
- Why fence the old writer before promotion?
- What is the no-cost equivalent exercise?
Review the answers
1. CCR preserves a single leader for the replicated index so follower operations do not create unresolved multi-master conflicts.
2. No; templates and many cluster/security/lifecycle artifacts must be deployed separately.
3. Measured follower checkpoint/operation lag compared with a declared business budget.
4. To prevent divergent writes and split-brain at the application/data-ownership layer.
5. Use the deterministic state/DR runbook here and execute the bundled OpenSearch CCR path in Lesson 4.
6. Production judgment
Elastic CCR is valuable when AtlasMart can name a leader, accept asynchronous lag, and independently manage target-cluster configuration. It is not a snapshot replacement: snapshots protect against corruption/operator error and provide point-in-time recovery, while CCR quickly copies good and bad leader operations. Use both when recovery requirements justify both.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic remote clusters — Current connection, security-model, managed-environment, and remote-cluster guidance.
- Elastic remote-cluster connection modes — Sniff versus proxy semantics and network reachability.
-
Elastic remote-cluster settings
—
skip_unavailable, roles, seeds, proxy settings, and defaults. - Elastic cross-cluster search — Search syntax, optional clusters, response metadata, and supported APIs.
- Elastic cross-cluster replication — Active-passive mechanics, limitations, and DR patterns.
- Elastic auto-follow patterns — Rolling-index/data-stream follower automation.
- Elastic subscriptions — Current cross-cluster licensing boundaries; verify the deployed license before labs.
- OpenSearch cross-cluster search — Security flow, remote connections, permissions, and examples.
- OpenSearch cross-cluster replication — Replication plugin model and operational behavior.
- OpenSearch CCR getting started — Sniff/proxy connectivity, roles, follower startup, and pull replication.
- OpenSearch auto-follow — Pattern-based automatic replication.
- Amazon OpenSearch Service cross-cluster search — Managed-service connection and IAM/network restrictions.
- Amazon OpenSearch Service cross-cluster replication — Managed CCR constraints and domain connection workflow.