Use Elastic cross-cluster replication as active-passive leader/follower replication, measure follower lag, document non-replicated configuration, and rehearse safe promotion boundaries.

Elastic Cross-Cluster Replication: Leader/Follower, Auto-Follow, Disaster Recovery, and Non-Replicated Configuration

Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.

Intermediate → Advanced120–165 minutesElastic CCR/DR design lab · Chapter 22 · Lesson 03Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

AtlasMart wants Region B to continue serving catalog searches when Region A is unavailable. Elastic CCR can maintain read-only follower indices, but it is an active-passive replication mechanism—not multi-master conflict resolution and not full-cluster configuration replication.

01

Explain leader/follower mechanics, remote recovery, retention leases, and read-only follower behavior.

02

Configure or reason about a follower and auto-follow pattern with explicit subscription/version prerequisites.

03

Measure follower checkpoints/lag and derive a real RPO from evidence.

04

Document configuration that CCR does not replicate and must be promoted separately.

05

Rehearse failover/promotion without allowing simultaneous writes to the same logical dataset.

Pinned multi-cluster baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course's existing local TLS/auth assumptions. Cross-cluster behavior is distribution-, license-, network-, and managed-service-sensitive, so every exercise starts by recording GET /, license/plugin state, remote-cluster settings, TLS trust, and the exact feature path being tested. Elastic's advanced API-key remote-cluster model and CCR have subscription boundaries; OpenSearch's replication plugin is bundled in the standard distribution but managed services can impose different connection/IAM constraints.

Elastic subscription boundary

Cross-cluster replication is not a Basic/free feature in the current Elastic subscription matrix. The mandatory no-cost path in this lesson is the deterministic follower-state simulation plus OpenSearch's bundled CCR lab in Lesson 4. If your Elastic license supports CCR, run the APIs against disposable clusters and record GET /_license.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Active-passive means one write owner per replicated index

The leader index accepts application writes. A follower cluster creates a follower index that pulls operations from the leader. The follower is read-only while following. Initial creation uses remote recovery to copy Lucene segment files, then replication follows operation history retained through soft-delete retention leases.

Elastic follower lifecycle
# On follower cluster; remote alias already configured
PUT /atlasmart-products-dr-v1/_ccr/follow
{
  "remote_cluster": "region_a",
  "leader_index": "atlasmart-products-v1"
}

GET /atlasmart-products-dr-v1/_ccr/info
GET /atlasmart-products-dr-v1/_ccr/stats

# Direct application writes to the follower should fail while following.
POST /atlasmart-products-dr-v1/_doc/test
{"name":"should-not-write"}

A follower that falls behind beyond available operation history can require recreation. That is a recovery event, not a reason to expand retention blindly; size history and RPO together.

2. Auto-follow automates follower creation, not all DR configuration

Elastic auto-follow patterns create follower indices for newly created matching leader indices and can follow rolling data-stream backing indices. Existing followers remain when a pattern changes. Auto-follow does not turn templates, ILM/SLM policies, roles, cluster settings, repositories, or system indices into replicated state.

Auto-follow intent
PUT /_ccr/auto_follow/atlasmart-catalog
{
  "remote_cluster": "region_a",
  "leader_index_patterns": ["atlasmart-products-*"],
  "follow_index_pattern": "{{leader_index}}"
}
GET /_ccr/auto_follow/atlasmart-catalog

Promote the non-data configuration through infrastructure-as-code or an explicit configuration release. Treat that release as independently testable and versioned.

3. What CCR replicates—and what it does not

Artifact CCR behavior DR action
User-generated index documents/operations Replicated to follower Monitor lag and follower health
Mappings and supported leader index changes Propagate with follower mechanics Test compatibility before leader change
Aliases Leader alias changes replicate, but write-index semantics cannot make follower writable Validate read aliases after failover
Security users/roles Not replicated Deploy independently
Index templates Not replicated Promote via config pipeline
ILM/SLM policies Not replicated Install/validate separately
Snapshot repository settings Not replicated Provision independently
Cluster settings/system indices/searchable snapshots Not a complete CCR payload Use documented product-specific recovery/config processes

Because security is independent, DR cannot be declared ready just because follower checkpoints are current. Application principals, TLS trust, templates, lifecycle rules, ingest pipelines, and routing expectations must be validated on the target cluster.

4. Measure lag as RPO evidence

RPO is the maximum acceptable data loss expressed in time or operations. CCR stats expose leader/follower checkpoints and operational state. Define a business RPO, then alarm before lag consumes that budget. Do not equate “follower task running” with “RPO met.”

Follower evidence template
GET /atlasmart-products-dr-v1/_ccr/stats

record:
leader_global_checkpoint=MEASURED
follower_global_checkpoint=MEASURED
operations_behind=MEASURED
last_successful_read=MEASURED
remote_recovery_state=MEASURED

rpo_gate: operations_behind <= DECLARED_BUDGET and follower is healthy

5. Failover is an application and write-ownership procedure

If Region A fails, do not immediately point writers at the follower while the old leader might still accept writes. First fence the old write path, verify lag against RPO, stop/pause the follow relationship as required, convert/unfollow according to the supported workflow, validate the target, then switch writes. The exact promotion sequence depends on version and topology and must be rehearsed.

DR decision gates
1. Detect leader-region failure.
2. Fence old write endpoint / prove it cannot accept writes.
3. Read follower CCR stats; compare lag to RPO.
4. Validate security/config/templates/application smoke tests.
5. Promote/unfollow using supported APIs only.
6. Switch writes and reads.
7. Record data-loss window and RTO.
8. Rebuild replication before allowing a second failure domain.
Bi-directional CCR is not same-index active-active.

You can design two directions for different leader indices, but replicated followers remain read-only. If two regions independently write the same logical entity set without a conflict model, CCR does not merge those conflicts for you.

Check your understanding

  1. Why is a follower read-only?
  2. Does auto-follow replicate index templates?
  3. What proves the RPO?
  4. Why fence the old writer before promotion?
  5. What is the no-cost equivalent exercise?
Review the answers

1. CCR preserves a single leader for the replicated index so follower operations do not create unresolved multi-master conflicts.

2. No; templates and many cluster/security/lifecycle artifacts must be deployed separately.

3. Measured follower checkpoint/operation lag compared with a declared business budget.

4. To prevent divergent writes and split-brain at the application/data-ownership layer.

5. Use the deterministic state/DR runbook here and execute the bundled OpenSearch CCR path in Lesson 4.

6. Production judgment

Elastic CCR is valuable when AtlasMart can name a leader, accept asynchronous lag, and independently manage target-cluster configuration. It is not a snapshot replacement: snapshots protect against corruption/operator error and provide point-in-time recovery, while CCR quickly copies good and bad leader operations. Use both when recovery requirements justify both.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.