Configure OpenSearch cross-cluster search and replication with Security-plugin roles, auto-follow, managed-service differences, and a free/local measured follower lab.

OpenSearch Cross-Cluster Replication/Search Concepts and Managed-Service Constraints

Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.

Intermediate → Advanced120–165 minutesOpenSearch CCR free/local lab · Chapter 22 · Lesson 04Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

OpenSearch offers cross-cluster search plus a bundled cross-cluster replication plugin. The architecture resembles Elastic's leader/follower model but the APIs, Security-plugin roles, auto-follow semantics, and managed-service workflow differ enough that operators must maintain product-specific runbooks.

01

Configure OpenSearch remote connectivity for CCS/CCR with Security-plugin boundaries.

02

Start, observe, pause/resume, and stop follower replication using the replication plugin APIs.

03

Use auto-follow deliberately and distinguish its semantics from Elastic auto-follow.

04

Explain Amazon OpenSearch Service connection/IAM/version restrictions separately from self-managed OpenSearch.

05

Build a free/local AtlasMart replication experiment with measured lag and a controlled outage.

Pinned multi-cluster baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course's existing local TLS/auth assumptions. Cross-cluster behavior is distribution-, license-, network-, and managed-service-sensitive, so every exercise starts by recording GET /, license/plugin state, remote-cluster settings, TLS trust, and the exact feature path being tested. Elastic's advanced API-key remote-cluster model and CCR have subscription boundaries; OpenSearch's replication plugin is bundled in the standard distribution but managed services can impose different connection/IAM constraints.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. OpenSearch CCR plugin is pull-based active-passive replication

The follower cluster connects to a leader cluster and creates a new read-only follower index. The standard OpenSearch distribution bundles the cross-cluster-replication plugin; the minimal distribution may not. If node.roles is explicitly set on follower nodes, include remote_cluster_client. Security should be enabled on both sides with node-to-node encryption in production.

OpenSearch: establish a remote and start replication
# Follower cluster: configure leader connection
PUT /_cluster/settings
{
  "persistent": {
    "cluster.remote.region_a.seeds": ["region-a-node:9300"]
  }
}

# Start follower replication
PUT /_plugins/_replication/atlasmart-products-dr-v1/_start
{
  "leader_alias": "region_a",
  "leader_index": "atlasmart-products-v1",
  "use_roles": {
    "leader_cluster_role": "atlasmart_ccr_leader",
    "follower_cluster_role": "atlasmart_ccr_follower"
  }
}
GET /_plugins/_replication/atlasmart-products-dr-v1/_status

2. Replication security is a two-cluster contract

The Security plugin includes replication permissions and built-in role concepts, but AtlasMart should create scoped roles rather than reusing all_access. The leader role should expose only indices intended for replication; the follower role should manage only required follower resources. CCS authorization and CCR authorization are related but distinct.

Operational controls
POST /_plugins/_replication/atlasmart-products-dr-v1/_pause
POST /_plugins/_replication/atlasmart-products-dr-v1/_resume
POST /_plugins/_replication/atlasmart-products-dr-v1/_stop

GET /_plugins/_replication/leader_stats
GET /_plugins/_replication/follower_stats
GET /_plugins/_replication/autofollow_stats

Stopping replication changes the follower's relationship; validate the exact current API semantics before using stop as a failover step. Treat promotion as a rehearsed procedure, not a generic “make writable” button.

3. Auto-follow semantics differ from Elastic

OpenSearch replication rules use wildcard patterns. Current OpenSearch documentation states that when a rule is created it begins by replicating existing matching indices as well as automatically following future matching indices. Elastic auto-follow patterns, by contrast, are described around newly created matching indices. This is a concrete portability trap: the same mental model can unexpectedly create more followers on OpenSearch.

OpenSearch auto-follow rule shape
POST /_plugins/_replication/_autofollow
{
  "leader_alias": "region_a",
  "name": "atlasmart-products",
  "pattern": "atlasmart-products-*",
  "use_roles": {
    "leader_cluster_role": "atlasmart_ccr_leader",
    "follower_cluster_role": "atlasmart_ccr_follower"
  }
}
GET /_plugins/_replication/autofollow_stats
Verify endpoint syntax on the pinned build.

OpenSearch replication APIs evolve independently from Elasticsearch CCR. Keep the course baseline at 3.8.0 and check the live plugin API documentation before automation. Never translate /_ccr/* paths mechanically to OpenSearch.

4. Managed Amazon OpenSearch Service is a different control plane

Amazon OpenSearch Service uses domain connection workflows and IAM/fine-grained access control rather than the self-managed cURL-only setup. Current service docs state that CCR is active-passive and incurs normal inter-domain data-transfer costs. A managed cross-cluster connection can have compatibility and account/Region approval constraints, and a connection previously created as search-only may need recreation for replication.

Concern Self-managed OpenSearch Amazon OpenSearch Service
Connection setup Cluster settings / transport connectivity Service domain connection request/acceptance
Authentication Security plugin / TLS / roles IAM + fine-grained access + service connection
Self-managed interop Directly possible when versions/security/network permit Service CCR/CCS docs restrict supported domain combinations
Cost Infrastructure/WAN you operate AWS domain + standard data-transfer charges
Upgrade control You sequence nodes/plugins Managed domain compatibility/upgrade workflow

AWS cross-cluster search documentation also carries service-specific limits such as connection-count/version restrictions and disallows connecting managed domains to self-managed clusters for that service feature. Use AWS documentation, not upstream OpenSearch setup steps, for managed domains.

5. Free/local AtlasMart replication lab

Use two disposable OpenSearch 3.8.0 clusters on one Docker network with the Security plugin enabled and a shared trusted node CA. If your machine cannot host both clusters, use the deterministic trace; do not weaken production conclusions by disabling all security.

Lab sequence and measurements
A. Leader: create atlasmart-products-v1 and index P-1001..P-1005.
B. Follower: configure region_a remote connection.
C. Start atlasmart-products-dr-v1 replication.
D. Wait until follower count == leader count.
E. Insert P-1006 on leader; record visibility delay on follower.
F. Stop only leader container/network path; record follower read behavior.
G. Restore leader; verify replication catches up.

leader_docs=MEASURED
follower_docs=MEASURED
replication_lag_ms_p95=MEASURED
outage_detection_ms=MEASURED
catchup_ms=MEASURED

Expected invariants: direct writes to the follower fail while replication is active; lag is measurable and non-negative; follower reads can continue during a leader outage if the follower cluster itself is healthy; configuration/security outside the replicated index must still be validated independently.

Check your understanding

  1. Why are OpenSearch CCR APIs not interchangeable with Elastic CCR APIs?
  2. What is a notable auto-follow difference?
  3. Why is all_access a bad production replication role?
  4. Why do AWS managed-domain instructions differ?
  5. What must the lab measure?
Review the answers

1. They are separate implementations with different plugins, endpoints, security models, and operational semantics.

2. OpenSearch documents rules as replicating existing matching indices when the rule is created, not only future matches.

3. It expands blast radius far beyond the indices/actions needed for replication.

4. The service adds its own connection approval, IAM, networking, version, and service-limit control plane.

5. Follower data parity, replication lag, outage detection, catch-up time, and application read behavior.

6. Production judgment

OpenSearch CCR is the practical no-cost replication mechanism for this course, but “free” does not mean operationally free: WAN transfer, follower capacity, observability, PKI, runbooks, and failure rehearsals still cost engineering time and infrastructure. Select only indices whose RPO/RTO justify continuous replication.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.