Connect AtlasMart clusters with explicit sniff/proxy-like network semantics, TLS/auth trust, required/optional remote behavior, and observable WAN failure handling.
Remote Cluster Connectivity, Sniff/Proxy-Like Modes, Security, Latency, and Failure Semantics
Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.
Learning outcomes
AtlasMart has a catalog cluster in Region A and a reporting cluster in Region B. A request to “just connect them” is incomplete: the platform team must decide who initiates the connection, which transport endpoint is reachable, how identity is propagated, whether the remote is required or optional, and what happens when the WAN is slow or unavailable.
Distinguish remote-cluster connectivity from ordinary HTTP client traffic and define local, remote, gateway, coordinating, and remote-cluster-client roles.
Choose sniff or proxy-like connectivity from network topology rather than convenience.
Build explicit TLS/authentication trust and least-privilege authorization across cluster boundaries.
Observe connection health and failure semantics instead of inferring them from successful local searches.
Define latency, availability, and security gates before a remote cluster is allowed into a production query path.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep
the course's existing local TLS/auth assumptions.
Cross-cluster behavior is distribution-, license-, network-,
and managed-service-sensitive, so every exercise starts by
recording GET /, license/plugin state,
remote-cluster settings, TLS trust, and the exact feature path
being tested. Elastic's advanced API-key remote-cluster model
and CCR have subscription boundaries; OpenSearch's replication
plugin is bundled in the standard distribution but managed
services can impose different connection/IAM constraints.
Every remote hop adds a separate failure domain, trust boundary, version-compatibility boundary, and WAN latency distribution. The coordinating cluster can only return results that remote clusters actually delivered under the configured failure policy.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Connection mechanism: local cluster, remote cluster, and gateway path
A remote cluster is another Elasticsearch or
OpenSearch cluster registered under an alias such as
region_b. Requests originate on the
local/querying cluster. Nodes that accept the
user request coordinate work and must be permitted to act as
remote-cluster clients. In Elasticsearch this is the
remote_cluster_client role; OpenSearch uses the
same role name when roles are explicitly overridden.
Sniff mode starts from seed transport addresses, discovers eligible remote nodes, and then opens direct transport connections to their published addresses. It therefore fails behind NAT/firewalls when discovered publish addresses are not reachable. Proxy mode pins connections to a layer-4 proxy/load-balancer address and ignores discovered publish addresses. Proxy is usually the practical choice across Kubernetes, managed environments, or segmented WANs.
| Decision | Sniff | Proxy |
|---|---|---|
| Reachability | Local remote-client nodes must reach remote publish addresses | Local nodes reach one proxy endpoint |
| Discovery | Seed then discover gateway nodes | No direct use of remote publish addresses |
| Operational fit | Flat/self-managed networks | NAT, firewalls, orchestration, managed endpoints |
| Failure surface | Seed/gateway reachability plus advertised addresses | Proxy health/routing plus remote target health |
PUT /_cluster/settings
{
"persistent": {
"cluster.remote.region_b.mode": "proxy",
"cluster.remote.region_b.proxy_address": "region-b-gateway.example:9443",
"cluster.remote.region_b.skip_unavailable": false
}
}
GET /_remote/info
GET /_resolve/cluster/region_b:atlasmart-*
With Elasticsearch's current API-key security model, incoming cross-cluster traffic uses the remote-cluster server interface, whose default port is 9443 when enabled. Do not expose it until its TLS certificate, bind/publish addresses, and cross-cluster API key scope are deliberate. The older certificate security model still exists but is deprecated for new Elastic remote-cluster setups.
2. OpenSearch connectivity is similar in shape, not identical in security semantics
OpenSearch CCS/CCR also supports sniff and proxy connection modes. In a self-managed lab the common path uses transport port 9300 and Security-plugin node-to-node TLS. Authentication for CCS is evaluated on the coordinating cluster and then on the remote cluster using the forwarded identity/backend roles. Replication has its own leader/follower role requirements. Never paste Elasticsearch remote-cluster credentials or privilege JSON into OpenSearch and assume equivalence.
PUT /_cluster/settings
{
"persistent": {
"cluster.remote.region_b.seeds": ["region-b-node:9300"],
"cluster.remote.region_b.skip_unavailable": false
}
}
GET /_remote/info
GET /region_b:atlasmart-products-v1/_search?size=0
For OpenSearch 3.8, skip_unavailable defaults to
false. For Elasticsearch 9.5.3 it defaults to
true because Elasticsearch changed that default
in 8.15. This is a migration-sensitive difference: an unset
field can mean “required” on one product and “optional” on the
other.
3. Security: trust is narrower than connectivity
A TCP connection proving that two clusters can talk does not prove the caller may read data. Define three layers separately: transport trust (TLS validates peers and encrypts the channel), authentication (which principal is represented), and authorization (which remote indices/actions that principal may perform). Use different credentials per region and function; a stolen reporting credential should not become a global administrative credential.
# Both products
GET /
GET /_remote/info
GET /_nodes?filter_path=nodes.*.roles,nodes.*.version
# Elasticsearch: resolve connectivity + authorization to target expression
GET /_resolve/cluster/region_b:atlasmart-*
# OpenSearch: verify Security-plugin identity on each side with least-privilege users
GET /_plugins/_security/authinfo
-k everywhere.
Disabling certificate verification can make a lab “work” while
deleting server identity from the security model. Keep
-k only for the already-documented disposable
demo-certificate path. Production uses a trusted CA and
hostname/SAN verification.
4. Latency and failure semantics are part of the API contract
Measure remote search p50/p95/p99 separately from local search. A required remote cluster can turn a regional network impairment into user-visible request failure; an optional remote can return success while silently omitting that region. Therefore the application must read cluster metadata, not just HTTP 200.
| Signal | Question it answers | What it does not prove |
|---|---|---|
connected in remote info |
At least one remote connection is open | That a target index is authorized or healthy |
CCS _clusters metadata |
Which clusters succeeded/skipped/failed | That omitted data is acceptable to the business |
| Replication checkpoint/lag | Follower progress relative to leader | That config/security is synchronized |
| WAN RTT/packet loss | Network condition | End-to-end query cost by itself |
Explicitly configure required/optional behavior. On
Elasticsearch 9.5.3 an unset skip_unavailable is
optional by default; on OpenSearch 3.8 it is required by
default. AtlasMart sets it explicitly in every environment to
prevent a version migration from changing behavior silently.
5. Controlled failure lab
The free/local mandatory path uses two disposable OpenSearch 3.8
clusters or, when cross-cluster TLS is not available on the
learner's machine, the deterministic trace below. Do not disable
security to make a production-like conclusion. Record the remote
alias, transport mode, certificate trust, version, and
skip_unavailable value.
# 1. Healthy baseline
GET /_remote/info
GET /region_b:atlasmart-products-v1/_search?size=0
# 2. Stop only the disposable Region B container/network path.
# 3. Repeat the search with region_b explicitly required.
# 4. Set region_b skip_unavailable=true and repeat.
# 5. Restore Region B; verify reconnection and application correctness.
# Record, do not fabricate:
wan_rtt_p95_ms=MEASURED
ccs_p95_required_ms=MEASURED
ccs_p95_optional_ms=MEASURED
failure_detection_ms=MEASURED
Expected invariant: when required, remote failure must be visible as request failure or failed-cluster metadata according to the API; when optional, the request may succeed but the remote cluster must be marked skipped/partial. The lab passes only if the application distinguishes those outcomes.
Check your understanding
- Why can sniff mode work on a flat lab network but fail after Kubernetes deployment?
- Why set
skip_unavailableexplicitly? - What does a green local cluster prove about a remote cluster?
- Why are remote credentials region-scoped?
- What is the bridge to Lesson 2?
Review the answers
1. Because discovered remote publish addresses must be directly reachable; orchestration/NAT often makes them private or unstable.
2. Its defaults differ between Elasticsearch 9.5.3 and OpenSearch 3.8 and it directly changes completeness/failure semantics.
3. Nothing about remote reachability, authorization, health, or data freshness.
4. To reduce blast radius and avoid one compromised credential granting broad multi-region access.
5. Connectivity is only the substrate; next we examine how a federated search fans out, aggregates results, and reports partial clusters.
6. Production judgment
Choose a remote-cluster mode from network topology, not habit. Require explicit trust, version compatibility, required/optional semantics, and latency/error budgets. A remote search path is ready only when operators can distinguish “remote unavailable,” “remote unauthorized,” “remote slow,” “remote skipped,” and “remote data stale.” Cross-region design begins with observable failure semantics.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic remote clusters — Current connection, security-model, managed-environment, and remote-cluster guidance.
- Elastic remote-cluster connection modes — Sniff versus proxy semantics and network reachability.
-
Elastic remote-cluster settings
—
skip_unavailable, roles, seeds, proxy settings, and defaults. - Elastic cross-cluster search — Search syntax, optional clusters, response metadata, and supported APIs.
- Elastic cross-cluster replication — Active-passive mechanics, limitations, and DR patterns.
- Elastic auto-follow patterns — Rolling-index/data-stream follower automation.
- Elastic subscriptions — Current cross-cluster licensing boundaries; verify the deployed license before labs.
- OpenSearch cross-cluster search — Security flow, remote connections, permissions, and examples.
- OpenSearch cross-cluster replication — Replication plugin model and operational behavior.
- OpenSearch CCR getting started — Sniff/proxy connectivity, roles, follower startup, and pull replication.
- OpenSearch auto-follow — Pattern-based automatic replication.
- Amazon OpenSearch Service cross-cluster search — Managed-service connection and IAM/network restrictions.
- Amazon OpenSearch Service cross-cluster replication — Managed CCR constraints and domain connection workflow.