Connect AtlasMart clusters with explicit sniff/proxy-like network semantics, TLS/auth trust, required/optional remote behavior, and observable WAN failure handling.

Remote Cluster Connectivity, Sniff/Proxy-Like Modes, Security, Latency, and Failure Semantics

Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.

Intermediate → Advanced115–150 minutesRemote connectivity & failure-semantics lab · Chapter 22 · Lesson 01Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

AtlasMart has a catalog cluster in Region A and a reporting cluster in Region B. A request to “just connect them” is incomplete: the platform team must decide who initiates the connection, which transport endpoint is reachable, how identity is propagated, whether the remote is required or optional, and what happens when the WAN is slow or unavailable.

01

Distinguish remote-cluster connectivity from ordinary HTTP client traffic and define local, remote, gateway, coordinating, and remote-cluster-client roles.

02

Choose sniff or proxy-like connectivity from network topology rather than convenience.

03

Build explicit TLS/authentication trust and least-privilege authorization across cluster boundaries.

04

Observe connection health and failure semantics instead of inferring them from successful local searches.

05

Define latency, availability, and security gates before a remote cluster is allowed into a production query path.

Pinned multi-cluster baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course's existing local TLS/auth assumptions. Cross-cluster behavior is distribution-, license-, network-, and managed-service-sensitive, so every exercise starts by recording GET /, license/plugin state, remote-cluster settings, TLS trust, and the exact feature path being tested. Elastic's advanced API-key remote-cluster model and CCR have subscription boundaries; OpenSearch's replication plugin is bundled in the standard distribution but managed services can impose different connection/IAM constraints.

Do not treat a remote cluster as a local shard set.

Every remote hop adds a separate failure domain, trust boundary, version-compatibility boundary, and WAN latency distribution. The coordinating cluster can only return results that remote clusters actually delivered under the configured failure policy.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Connection mechanism: local cluster, remote cluster, and gateway path

A remote cluster is another Elasticsearch or OpenSearch cluster registered under an alias such as region_b. Requests originate on the local/querying cluster. Nodes that accept the user request coordinate work and must be permitted to act as remote-cluster clients. In Elasticsearch this is the remote_cluster_client role; OpenSearch uses the same role name when roles are explicitly overridden.

Sniff mode starts from seed transport addresses, discovers eligible remote nodes, and then opens direct transport connections to their published addresses. It therefore fails behind NAT/firewalls when discovered publish addresses are not reachable. Proxy mode pins connections to a layer-4 proxy/load-balancer address and ignores discovered publish addresses. Proxy is usually the practical choice across Kubernetes, managed environments, or segmented WANs.

Decision Sniff Proxy
Reachability Local remote-client nodes must reach remote publish addresses Local nodes reach one proxy endpoint
Discovery Seed then discover gateway nodes No direct use of remote publish addresses
Operational fit Flat/self-managed networks NAT, firewalls, orchestration, managed endpoints
Failure surface Seed/gateway reachability plus advertised addresses Proxy health/routing plus remote target health
Elastic: configure and inspect a remote cluster
PUT /_cluster/settings
{
  "persistent": {
    "cluster.remote.region_b.mode": "proxy",
    "cluster.remote.region_b.proxy_address": "region-b-gateway.example:9443",
    "cluster.remote.region_b.skip_unavailable": false
  }
}
GET /_remote/info
GET /_resolve/cluster/region_b:atlasmart-*

With Elasticsearch's current API-key security model, incoming cross-cluster traffic uses the remote-cluster server interface, whose default port is 9443 when enabled. Do not expose it until its TLS certificate, bind/publish addresses, and cross-cluster API key scope are deliberate. The older certificate security model still exists but is deprecated for new Elastic remote-cluster setups.

2. OpenSearch connectivity is similar in shape, not identical in security semantics

OpenSearch CCS/CCR also supports sniff and proxy connection modes. In a self-managed lab the common path uses transport port 9300 and Security-plugin node-to-node TLS. Authentication for CCS is evaluated on the coordinating cluster and then on the remote cluster using the forwarded identity/backend roles. Replication has its own leader/follower role requirements. Never paste Elasticsearch remote-cluster credentials or privilege JSON into OpenSearch and assume equivalence.

OpenSearch: seed a remote connection
PUT /_cluster/settings
{
  "persistent": {
    "cluster.remote.region_b.seeds": ["region-b-node:9300"],
    "cluster.remote.region_b.skip_unavailable": false
  }
}
GET /_remote/info
GET /region_b:atlasmart-products-v1/_search?size=0

For OpenSearch 3.8, skip_unavailable defaults to false. For Elasticsearch 9.5.3 it defaults to true because Elasticsearch changed that default in 8.15. This is a migration-sensitive difference: an unset field can mean “required” on one product and “optional” on the other.

3. Security: trust is narrower than connectivity

A TCP connection proving that two clusters can talk does not prove the caller may read data. Define three layers separately: transport trust (TLS validates peers and encrypts the channel), authentication (which principal is represented), and authorization (which remote indices/actions that principal may perform). Use different credentials per region and function; a stolen reporting credential should not become a global administrative credential.

Evidence to capture before enabling production CCS
# Both products
GET /
GET /_remote/info
GET /_nodes?filter_path=nodes.*.roles,nodes.*.version

# Elasticsearch: resolve connectivity + authorization to target expression
GET /_resolve/cluster/region_b:atlasmart-*

# OpenSearch: verify Security-plugin identity on each side with least-privilege users
GET /_plugins/_security/authinfo
Unsafe shortcut: -k everywhere.

Disabling certificate verification can make a lab “work” while deleting server identity from the security model. Keep -k only for the already-documented disposable demo-certificate path. Production uses a trusted CA and hostname/SAN verification.

4. Latency and failure semantics are part of the API contract

Measure remote search p50/p95/p99 separately from local search. A required remote cluster can turn a regional network impairment into user-visible request failure; an optional remote can return success while silently omitting that region. Therefore the application must read cluster metadata, not just HTTP 200.

Signal Question it answers What it does not prove
connected in remote info At least one remote connection is open That a target index is authorized or healthy
CCS _clusters metadata Which clusters succeeded/skipped/failed That omitted data is acceptable to the business
Replication checkpoint/lag Follower progress relative to leader That config/security is synchronized
WAN RTT/packet loss Network condition End-to-end query cost by itself

Explicitly configure required/optional behavior. On Elasticsearch 9.5.3 an unset skip_unavailable is optional by default; on OpenSearch 3.8 it is required by default. AtlasMart sets it explicitly in every environment to prevent a version migration from changing behavior silently.

5. Controlled failure lab

The free/local mandatory path uses two disposable OpenSearch 3.8 clusters or, when cross-cluster TLS is not available on the learner's machine, the deterministic trace below. Do not disable security to make a production-like conclusion. Record the remote alias, transport mode, certificate trust, version, and skip_unavailable value.

Failure-injection sequence
# 1. Healthy baseline
GET /_remote/info
GET /region_b:atlasmart-products-v1/_search?size=0

# 2. Stop only the disposable Region B container/network path.
# 3. Repeat the search with region_b explicitly required.
# 4. Set region_b skip_unavailable=true and repeat.
# 5. Restore Region B; verify reconnection and application correctness.

# Record, do not fabricate:
wan_rtt_p95_ms=MEASURED
ccs_p95_required_ms=MEASURED
ccs_p95_optional_ms=MEASURED
failure_detection_ms=MEASURED

Expected invariant: when required, remote failure must be visible as request failure or failed-cluster metadata according to the API; when optional, the request may succeed but the remote cluster must be marked skipped/partial. The lab passes only if the application distinguishes those outcomes.

Check your understanding

  1. Why can sniff mode work on a flat lab network but fail after Kubernetes deployment?
  2. Why set skip_unavailable explicitly?
  3. What does a green local cluster prove about a remote cluster?
  4. Why are remote credentials region-scoped?
  5. What is the bridge to Lesson 2?
Review the answers

1. Because discovered remote publish addresses must be directly reachable; orchestration/NAT often makes them private or unstable.

2. Its defaults differ between Elasticsearch 9.5.3 and OpenSearch 3.8 and it directly changes completeness/failure semantics.

3. Nothing about remote reachability, authorization, health, or data freshness.

4. To reduce blast radius and avoid one compromised credential granting broad multi-region access.

5. Connectivity is only the substrate; next we examine how a federated search fans out, aggregates results, and reports partial clusters.

6. Production judgment

Choose a remote-cluster mode from network topology, not habit. Require explicit trust, version compatibility, required/optional semantics, and latency/error budgets. A remote search path is ready only when operators can distinguish “remote unavailable,” “remote unauthorized,” “remote slow,” “remote skipped,” and “remote data stale.” Cross-region design begins with observable failure semantics.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.