Run federated AtlasMart searches across local and remote indices/data streams, inspect cluster-completeness metadata, and choose CCS only when WAN/query tradeoffs are acceptable.
Cross-Cluster Search Across Indices/Data Streams and Federated Query Tradeoffs
Design multi-cluster and multi-region search/replication with explicit latency, security, write ownership, failure semantics, and measured recovery objectives.
Learning outcomes
AtlasMart product search is now federated across Region A and Region B. The next question is not “does the query return hits?” but “what work happened on each cluster, how much WAN coordination was added, and can the result be considered complete when one cluster is slow or missing?”
Explain CCS fan-out and merge without pretending remote shards are local.
Use remote-alias index/data-stream expressions correctly and inspect cluster-level response metadata.
Control required/optional cluster semantics and understand minimized-round-trip tradeoffs.
Diagnose WAN-dominated tail latency and fetch/aggregation amplification.
Choose between federated search, replication, and upstream aggregation for a concrete workload.
Examples are reviewed against
Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep
the course's existing local TLS/auth assumptions.
Cross-cluster behavior is distribution-, license-, network-,
and managed-service-sensitive, so every exercise starts by
recording GET /, license/plugin state,
remote-cluster settings, TLS trust, and the exact feature path
being tested. Elastic's advanced API-key remote-cluster model
and CCR have subscription boundaries; OpenSearch's replication
plugin is bundled in the standard distribution but managed
services can impose different connection/IAM constraints.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Federated query flow
A cross-cluster search starts on one coordinating cluster. It resolves local and remote index expressions, contacts configured remotes, performs distributed search in each participating cluster, and merges results. The WAN boundary sits between cluster-level coordinators, so local shard count and remote shard count are not the only costs. Serialized requests/results, remote queuing, TLS, and long-tail network delay matter.
# Elasticsearch and OpenSearch shape
GET /atlasmart-products-v1,region_b:atlasmart-products-v1/_search
{
"size": 20,
"query": {"bool": {"filter": [{"term": {"status": "active"}}]}},
"sort": [{"updated_at": "desc"}, {"product_id": "asc"}]
}
# Search only remote data
GET /region_b:atlasmart-products-v1/_search?size=0
Remote aliases are part of the index expression, not hostnames. Permissions must succeed both for the caller's local role and the remote authorization model. Data streams follow the same principle: target the remote alias plus stream expression supported by the product/version.
2. Read _clusters before trusting totals
Elasticsearch CCS responses expose cluster metadata describing
total, successful, skipped, failed, and sometimes partial
clusters. OpenSearch also surfaces remote failures through CCS
behavior; managed products may add service-specific connection
state. An application that only reads
hits.total can mistake a partial global view for a
complete one.
result = ccs_search(...)
assert result.http_status == 200
# Product-specific response parsing:
if result.clusters.failed > 0:
reject("global search incomplete")
if result.clusters.skipped > 0 and business_requires_all_regions:
reject("required region skipped")
publish(result, completeness="complete" or "degraded")
With skip_unavailable=true, Elasticsearch can
ignore remote connectivity and shard/index errors and still
return HTTP 200. The application must explicitly decide
whether a degraded result is acceptable.
3. Minimize round trips—but measure
CCS implementations can minimize cross-cluster round trips so
the local cluster exchanges fewer WAN messages with each remote.
That often reduces latency when networks are expensive, but it
changes when shard discovery and intermediate reduction happen.
OpenSearch also exposes the
ccs_minimize_roundtrips search parameter; some
permission paths require additional shard-search permissions
when it is disabled. Managed services may restrict combinations.
| Workload | Federated CCS fit | Alternative |
|---|---|---|
| Interactive global search, modest result set | Often good if regions are reachable and partial semantics are explicit | Replicate hot search index closer to users when WAN tail dominates |
| Huge export across regions | Poor: WAN result transfer/coordinator pressure | Batch ETL/object storage/analytics pipeline |
| Regional dashboards with tolerant partial data | Good with explicit degraded-state UI | Central replicated reporting cluster |
| Strict write-after-read global consistency | CCS does not solve write ownership/replication freshness | Single write owner + replication/freshness contract |
4. Relevance and aggregation boundaries
BM25 remains query-relative and shard/segment statistics matter.
Cross-cluster search does not make score values globally
portable across unrelated corpora. For aggregations, each
cluster/shard computes partial state that the coordinator
reduces. High-cardinality buckets, large size, or
returning many source fields can multiply WAN bytes and
coordinator memory. Treat profile as a diagnostic,
not a production benchmark.
# Baseline local
GET /atlasmart-products-v1/_search?filter_path=took,_shards,hits.total
{ "size": 0, "aggs": {"brands": {"terms": {"field": "brand", "size": 20}}}}
# Federated
GET /atlasmart-products-v1,region_b:atlasmart-products-v1/_search?filter_path=took,_clusters,_shards,hits.total
{ "size": 0, "aggs": {"brands": {"terms": {"field": "brand", "size": 20}}}}
# Measure client-side bytes and p50/p95/p99 outside the response too.
5. Lab: remote outage and concurrent data change
Index a deterministic AtlasMart document set in both clusters with one region-specific marker. Run a federated query, record cluster metadata and counts, then stop Region B. Repeat once with Region B required and once optional. Restore Region B, insert one new document there, and rerun to prove that CCS freshness is simply the remote cluster's current searchable state—it is not replication.
scenario,required,remote_state,http,clusters_success,clusters_skipped,clusters_failed,hits,p95_ms
baseline,true,up,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED
outage,true,down,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED
outage,false,down,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED
restored,true,up,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED,MEASURED
Do not pre-fill timings. The point is to prove whether WAN and remote-cluster availability violate AtlasMart's p95/p99 and completeness objectives.
Check your understanding
- Why is HTTP 200 insufficient for global-search correctness?
- When does replication beat CCS?
- Why can large aggregations be expensive across clusters?
- Does CCS replicate data?
- What is Lesson 3 about?
Review the answers
1. An optional remote may be skipped or partial while the overall request still succeeds.
2. When serving queries locally with bounded replication lag is preferable to paying WAN/query availability on every request.
3. Partial aggregation state and final reductions consume remote CPU, WAN bandwidth, and coordinator memory.
4. No. It queries data in place; freshness is whatever each remote cluster can currently search.
5. Elastic CCR, where one cluster owns writes and followers pull changes for local reads/DR.
6. Production judgment
Federated search is a query architecture, not a DR plan. Use it when live remote data is worth the WAN dependency. If user-facing latency or strict availability cannot tolerate a remote hop, replicate selected data or build a dedicated reporting plane. Always surface completeness state to callers.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic remote clusters — Current connection, security-model, managed-environment, and remote-cluster guidance.
- Elastic remote-cluster connection modes — Sniff versus proxy semantics and network reachability.
-
Elastic remote-cluster settings
—
skip_unavailable, roles, seeds, proxy settings, and defaults. - Elastic cross-cluster search — Search syntax, optional clusters, response metadata, and supported APIs.
- Elastic cross-cluster replication — Active-passive mechanics, limitations, and DR patterns.
- Elastic auto-follow patterns — Rolling-index/data-stream follower automation.
- Elastic subscriptions — Current cross-cluster licensing boundaries; verify the deployed license before labs.
- OpenSearch cross-cluster search — Security flow, remote connections, permissions, and examples.
- OpenSearch cross-cluster replication — Replication plugin model and operational behavior.
- OpenSearch CCR getting started — Sniff/proxy connectivity, roles, follower startup, and pull replication.
- OpenSearch auto-follow — Pattern-based automatic replication.
- Amazon OpenSearch Service cross-cluster search — Managed-service connection and IAM/network restrictions.
- Amazon OpenSearch Service cross-cluster replication — Managed CCR constraints and domain connection workflow.