Compare off-region snapshot strategy, Elastic searchable snapshots, and OpenSearch remote-backed storage without confusing backup, primary storage, cache, and recovery semantics.
Cross-Region Repository Strategy, Searchable Snapshot/Remote Store Concepts, and Cost/Latency Tradeoffs
Design equivalent AtlasMart retention intent in Elastic and OpenSearch while documenting non-equivalent lifecycle, tiering, policy-update, simulation, and managed-service behavior.
Learning outcomes
AtlasMart’s operations team wants three outcomes at once: survive a regional outage, keep older data searchable at low cost, and recover quickly after node loss. A vendor slide labels all three as “remote storage,” but these are different mechanisms. This lesson separates backup repository copies, Elastic searchable snapshots, and OpenSearch remote-backed storage/remote snapshot restore.
Design a cross-region repository strategy with one writer, read-only consumers and storage-layer replication/immutability.
Explain Elastic searchable snapshots as read-only indices backed by an underlying snapshot repository and identify their license dependency.
Explain OpenSearch remote-backed storage as remote segment/translog persistence distinct from ordinary snapshots.
Compare local cache, remote read latency, recovery speed and storage cost without treating cache as backup.
Choose mechanisms by failure mode and prove assumptions with cost/latency/recovery measurements.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0, reviewed 11 September 2026. AtlasMart keeps the Chapter 01
TLS/auth conventions: Elasticsearch at
https://localhost:9200 with CA verification and
OpenSearch at https://localhost:9201 with the
upstream demo certificate only in the disposable lab. Existing
containers remain atlasmart-es and
atlasmart-os. The snapshot lab uses one primary
shard and zero replicas only because the local environment is
single-node; that is not production guidance. A filesystem
repository requires path.repo to be configured
before node start and the repository path to be reachable by
every master/data node that participates. If your earlier
containers were not created with that setting, recreate
disposable lab containers rather than editing production-like
nodes in place.
The required Chapter 18 lab uses ordinary filesystem snapshots and restore. Elastic searchable snapshots currently require an Enterprise license, so they are discussed and can be simulated conceptually with a repository/cache cost model. OpenSearch remote-backed storage requires a separately configured remote-backed cluster, so it is optional; the free/local learning path remains standard snapshots plus measured restore.
The generation environment does not run the AtlasMart Docker clusters, so commands are reproducible lab instructions and response fragments are labeled expected shapes/invariants rather than fabricated measurements. Record your own snapshot duration, bytes transferred, repository latency, restore throughput, p95/p99 application latency, RPO and RTO.
1. Start with failure domains, not product names
| Failure | Mechanism that helps | Mechanism that does not automatically solve it |
|---|---|---|
| Single node/disk loss | Replica shards; OpenSearch remote-backed recovery can also help | A slow off-region snapshot may have higher RTO. |
| Whole cluster deletion/corruption | Independent snapshot recovery point | Replicas mirror the damage. |
| Region outage | Repository/copy accessible from another region + target capacity/runbook | Same-region repository alone. |
| Low-cost read-only historical search | Elastic searchable snapshots or product-specific remote-search modes | A plain snapshot is not automatically queryable as an index. |
| Ransomware/admin compromise | Separate credentials, immutable/off-host copy, retention controls | A repository writable by the same compromised principal. |
2. Cross-region repository strategy: replicate storage, not writers
The safest pattern is one authoritative cluster writing a repository/base path. If another cluster needs to inspect or restore those snapshots, register the replicated/source repository read-only. Cross-region durability is normally achieved at the object-storage layer or by copying recovery artifacts into an independent account/region under a controlled process—not by allowing two clusters to write the same repository namespace.
repository_id: atlasmart-dr-primary
writer_cluster: atlasmart-search-prod-a
writer_region: region-a
secondary_copy_region: region-b
secondary_cluster_registration: read_only
immutability: object-lock/versioning policy documented separately
credentials:
writer: create/list/read/delete under controlled retention role
dr-reader: list/read only
validation:
- verify repository from target
- restore isolated sample monthly
- record observed copy lag
RPO includes replication lag. If a snapshot finishes in region A at 10:00 but the immutable copy becomes available in region B at 10:12, the regional-disaster recovery point cannot be newer than that replicated copy.
3. Elastic searchable snapshots: repository is authoritative for read-only data
Elastic searchable snapshots mount snapshot-backed indices so infrequently accessed data can remain searchable with reduced local storage. Current Elastic documentation marks searchable snapshots as an Enterprise-license feature. The underlying snapshot is the durable source of the mounted index; local cache is an optimization. Deleting the underlying snapshot can therefore destroy the data even if some blocks happen to be cached locally.
| Property | Searchable snapshot implication |
|---|---|
| Mutability | Designed for read-only snapshot-backed indices, not normal hot-write semantics. |
| Repository | Must remain available and protected; it is part of the serving path. |
| Local disk | Cache can reduce remote reads but is not a complete recovery copy. |
| Latency | Cold/cache-miss search can pay repository/network latency; measure p95/p99, not only warm-cache medians. |
| Backup chain | Snapshot of a searchable-snapshot index may mostly reference its original snapshot; protect the original recovery artifact. |
A demo that searches the same query repeatedly after cache warm-up hides first-read and cache-miss costs. Report cold-ish and warm runs separately, disclose cache state, and include repository egress/request cost where applicable.
4. OpenSearch remote-backed storage: remote segments + translog are a different contract
OpenSearch remote-backed storage can persist index segments and
translogs to configured remote repositories. It is designed to
improve durability/recovery characteristics of the live index
path. OpenSearch also exposes restore behavior for remote-backed
clusters and a storage_type: remote_snapshot mode
where remote snapshot data remains authoritative and is
fetched/cached for search. These are OpenSearch-specific
mechanics; they are not “OpenSearch searchable snapshots with
Elastic syntax.”
POST /_snapshot/source-cluster-snapshots/snapshot-1/_restore
{
"indices": "my-remote-index",
"storage_type": "remote_snapshot"
}
Current OpenSearch restore documentation requires the
appropriate remote-backed cluster configuration and a node with
the search role for remote_snapshot. Cross-cluster
remote-backed restore can also require the source remote
segment/translog repositories to be registered read-only. Treat
all of this as a topology-specific feature contract.
5. Remote-backed durability still needs DR reasoning
Remote store can improve recovery after node failure and can preserve acknowledged data under its documented durability conditions. It does not automatically create historical recovery points, legal retention, immutable copies or clean rollback after application corruption. A bad delete written to the live index is still a valid new state. Snapshots preserve older points if retention kept them.
| Question | Snapshot repository | Remote-backed store |
|---|---|---|
| Historical recovery points? | Yes, named snapshots retained by policy | Not inherently; represents live remote durability. |
| Supports rollback after bad write? | Yes, to retained pre-error snapshot | Not by itself. |
| Can be primary serving path? | Ordinary snapshot: no; Elastic searchable snapshots: yes for mounted read-only data | Yes in OpenSearch remote-backed architecture, with product-specific behavior. |
| RPO semantics | Bounded by last completed/replicated snapshot | Can approach acknowledged-write durability for documented configurations, but test failure mode. |
| Cross-product portable? | No generic ES↔OS guarantee | No. Product-specific. |
6. Cost/latency model for AtlasMart
dataset_bytes=...
monthly_growth_bytes=...
snapshot_new_segment_bytes_per_day=...
repository_storage_price=...
repository_request_price=...
cross_region_replication_price=...
restore_bandwidth_bytes_per_sec=...
restore_observed_rto_seconds=...
search_cache_hit_ratio=...
remote_read_p95_ms=...
remote_read_p99_ms=...
monthly_cost = storage + requests + egress + target_capacity + operations
Do not reduce the comparison to storage price per GB. Recovery can require temporarily doubling capacity, moving many terabytes through constrained links, warming caches, recreating templates/security, and running application validation. The least expensive repository is useless if it misses RTO.
7. Cross-region restore acceptance criteria
| Gate | Pass condition |
|---|---|
| Isolation | DR cluster cannot accidentally write the source repository. |
| Replica lag | Newest replicated recovery point is inside regional RPO. |
| Compatibility | Exact target version/product/plugins checked against official docs. |
| Capacity | Target has disk/heap/network/shard headroom to restore. |
| Security | Repository credentials are least privilege; no secrets embedded in repository JSON. |
| Performance | Restore and application warm-up finish inside RTO; p95/p99 smoke tests pass. |
| Evidence | Drill timestamp, snapshot ID, counts/query fixtures and cutover/rollback actions retained. |
8. Production judgment and bridge to the restore drill
AtlasMart should use replicas for fast local availability, snapshots for retained recovery points, independent repository copies for regional/administrative failure, and product-specific remote/searchable storage only where workload economics justify them. Chapter 18’s final lesson now proves the design with a timed restore drill.
Check your understanding
- Why should a DR cluster usually register the source repository read-only?
- What makes Elastic searchable snapshots different from plain snapshots?
- What does OpenSearch remote-backed storage persist?
- Why is a local cache not a backup?
- Which number must cross-region RPO include?
Review the answers
1. To enforce a single writer and prevent repository corruption or unintended mutation while still allowing restore.
2. They mount read-only snapshot-backed data for search, with the repository remaining authoritative; current Elastic docs require Enterprise licensing.
3. Product-specific remote copies of segments and translogs for live-index durability/recovery; it is not simply a historical snapshot schedule.
4. It can be incomplete/evictable and depends on the authoritative remote repository or index state.
5. The lag until a completed recovery point is available in the independent DR region/account, not merely local snapshot completion time.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic snapshot and restore — Snapshot contents, incremental segment reuse, compatibility, and deployment-specific behavior.
- Elastic create/monitor snapshots and SLM — SLM schedules, retention, snapshot status, feature-state backup, and operational monitoring.
- Elastic restore a snapshot — Restore prerequisites, rename-on-restore, capacity, compatibility, and cross-cluster cautions.
- Elastic searchable snapshots — Enterprise-licensed searchable-snapshot behavior and repository dependency.
- OpenSearch snapshot and restore — Incremental snapshots, repository types, restore compatibility, and security constraints.
- OpenSearch Snapshot Management — Scheduled snapshot creation/deletion, failure metadata, plugin and security requirements.
- OpenSearch Snapshot Management API — SM policy schema, explain state, start/stop, schedules, retention and OCC updates.
- OpenSearch remote-backed storage — Remote segment/translog storage and remote-store recovery concepts.
- OpenSearch release artifacts — Pinned OpenSearch release baseline.
- Elasticsearch downloads — Pinned Elasticsearch release baseline.