Understand snapshot repositories as shared backup systems built from immutable Lucene segments, with explicit consistency, access-control, compatibility, and single-writer rules.
Snapshot Repository Architecture, Incremental Segment Reuse, Consistency, and Repository Access Control
Design equivalent AtlasMart retention intent in Elastic and OpenSearch while documenting non-equivalent lifecycle, tiering, policy-update, simulation, and managed-service behavior.
Learning outcomes
AtlasMart’s search cluster is green because every primary shard has a replica, so an engineer proposes that “backup is already solved.” The design fails the moment an operator deletes an index, credentials are compromised, a bad migration rewrites documents, or an entire failure domain disappears: replicas faithfully reproduce the current cluster state, including destructive mistakes. A snapshot is a separate recovery artifact stored in a snapshot repository outside the live shard set.
Explain why replicas, remote-backed storage and snapshots solve different failure classes.
Trace how immutable Lucene segments make snapshots incremental and deduplicated without making old snapshots dependent on mutable files.
Register and verify isolated filesystem repositories with single-writer and least-privilege rules.
Reason about snapshot consistency, primary-shard participation, start/end windows and application-level point-in-time requirements.
Design repository access, credentials, immutability/off-host copies and compatibility evidence before calling a backup recoverable.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0, reviewed 11 September 2026. AtlasMart keeps the Chapter 01
TLS/auth conventions: Elasticsearch at
https://localhost:9200 with CA verification and
OpenSearch at https://localhost:9201 with the
upstream demo certificate only in the disposable lab. Existing
containers remain atlasmart-es and
atlasmart-os. The snapshot lab uses one primary
shard and zero replicas only because the local environment is
single-node; that is not production guidance. A filesystem
repository requires path.repo to be configured
before node start and the repository path to be reachable by
every master/data node that participates. If your earlier
containers were not created with that setting, recreate
disposable lab containers rather than editing production-like
nodes in place.
A replica is an online shard copy used for availability and read capacity. It follows the primary and therefore follows deletions, mapping changes and bad writes. A snapshot is an independently retained recovery point in an external repository. OpenSearch remote-backed storage and Elastic searchable snapshots are additional storage/recovery mechanisms, but neither should be silently substituted for an independently tested backup policy.
The generation environment does not run the AtlasMart Docker clusters, so commands are reproducible lab instructions and response fragments are labeled expected shapes/invariants rather than fabricated measurements. Record your own snapshot duration, bytes transferred, repository latency, restore throughput, p95/p99 application latency, RPO and RTO.
1. Repository architecture: control plane, metadata and immutable segment blobs
Both Elasticsearch and OpenSearch ultimately snapshot Lucene shard contents. Lucene writes immutable segments: once a segment is published, later document updates are represented by new segments plus deletion metadata rather than in-place mutation. Snapshotting can therefore copy segment blobs that are not already present in the repository and reuse existing blobs across later snapshots. The repository also stores snapshot and index metadata that tells the cluster which immutable blobs form each logical recovery point.
| Layer | What it owns | Failure implication |
|---|---|---|
| Live primary/replica shards | Current searchable/indexable shard copies | Node loss can be tolerated while enough copies remain; logical corruption is replicated. |
| Snapshot repository | Snapshot metadata plus immutable segment objects | Provides retained recovery points outside live shard allocation. |
| Repository credentials / ACL | Who may read, write, delete or enumerate backups | Compromise can turn backup into a ransomware target; separate duties and audit access. |
| Repository catalog/metadata | Which snapshots reference which blobs | Do not manually delete repository files; APIs preserve shared-blob references. |
| Restore target cluster | Capacity, version compatibility, templates/security/runtime config | A usable snapshot still needs a compatible and sufficiently sized restore environment. |
Incremental means later snapshots generally transfer only new segment data; it does not mean “snapshot B is a patch that cannot survive without snapshot A.” Each snapshot is logically independent because repository metadata tracks shared segments. When the product deletes a snapshot through its API, it removes only blobs no remaining snapshot needs.
2. Consistency is shard-consistent, not a magical cluster-wide transaction
A snapshot has a start and end time. A primary shard contributes a stable segment view from some point inside that interval. Elasticsearch explicitly documents that a snapshot is not a precise cluster-wide point-in-time view; OpenSearch likewise warns that different shards can be captured at different times. This matters when AtlasMart must restore several indices whose records have cross-index business invariants.
| Requirement | Snapshot alone? | Safer design |
|---|---|---|
| Recover one independently searchable product index | Usually sufficient | Measure snapshot completion and validate counts/query fixtures after restore. |
| Recover order + payment indices at one business transaction boundary | Not guaranteed by shard snapshot timing | Coordinate application writes, use transaction/event-log replay, or record a consistent external checkpoint. |
| Protect against node loss | Snapshot is useful but slower than replica failover | Use replicas/failure domains for HA and snapshots for DR. |
| Protect against accidental delete | Yes, if retention preserves a pre-delete recovery point | Keep repository inaccessible to ordinary app credentials and test rename-on-restore. |
Snapshot creation depends on available primary shards. If required primaries are unavailable, the operation can fail or become partial only when you explicitly allow partial behavior. A green replica set before snapshotting improves availability, but the snapshot still reads from primaries and is not created from “any copy whatsoever.”
3. AtlasMart repository prerequisite and registration
The mandatory lab uses separate filesystem repositories so
Elasticsearch and OpenSearch never share writable repository
contents. In a multi-node cluster the path must be a real shared
filesystem mounted consistently. In the one-node lab a Docker
volume or host bind mount is enough. Configure
path.repo before startup; repository registration
is cluster metadata, not a way to grant the process arbitrary
filesystem access.
# Elasticsearch
curl --cacert ./atlasmart-es-http-ca -u "elastic:$ELASTIC_PASSWORD" \
https://localhost:9200/
curl --cacert ./atlasmart-es-http-ca -u "elastic:$ELASTIC_PASSWORD" \
'https://localhost:9200/_nodes/jvm,settings?filter_path=nodes.*.jvm.version,nodes.*.settings.path.repo'
# OpenSearch (demo TLS only; -k is NOT production practice)
curl -k -u "admin:$OPENSEARCH_INITIAL_ADMIN_PASSWORD" \
https://localhost:9201/
curl -k -u "admin:$OPENSEARCH_INITIAL_ADMIN_PASSWORD" \
'https://localhost:9201/_nodes/jvm,settings?filter_path=nodes.*.jvm.version,nodes.*.settings.path.repo'
# Elasticsearch
PUT _snapshot/atlasmart-es-fs
{
"type": "fs",
"settings": {
"location": "/mnt/snapshots/elastic",
"compress": true
}
}
# OpenSearch
PUT _snapshot/atlasmart-os-fs
{
"type": "fs",
"settings": {
"location": "/mnt/snapshots/opensearch",
"compress": true
}
}
For object-store repositories, keep secret material in the product’s supported secure credential mechanism (for example a keystore or workload identity), not in lesson files, Git, curl history or repository JSON. The repository definition should point to a bucket/base path; it should not become a secret vault.
4. Verify connectivity before trusting a snapshot
POST _snapshot/atlasmart-es-fs/_verify
POST _snapshot/atlasmart-os-fs/_verify
Expected invariant: the response names the nodes that can access the repository. That proves basic repository reachability and common operations from those nodes; it does not prove that every future snapshot is restorable, that a third-party object store implements all required concurrency semantics, or that your restore target has capacity. Elasticsearch also offers repository analysis and an experimental deep integrity verification API for stronger diagnostics; treat those as additional evidence, not substitutes for restore drills.
5. Build a deterministic AtlasMart recovery fixture
PUT atlasmart-dr-products-v1
{
"settings": {"number_of_shards": 1, "number_of_replicas": 0},
"mappings": {
"dynamic": "strict",
"properties": {
"sku": {"type": "keyword"},
"name": {"type": "text", "fields": {"raw": {"type": "keyword"}}},
"category": {"type": "keyword"},
"price": {"type": "double"},
"updated_at": {"type": "date"}
}
}
}
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-dr-products-v1","_id":"P-1001"}}
{"sku":"AM-AU-100","name":"Wireless Noise Cancelling Headphones","category":"audio","price":199.99,"updated_at":"2026-09-11T08:00:00Z"}
{"index":{"_index":"atlasmart-dr-products-v1","_id":"P-1002"}}
{"sku":"AM-AU-200","name":"Wired Studio Headphones","category":"audio","price":89.99,"updated_at":"2026-09-11T08:01:00Z"}
{"index":{"_index":"atlasmart-dr-products-v1","_id":"P-1003"}}
{"sku":"AM-WB-300","name":"Portable Bluetooth Speaker","category":"audio","price":79.99,"updated_at":"2026-09-11T08:02:00Z"}
{"index":{"_index":"atlasmart-dr-products-v1","_id":"P-1004"}}
{"sku":"AM-WN-500","name":"Aluminum Laptop Stand","category":"office","price":49.00,"updated_at":"2026-09-11T08:03:00Z"}
{"index":{"_index":"atlasmart-dr-products-v1","_id":"P-1005"}}
{"sku":"AM-ST-600","name":"Lightweight Running Shoes","category":"sports","price":69.00,"updated_at":"2026-09-11T08:04:00Z"}
Record pre-snapshot evidence: count, mapping, settings, selected IDs, and a stable query result. Do not use an order-dependent whole-document checksum unless you canonicalize ordering and serialization first.
GET atlasmart-dr-products-v1/_count
GET atlasmart-dr-products-v1/_mapping
GET atlasmart-dr-products-v1/_settings?filter_path=*.settings.index.number_of_shards,*.settings.index.number_of_replicas
GET atlasmart-dr-products-v1/_search
{
"size": 0,
"aggs": {
"by_category": {"terms": {"field": "category", "size": 10}},
"price_sum": {"sum": {"field": "price"}}
}
}
6. Create snapshot 001 and inspect what it proves
# Elasticsearch
PUT _snapshot/atlasmart-es-fs/atlasmart-dr-2026.09.11-001?wait_for_completion=true
{
"indices": "atlasmart-dr-products-v1",
"include_global_state": false
}
# OpenSearch
PUT _snapshot/atlasmart-os-fs/atlasmart-dr-2026.09.11-001?wait_for_completion=true
{
"indices": "atlasmart-dr-products-v1",
"include_global_state": false
}
GET _snapshot/atlasmart-es-fs/atlasmart-dr-2026.09.11-001
GET _snapshot/atlasmart-es-fs/atlasmart-dr-2026.09.11-001/_status
GET _snapshot/atlasmart-os-fs/atlasmart-dr-2026.09.11-001
GET _snapshot/atlasmart-os-fs/atlasmart-dr-2026.09.11-001/_status
Expected invariant: state is successful, the intended index is listed, shard success/failure counts are explicit, and start/end timestamps exist. This proves the cluster completed the snapshot API workflow. It still does not prove that a different cluster/version can restore it or that the application will behave correctly after restore.
7. The deliberately wrong design: one repository, many writers
# WRONG conceptual design
cluster-a -> same bucket/base_path with write access
cluster-b -> same bucket/base_path with write access
manual cron -> rm old snapshot files directly in object storage
Repository metadata is not a generic folder where independent writers may safely append arbitrary files. Elasticsearch explicitly requires a single writer when the same repository is registered by multiple clusters; secondary clusters should register it read-only. OpenSearch likewise warns against conflicting repository use and manual deletion because snapshots share data. The repair is one authoritative writer, read-only consumers, API-driven deletion and immutable/off-host replication at the storage layer if a second failure domain is required.
Elasticsearch snapshots follow Elasticsearch snapshot/index compatibility rules. OpenSearch documents its own compatibility constraints. Shared historical ancestry does not create a supported arbitrary Elasticsearch ↔ OpenSearch snapshot-restore contract. For cross-product migration, plan a supported reindex/export pipeline and validate mappings, analyzers and application semantics.
8. Production judgment: define the recovery contract now
| Decision surface | Evidence to require |
|---|---|
| Backup frequency / RPO | Maximum acceptable loss window, observed snapshot schedule/finish times, write-log replay strategy. |
| RTO | Restore-throughput measurement, target-cluster provisioning time, shard recovery time, application smoke-test time. |
| Repository security | Single writer, least privilege, audit trail, secret storage, encryption, immutability/object lock where appropriate. |
| Regional failure | Repository copy/replication in an independent failure domain plus tested read-only registration from DR region. |
| Compatibility | Source version, target version, index creation versions, plugins/features, product distribution and restore test. |
| Capacity | Free disk, network/object-store bandwidth, recovery throttles, shard count, heap and filesystem cache headroom. |
Check your understanding
- Why can a green cluster still need snapshots?
- What makes snapshots incremental?
- Does a snapshot represent one exact cluster-wide instant?
- Why should only one cluster write a shared repository?
- What proves a backup is actually useful?
Review the answers
1. Replicas protect online availability, not independent historical recovery from deletions, bad writes, compromise or whole-cluster failure.
2. Immutable Lucene segments can be reused across snapshots, so only new segment data generally needs to be copied.
3. No. Each shard contributes a view from within the snapshot start/end window; strict cross-index consistency needs application-level coordination.
4. Concurrent independent writers can corrupt or invalidate repository metadata/assumptions; other clusters should be read-only.
5. A compatible restore drill plus data/schema/application validation and measured RPO/RTO—not a successful snapshot API response alone.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic snapshot and restore — Snapshot contents, incremental segment reuse, compatibility, and deployment-specific behavior.
- Elastic create/monitor snapshots and SLM — SLM schedules, retention, snapshot status, feature-state backup, and operational monitoring.
- Elastic restore a snapshot — Restore prerequisites, rename-on-restore, capacity, compatibility, and cross-cluster cautions.
- Elastic searchable snapshots — Enterprise-licensed searchable-snapshot behavior and repository dependency.
- OpenSearch snapshot and restore — Incremental snapshots, repository types, restore compatibility, and security constraints.
- OpenSearch Snapshot Management — Scheduled snapshot creation/deletion, failure metadata, plugin and security requirements.
- OpenSearch Snapshot Management API — SM policy schema, explain state, start/stop, schedules, retention and OCC updates.
- OpenSearch remote-backed storage — Remote segment/translog storage and remote-store recovery concepts.
- OpenSearch release artifacts — Pinned OpenSearch release baseline.
- Elasticsearch downloads — Pinned Elasticsearch release baseline.