Turn backup frequency and retention objectives into observable Elastic SLM and OpenSearch Snapshot Management workflows, then prove that scheduled snapshots actually complete.
Snapshot Lifecycle/Scheduled Backups, Retention, Verification, and Monitoring
Design equivalent AtlasMart retention intent in Elastic and OpenSearch while documenting non-equivalent lifecycle, tiering, policy-update, simulation, and managed-service behavior.
Learning outcomes
AtlasMart now has a repository and one successful manual snapshot. That is still an operator-dependent backup process. The next failure mode is predictable: the engineer who remembers the cron job is on leave, the last three scheduled snapshots failed, retention silently kept only recent bad recovery points, and nobody noticed. This lesson converts backup intent into scheduled, observable, retained recovery points.
Translate an RPO into snapshot cadence while separating start time, completion time and restore usefulness.
Configure Elastic SLM and OpenSearch Snapshot Management as product-specific schedulers rather than copied JSON.
Apply retention without deleting the only known-good recovery point or violating legal/compliance holds.
Monitor policy state, snapshot status, repository health and failure history instead of trusting schedule configuration.
Design restore verification as a recurring control, not a one-time preproduction exercise.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0, reviewed 11 September 2026. AtlasMart keeps the Chapter 01
TLS/auth conventions: Elasticsearch at
https://localhost:9200 with CA verification and
OpenSearch at https://localhost:9201 with the
upstream demo certificate only in the disposable lab. Existing
containers remain atlasmart-es and
atlasmart-os. The snapshot lab uses one primary
shard and zero replicas only because the local environment is
single-node; that is not production guidance. A filesystem
repository requires path.repo to be configured
before node start and the repository path to be reachable by
every master/data node that participates. If your earlier
containers were not created with that setting, recreate
disposable lab containers rather than editing production-like
nodes in place.
A scheduler can successfully start a snapshot and still end with a failed or unusable recovery point. Alert on completion/failure metadata, verify repository access, and regularly restore into an isolated target. An RPO is a business promise about recoverable data, not the cron expression printed in a policy.
The generation environment does not run the AtlasMart Docker clusters, so commands are reproducible lab instructions and response fragments are labeled expected shapes/invariants rather than fabricated measurements. Record your own snapshot duration, bytes transferred, repository latency, restore throughput, p95/p99 application latency, RPO and RTO.
1. From “every hour” to an RPO contract
Recovery point objective (RPO) is the maximum acceptable amount of data loss measured backward from the disaster. A 30-minute snapshot schedule does not automatically give a 30-minute RPO: snapshot duration, failures, repository outage, write timing and application replay options matter. Record the timestamp of the last completed and validated recovery point.
| Term | Operational meaning | AtlasMart evidence |
|---|---|---|
| Cadence | When a policy attempts a snapshot | Policy schedule plus synchronized node clocks. |
| Completion age | Age of most recent successful snapshot | Snapshot metadata/status; alert if older than RPO budget. |
| Retention | Which recovery points are eligible for deletion | Age + min/max count + hold/immutability policy. |
| Validation | Whether a recovery point restores and serves expected data | Recurring isolated restore drill with counts/mapping/query fixtures. |
| RPO breach | No validated recovery point inside allowed loss window | Page/incident; do not wait for disaster to discover it. |
2. Elastic SLM: schedule, retention and observability
Elasticsearch Snapshot Lifecycle Management (SLM) is the native policy engine for scheduled snapshots. The policy controls the cron schedule, repository, included data and retention rules. Retention execution is a separate cluster-level task; therefore “policy runs hourly” and “old snapshots are pruned hourly” are different statements.
PUT _slm/policy/atlasmart-dr-slm
{
"schedule": "0 0 * * * ?",
"name": "<atlasmart-dr-{now/d}>",
"repository": "atlasmart-es-fs",
"config": {
"indices": ["atlasmart-dr-products-v1"],
"include_global_state": false
},
"retention": {
"expire_after": "7d",
"min_count": 3,
"max_count": 48
}
}
The exact values above are a lab example, not universal
production policy. Choose cadence from RPO and
repository/cluster capacity. min_count protects
against age-based deletion eliminating every recovery point
after a prolonged failure; it does not guarantee that those
retained snapshots are healthy.
POST _slm/policy/atlasmart-dr-slm/_execute
GET _slm/status
GET _slm/policy/atlasmart-dr-slm
GET _slm/stats
GET _snapshot/atlasmart-es-fs/_current
GET _snapshot/atlasmart-es-fs/_all
Expected invariant: policy metadata distinguishes last success from last failure, snapshot metadata shows actual completion, and the repository contains named recovery points. A successful policy invocation means the snapshot process started; inspect completion separately.
3. OpenSearch Snapshot Management: separate product, separate state machine
OpenSearch Snapshot Management (SM) is provided through the
Index Management plugin and Job Scheduler. It has its own policy
document, creation and deletion schedules, execution metadata,
retries and permissions. Do not paste an Elastic SLM document
into _plugins/_sm.
POST _plugins/_sm/policies/atlasmart-dr-sm
{
"description": "AtlasMart Chapter 18 DR snapshots",
"creation": {
"schedule": {
"cron": {
"expression": "0 * * * *",
"timezone": "UTC"
}
},
"time_limit": "30m"
},
"deletion": {
"schedule": {
"cron": {
"expression": "15 * * * *",
"timezone": "UTC"
}
},
"condition": {
"max_age": "7d",
"min_count": 3,
"max_count": 48
},
"time_limit": "30m"
},
"snapshot_config": {
"date_format": "yyyy-MM-dd-HH-mm",
"timezone": "UTC",
"indices": "atlasmart-dr-products-v1",
"repository": "atlasmart-os-fs",
"ignore_unavailable": false,
"include_global_state": false,
"partial": false
}
}
GET _plugins/_sm/policies/atlasmart-dr-sm
GET _plugins/_sm/policies/atlasmart-dr-sm/_explain
POST _plugins/_sm/policies/atlasmart-dr-sm/_start
# ...after observation...
POST _plugins/_sm/policies/atlasmart-dr-sm/_stop
GET _snapshot/atlasmart-os-fs/_all
The SM explain response exposes the creation/deletion state machine and latest execution status such as in-progress, success, retrying, failed or time-limit exceeded. OpenSearch documents retries for failed snapshot operations. Treat those fields as monitoring signals, not implementation trivia.
4. Retention protects capacity—but can destroy your only clean point
Retention is a destructive control. “Delete anything older than seven days” may be correct for a disposable lab and disastrous for ransomware recovery, legal hold, monthly close, audit or slow corruption discovered weeks later. Separate operational retention from compliance/immutable retention and from a long-term archive schedule.
| Recovery class | Illustrative cadence | Purpose | Do not infer |
|---|---|---|---|
| Frequent | 15–60 min | Short RPO for recent operational mistakes | That every point is application-consistent. |
| Daily | Once/day | Longer rollback window | That daily is sufficient for high-write systems. |
| Monthly/quarterly | Business calendar | Audit/history/release checkpoints | That scheduler retention equals legal retention. |
| Immutable/off-host copy | Storage policy | Protect backup from cluster/admin compromise | That immutability proves restore compatibility. |
If automated deletion prunes older snapshots before the newest point has passed an isolated restore test, a latent repository or application-level problem can erase the last known-good option. Retention gates should incorporate monitoring and, where risk warrants, protected recovery points outside the scheduler’s ordinary deletion authority.
5. Monitor four layers, not one dashboard tile
| Layer | Signal | Question answered |
|---|---|---|
| Scheduler | SLM stats / SM explain | Did policy execution start and finish? |
| Snapshot | snapshot metadata/status, shard failures | Did the requested recovery point complete? |
| Repository | verify/analyze/integrity checks, object-store errors | Can cluster nodes read/write the repository correctly? |
| Restore | scheduled DR drill | Can a compatible target reconstruct usable data within RTO? |
# Elasticsearch
GET _slm/stats
GET _slm/policy/atlasmart-dr-slm
GET _snapshot/atlasmart-es-fs/_current
GET _snapshot/atlasmart-es-fs/_all
POST _snapshot/atlasmart-es-fs/_verify
# OpenSearch
GET _plugins/_sm/policies/atlasmart-dr-sm/_explain
GET _snapshot/atlasmart-os-fs/_current
GET _snapshot/atlasmart-os-fs/_all
POST _snapshot/atlasmart-os-fs/_verify
Alert on age of last completed snapshot, repeated failures, snapshot duration growth, repository access failures, unexpected partial snapshots, restore-test failures and capacity pressure. A low error rate means little if the last successful recovery point is older than the promised RPO.
6. Safe failure injection: unavailable repository, not corrupted blobs
To teach observability, simulate a missing mount or deny repository access on a disposable lab, then run verification/snapshot and capture the explicit failure. Do not hand-edit repository metadata, truncate blobs or corrupt a shared object-store bucket. The learning goal is to verify detection and runbook behavior without creating unrecoverable or ambiguous damage.
failure_id: CH18-SM-REPO-UNAVAILABLE
expected:
- repository verification fails clearly
- scheduled snapshot records failure/retry state
- alert fires before RPO budget is exhausted
- no source index is modified
repair:
- restore repository access/mount
- verify repository
- run one manual snapshot
- restore-test the new recovery point
7. Production judgment
Backup automation consumes I/O, network and repository metadata operations. Measure snapshot duration and impact during realistic indexing/search traffic. Separate backup traffic from user-facing p95/p99 latency budgets, and schedule around cluster headroom rather than folklore. Managed Elastic and Amazon OpenSearch offerings can add platform-managed snapshots with their own scope, region and retention semantics; validate the provider contract instead of assuming it matches self-managed SLM/SM.
Check your understanding
- Why is a cron interval not the same thing as an RPO?
- What is the main product difference between SLM and OpenSearch SM?
- Why keep a minimum snapshot count?
- Which monitoring layer proves application recoverability?
- What is a safe backup failure injection?
Review the answers
1. RPO is based on the age of the last completed, usable recovery point; duration, failure and validation can widen the actual loss window.
2. They are separate policy engines with different APIs, state/metadata, scheduling and permissions; common intent does not make JSON portable.
3. It reduces the chance that age-based retention removes every recovery point during an extended failure, though it does not prove the retained points are healthy.
4. A restore drill with schema/data/query/application validation, not scheduler or snapshot status alone.
5. Temporarily make a disposable repository unavailable or deny access, observe clear failure, then repair and verify—do not corrupt repository blobs manually.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic snapshot and restore — Snapshot contents, incremental segment reuse, compatibility, and deployment-specific behavior.
- Elastic create/monitor snapshots and SLM — SLM schedules, retention, snapshot status, feature-state backup, and operational monitoring.
- Elastic restore a snapshot — Restore prerequisites, rename-on-restore, capacity, compatibility, and cross-cluster cautions.
- Elastic searchable snapshots — Enterprise-licensed searchable-snapshot behavior and repository dependency.
- OpenSearch snapshot and restore — Incremental snapshots, repository types, restore compatibility, and security constraints.
- OpenSearch Snapshot Management — Scheduled snapshot creation/deletion, failure metadata, plugin and security requirements.
- OpenSearch Snapshot Management API — SM policy schema, explain state, start/stop, schedules, retention and OCC updates.
- OpenSearch remote-backed storage — Remote segment/translog storage and remote-store recovery concepts.
- OpenSearch release artifacts — Pinned OpenSearch release baseline.
- Elasticsearch downloads — Pinned Elasticsearch release baseline.