Run an evidence-driven AtlasMart disaster-recovery drill that measures RPO/RTO and validates data, mappings, aliases, templates, security assumptions, and application cutover.
Run a Restore Drill and Measure RPO/RTO, Data Completeness, Security State, Templates, and Application Cutover
Design equivalent AtlasMart retention intent in Elastic and OpenSearch while documenting non-equivalent lifecycle, tiering, policy-update, simulation, and managed-service behavior.
Learning outcomes
The only defensible statement that “AtlasMart can recover in four hours with at most one hour of data loss” is a measured restore drill with evidence. This lesson turns the chapter into a runbook: capture a recovery point, create a known post-snapshot change, simulate source loss, restore under an isolated name, validate data/configuration/security assumptions, cut an alias deliberately, measure RPO/RTO, and prove rollback.
Define start/stop timestamps and evidence needed to measure RPO and RTO rather than estimate them.
Run an isolated restore without deleting the original recovery evidence or blindly restoring global/security state.
Validate data completeness, mappings/settings, templates, aliases and representative application queries.
Execute an explicit application cutover and rollback using a stable alias or configuration boundary.
Produce a DR report that separates measured facts, assumptions, version/product constraints and remediation actions.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0, reviewed 11 September 2026. AtlasMart keeps the Chapter 01
TLS/auth conventions: Elasticsearch at
https://localhost:9200 with CA verification and
OpenSearch at https://localhost:9201 with the
upstream demo certificate only in the disposable lab. Existing
containers remain atlasmart-es and
atlasmart-os. The snapshot lab uses one primary
shard and zero replicas only because the local environment is
single-node; that is not production guidance. A filesystem
repository requires path.repo to be configured
before node start and the repository path to be reachable by
every master/data node that participates. If your earlier
containers were not created with that setting, recreate
disposable lab containers rather than editing production-like
nodes in place.
Run this drill only on the disposable
atlasmart-dr-* fixture or an isolated recovery
cluster. Do not delete production indices, mutate a real
snapshot repository, restore security/global state, or repoint
production clients merely to practice the runbook.
The generation environment does not run the AtlasMart Docker clusters, so commands are reproducible lab instructions and response fragments are labeled expected shapes/invariants rather than fabricated measurements. Record your own snapshot duration, bytes transferred, repository latency, restore throughput, p95/p99 application latency, RPO and RTO.
1. Define the clock before the incident
RPO measures recoverable data age at the failure boundary. RTO measures how long service recovery takes from declared incident until the agreed service acceptance gate passes. Decide what “service recovered” means: cluster green? index searchable? API smoke tests passing? full p95/p99 SLO restored? Different definitions produce different RTOs.
| Timestamp | Meaning |
|---|---|
| T_snap_start | Snapshot operation starts. |
| T_snap_end | Snapshot successfully completes. |
| T_last_recoverable_event | Newest business event proven present in that snapshot. |
| T_incident | Failure/data-loss incident declared. |
| T_restore_start | Restore request accepted. |
| T_data_ready | Restored index shards available for validation. |
| T_app_ready | Smoke tests/SLO gate passes after cutover. |
observed_rpo = T_incident - T_last_recoverable_event
observed_data_restore_time = T_data_ready - T_restore_start
observed_rto = T_app_ready - T_incident
# Do not substitute:
# schedule_interval for observed_rpo
# restore_API_ack for T_app_ready
2. Prepare recovery point and deliberate post-snapshot evidence
Use snapshot 002 from Lesson 3 or create a fresh snapshot. Then write one document after snapshot completion. That post-snapshot document acts as a visible RPO marker: an ordinary restore of the snapshot should not contain it.
PUT atlasmart-dr-products-v1/_doc/P-AFTER-SNAPSHOT?refresh=wait_for
{
"sku":"AM-DR-AFTER",
"name":"Post Snapshot Marker",
"category":"dr-test",
"price":1.00,
"updated_at":"2026-09-11T10:05:00Z"
}
GET atlasmart-dr-products-v1/_doc/P-AFTER-SNAPSHOT
Record the exact snapshot completion time and marker event time. In a real system, the newest recoverable business event should come from a canonical event/transaction sequence, not from wall-clock guessing alone.
3. Simulate source loss safely
The safest drill is to leave the source intact and pretend it is unavailable to the recovery operator. If you want deletion semantics, clone the fixture first and delete only the disposable clone. The goal is to practice restore, not to manufacture real data loss.
POST _aliases
{
"actions": [
{"add": {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-read"}},
{"add": {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-write", "is_write_index":true}}
]
}
GET _alias/atlasmart-dr-products-read,atlasmart-dr-products-write
Clients in the drill should use the aliases, not the physical index name. This gives the recovery procedure an explicit cutover boundary and lets rollback remain an alias operation instead of a destructive restore-overwrite.
4. Restore the snapshot under an isolated name and time it
# Record T_restore_start immediately before this request.
POST _snapshot/atlasmart-es-fs/atlasmart-dr-2026.09.11-002/_restore?wait_for_completion=true
{
"indices":"atlasmart-dr-products-v1",
"include_global_state":false,
"include_aliases":false,
"rename_pattern":"atlasmart-dr-products-v1",
"rename_replacement":"atlasmart-dr-products-drill-002"
}
# Then inspect recovery, health and index state.
GET _cluster/health/atlasmart-dr-products-drill-002
GET atlasmart-dr-products-drill-002/_recovery?active_only=false
GET atlasmart-dr-products-drill-002/_count
Expected invariant: restore completes without
mutating the source index or aliases. The exact response timing
is environment-specific.
wait_for_completion=true measures API completion
for the restore operation, but your RTO continues until
application validation and cutover gates pass.
5. Validate completeness and the intentional RPO gap
GET atlasmart-dr-products-drill-002/_doc/P-1001
GET atlasmart-dr-products-drill-002/_doc/P-1006
# The marker was written AFTER snapshot 002 and should be absent.
GET atlasmart-dr-products-drill-002/_doc/P-AFTER-SNAPSHOT
GET atlasmart-dr-products-drill-002/_mapping
GET atlasmart-dr-products-drill-002/_settings
GET atlasmart-dr-products-drill-002/_search
{
"size": 0,
"aggs": {
"by_category": {"terms": {"field":"category","size":10}},
"price_sum": {"sum": {"field":"price"}}
}
}
The missing post-snapshot marker is not a restore bug; it is a visible representation of the recovery point. If the documented RPO is smaller than that gap, the design needs more frequent snapshots, external change replay, or another durability mechanism.
6. Validate dependencies outside the restored index
| Dependency | Drill action | Why it matters |
|---|---|---|
| Index templates | GET expected templates and compare version/hash | Future rollover/new index must recreate the same schema contract. |
| Aliases | Inspect source and restored alias state before cutover | Snapshot restore should not unexpectedly hijack live traffic. |
| Ingest pipelines | Verify names/versions referenced by writers | Restored data may search fine while new writes fail or transform differently. |
| Lifecycle | Confirm policy attachment/timing | A recovered index must not be immediately deleted or moved unexpectedly. |
| Security | Verify service principal can read restored alias after controlled grant | Data recovery without application authorization is not service recovery. |
| TLS/keystore/config files | Recover from separate secure configuration backup/runbook | Snapshots do not replace node config/secret backup. |
Elasticsearch feature states and OpenSearch Security-plugin configuration require product-specific privileged handling. The mandatory lab intentionally does not restore them. Instead, record whether the DR environment obtains security configuration from infrastructure-as-code, identity provider integration, separate secure backups, or a product-supported security-state recovery process.
7. Atomic application cutover and smoke tests
POST _aliases
{
"actions": [
{"remove": {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-read"}},
{"add": {"index":"atlasmart-dr-products-drill-002", "alias":"atlasmart-dr-products-read"}}
]
}
GET _alias/atlasmart-dr-products-read
GET atlasmart-dr-products-read/_search
{
"query": {"match": {"name":"wireless headphones"}},
"size": 3
}
If the recovered index is intended to resume writes, do not simply repoint the write alias until you have decided how to reconcile post-snapshot writes, mapping/version constraints and write ownership. In many DR designs the restored cluster becomes authoritative only after an explicit incident command and application-level replay/catch-up step.
8. Rollback evidence
POST _aliases
{
"actions": [
{"remove": {"index":"atlasmart-dr-products-drill-002", "alias":"atlasmart-dr-products-read"}},
{"add": {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-read"}}
]
}
Rollback is only possible because the drill preserved the original source. In a real regional disaster the old cluster may be gone; “rollback” could instead mean abandon the failed target and retry restore into another target. Write that distinction into the runbook.
9. Disaster-recovery report
drill_id: atlasmart-ch18-2026-09-11
source_product_version: ...
target_product_version: ...
repository: ...
snapshot: ...
snapshot_start: ...
snapshot_end: ...
last_recoverable_business_event: ...
incident_time: ...
restore_start: ...
data_ready: ...
application_ready: ...
observed_rpo: ...
observed_rto: ...
source_count: ...
restored_count: ...
mapping_diff: PASS|FAIL
query_fixture: PASS|FAIL
template_check: PASS|FAIL
alias_cutover: PASS|FAIL
security_assumption: documented
post_snapshot_marker_absent: EXPECTED|UNEXPECTED
rollback_test: PASS|FAIL
open_actions:
- ...
| If this fails… | Likely design action |
|---|---|
| RPO too large | Increase snapshot cadence, reduce failures/replication lag, add event-log replay or change durability architecture. |
| Restore too slow | Increase restore bandwidth/target capacity, reduce shard/pathology, pre-stage DR infrastructure, evaluate product-specific remote/searchable options. |
| Schema/query mismatch | Version templates/analyzers and include application-level regression tests in backup validation. |
| Security cutover fails | Treat identity/roles/certs/API keys as a first-class DR dependency with separate tested recovery. |
| Repository unavailable | Add independent repository copies/credentials/network paths and regularly verify from DR environment. |
10. Chapter completion criteria
Chapter 18 is complete only when AtlasMart can point to a named snapshot, prove its repository is reachable, restore it into an isolated target, explain exactly what global/security state was excluded, demonstrate the expected RPO gap, pass data/schema/application checks, execute cutover and rollback, and report measured RPO/RTO. This hands off naturally to Chapter 19: the recovered service must now be secured with correct TLS, identities, roles, API keys and audit controls.
Check your understanding
- What starts and stops the RTO clock?
- Why write a marker after the snapshot?
- Why keep read/write aliases during the drill?
- Why is security state validated separately?
- What is the strongest evidence that backup works?
Review the answers
1. Use a documented incident start and stop only when the agreed service acceptance gate passes; restore API acknowledgement alone is not service recovery.
2. It makes the recovery boundary observable: its expected absence after restore demonstrates what data falls outside that recovery point.
3. They provide an explicit, atomic application cutover boundary and make rollback testable without destructive renaming.
4. Restoring old authority can reintroduce revoked privilege or lock operators out; Elastic/OpenSearch security recovery mechanisms also differ.
5. A repeated, compatible restore drill with data/schema/dependency/application validation and measured RPO/RTO.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic snapshot and restore — Snapshot contents, incremental segment reuse, compatibility, and deployment-specific behavior.
- Elastic create/monitor snapshots and SLM — SLM schedules, retention, snapshot status, feature-state backup, and operational monitoring.
- Elastic restore a snapshot — Restore prerequisites, rename-on-restore, capacity, compatibility, and cross-cluster cautions.
- Elastic searchable snapshots — Enterprise-licensed searchable-snapshot behavior and repository dependency.
- OpenSearch snapshot and restore — Incremental snapshots, repository types, restore compatibility, and security constraints.
- OpenSearch Snapshot Management — Scheduled snapshot creation/deletion, failure metadata, plugin and security requirements.
- OpenSearch Snapshot Management API — SM policy schema, explain state, start/stop, schedules, retention and OCC updates.
- OpenSearch remote-backed storage — Remote segment/translog storage and remote-store recovery concepts.
- OpenSearch release artifacts — Pinned OpenSearch release baseline.
- Elasticsearch downloads — Pinned Elasticsearch release baseline.