Run an evidence-driven AtlasMart disaster-recovery drill that measures RPO/RTO and validates data, mappings, aliases, templates, security assumptions, and application cutover.

Run a Restore Drill and Measure RPO/RTO, Data Completeness, Security State, Templates, and Application Cutover

Design equivalent AtlasMart retention intent in Elastic and OpenSearch while documenting non-equivalent lifecycle, tiering, policy-update, simulation, and managed-service behavior.

Intermediate → Advanced120–160 minutesDisaster-recovery drill & RPO/RTO evidenceElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

The only defensible statement that “AtlasMart can recover in four hours with at most one hour of data loss” is a measured restore drill with evidence. This lesson turns the chapter into a runbook: capture a recovery point, create a known post-snapshot change, simulate source loss, restore under an isolated name, validate data/configuration/security assumptions, cut an alias deliberately, measure RPO/RTO, and prove rollback.

01

Define start/stop timestamps and evidence needed to measure RPO and RTO rather than estimate them.

02

Run an isolated restore without deleting the original recovery evidence or blindly restoring global/security state.

03

Validate data completeness, mappings/settings, templates, aliases and representative application queries.

04

Execute an explicit application cutover and rollback using a stable alias or configuration boundary.

05

Produce a DR report that separates measured facts, assumptions, version/product constraints and remediation actions.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0, reviewed 11 September 2026. AtlasMart keeps the Chapter 01 TLS/auth conventions: Elasticsearch at https://localhost:9200 with CA verification and OpenSearch at https://localhost:9201 with the upstream demo certificate only in the disposable lab. Existing containers remain atlasmart-es and atlasmart-os. The snapshot lab uses one primary shard and zero replicas only because the local environment is single-node; that is not production guidance. A filesystem repository requires path.repo to be configured before node start and the repository path to be reachable by every master/data node that participates. If your earlier containers were not created with that setting, recreate disposable lab containers rather than editing production-like nodes in place.

Safety boundary

Run this drill only on the disposable atlasmart-dr-* fixture or an isolated recovery cluster. Do not delete production indices, mutate a real snapshot repository, restore security/global state, or repoint production clients merely to practice the runbook.

Execution note

The generation environment does not run the AtlasMart Docker clusters, so commands are reproducible lab instructions and response fragments are labeled expected shapes/invariants rather than fabricated measurements. Record your own snapshot duration, bytes transferred, repository latency, restore throughput, p95/p99 application latency, RPO and RTO.

1. Define the clock before the incident

RPO measures recoverable data age at the failure boundary. RTO measures how long service recovery takes from declared incident until the agreed service acceptance gate passes. Decide what “service recovered” means: cluster green? index searchable? API smoke tests passing? full p95/p99 SLO restored? Different definitions produce different RTOs.

Timestamp Meaning
T_snap_start Snapshot operation starts.
T_snap_end Snapshot successfully completes.
T_last_recoverable_event Newest business event proven present in that snapshot.
T_incident Failure/data-loss incident declared.
T_restore_start Restore request accepted.
T_data_ready Restored index shards available for validation.
T_app_ready Smoke tests/SLO gate passes after cutover.
Measured objectives
observed_rpo = T_incident - T_last_recoverable_event
observed_data_restore_time = T_data_ready - T_restore_start
observed_rto = T_app_ready - T_incident

# Do not substitute:
# schedule_interval for observed_rpo
# restore_API_ack for T_app_ready

2. Prepare recovery point and deliberate post-snapshot evidence

Use snapshot 002 from Lesson 3 or create a fresh snapshot. Then write one document after snapshot completion. That post-snapshot document acts as a visible RPO marker: an ordinary restore of the snapshot should not contain it.

Create the RPO marker after the recovery point
PUT atlasmart-dr-products-v1/_doc/P-AFTER-SNAPSHOT?refresh=wait_for
{
  "sku":"AM-DR-AFTER",
  "name":"Post Snapshot Marker",
  "category":"dr-test",
  "price":1.00,
  "updated_at":"2026-09-11T10:05:00Z"
}

GET atlasmart-dr-products-v1/_doc/P-AFTER-SNAPSHOT

Record the exact snapshot completion time and marker event time. In a real system, the newest recoverable business event should come from a canonical event/transaction sequence, not from wall-clock guessing alone.

3. Simulate source loss safely

The safest drill is to leave the source intact and pretend it is unavailable to the recovery operator. If you want deletion semantics, clone the fixture first and delete only the disposable clone. The goal is to practice restore, not to manufacture real data loss.

Create stable application aliases before cutover test
POST _aliases
{
  "actions": [
    {"add": {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-read"}},
    {"add": {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-write", "is_write_index":true}}
  ]
}

GET _alias/atlasmart-dr-products-read,atlasmart-dr-products-write

Clients in the drill should use the aliases, not the physical index name. This gives the recovery procedure an explicit cutover boundary and lets rollback remain an alias operation instead of a destructive restore-overwrite.

4. Restore the snapshot under an isolated name and time it

Timed restore request
# Record T_restore_start immediately before this request.
POST _snapshot/atlasmart-es-fs/atlasmart-dr-2026.09.11-002/_restore?wait_for_completion=true
{
  "indices":"atlasmart-dr-products-v1",
  "include_global_state":false,
  "include_aliases":false,
  "rename_pattern":"atlasmart-dr-products-v1",
  "rename_replacement":"atlasmart-dr-products-drill-002"
}

# Then inspect recovery, health and index state.
GET _cluster/health/atlasmart-dr-products-drill-002
GET atlasmart-dr-products-drill-002/_recovery?active_only=false
GET atlasmart-dr-products-drill-002/_count

Expected invariant: restore completes without mutating the source index or aliases. The exact response timing is environment-specific. wait_for_completion=true measures API completion for the restore operation, but your RTO continues until application validation and cutover gates pass.

5. Validate completeness and the intentional RPO gap

Data and behavior checks
GET atlasmart-dr-products-drill-002/_doc/P-1001
GET atlasmart-dr-products-drill-002/_doc/P-1006

# The marker was written AFTER snapshot 002 and should be absent.
GET atlasmart-dr-products-drill-002/_doc/P-AFTER-SNAPSHOT

GET atlasmart-dr-products-drill-002/_mapping
GET atlasmart-dr-products-drill-002/_settings

GET atlasmart-dr-products-drill-002/_search
{
  "size": 0,
  "aggs": {
    "by_category": {"terms": {"field":"category","size":10}},
    "price_sum": {"sum": {"field":"price"}}
  }
}

The missing post-snapshot marker is not a restore bug; it is a visible representation of the recovery point. If the documented RPO is smaller than that gap, the design needs more frequent snapshots, external change replay, or another durability mechanism.

6. Validate dependencies outside the restored index

Dependency Drill action Why it matters
Index templates GET expected templates and compare version/hash Future rollover/new index must recreate the same schema contract.
Aliases Inspect source and restored alias state before cutover Snapshot restore should not unexpectedly hijack live traffic.
Ingest pipelines Verify names/versions referenced by writers Restored data may search fine while new writes fail or transform differently.
Lifecycle Confirm policy attachment/timing A recovered index must not be immediately deleted or moved unexpectedly.
Security Verify service principal can read restored alias after controlled grant Data recovery without application authorization is not service recovery.
TLS/keystore/config files Recover from separate secure configuration backup/runbook Snapshots do not replace node config/secret backup.

Elasticsearch feature states and OpenSearch Security-plugin configuration require product-specific privileged handling. The mandatory lab intentionally does not restore them. Instead, record whether the DR environment obtains security configuration from infrastructure-as-code, identity provider integration, separate secure backups, or a product-supported security-state recovery process.

7. Atomic application cutover and smoke tests

Cut the read alias to the restored index
POST _aliases
{
  "actions": [
    {"remove": {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-read"}},
    {"add":    {"index":"atlasmart-dr-products-drill-002", "alias":"atlasmart-dr-products-read"}}
  ]
}

GET _alias/atlasmart-dr-products-read
GET atlasmart-dr-products-read/_search
{
  "query": {"match": {"name":"wireless headphones"}},
  "size": 3
}

If the recovered index is intended to resume writes, do not simply repoint the write alias until you have decided how to reconcile post-snapshot writes, mapping/version constraints and write ownership. In many DR designs the restored cluster becomes authoritative only after an explicit incident command and application-level replay/catch-up step.

8. Rollback evidence

Rollback the read alias
POST _aliases
{
  "actions": [
    {"remove": {"index":"atlasmart-dr-products-drill-002", "alias":"atlasmart-dr-products-read"}},
    {"add":    {"index":"atlasmart-dr-products-v1", "alias":"atlasmart-dr-products-read"}}
  ]
}

Rollback is only possible because the drill preserved the original source. In a real regional disaster the old cluster may be gone; “rollback” could instead mean abandon the failed target and retry restore into another target. Write that distinction into the runbook.

9. Disaster-recovery report

Minimum evidence record
drill_id: atlasmart-ch18-2026-09-11
source_product_version: ...
target_product_version: ...
repository: ...
snapshot: ...
snapshot_start: ...
snapshot_end: ...
last_recoverable_business_event: ...
incident_time: ...
restore_start: ...
data_ready: ...
application_ready: ...
observed_rpo: ...
observed_rto: ...
source_count: ...
restored_count: ...
mapping_diff: PASS|FAIL
query_fixture: PASS|FAIL
template_check: PASS|FAIL
alias_cutover: PASS|FAIL
security_assumption: documented
post_snapshot_marker_absent: EXPECTED|UNEXPECTED
rollback_test: PASS|FAIL
open_actions:
  - ...
If this fails… Likely design action
RPO too large Increase snapshot cadence, reduce failures/replication lag, add event-log replay or change durability architecture.
Restore too slow Increase restore bandwidth/target capacity, reduce shard/pathology, pre-stage DR infrastructure, evaluate product-specific remote/searchable options.
Schema/query mismatch Version templates/analyzers and include application-level regression tests in backup validation.
Security cutover fails Treat identity/roles/certs/API keys as a first-class DR dependency with separate tested recovery.
Repository unavailable Add independent repository copies/credentials/network paths and regularly verify from DR environment.

10. Chapter completion criteria

Chapter 18 is complete only when AtlasMart can point to a named snapshot, prove its repository is reachable, restore it into an isolated target, explain exactly what global/security state was excluded, demonstrate the expected RPO gap, pass data/schema/application checks, execute cutover and rollback, and report measured RPO/RTO. This hands off naturally to Chapter 19: the recovered service must now be secured with correct TLS, identities, roles, API keys and audit controls.

Check your understanding

  1. What starts and stops the RTO clock?
  2. Why write a marker after the snapshot?
  3. Why keep read/write aliases during the drill?
  4. Why is security state validated separately?
  5. What is the strongest evidence that backup works?
Review the answers

1. Use a documented incident start and stop only when the agreed service acceptance gate passes; restore API acknowledgement alone is not service recovery.

2. It makes the recovery boundary observable: its expected absence after restore demonstrates what data falls outside that recovery point.

3. They provide an explicit, atomic application cutover boundary and make rollback testable without destructive renaming.

4. Restoring old authority can reintroduce revoked privilege or lock operators out; Elastic/OpenSearch security recovery mechanisms also differ.

5. A repeated, compatible restore drill with data/schema/dependency/application validation and measured RPO/RTO.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.