Chapter 16 · Backup, Restore, Transaction Logs, Disaster Recovery, and Upgrade Safety

Run a Disaster-Recovery Drill and Measure RPO, RTO, Data Integrity, Application Recovery, and Operator Steps

Run an AtlasMart disaster-recovery exercise end to end, measure RPO/RTO without fabricated numbers, validate data and application behavior, and turn operator actions into an auditable runbook.

Advanced240–320 minutesMeasured disaster-recovery drillNeo4j 2026.07.1 Community baseline · Cypher 25Enterprise online backup/PITR explicitly separatedJava 21/25 · Python driver 6.3Last reviewed: September 2026

Learning outcomes

01

Execute an end-to-end Community recovery drill without overwriting the source.

02

Measure observed RPO and end-to-end RTO from timestamps and a known recovery marker.

03

Reconcile database structure, graph invariants and application behavior after recovery.

04

Record operator steps, dependencies and failure points as a reusable runbook.

05

Use drill evidence to choose concrete improvements in cadence, retention, automation and upgrade rollback.

Reproducible AtlasMart setup

Chapter 16 baseline · reviewed 9 September 2026

The continuity lab remains Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 where language behavior matters, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic credential neo4j / atlasmart-course-2026, Java 21 or 25, Python driver 6.3, and named volume atlasmart-neo4j-data. Current 5.26 LTS is 5.26.30. APOC Core 2026.07.1 and GDS 2026.07.0 are compatibility references only; neither is mandatory for this chapter's Community drill.

Edition / platform boundary

Community provides offline neo4j-admin database dump, load, and consistency checking. Enterprise adds online full/differential backup chains, backup metadata/aggregation, TLS-capable backup service, and transaction-log replay during restore. The self-managed neo4j-admin database backup command is not an Aura workflow; Aura uses managed backup/recovery features whose retention and restore controls depend on the service tier. Do not claim that a replica, cluster member, filesystem snapshot, or Aura copy is automatically equivalent to a validated backup.

Safety boundary

Every destructive command in this chapter targets disposable course containers/volumes or a new isolated restore volume. Never overwrite the only known-good store during a drill. Capture a hash and inventory first, restore to a separate target, validate it, and only then decide whether promotion is safe.

The drill adds one tiny course-owned marker so recovery correctness has a deterministic invariant even if your earlier AtlasMart graph has grown. The marker is not a substitute for business reconciliation; it is a canary.

Cypher · deterministic recovery canary
CYPHER 25
CREATE CONSTRAINT recovery_marker_id IF NOT EXISTS
FOR (m:RecoveryMarker) REQUIRE m.markerId IS UNIQUE;
MERGE (c:Customer {customerId:'C-1001'})
ON CREATE SET c.name = 'Mina Rahimi'
MERGE (m:RecoveryMarker {markerId:'DR-CH16-001'})
SET m.createdAt = datetime('2026-09-09T17:00:00Z'),
    m.expectedState = 'before-dump',
    m.exercise = 'chapter-16'
MERGE (c)-[:HAS_RECOVERY_MARKER]->(m);
MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN m.markerId, m.expectedState;
Assumption Pinned value / rule
server Neo4j Community 2026.07.1; 5.26.30 is the current LTS comparison line
Java 21 or 25 for current 2026.07 server
database / volume neo4j / atlasmart-neo4j-data
transport loopback Bolt without TLS only for this disposable local lab
backup destination new host directory ./neo4j-dr/backups; production must be off-host and access-controlled
plugins none required; if installed, inventory exact APOC/GDS versions before upgrade/restore
measurement record real start/end timestamps and artifact hashes; no invented RPO/RTO numbers

1. Disaster recovery is a timed system test

The drill scenario is intentionally narrow: assume the AtlasMart primary data volume is unavailable at incident time. You may use only the previously created neo4j.dump, pinned container image, documented configuration and synthetic credentials. The original volume remains untouched so the exercise is reversible.

Clock point Record
T0 incident/drill declaration
Tbackup timestamp of last proven recoverable marker/artifact
Tdb restored database answers verified query
Tapp critical application smoke tests meet recovery acceptance criteria

Observed RPO = T0 − Tbackup. Database recovery duration = Tdb − T0. End-to-end RTO = Tapp − T0. Use your real timestamps; no course-provided number is a benchmark.

2. Preflight: freeze evidence, do not mutate the source

Bash · DR preflight
set -euo pipefail
mkdir -p ./neo4j-dr/run-001
 date -u +'%Y-%m-%dT%H:%M:%SZ' | tee ./neo4j-dr/run-001/T0.txt
sha256sum -c ./neo4j-dr/evidence/neo4j.dump.sha256
cp ./neo4j-dr/evidence/manifest.txt ./neo4j-dr/run-001/
docker image inspect neo4j:2026.07.1 --format '{{.Id}}' | tee ./neo4j-dr/run-001/image-id.txt
Evidence rule

If the hash does not match, stop. A DR drill should not silently “use whatever file is there.”

3. Restore into the isolated target and time it

Bash · timed restore
docker rm -f atlasmart-neo4j-restore 2>/dev/null || true
docker volume rm atlasmart-neo4j-restore-data 2>/dev/null || true
docker volume create atlasmart-neo4j-restore-data

date -u +'%Y-%m-%dT%H:%M:%SZ' | tee ./neo4j-dr/run-001/restore-start.txt
docker run --rm \
  -v atlasmart-neo4j-restore-data:/data \
  -v "$PWD/neo4j-dr/backups:/backups" \
  neo4j/neo4j-admin:2026.07.1 \
  neo4j-admin database load neo4j --from-path=/backups --overwrite-destination=true

docker run -d --name atlasmart-neo4j-restore \
  -p 127.0.0.1:7475:7474 -p 127.0.0.1:7688:7687 \
  -v atlasmart-neo4j-restore-data:/data \
  -e NEO4J_AUTH=neo4j/atlasmart-course-2026 \
  neo4j:2026.07.1

4. Gate 1: database readiness and graph integrity

Do not declare success when the container status becomes “running.” The database must authenticate and return the expected recovery point.

Bash · query canary and record Tdb
until docker exec atlasmart-neo4j-restore cypher-shell \
  -u neo4j -p atlasmart-course-2026 \
  "RETURN 1 AS ready" >/dev/null 2>&1; do sleep 2; done

docker exec atlasmart-neo4j-restore cypher-shell \
  -u neo4j -p atlasmart-course-2026 \
  "CYPHER 25 MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'}) RETURN m.expectedState;"
date -u +'%Y-%m-%dT%H:%M:%SZ' | tee ./neo4j-dr/run-001/Tdb.txt
Cypher · invariants to capture
CYPHER 25
MATCH (n) RETURN count(n) AS nodes;
MATCH ()-[r]->() RETURN count(r) AS relationships;
MATCH (c:Customer {customerId:'C-1001'})-[:HAS_RECOVERY_MARKER]->(m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN c.customerId, m.expectedState;
SHOW CONSTRAINTS;
SHOW INDEXES;

5. Gate 2: consistency and application recovery

For a direct store consistency check, stop the isolated restore DBMS, run database check, then restart it. Application recovery then verifies the actual repository/service layer from Chapter 13/14 against port 7688: authentication, a parameterized read, a safe idempotent write, result consumption, timeout/retry behavior and any external dependency needed by the chosen SLO.

Bash · store consistency gate
docker stop atlasmart-neo4j-restore
docker run --rm -v atlasmart-neo4j-restore-data:/data \
  neo4j/neo4j-admin:2026.07.1 \
  neo4j-admin database check neo4j
docker start atlasmart-neo4j-restore
Application Tapp

Record Tapp only after the chosen critical application checks pass. If the graph is healthy but DNS/secrets/certificates/driver configuration is wrong, the business RTO has not ended.

6. Failure injection: remove a dependency, not data

Safe drills should expose runbook weaknesses without corrupting the source. One reversible test is to start the restored service with a wrong synthetic password, confirm authentication failure, then correct the secret reference. Another is to rename the dump file before preflight and prove the runbook stops on missing/hash evidence.

Injected failure Expected detection Safe repair
wrong restored-service password driver auth error before business query rotate/reload correct synthetic secret
missing dump filename preflight/hash step fails select exact manifest artifact; do not guess
wrong server/plugin image compatibility/inventory gate fails use pinned tested image/plugin set
insufficient restore disk load/check fails and disk metrics alert provision headroom; do not delete source/only backup

7. Measure and interpret—do not celebrate one stopwatch

Metric Record alongside it
RPO artifact/marker time, incident time, whether external event replay can close the gap
database restore duration dump size, store size, disk/CPU, image cache state, consistency-check duration
application RTO DNS/endpoint/secret/cert/driver changes, smoke-test duration, operator waits
operator error count step, reason, workaround, automation opportunity
data reconciliation source recovery-point counts/invariants and any known acceptable difference

8. Runbook result and improvement backlog

A useful drill ends with decisions. Examples: increase backup frequency to meet RPO; move artifacts to a separate account/site; shorten restore time by adjusting full/differential cadence in Enterprise; automate hash/metadata checks; pre-stage images; reduce manual secret/DNS steps; or add application-level reconciliation for external payments/orders.

Finding Example corrective action
RPO exceeds objective shorter backup cadence or proven event replay
RTO dominated by artifact transfer closer protected replica of backup artifacts / larger bandwidth / pre-staging
restore passes but app fails include secrets/cert/config/dependency recovery in runbook
operators choose wrong artifact machine-readable manifest/catalog + policy selection
upgrade rollback untested schedule pre-upgrade restore rehearsal on target/prior stacks

Check your understanding

  1. What marks the end of database recovery?
  2. What marks the end of business RTO?
  3. Why keep source and restore volumes separate?
  4. Should a drill invent a target p99/RTO result?
  5. What is the most valuable output after the drill?
Review the answers

1. A verified database query/invariant, not merely process startup.

2. Critical application behavior meets the defined recovery acceptance/SLO.

3. To make the drill reversible and prevent accidental destruction of the source.

4. No. Measure real environment results and record conditions.

5. Validated recovery evidence plus a prioritized runbook/architecture improvement backlog.

9. Cleanup/reset

Bash · remove drill-only restored database
docker rm -f atlasmart-neo4j-restore 2>/dev/null || true
docker volume rm atlasmart-neo4j-restore-data 2>/dev/null || true
# Keep backup/evidence if you want to continue the recovery exercises.
docker start atlasmart-neo4j 2>/dev/null || true

Production judgment

Review area Decision evidence
graph/workload fit recovery scope includes every database and external dependency needed to make AtlasMart useful, not only graph files
correctness restored node/relationship/business invariants and application smoke tests; a successful command exit is insufficient
RPO/RTO measured from real cadence, last recoverable point, restore duration and operator/application recovery steps
transactions/concurrency backup method preserves a consistent recoverable state; log retention covers required differential/PITR window
memory/CPU/disk/network backup, restore and consistency-check resource use measured separately from normal workload
indexes/constraints index/constraint state reconciled after restore; rebuild/population time included in RTO if applicable
driver/service pool/retry/bookmark behavior revalidated after endpoint/version changes; ambiguous writes reconciled
security backup files encrypted/protected by platform controls, least-privilege access, secret/certificate handling and deletion policy
observability backup age, artifact chain, failures, restore drills, disk pressure and operator actions are monitored/audited
version/edition server, store format, Java, Cypher, driver, APOC/GDS and Aura/self-managed boundaries captured before change
rollback pre-upgrade artifact remains immutable and compatible with the rollback server; rollback trigger and owner are explicit
cost/governance retention, egress/object-lock/license cost balanced against business RPO/RTO and compliance requirements

Summary and next step

Chapter 16 has turned backup from a file-creation task into measured recoverability and upgrade rollback evidence. Chapter 17 moves from recovery of a failed deployment to keeping service available during member/network failures through Neo4j clustering, consensus, routing and failure-domain design.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.