Chapter 16 · Backup, Restore, Transaction Logs, Disaster Recovery, and Upgrade Safety
Run a Disaster-Recovery Drill and Measure RPO, RTO, Data Integrity, Application Recovery, and Operator Steps
Run an AtlasMart disaster-recovery exercise end to end, measure RPO/RTO without fabricated numbers, validate data and application behavior, and turn operator actions into an auditable runbook.
Learning outcomes
Execute an end-to-end Community recovery drill without overwriting the source.
Measure observed RPO and end-to-end RTO from timestamps and a known recovery marker.
Reconcile database structure, graph invariants and application behavior after recovery.
Record operator steps, dependencies and failure points as a reusable runbook.
Use drill evidence to choose concrete improvements in cadence, retention, automation and upgrade rollback.
Reproducible AtlasMart setup
The continuity lab remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language behavior matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic credential
neo4j / atlasmart-course-2026, Java 21 or 25,
Python driver 6.3, and named volume
atlasmart-neo4j-data. Current 5.26 LTS is
5.26.30. APOC Core 2026.07.1 and GDS 2026.07.0 are
compatibility references only; neither is mandatory for this
chapter's Community drill.
Community provides offline
neo4j-admin database dump, load, and
consistency checking. Enterprise adds online full/differential
backup chains, backup metadata/aggregation, TLS-capable backup
service, and transaction-log replay during restore. The
self-managed neo4j-admin database backup command
is not an Aura workflow; Aura uses managed backup/recovery
features whose retention and restore controls depend on the
service tier. Do not claim that a replica, cluster member,
filesystem snapshot, or Aura copy is automatically equivalent
to a validated backup.
Every destructive command in this chapter targets disposable course containers/volumes or a new isolated restore volume. Never overwrite the only known-good store during a drill. Capture a hash and inventory first, restore to a separate target, validate it, and only then decide whether promotion is safe.
The drill adds one tiny course-owned marker so recovery correctness has a deterministic invariant even if your earlier AtlasMart graph has grown. The marker is not a substitute for business reconciliation; it is a canary.
CYPHER 25
CREATE CONSTRAINT recovery_marker_id IF NOT EXISTS
FOR (m:RecoveryMarker) REQUIRE m.markerId IS UNIQUE;
MERGE (c:Customer {customerId:'C-1001'})
ON CREATE SET c.name = 'Mina Rahimi'
MERGE (m:RecoveryMarker {markerId:'DR-CH16-001'})
SET m.createdAt = datetime('2026-09-09T17:00:00Z'),
m.expectedState = 'before-dump',
m.exercise = 'chapter-16'
MERGE (c)-[:HAS_RECOVERY_MARKER]->(m);
MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN m.markerId, m.expectedState;
| Assumption | Pinned value / rule |
|---|---|
| server | Neo4j Community 2026.07.1; 5.26.30 is the current LTS comparison line |
| Java | 21 or 25 for current 2026.07 server |
| database / volume | neo4j / atlasmart-neo4j-data |
| transport | loopback Bolt without TLS only for this disposable local lab |
| backup destination |
new host directory ./neo4j-dr/backups;
production must be off-host and access-controlled
|
| plugins | none required; if installed, inventory exact APOC/GDS versions before upgrade/restore |
| measurement | record real start/end timestamps and artifact hashes; no invented RPO/RTO numbers |
1. Disaster recovery is a timed system test
The drill scenario is intentionally narrow: assume the AtlasMart
primary data volume is unavailable at incident time. You may use
only the previously created neo4j.dump, pinned
container image, documented configuration and synthetic
credentials. The original volume remains untouched so the
exercise is reversible.
| Clock point | Record |
|---|---|
| T0 | incident/drill declaration |
| Tbackup | timestamp of last proven recoverable marker/artifact |
| Tdb | restored database answers verified query |
| Tapp | critical application smoke tests meet recovery acceptance criteria |
Observed RPO = T0 − Tbackup. Database recovery duration = Tdb − T0. End-to-end RTO = Tapp − T0. Use your real timestamps; no course-provided number is a benchmark.
2. Preflight: freeze evidence, do not mutate the source
set -euo pipefail
mkdir -p ./neo4j-dr/run-001
date -u +'%Y-%m-%dT%H:%M:%SZ' | tee ./neo4j-dr/run-001/T0.txt
sha256sum -c ./neo4j-dr/evidence/neo4j.dump.sha256
cp ./neo4j-dr/evidence/manifest.txt ./neo4j-dr/run-001/
docker image inspect neo4j:2026.07.1 --format '{{.Id}}' | tee ./neo4j-dr/run-001/image-id.txt
If the hash does not match, stop. A DR drill should not silently “use whatever file is there.”
3. Restore into the isolated target and time it
docker rm -f atlasmart-neo4j-restore 2>/dev/null || true
docker volume rm atlasmart-neo4j-restore-data 2>/dev/null || true
docker volume create atlasmart-neo4j-restore-data
date -u +'%Y-%m-%dT%H:%M:%SZ' | tee ./neo4j-dr/run-001/restore-start.txt
docker run --rm \
-v atlasmart-neo4j-restore-data:/data \
-v "$PWD/neo4j-dr/backups:/backups" \
neo4j/neo4j-admin:2026.07.1 \
neo4j-admin database load neo4j --from-path=/backups --overwrite-destination=true
docker run -d --name atlasmart-neo4j-restore \
-p 127.0.0.1:7475:7474 -p 127.0.0.1:7688:7687 \
-v atlasmart-neo4j-restore-data:/data \
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 \
neo4j:2026.07.1
4. Gate 1: database readiness and graph integrity
Do not declare success when the container status becomes “running.” The database must authenticate and return the expected recovery point.
until docker exec atlasmart-neo4j-restore cypher-shell \
-u neo4j -p atlasmart-course-2026 \
"RETURN 1 AS ready" >/dev/null 2>&1; do sleep 2; done
docker exec atlasmart-neo4j-restore cypher-shell \
-u neo4j -p atlasmart-course-2026 \
"CYPHER 25 MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'}) RETURN m.expectedState;"
date -u +'%Y-%m-%dT%H:%M:%SZ' | tee ./neo4j-dr/run-001/Tdb.txt
CYPHER 25
MATCH (n) RETURN count(n) AS nodes;
MATCH ()-[r]->() RETURN count(r) AS relationships;
MATCH (c:Customer {customerId:'C-1001'})-[:HAS_RECOVERY_MARKER]->(m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN c.customerId, m.expectedState;
SHOW CONSTRAINTS;
SHOW INDEXES;
5. Gate 2: consistency and application recovery
For a direct store consistency check, stop the isolated restore
DBMS, run database check, then restart it.
Application recovery then verifies the actual repository/service
layer from Chapter 13/14 against port 7688: authentication, a
parameterized read, a safe idempotent write, result consumption,
timeout/retry behavior and any external dependency needed by the
chosen SLO.
docker stop atlasmart-neo4j-restore
docker run --rm -v atlasmart-neo4j-restore-data:/data \
neo4j/neo4j-admin:2026.07.1 \
neo4j-admin database check neo4j
docker start atlasmart-neo4j-restore
Record Tapp only after the chosen critical
application checks pass. If the graph is healthy but
DNS/secrets/certificates/driver configuration is wrong, the
business RTO has not ended.
6. Failure injection: remove a dependency, not data
Safe drills should expose runbook weaknesses without corrupting the source. One reversible test is to start the restored service with a wrong synthetic password, confirm authentication failure, then correct the secret reference. Another is to rename the dump file before preflight and prove the runbook stops on missing/hash evidence.
| Injected failure | Expected detection | Safe repair |
|---|---|---|
| wrong restored-service password | driver auth error before business query | rotate/reload correct synthetic secret |
| missing dump filename | preflight/hash step fails | select exact manifest artifact; do not guess |
| wrong server/plugin image | compatibility/inventory gate fails | use pinned tested image/plugin set |
| insufficient restore disk | load/check fails and disk metrics alert | provision headroom; do not delete source/only backup |
7. Measure and interpret—do not celebrate one stopwatch
| Metric | Record alongside it |
|---|---|
| RPO | artifact/marker time, incident time, whether external event replay can close the gap |
| database restore duration | dump size, store size, disk/CPU, image cache state, consistency-check duration |
| application RTO | DNS/endpoint/secret/cert/driver changes, smoke-test duration, operator waits |
| operator error count | step, reason, workaround, automation opportunity |
| data reconciliation | source recovery-point counts/invariants and any known acceptable difference |
8. Runbook result and improvement backlog
A useful drill ends with decisions. Examples: increase backup frequency to meet RPO; move artifacts to a separate account/site; shorten restore time by adjusting full/differential cadence in Enterprise; automate hash/metadata checks; pre-stage images; reduce manual secret/DNS steps; or add application-level reconciliation for external payments/orders.
| Finding | Example corrective action |
|---|---|
| RPO exceeds objective | shorter backup cadence or proven event replay |
| RTO dominated by artifact transfer | closer protected replica of backup artifacts / larger bandwidth / pre-staging |
| restore passes but app fails | include secrets/cert/config/dependency recovery in runbook |
| operators choose wrong artifact | machine-readable manifest/catalog + policy selection |
| upgrade rollback untested | schedule pre-upgrade restore rehearsal on target/prior stacks |
Check your understanding
- What marks the end of database recovery?
- What marks the end of business RTO?
- Why keep source and restore volumes separate?
- Should a drill invent a target p99/RTO result?
- What is the most valuable output after the drill?
Review the answers
1. A verified database query/invariant, not merely process startup.
2. Critical application behavior meets the defined recovery acceptance/SLO.
3. To make the drill reversible and prevent accidental destruction of the source.
4. No. Measure real environment results and record conditions.
5. Validated recovery evidence plus a prioritized runbook/architecture improvement backlog.
9. Cleanup/reset
docker rm -f atlasmart-neo4j-restore 2>/dev/null || true
docker volume rm atlasmart-neo4j-restore-data 2>/dev/null || true
# Keep backup/evidence if you want to continue the recovery exercises.
docker start atlasmart-neo4j 2>/dev/null || true
Production judgment
| Review area | Decision evidence |
|---|---|
| graph/workload fit | recovery scope includes every database and external dependency needed to make AtlasMart useful, not only graph files |
| correctness | restored node/relationship/business invariants and application smoke tests; a successful command exit is insufficient |
| RPO/RTO | measured from real cadence, last recoverable point, restore duration and operator/application recovery steps |
| transactions/concurrency | backup method preserves a consistent recoverable state; log retention covers required differential/PITR window |
| memory/CPU/disk/network | backup, restore and consistency-check resource use measured separately from normal workload |
| indexes/constraints | index/constraint state reconciled after restore; rebuild/population time included in RTO if applicable |
| driver/service | pool/retry/bookmark behavior revalidated after endpoint/version changes; ambiguous writes reconciled |
| security | backup files encrypted/protected by platform controls, least-privilege access, secret/certificate handling and deletion policy |
| observability | backup age, artifact chain, failures, restore drills, disk pressure and operator actions are monitored/audited |
| version/edition | server, store format, Java, Cypher, driver, APOC/GDS and Aura/self-managed boundaries captured before change |
| rollback | pre-upgrade artifact remains immutable and compatible with the rollback server; rollback trigger and owner are explicit |
| cost/governance | retention, egress/object-lock/license cost balanced against business RPO/RTO and compliance requirements |
Summary and next step
Chapter 16 has turned backup from a file-creation task into measured recoverability and upgrade rollback evidence. Chapter 17 moves from recovery of a failed deployment to keeping service available during member/network failures through Neo4j clustering, consensus, routing and failure-domain design.
Authoritative references
- Current Neo4j versions — Current server and 5.26 LTS release snapshot.
- Backup and restore — Edition-aware entry point for dump/load, online backup, restore and planning.
- Backup and restore planning — RPO/RTO, backup mode, storage location, cadence and retention planning.
- Back up an offline database — neo4j-admin database dump semantics and Community offline boundary.
- Restore a database dump — neo4j-admin database load semantics, overwrite rules and edition differences.
- Back up an online database — Enterprise full/differential backup artifacts and chain semantics.
- Restore a database backup — Enterprise recovery of backup chains and restore-until predicates.
- Check database consistency — neo4j-admin database check for stores, dumps and recovered full backups.
- Transaction logging — Transaction-log retention, checkpointing and pruning behavior.
- Store formats — Current aligned/block formats, limits and legacy-format deprecation.
- Migrate a database — neo4j-admin database migrate and store-format migration boundaries.
- System requirements — Supported Java/runtime and platform requirements for current Neo4j.
- APOC installation — APOC/server release compatibility and restart/deployment coupling.
- GDS compatibility — Graph Data Science and Neo4j version compatibility matrix.
- Python driver installation — Current 6.x driver/server compatibility baseline.