Chapter 16 · Backup, Restore, Transaction Logs, Disaster Recovery, and Upgrade Safety
Restore Workflows, Validation, Point-in-Time/Transaction-Log Reasoning, and Isolated Recovery Testing
Restore AtlasMart into an isolated target, validate the archive before trusting it, reason correctly about transaction-log/PITR boundaries, and reconcile graph invariants after recovery.
Learning outcomes
Restore a dump into an isolated data volume instead of overwriting the source.
Validate archive metadata/consistency and reconcile graph invariants after load.
Explain point-in-time restore semantics from Enterprise transaction-log backup chains.
Separate database restore completion from application-ready RTO.
Recognize ambiguous or incomplete recovery evidence and repair the runbook.
Reproducible AtlasMart setup
The continuity lab remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language behavior matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic credential
neo4j / atlasmart-course-2026, Java 21 or 25,
Python driver 6.3, and named volume
atlasmart-neo4j-data. Current 5.26 LTS is
5.26.30. APOC Core 2026.07.1 and GDS 2026.07.0 are
compatibility references only; neither is mandatory for this
chapter's Community drill.
Community provides offline
neo4j-admin database dump, load, and
consistency checking. Enterprise adds online full/differential
backup chains, backup metadata/aggregation, TLS-capable backup
service, and transaction-log replay during restore. The
self-managed neo4j-admin database backup command
is not an Aura workflow; Aura uses managed backup/recovery
features whose retention and restore controls depend on the
service tier. Do not claim that a replica, cluster member,
filesystem snapshot, or Aura copy is automatically equivalent
to a validated backup.
Every destructive command in this chapter targets disposable course containers/volumes or a new isolated restore volume. Never overwrite the only known-good store during a drill. Capture a hash and inventory first, restore to a separate target, validate it, and only then decide whether promotion is safe.
The drill adds one tiny course-owned marker so recovery correctness has a deterministic invariant even if your earlier AtlasMart graph has grown. The marker is not a substitute for business reconciliation; it is a canary.
CYPHER 25
CREATE CONSTRAINT recovery_marker_id IF NOT EXISTS
FOR (m:RecoveryMarker) REQUIRE m.markerId IS UNIQUE;
MERGE (c:Customer {customerId:'C-1001'})
ON CREATE SET c.name = 'Mina Rahimi'
MERGE (m:RecoveryMarker {markerId:'DR-CH16-001'})
SET m.createdAt = datetime('2026-09-09T17:00:00Z'),
m.expectedState = 'before-dump',
m.exercise = 'chapter-16'
MERGE (c)-[:HAS_RECOVERY_MARKER]->(m);
MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN m.markerId, m.expectedState;
| Assumption | Pinned value / rule |
|---|---|
| server | Neo4j Community 2026.07.1; 5.26.30 is the current LTS comparison line |
| Java | 21 or 25 for current 2026.07 server |
| database / volume | neo4j / atlasmart-neo4j-data |
| transport | loopback Bolt without TLS only for this disposable local lab |
| backup destination |
new host directory ./neo4j-dr/backups;
production must be off-host and access-controlled
|
| plugins | none required; if installed, inventory exact APOC/GDS versions before upgrade/restore |
| measurement | record real start/end timestamps and artifact hashes; no invented RPO/RTO numbers |
1. Restore is the test that turns a file into evidence
AtlasMart’s backup is not “verified” until operators can build a
usable database from it without relying on the original store.
The safest drill creates a new Docker volume and new ports so no
command can silently overwrite
atlasmart-neo4j-data.
Running load --overwrite-destination directly
against the only production volume during a drill converts a
test into an outage and may destroy the last good state.
2. Inspect the archive before load
neo4j-admin database load --info reads archive
information without loading it. Keep the previously recorded
hash beside this metadata.
docker run --rm \
-v "$PWD/neo4j-dr/backups:/backups" \
neo4j/neo4j-admin:2026.07.1 \
neo4j-admin database load --info --from-path=/backups neo4j
sha256sum ./neo4j-dr/backups/neo4j.dump
3. Load into a new volume
Use a separate restore volume. The admin image writes the recovered store there; afterward a separate Neo4j container exposes it on alternate loopback ports. The local credential is freshly supplied for the disposable restored DBMS; do not assume a database dump is a complete identity/secrets backup.
docker rm -f atlasmart-neo4j-restore 2>/dev/null || true
docker volume rm atlasmart-neo4j-restore-data 2>/dev/null || true
docker volume create atlasmart-neo4j-restore-data
docker run --rm \
-v atlasmart-neo4j-restore-data:/data \
-v "$PWD/neo4j-dr/backups:/backups" \
neo4j/neo4j-admin:2026.07.1 \
neo4j-admin database load neo4j --from-path=/backups --overwrite-destination=true
docker run -d --name atlasmart-neo4j-restore \
-p 127.0.0.1:7475:7474 -p 127.0.0.1:7688:7687 \
-v atlasmart-neo4j-restore-data:/data \
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 \
neo4j:2026.07.1
4. Reconcile state, not just process health
Wait until Bolt is ready, then query the deterministic recovery canary and capture broad counts. In a real system, compare the restored graph to independently recorded source totals/orders/payments and external-system checkpoints.
docker exec atlasmart-neo4j-restore cypher-shell \
-a bolt://127.0.0.1:7687 -u neo4j -p atlasmart-course-2026 \
"CYPHER 25 MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'}) RETURN m.expectedState;"
CYPHER 25
MATCH (n) RETURN count(n) AS nodes;
MATCH ()-[r]->() RETURN count(r) AS relationships;
MATCH (c:Customer {customerId:'C-1001'})-[:HAS_RECOVERY_MARKER]->(m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN c.customerId, m.expectedState;
SHOW CONSTRAINTS;
The canary relationship exists and
expectedState = "before-dump". Exact total
node/relationship counts depend on which earlier chapter labs
you retained, so record the source counts immediately before
the dump and compare them to the restored counts rather than
copying a hard-coded number.
5. Consistency check and application smoke test are different layers
Stop the restore container before checking its store directly. A clean store check does not validate API permissions, driver routing, Cypher compatibility or business behavior.
docker stop atlasmart-neo4j-restore
docker run --rm \
-v atlasmart-neo4j-restore-data:/data \
neo4j/neo4j-admin:2026.07.1 \
neo4j-admin database check neo4j
| Layer | Pass condition |
|---|---|
| artifact | hash/metadata match intended backup |
| store | consistency checker exits cleanly |
| graph | counts/constraints/business invariants match recovery point |
| driver | authentication/TLS/routing/query smoke tests pass |
| application | critical AtlasMart read/write paths and dependent services behave correctly |
6. Point-in-time reasoning belongs to the Enterprise backup chain
A differential backup artifact contains transaction logs that
can be replayed onto the full store during restore. Current
Enterprise restore accepts --restore-until by
transaction ID or UTC timestamp. The boundary is precise:
transaction-ID recovery stops before the specified transaction.
# Licensed/self-managed Enterprise example only:
neo4j-admin database restore \
--from-path=/srv/neo4j-backups \
--restore-until="2026-09-09 16:45:00" \
atlasmart
Point-in-time recovery is only possible for transaction history actually present in a valid recoverable backup chain. A timestamp outside retained coverage cannot be recreated from wishful configuration.
7. Measure RPO and RTO from evidence
| Metric | How to calculate during the drill |
|---|---|
| observed RPO | failure/cutover decision time minus latest business transaction proven recoverable |
| database restore time | start of restore operation to DB ready for verified queries |
| application RTO | incident declaration/failover start to application SLO restored and critical smoke tests passing |
| operator steps | count/manual decisions from runbook; record errors/rework separately |
This course provides commands and formulas, not pretend seconds/minutes. Your disk, graph size, cache, CPU, artifact location and operator workflow determine actual results.
Check your understanding
- Why restore to a new volume?
- Can a consistency check prove an application is recovered?
- What does
--restore-untilrequire? - What should restored counts be compared with?
- When does RTO end?
Review the answers
1. It isolates the drill, protects the source, and proves the artifact is sufficient to build a separate store.
2. No. It checks supported store structures, not end-to-end application semantics.
3. A suitable Enterprise full/differential backup chain containing the required transaction history.
4. Counts/invariants captured from the intended source recovery point, not a generic course number.
5. When the required service/application SLO is restored, not merely when files finish loading.
8. Cleanup/reset
docker rm -f atlasmart-neo4j-restore 2>/dev/null || true
docker volume rm atlasmart-neo4j-restore-data
# Keep the source volume and backup artifact for the next lessons.
docker start atlasmart-neo4j
Production judgment
| Review area | Decision evidence |
|---|---|
| graph/workload fit | recovery scope includes every database and external dependency needed to make AtlasMart useful, not only graph files |
| correctness | restored node/relationship/business invariants and application smoke tests; a successful command exit is insufficient |
| RPO/RTO | measured from real cadence, last recoverable point, restore duration and operator/application recovery steps |
| transactions/concurrency | backup method preserves a consistent recoverable state; log retention covers required differential/PITR window |
| memory/CPU/disk/network | backup, restore and consistency-check resource use measured separately from normal workload |
| indexes/constraints | index/constraint state reconciled after restore; rebuild/population time included in RTO if applicable |
| driver/service | pool/retry/bookmark behavior revalidated after endpoint/version changes; ambiguous writes reconciled |
| security | backup files encrypted/protected by platform controls, least-privilege access, secret/certificate handling and deletion policy |
| observability | backup age, artifact chain, failures, restore drills, disk pressure and operator actions are monitored/audited |
| version/edition | server, store format, Java, Cypher, driver, APOC/GDS and Aura/self-managed boundaries captured before change |
| rollback | pre-upgrade artifact remains immutable and compatible with the rollback server; rollback trigger and owner are explicit |
| cost/governance | retention, egress/object-lock/license cost balanced against business RPO/RTO and compliance requirements |
Summary and next step
A restore drill proves much more than archive creation. Lesson 3 treats the artifact itself as high-value sensitive data and builds off-host, retention, immutability and chain-of-custody evidence around it.
Authoritative references
- Current Neo4j versions — Current server and 5.26 LTS release snapshot.
- Backup and restore — Edition-aware entry point for dump/load, online backup, restore and planning.
- Backup and restore planning — RPO/RTO, backup mode, storage location, cadence and retention planning.
- Back up an offline database — neo4j-admin database dump semantics and Community offline boundary.
- Restore a database dump — neo4j-admin database load semantics, overwrite rules and edition differences.
- Back up an online database — Enterprise full/differential backup artifacts and chain semantics.
- Restore a database backup — Enterprise recovery of backup chains and restore-until predicates.
- Check database consistency — neo4j-admin database check for stores, dumps and recovered full backups.
- Transaction logging — Transaction-log retention, checkpointing and pruning behavior.
- Store formats — Current aligned/block formats, limits and legacy-format deprecation.
- Migrate a database — neo4j-admin database migrate and store-format migration boundaries.
- System requirements — Supported Java/runtime and platform requirements for current Neo4j.
- APOC installation — APOC/server release compatibility and restart/deployment coupling.
- GDS compatibility — Graph Data Science and Neo4j version compatibility matrix.
- Python driver installation — Current 6.x driver/server compatibility baseline.