Chapter 16 · Backup, Restore, Transaction Logs, Disaster Recovery, and Upgrade Safety
Backup Types, Online Backup Concepts, Transaction Logs, Retention, and Recovery Objectives
Design an AtlasMart recovery strategy from RPO/RTO first principles, distinguish offline Community dumps from Enterprise online backup chains, and prove what transaction-log retention can and cannot recover.
Learning outcomes
Distinguish replication/high availability from an independently recoverable backup.
Define RPO and RTO and derive backup cadence from business loss/downtime tolerance.
Explain Community offline dumps versus Enterprise online full/differential backup chains.
Connect transaction-log retention and checkpoints to differential/PITR recoverability without treating logs as a backup by themselves.
Create, hash, inspect and consistency-check a reproducible AtlasMart dump artifact.
Reproducible AtlasMart setup
The continuity lab remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language behavior matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic credential
neo4j / atlasmart-course-2026, Java 21 or 25,
Python driver 6.3, and named volume
atlasmart-neo4j-data. Current 5.26 LTS is
5.26.30. APOC Core 2026.07.1 and GDS 2026.07.0 are
compatibility references only; neither is mandatory for this
chapter's Community drill.
Community provides offline
neo4j-admin database dump, load, and
consistency checking. Enterprise adds online full/differential
backup chains, backup metadata/aggregation, TLS-capable backup
service, and transaction-log replay during restore. The
self-managed neo4j-admin database backup command
is not an Aura workflow; Aura uses managed backup/recovery
features whose retention and restore controls depend on the
service tier. Do not claim that a replica, cluster member,
filesystem snapshot, or Aura copy is automatically equivalent
to a validated backup.
Every destructive command in this chapter targets disposable course containers/volumes or a new isolated restore volume. Never overwrite the only known-good store during a drill. Capture a hash and inventory first, restore to a separate target, validate it, and only then decide whether promotion is safe.
The drill adds one tiny course-owned marker so recovery correctness has a deterministic invariant even if your earlier AtlasMart graph has grown. The marker is not a substitute for business reconciliation; it is a canary.
CYPHER 25
CREATE CONSTRAINT recovery_marker_id IF NOT EXISTS
FOR (m:RecoveryMarker) REQUIRE m.markerId IS UNIQUE;
MERGE (c:Customer {customerId:'C-1001'})
ON CREATE SET c.name = 'Mina Rahimi'
MERGE (m:RecoveryMarker {markerId:'DR-CH16-001'})
SET m.createdAt = datetime('2026-09-09T17:00:00Z'),
m.expectedState = 'before-dump',
m.exercise = 'chapter-16'
MERGE (c)-[:HAS_RECOVERY_MARKER]->(m);
MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN m.markerId, m.expectedState;
| Assumption | Pinned value / rule |
|---|---|
| server | Neo4j Community 2026.07.1; 5.26.30 is the current LTS comparison line |
| Java | 21 or 25 for current 2026.07 server |
| database / volume | neo4j / atlasmart-neo4j-data |
| transport | loopback Bolt without TLS only for this disposable local lab |
| backup destination |
new host directory ./neo4j-dr/backups;
production must be off-host and access-controlled
|
| plugins | none required; if installed, inventory exact APOC/GDS versions before upgrade/restore |
| measurement | record real start/end timestamps and artifact hashes; no invented RPO/RTO numbers |
1. Start with the recovery objective, not the command
AtlasMart can survive a server restart and still fail a disaster: an operator can delete data correctly on every cluster member, ransomware can encrypt every writable replica, or an application bug can commit bad state. Replication keeps copies available; a backup preserves an independent recoverable state. The design question is therefore: how much committed work may the business lose, and how long may recovery take?
| Term | Mechanism | AtlasMart question |
|---|---|---|
| RPO | maximum acceptable gap between failure and last recoverable state | Can orders committed in the last 15 minutes be reconstructed from another system? |
| RTO | maximum acceptable time to restore service | Does the runbook include download, load, consistency checks, DNS/driver cutover and application verification? |
| full backup | complete store state | How often can a full artifact be created without exceeding the recovery window/cost? |
| differential backup | Enterprise transaction-log artifact chained to a recoverable full | Is every required transaction interval contiguous and retained? |
| transaction log | durable record of committed changes used by recovery/backup machinery | Will pruning remove logs needed by the planned backup chain? |
2. Community offline dump vs Enterprise online backup
neo4j-admin database dump requires the Community
DBMS/database to be offline. It produces a single archive of
database contents. It is operationally simple and free, but
downtime contributes directly to RTO and the artifact is a
point-in-time full copy rather than a continuously replayable
backup chain.
Enterprise neo4j-admin database backup talks to a
configured online backup service. The first artifact is full;
subsequent runs can produce differential artifacts containing
transaction logs. A valid chain is the full plus contiguous
differentials. Since Neo4j 2026.02, the first differential may
overlap the full, but subsequent parent/child coverage still
matters.
| Capability | Community | Enterprise self-managed | Aura |
|---|---|---|---|
| offline dump/load | yes | yes | not the self-managed operational path |
| online backup command | no | yes |
neo4j-admin database backup not supported
|
| differential chain | no | yes | managed service behavior/tier-specific |
| restore-until tx/time | not from a dump chain | yes, from suitable backup chain | managed service controls/tier-specific |
| consistency check | yes on offline store/dump | yes on supported store/dump/recovered backup | provider-managed plus application reconciliation |
3. Transaction-log retention is a recoverability dependency
db.tx_log.rotation.retention_policy controls how
long logical transaction logs remain before safe pruning after
checkpoints. The current default is 2 days 2G. The
setting is dynamic. Manual deletion of transaction-log files is
unsupported because pruning must respect checkpoint/recovery
state.
SHOW SETTINGS YIELD name, value, dynamic
WHERE name = 'db.tx_log.rotation.retention_policy'
RETURN name, value, dynamic;
Setting keep_none may save disk but can destroy
the transaction history an Enterprise differential-backup
schedule needs. Conversely, keep_all can grow
storage without bound. Choose retention from measured backup
cadence and disk headroom, not folklore.
4. Produce the Community artifact without touching the source store
Stop the course container first. Then mount its data volume read/write into the official admin image and write the dump to a host backup directory. Exact Docker filesystem ownership varies by host; fix permissions deliberately rather than using world-writable backup directories.
mkdir -p ./neo4j-dr/backups
docker stop atlasmart-neo4j
docker run --rm \
-v atlasmart-neo4j-data:/data \
-v "$PWD/neo4j-dr/backups:/backups" \
neo4j/neo4j-admin:2026.07.1 \
neo4j-admin database dump neo4j --to-path=/backups
ls -lh ./neo4j-dr/backups/neo4j.dump
sha256sum ./neo4j-dr/backups/neo4j.dump
A neo4j.dump file exists and has a stable SHA-256
for those exact bytes. Its size/hash are machine/data
dependent, so this lesson deliberately does not invent them.
5. Inspect and consistency-check before calling it “good”
An artifact that merely exists can still be unusable or incomplete. Record its hash and run a consistency check while it is still isolated.
docker run --rm \
-v "$PWD/neo4j-dr/backups:/backups" \
neo4j/neo4j-admin:2026.07.1 \
neo4j-admin database check --from-path=/backups/neo4j.dump neo4j
A clean consistency check increases confidence in the archive’s internal store consistency. It does not prove that the backup contains the intended business point, that every dependent database/secret/plugin exists, or that the application can start from it.
6. The misleading shortcut: “the cluster is my backup”
A replicated topology can protect availability from a member failure, but it normally replicates legitimate destructive transactions too. A copied live data directory can also be crash-inconsistent unless the mechanism is designed for it. The repair is independent, access-controlled artifacts plus routine restore drills.
| Failure | Replica helps? | Independent validated backup helps? |
|---|---|---|
| single disk/server loss | often | yes |
| accidental DELETE committed everywhere | no | yes, if recovery point predates it |
| ransomware/admin compromise | maybe not | yes only if off-host/immutable controls resist the same compromise |
| bad upgrade/store migration | not necessarily | yes if pre-upgrade artifact remains compatible |
7. Recovery-objective worksheet
| Evidence to capture | Why |
|---|---|
| last successful backup time + highest tx/time covered | establishes recoverable point |
| full/differential chain IDs | proves chain continuity |
| artifact hash/size/location | detects accidental substitution/corruption |
| restore start → database ready → application ready timestamps | separates DB restore from end-to-end RTO |
| business reconciliation queries | proves domain state rather than filesystem success |
Check your understanding
- Is a three-member cluster automatically a backup?
- What drives differential cadence?
-
Can Community perform online
neo4j-admin database backup? - Why should transaction logs not be deleted manually?
- Does a zero exit code prove recoverability?
Review the answers
1. No. Replication/HA and independently recoverable backup solve different failure classes.
2. The required RPO, constrained by transaction-log availability, operational cost and backup duration.
3. No; the free reproducible path is offline dump/load.
4. Safe pruning depends on checkpoints/recovery state; manual deletion is unsupported.
5. No. You still need artifact evidence, restore validation, consistency/business checks and application recovery tests.
Production judgment
| Review area | Decision evidence |
|---|---|
| graph/workload fit | recovery scope includes every database and external dependency needed to make AtlasMart useful, not only graph files |
| correctness | restored node/relationship/business invariants and application smoke tests; a successful command exit is insufficient |
| RPO/RTO | measured from real cadence, last recoverable point, restore duration and operator/application recovery steps |
| transactions/concurrency | backup method preserves a consistent recoverable state; log retention covers required differential/PITR window |
| memory/CPU/disk/network | backup, restore and consistency-check resource use measured separately from normal workload |
| indexes/constraints | index/constraint state reconciled after restore; rebuild/population time included in RTO if applicable |
| driver/service | pool/retry/bookmark behavior revalidated after endpoint/version changes; ambiguous writes reconciled |
| security | backup files encrypted/protected by platform controls, least-privilege access, secret/certificate handling and deletion policy |
| observability | backup age, artifact chain, failures, restore drills, disk pressure and operator actions are monitored/audited |
| version/edition | server, store format, Java, Cypher, driver, APOC/GDS and Aura/self-managed boundaries captured before change |
| rollback | pre-upgrade artifact remains immutable and compatible with the rollback server; rollback trigger and owner are explicit |
| cost/governance | retention, egress/object-lock/license cost balanced against business RPO/RTO and compliance requirements |
Summary and next step
Backup design begins with RPO/RTO and failure classes, not with a command. Lesson 2 restores the Community dump into an isolated volume, validates the graph and explains exactly where Enterprise point-in-time recovery changes the mechanism.
Authoritative references
- Current Neo4j versions — Current server and 5.26 LTS release snapshot.
- Backup and restore — Edition-aware entry point for dump/load, online backup, restore and planning.
- Backup and restore planning — RPO/RTO, backup mode, storage location, cadence and retention planning.
- Back up an offline database — neo4j-admin database dump semantics and Community offline boundary.
- Restore a database dump — neo4j-admin database load semantics, overwrite rules and edition differences.
- Back up an online database — Enterprise full/differential backup artifacts and chain semantics.
- Restore a database backup — Enterprise recovery of backup chains and restore-until predicates.
- Check database consistency — neo4j-admin database check for stores, dumps and recovered full backups.
- Transaction logging — Transaction-log retention, checkpointing and pruning behavior.
- Store formats — Current aligned/block formats, limits and legacy-format deprecation.
- Migrate a database — neo4j-admin database migrate and store-format migration boundaries.
- System requirements — Supported Java/runtime and platform requirements for current Neo4j.
- APOC installation — APOC/server release compatibility and restart/deployment coupling.
- GDS compatibility — Graph Data Science and Neo4j version compatibility matrix.
- Python driver installation — Current 6.x driver/server compatibility baseline.