Chapter 16 · Backup, Restore, Transaction Logs, Disaster Recovery, and Upgrade Safety

Backup Types, Online Backup Concepts, Transaction Logs, Retention, and Recovery Objectives

Design an AtlasMart recovery strategy from RPO/RTO first principles, distinguish offline Community dumps from Enterprise online backup chains, and prove what transaction-log retention can and cannot recover.

Advanced210–270 minutesBackup strategy and offline dump labNeo4j 2026.07.1 Community baseline · Cypher 25Enterprise online backup/PITR explicitly separatedJava 21/25 · Python driver 6.3Last reviewed: September 2026

Learning outcomes

01

Distinguish replication/high availability from an independently recoverable backup.

02

Define RPO and RTO and derive backup cadence from business loss/downtime tolerance.

03

Explain Community offline dumps versus Enterprise online full/differential backup chains.

04

Connect transaction-log retention and checkpoints to differential/PITR recoverability without treating logs as a backup by themselves.

05

Create, hash, inspect and consistency-check a reproducible AtlasMart dump artifact.

Reproducible AtlasMart setup

Chapter 16 baseline · reviewed 9 September 2026

The continuity lab remains Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 where language behavior matters, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic credential neo4j / atlasmart-course-2026, Java 21 or 25, Python driver 6.3, and named volume atlasmart-neo4j-data. Current 5.26 LTS is 5.26.30. APOC Core 2026.07.1 and GDS 2026.07.0 are compatibility references only; neither is mandatory for this chapter's Community drill.

Edition / platform boundary

Community provides offline neo4j-admin database dump, load, and consistency checking. Enterprise adds online full/differential backup chains, backup metadata/aggregation, TLS-capable backup service, and transaction-log replay during restore. The self-managed neo4j-admin database backup command is not an Aura workflow; Aura uses managed backup/recovery features whose retention and restore controls depend on the service tier. Do not claim that a replica, cluster member, filesystem snapshot, or Aura copy is automatically equivalent to a validated backup.

Safety boundary

Every destructive command in this chapter targets disposable course containers/volumes or a new isolated restore volume. Never overwrite the only known-good store during a drill. Capture a hash and inventory first, restore to a separate target, validate it, and only then decide whether promotion is safe.

The drill adds one tiny course-owned marker so recovery correctness has a deterministic invariant even if your earlier AtlasMart graph has grown. The marker is not a substitute for business reconciliation; it is a canary.

Cypher · deterministic recovery canary
CYPHER 25
CREATE CONSTRAINT recovery_marker_id IF NOT EXISTS
FOR (m:RecoveryMarker) REQUIRE m.markerId IS UNIQUE;
MERGE (c:Customer {customerId:'C-1001'})
ON CREATE SET c.name = 'Mina Rahimi'
MERGE (m:RecoveryMarker {markerId:'DR-CH16-001'})
SET m.createdAt = datetime('2026-09-09T17:00:00Z'),
    m.expectedState = 'before-dump',
    m.exercise = 'chapter-16'
MERGE (c)-[:HAS_RECOVERY_MARKER]->(m);
MATCH (m:RecoveryMarker {markerId:'DR-CH16-001'})
RETURN m.markerId, m.expectedState;
Assumption Pinned value / rule
server Neo4j Community 2026.07.1; 5.26.30 is the current LTS comparison line
Java 21 or 25 for current 2026.07 server
database / volume neo4j / atlasmart-neo4j-data
transport loopback Bolt without TLS only for this disposable local lab
backup destination new host directory ./neo4j-dr/backups; production must be off-host and access-controlled
plugins none required; if installed, inventory exact APOC/GDS versions before upgrade/restore
measurement record real start/end timestamps and artifact hashes; no invented RPO/RTO numbers

1. Start with the recovery objective, not the command

AtlasMart can survive a server restart and still fail a disaster: an operator can delete data correctly on every cluster member, ransomware can encrypt every writable replica, or an application bug can commit bad state. Replication keeps copies available; a backup preserves an independent recoverable state. The design question is therefore: how much committed work may the business lose, and how long may recovery take?

Term Mechanism AtlasMart question
RPO maximum acceptable gap between failure and last recoverable state Can orders committed in the last 15 minutes be reconstructed from another system?
RTO maximum acceptable time to restore service Does the runbook include download, load, consistency checks, DNS/driver cutover and application verification?
full backup complete store state How often can a full artifact be created without exceeding the recovery window/cost?
differential backup Enterprise transaction-log artifact chained to a recoverable full Is every required transaction interval contiguous and retained?
transaction log durable record of committed changes used by recovery/backup machinery Will pruning remove logs needed by the planned backup chain?

2. Community offline dump vs Enterprise online backup

neo4j-admin database dump requires the Community DBMS/database to be offline. It produces a single archive of database contents. It is operationally simple and free, but downtime contributes directly to RTO and the artifact is a point-in-time full copy rather than a continuously replayable backup chain.

Enterprise neo4j-admin database backup talks to a configured online backup service. The first artifact is full; subsequent runs can produce differential artifacts containing transaction logs. A valid chain is the full plus contiguous differentials. Since Neo4j 2026.02, the first differential may overlap the full, but subsequent parent/child coverage still matters.

Capability Community Enterprise self-managed Aura
offline dump/load yes yes not the self-managed operational path
online backup command no yes neo4j-admin database backup not supported
differential chain no yes managed service behavior/tier-specific
restore-until tx/time not from a dump chain yes, from suitable backup chain managed service controls/tier-specific
consistency check yes on offline store/dump yes on supported store/dump/recovered backup provider-managed plus application reconciliation

3. Transaction-log retention is a recoverability dependency

db.tx_log.rotation.retention_policy controls how long logical transaction logs remain before safe pruning after checkpoints. The current default is 2 days 2G. The setting is dynamic. Manual deletion of transaction-log files is unsupported because pruning must respect checkpoint/recovery state.

Cypher · inspect the current retention setting
SHOW SETTINGS YIELD name, value, dynamic
WHERE name = 'db.tx_log.rotation.retention_policy'
RETURN name, value, dynamic;
Boundary case

Setting keep_none may save disk but can destroy the transaction history an Enterprise differential-backup schedule needs. Conversely, keep_all can grow storage without bound. Choose retention from measured backup cadence and disk headroom, not folklore.

4. Produce the Community artifact without touching the source store

Stop the course container first. Then mount its data volume read/write into the official admin image and write the dump to a host backup directory. Exact Docker filesystem ownership varies by host; fix permissions deliberately rather than using world-writable backup directories.

Bash · offline dump to a separate host path
mkdir -p ./neo4j-dr/backups
docker stop atlasmart-neo4j
docker run --rm \
  -v atlasmart-neo4j-data:/data \
  -v "$PWD/neo4j-dr/backups:/backups" \
  neo4j/neo4j-admin:2026.07.1 \
  neo4j-admin database dump neo4j --to-path=/backups
ls -lh ./neo4j-dr/backups/neo4j.dump
sha256sum ./neo4j-dr/backups/neo4j.dump
Expected evidence

A neo4j.dump file exists and has a stable SHA-256 for those exact bytes. Its size/hash are machine/data dependent, so this lesson deliberately does not invent them.

5. Inspect and consistency-check before calling it “good”

An artifact that merely exists can still be unusable or incomplete. Record its hash and run a consistency check while it is still isolated.

Bash · consistency-check the dump
docker run --rm \
  -v "$PWD/neo4j-dr/backups:/backups" \
  neo4j/neo4j-admin:2026.07.1 \
  neo4j-admin database check --from-path=/backups/neo4j.dump neo4j
What this proves

A clean consistency check increases confidence in the archive’s internal store consistency. It does not prove that the backup contains the intended business point, that every dependent database/secret/plugin exists, or that the application can start from it.

6. The misleading shortcut: “the cluster is my backup”

A replicated topology can protect availability from a member failure, but it normally replicates legitimate destructive transactions too. A copied live data directory can also be crash-inconsistent unless the mechanism is designed for it. The repair is independent, access-controlled artifacts plus routine restore drills.

Failure Replica helps? Independent validated backup helps?
single disk/server loss often yes
accidental DELETE committed everywhere no yes, if recovery point predates it
ransomware/admin compromise maybe not yes only if off-host/immutable controls resist the same compromise
bad upgrade/store migration not necessarily yes if pre-upgrade artifact remains compatible

7. Recovery-objective worksheet

Evidence to capture Why
last successful backup time + highest tx/time covered establishes recoverable point
full/differential chain IDs proves chain continuity
artifact hash/size/location detects accidental substitution/corruption
restore start → database ready → application ready timestamps separates DB restore from end-to-end RTO
business reconciliation queries proves domain state rather than filesystem success

Check your understanding

  1. Is a three-member cluster automatically a backup?
  2. What drives differential cadence?
  3. Can Community perform online neo4j-admin database backup?
  4. Why should transaction logs not be deleted manually?
  5. Does a zero exit code prove recoverability?
Review the answers

1. No. Replication/HA and independently recoverable backup solve different failure classes.

2. The required RPO, constrained by transaction-log availability, operational cost and backup duration.

3. No; the free reproducible path is offline dump/load.

4. Safe pruning depends on checkpoints/recovery state; manual deletion is unsupported.

5. No. You still need artifact evidence, restore validation, consistency/business checks and application recovery tests.

Production judgment

Review area Decision evidence
graph/workload fit recovery scope includes every database and external dependency needed to make AtlasMart useful, not only graph files
correctness restored node/relationship/business invariants and application smoke tests; a successful command exit is insufficient
RPO/RTO measured from real cadence, last recoverable point, restore duration and operator/application recovery steps
transactions/concurrency backup method preserves a consistent recoverable state; log retention covers required differential/PITR window
memory/CPU/disk/network backup, restore and consistency-check resource use measured separately from normal workload
indexes/constraints index/constraint state reconciled after restore; rebuild/population time included in RTO if applicable
driver/service pool/retry/bookmark behavior revalidated after endpoint/version changes; ambiguous writes reconciled
security backup files encrypted/protected by platform controls, least-privilege access, secret/certificate handling and deletion policy
observability backup age, artifact chain, failures, restore drills, disk pressure and operator actions are monitored/audited
version/edition server, store format, Java, Cypher, driver, APOC/GDS and Aura/self-managed boundaries captured before change
rollback pre-upgrade artifact remains immutable and compatible with the rollback server; rollback trigger and owner are explicit
cost/governance retention, egress/object-lock/license cost balanced against business RPO/RTO and compliance requirements

Summary and next step

Backup design begins with RPO/RTO and failure classes, not with a command. Lesson 2 restores the Community dump into an isolated volume, validates the graph and explains exactly where Enterprise point-in-time recovery changes the mechanism.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.