Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j

Load-Test, Profile, Secure, Back Up, Restore, Fail Over, Monitor, Upgrade, and Execute Incident Runbooks

Test the capstone as an operated system: load distributions, Community security boundaries, backup/restore proof, controlled outage recovery, monitoring evidence, upgrade gates and incident runbooks.

Advanced420–540 minutesLoad · security · recovery · operationsNeo4j 2026.07.1 · Community mandatoryOffline dump/load · isolated restoreCommunity outage ≠ Enterprise failoverEnterprise HA/RBAC/online backup optionalLast reviewed: September 2026

Learning outcomes

01

Run a reproducible mixed workload with percentile latency, throughput, failures and environment metadata instead of publishing one context-free QPS number.

02

Separate Community-reproducible auth/network/process recovery from Enterprise-only least-privilege RBAC and true cluster failover.

03

Create a Community offline dump, verify its artifact hash, load it into an isolated restore container, and reconcile the restored graph before calling the backup usable.

04

Collect query/transaction/log/OS evidence and turn observed symptoms into bounded incident actions with explicit rollback/reset steps.

05

Gate upgrades on compatibility, backups, plugin/driver/Cypher checks and rollback evidence rather than replacing binaries first and diagnosing later.

Execution and safety note

Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.

1. AtlasMart problem: “works on my laptop” is not operability

The capstone is now functional, but production readiness requires evidence under load and failure. Tail latency (p95/p99), saturation (where more concurrency stops increasing useful throughput), recovery evidence, and security tests describe different properties. One green smoke test cannot substitute for them.

Dimension Chapter 27 reproducible assumption
Neo4j 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled.
Cypher Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default.
Java Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK.
Driver Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions.
Database / auth Database neo4j; local user neo4j; synthetic password atlasmart-course-2026. Never reuse these lab credentials in production.
Network / TLS Loopback-only HTTP/Bolt for the disposable lab: 127.0.0.1:27474→7474 and 127.0.0.1:27687→7687. No TLS only because traffic stays on localhost; production/remote connections require a real TLS policy.
Plugins No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line.
Edition boundary Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries.
Evidence rule This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally.
Continuity This lesson may stop/restart only the disposable capstone container and may create a separate restore volume/container. It never manipulates host firewall/clock or unrelated Neo4j data.
Run a small mixed-load harness and capture distributions
# pip install neo4j==6.3.0
from concurrent.futures import ThreadPoolExecutor, as_completed
from time import perf_counter
from neo4j import GraphDatabase
import math, random

URI="bolt://localhost:27687"; AUTH=("neo4j","atlasmart-course-2026"); DB="neo4j"
Q={
 "history":("MATCH (c:Customer {customerId:$id})-[:PLACED]->(o:Order)-[:CONTAINS]->(p:Product) RETURN o.orderId,p.productId",{"id":"C-1001"}),
 "copurchase":("MATCH (:Product {productId:$id})<-[:CONTAINS]-(o:Order)-[:CONTAINS]->(p:Product) RETURN p.productId,count(DISTINCT o) AS n ORDER BY n DESC LIMIT 5",{"id":"P-2001"}),
 "write":("MERGE (m:CapstoneProbe {id:$id}) SET m.lastSeen=datetime(),m.capstone=true RETURN m.id",{"id":"load-probe"}),
}
def pct(xs,p):
    ys=sorted(xs); k=(len(ys)-1)*p/100; lo=math.floor(k); hi=math.ceil(k)
    return ys[lo] if lo==hi else ys[lo]*(hi-k)+ys[hi]*(k-lo)
def one(driver,name):
    cypher,params=Q[name]; t=perf_counter()
    try:
        records,summary,keys=driver.execute_query(cypher,parameters_=params,database_=DB)
        return (perf_counter()-t)*1000, True
    except Exception:
        return (perf_counter()-t)*1000, False

def run(concurrency,ops=200,seed=27):
    rnd=random.Random(seed); mix=["history"]*55+["copurchase"]*35+["write"]*10
    lat=[]; failures=0; t0=perf_counter()
    with GraphDatabase.driver(URI,auth=AUTH,max_connection_pool_size=max(4,concurrency)) as driver:
        driver.verify_connectivity()
        with ThreadPoolExecutor(max_workers=concurrency) as ex:
            fs=[ex.submit(one,driver,rnd.choice(mix)) for _ in range(ops)]
            for f in as_completed(fs):
                ms,ok=f.result(); lat.append(ms); failures += (not ok)
    sec=perf_counter()-t0
    print({"c":concurrency,"ops_s":round(ops/sec,2),"p50_ms":round(pct(lat,50),2),
           "p95_ms":round(pct(lat,95),2),"p99_ms":round(pct(lat,99),2),"failures":failures})

for c in (1,2,4,8): run(c)

# Capture your local outputs. No latency/QPS number in this lesson is asserted as universal.

2. Security test: prove only what Community actually enforces

Community can prove that authentication is enabled, the lab listener is loopback-bound, and secrets are not embedded in source artifacts. It cannot reproduce Enterprise fine-grained RBAC. Current Community users effectively have administrator-level authority, so creating a second user does not prove least privilege. An Enterprise deployment should add roles/privileges and explicit denied-path tests; Aura security controls must be verified against the selected tier.

Inspect current identity and network exposure
# Application-visible identity
$cypher = "SHOW CURRENT USER YIELD user,roles RETURN user,roles;"
docker exec atlasmart-neo4j-capstone cypher-shell -u neo4j -p atlasmart-course-2026 -d system $cypher

# Host-side exposure: ports should be published only on 127.0.0.1.
docker port atlasmart-neo4j-capstone

# Do not commit the lab password. In production use an approved secret store and TLS policy.

3. Backup proof: artifact → isolated restore → semantic reconciliation

A backup file existing is not recovery evidence. Community neo4j-admin database dump requires the database to be offline; the exercise therefore stops only the disposable source container, uses a separate admin container/volume path, hashes the artifact, then restores into an isolated volume and alternate ports. The restored graph must pass canary/count checks before the source is restarted for normal work.

Create a Community offline dump and hash it
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Backup = Join-Path $Root "backup"
New-Item -ItemType Directory -Force -Path $Backup | Out-Null

docker stop atlasmart-neo4j-capstone

docker run --rm `
  -v atlasmart-neo4j-capstone-data:/data `
  --mount type=bind,source="$Backup",target=/backups `
  neo4j/neo4j-admin:2026.07.1 `
  neo4j-admin database dump neo4j --to-path=/backups --overwrite-destination=true

Get-FileHash (Join-Path $Backup 'neo4j.dump') -Algorithm SHA256

# Hash proves byte identity of this artifact, not semantic correctness or confidentiality.
Load into an isolated restore volume and verify
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Backup = Join-Path $Root "backup"
docker rm -f atlasmart-neo4j-capstone-restore 2>$null
docker volume rm atlasmart-neo4j-capstone-restore-data 2>$null

docker run --rm `
  -v atlasmart-neo4j-capstone-restore-data:/data `
  --mount type=bind,source="$Backup",target=/backups `
  neo4j/neo4j-admin:2026.07.1 `
  neo4j-admin database load neo4j --from-path=/backups --overwrite-destination=true

docker run -d --name atlasmart-neo4j-capstone-restore `
  -p 127.0.0.1:28474:7474 -p 127.0.0.1:28687:7687 `
  -e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
  -v atlasmart-neo4j-capstone-restore-data:/data `
  neo4j:2026.07.1

# After the restore DBMS is ready, verify semantic state:
docker exec atlasmart-neo4j-capstone-restore cypher-shell `
  -u neo4j -p atlasmart-course-2026 -d neo4j `
  "MATCH (c:Customer) WHERE c.capstone=true RETURN count(c) AS customers;"

# Expected invariant: customers=5. Repeat the full reconciliation query from Lesson 2.

# Restore normal source state for the next bounded outage exercise.
docker start atlasmart-neo4j-capstone
# Wait for readiness, then verify source connectivity and the same reconciliation counts.

4. Controlled outage: driver recovery is not cluster failover

Stop the Community container while a client performs a harmless idempotent read. Observe connection errors/timeouts, restart the same container, then verify connectivity and data. This teaches application outage handling and recovery. It does not prove Enterprise writer election, quorum behavior, zero-downtime failover or ambiguous commit handling.

Inject and recover a bounded Community process outage
# Blast radius: only atlasmart-neo4j-capstone.
docker stop atlasmart-neo4j-capstone
# Run the read client now and record its exception/timeout category.
docker start atlasmart-neo4j-capstone
# Wait for readiness, then re-run driver.verify_connectivity() and the reconciliation query.

docker logs --since 5m atlasmart-neo4j-capstone

docker stats --no-stream atlasmart-neo4j-capstone

# Enterprise-only extension: repeat with a licensed multi-primary cluster, route via neo4j://,
# remove the current writer safely, and record writer change, quorum, bookmarks,
# retry/ambiguous outcomes and alerts. Do not label this Community restart as failover.

5. Observability and incident runbook

Symptom Evidence first Hypothesis Bounded action / verification
p99 rises while throughput plateaus Driver timings, pool wait, Docker CPU/I/O, PROFILE representative query. Saturation, fan-out, pool queue or store I/O. Change one factor; rerun same workload; compare p50/p95/p99 and failures.
Writes time out SHOW TRANSACTIONS, logs, hot-node pattern, client retry errors. Lock contention or overloaded server/client queue. Remove synthetic contention, reduce concurrency/batch, verify idempotent retries.
Process unavailable Client status codes/exceptions, container state/logs. DBMS stopped/crashed/network boundary. Restart isolated process; verify connectivity + semantic canary; escalate if repeated.
Restore succeeds but app fails Counts, constraints/index state, auth config, driver database/URI. Recovery incomplete at application boundary. Run full smoke test; restore/RTO ends only when application is useful.

6. Upgrade gate

Do not “test” upgrades by pointing a new image at the only copy of a store. Record current server/store/Cypher/driver/APOC/GDS versions, read the target upgrade notes, verify Java/platform compatibility, create and restore-test a pre-upgrade backup, run staging application tests, and define rollback. Neo4j 5.26 LTS is the checkpoint for upgrades from older 5.x-era versions; current 2025–2026 upgrades have their own compatibility guide. If a store migration is not backward-compatible, rollback means restoring a validated pre-upgrade artifact into a compatible prior stack—not merely reinstalling an older binary.

7. Deliberately wrong approach: “HA means no backup drill”

Replication can keep serving after a node failure, but it can also faithfully replicate deletion/corruption. Conversely, a backup may be healthy while the application still exceeds RTO because credentials, indexes, routing or smoke tests were omitted.

Repair: test availability/failure handling and backup/restore as independent controls. Measure database recovery and application recovery separately. Keep an evidence bundle with artifact hash, version manifest, restore steps, canary/counts, errors and operator timestamps.

8. Hands-on lab: production-readiness evidence bundle

Setup: start from the Lesson 2/3 source state, capture a version/index/count manifest, run the mixed-load harness, then perform the offline dump/isolated restore and the bounded source-process outage. Enterprise cluster/RBAC/online-backup extensions remain separate optional exercises.

Verification checklist

  • The load report records concurrency, operations, p50/p95/p99, throughput and failures together with the machine/container envelope.
  • Security evidence shows loopback-only published ports and does not claim Community fine-grained RBAC.
  • The dump has a recorded SHA-256 hash and the isolated restore reproduces all Lesson 2 reconciliation counts.
  • After the source process restart, the official driver reconnects and semantic canaries pass; the report labels this process recovery, not cluster failover.
  • The upgrade checklist includes target-version notes, Java/Cypher/driver/plugin compatibility, a restore-tested pre-upgrade backup and rollback criteria.

Cleanup/reset: remove the isolated restore container/volume after recovery evidence is captured. Keep or restart the source capstone container for Lesson 5. Never delete the only backup artifact until another validated copy exists.

Production judgment

Load, security, recovery and upgrade decisions interact. Larger result sets can dominate network/client time; more heap can starve page-cache/OS headroom; retries can amplify load unless idempotent; a cluster changes failure modes and licensing; observability can add cost; and recovery must include credentials/config/plugins as well as graph data. Capacity headroom is the distance from saturation under the actual mixed workload, not a fixed CPU percentage copied from another system.

Lesson 5 turns all this evidence into an architecture decision package with known limits, costs, exit criteria, and explicit cases where Neo4j should not be the chosen system.

Check your understanding

  1. Why are p95/p99 required alongside throughput?
  2. What does the Community outage lab prove?
  3. Why hash a dump if hashing is not enough?
  4. When is RTO complete?
  5. Why must rollback be designed before upgrade?
Review the answers

1. Throughput can stay flat while requests queue and tail latency/failures become unacceptable; percentiles expose service quality under load.

2. It proves client-visible outage/recovery behavior for one process; it does not prove Enterprise cluster failover/quorum.

3. The hash records artifact byte identity/integrity; isolated restore and semantic checks prove recoverability of the graph state.

4. When the required application service is useful again, not merely when database files load.

5. Store/config/plugin changes may not be backward-compatible; a validated pre-upgrade recovery path is needed before mutation.

Summary and next step

Load-Test, Profile, Secure, Back Up, Restore, Fail Over, Monitor, Upgrade, and Execute Incident Runbooks is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Present the Architecture with Evidence, Known Limits, Cost/Capacity Assumptions, and Clear Criteria for When Neo4j Is Not the Right Tool. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.