Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j
Load-Test, Profile, Secure, Back Up, Restore, Fail Over, Monitor, Upgrade, and Execute Incident Runbooks
Test the capstone as an operated system: load distributions, Community security boundaries, backup/restore proof, controlled outage recovery, monitoring evidence, upgrade gates and incident runbooks.
Learning outcomes
Run a reproducible mixed workload with percentile latency, throughput, failures and environment metadata instead of publishing one context-free QPS number.
Separate Community-reproducible auth/network/process recovery from Enterprise-only least-privilege RBAC and true cluster failover.
Create a Community offline dump, verify its artifact hash, load it into an isolated restore container, and reconcile the restored graph before calling the backup usable.
Collect query/transaction/log/OS evidence and turn observed symptoms into bounded incident actions with explicit rollback/reset steps.
Gate upgrades on compatibility, backups, plugin/driver/Cypher checks and rollback evidence rather than replacing binaries first and diagnosing later.
Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.
1. AtlasMart problem: “works on my laptop” is not operability
The capstone is now functional, but production readiness requires evidence under load and failure. Tail latency (p95/p99), saturation (where more concurrency stops increasing useful throughput), recovery evidence, and security tests describe different properties. One green smoke test cannot substitute for them.
| Dimension | Chapter 27 reproducible assumption |
|---|---|
| Neo4j | 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled. |
| Cypher | Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default. |
| Java | Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK. |
| Driver | Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions. |
| Database / auth |
Database neo4j; local user
neo4j; synthetic password
atlasmart-course-2026. Never reuse these lab
credentials in production.
|
| Network / TLS |
Loopback-only HTTP/Bolt for the disposable lab:
127.0.0.1:27474→7474 and
127.0.0.1:27687→7687. No TLS only because
traffic stays on localhost; production/remote connections
require a real TLS policy.
|
| Plugins | No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line. |
| Edition boundary | Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries. |
| Evidence rule | This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally. |
| Continuity | This lesson may stop/restart only the disposable capstone container and may create a separate restore volume/container. It never manipulates host firewall/clock or unrelated Neo4j data. |
# pip install neo4j==6.3.0
from concurrent.futures import ThreadPoolExecutor, as_completed
from time import perf_counter
from neo4j import GraphDatabase
import math, random
URI="bolt://localhost:27687"; AUTH=("neo4j","atlasmart-course-2026"); DB="neo4j"
Q={
"history":("MATCH (c:Customer {customerId:$id})-[:PLACED]->(o:Order)-[:CONTAINS]->(p:Product) RETURN o.orderId,p.productId",{"id":"C-1001"}),
"copurchase":("MATCH (:Product {productId:$id})<-[:CONTAINS]-(o:Order)-[:CONTAINS]->(p:Product) RETURN p.productId,count(DISTINCT o) AS n ORDER BY n DESC LIMIT 5",{"id":"P-2001"}),
"write":("MERGE (m:CapstoneProbe {id:$id}) SET m.lastSeen=datetime(),m.capstone=true RETURN m.id",{"id":"load-probe"}),
}
def pct(xs,p):
ys=sorted(xs); k=(len(ys)-1)*p/100; lo=math.floor(k); hi=math.ceil(k)
return ys[lo] if lo==hi else ys[lo]*(hi-k)+ys[hi]*(k-lo)
def one(driver,name):
cypher,params=Q[name]; t=perf_counter()
try:
records,summary,keys=driver.execute_query(cypher,parameters_=params,database_=DB)
return (perf_counter()-t)*1000, True
except Exception:
return (perf_counter()-t)*1000, False
def run(concurrency,ops=200,seed=27):
rnd=random.Random(seed); mix=["history"]*55+["copurchase"]*35+["write"]*10
lat=[]; failures=0; t0=perf_counter()
with GraphDatabase.driver(URI,auth=AUTH,max_connection_pool_size=max(4,concurrency)) as driver:
driver.verify_connectivity()
with ThreadPoolExecutor(max_workers=concurrency) as ex:
fs=[ex.submit(one,driver,rnd.choice(mix)) for _ in range(ops)]
for f in as_completed(fs):
ms,ok=f.result(); lat.append(ms); failures += (not ok)
sec=perf_counter()-t0
print({"c":concurrency,"ops_s":round(ops/sec,2),"p50_ms":round(pct(lat,50),2),
"p95_ms":round(pct(lat,95),2),"p99_ms":round(pct(lat,99),2),"failures":failures})
for c in (1,2,4,8): run(c)
# Capture your local outputs. No latency/QPS number in this lesson is asserted as universal.
2. Security test: prove only what Community actually enforces
Community can prove that authentication is enabled, the lab listener is loopback-bound, and secrets are not embedded in source artifacts. It cannot reproduce Enterprise fine-grained RBAC. Current Community users effectively have administrator-level authority, so creating a second user does not prove least privilege. An Enterprise deployment should add roles/privileges and explicit denied-path tests; Aura security controls must be verified against the selected tier.
# Application-visible identity
$cypher = "SHOW CURRENT USER YIELD user,roles RETURN user,roles;"
docker exec atlasmart-neo4j-capstone cypher-shell -u neo4j -p atlasmart-course-2026 -d system $cypher
# Host-side exposure: ports should be published only on 127.0.0.1.
docker port atlasmart-neo4j-capstone
# Do not commit the lab password. In production use an approved secret store and TLS policy.
3. Backup proof: artifact → isolated restore → semantic reconciliation
A backup file existing is not recovery evidence. Community
neo4j-admin database dump requires the database to
be offline; the exercise therefore stops only the disposable
source container, uses a separate admin container/volume path,
hashes the artifact, then restores into an isolated volume and
alternate ports. The restored graph must pass canary/count
checks before the source is restarted for normal work.
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Backup = Join-Path $Root "backup"
New-Item -ItemType Directory -Force -Path $Backup | Out-Null
docker stop atlasmart-neo4j-capstone
docker run --rm `
-v atlasmart-neo4j-capstone-data:/data `
--mount type=bind,source="$Backup",target=/backups `
neo4j/neo4j-admin:2026.07.1 `
neo4j-admin database dump neo4j --to-path=/backups --overwrite-destination=true
Get-FileHash (Join-Path $Backup 'neo4j.dump') -Algorithm SHA256
# Hash proves byte identity of this artifact, not semantic correctness or confidentiality.
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Backup = Join-Path $Root "backup"
docker rm -f atlasmart-neo4j-capstone-restore 2>$null
docker volume rm atlasmart-neo4j-capstone-restore-data 2>$null
docker run --rm `
-v atlasmart-neo4j-capstone-restore-data:/data `
--mount type=bind,source="$Backup",target=/backups `
neo4j/neo4j-admin:2026.07.1 `
neo4j-admin database load neo4j --from-path=/backups --overwrite-destination=true
docker run -d --name atlasmart-neo4j-capstone-restore `
-p 127.0.0.1:28474:7474 -p 127.0.0.1:28687:7687 `
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
-v atlasmart-neo4j-capstone-restore-data:/data `
neo4j:2026.07.1
# After the restore DBMS is ready, verify semantic state:
docker exec atlasmart-neo4j-capstone-restore cypher-shell `
-u neo4j -p atlasmart-course-2026 -d neo4j `
"MATCH (c:Customer) WHERE c.capstone=true RETURN count(c) AS customers;"
# Expected invariant: customers=5. Repeat the full reconciliation query from Lesson 2.
# Restore normal source state for the next bounded outage exercise.
docker start atlasmart-neo4j-capstone
# Wait for readiness, then verify source connectivity and the same reconciliation counts.
4. Controlled outage: driver recovery is not cluster failover
Stop the Community container while a client performs a harmless idempotent read. Observe connection errors/timeouts, restart the same container, then verify connectivity and data. This teaches application outage handling and recovery. It does not prove Enterprise writer election, quorum behavior, zero-downtime failover or ambiguous commit handling.
# Blast radius: only atlasmart-neo4j-capstone.
docker stop atlasmart-neo4j-capstone
# Run the read client now and record its exception/timeout category.
docker start atlasmart-neo4j-capstone
# Wait for readiness, then re-run driver.verify_connectivity() and the reconciliation query.
docker logs --since 5m atlasmart-neo4j-capstone
docker stats --no-stream atlasmart-neo4j-capstone
# Enterprise-only extension: repeat with a licensed multi-primary cluster, route via neo4j://,
# remove the current writer safely, and record writer change, quorum, bookmarks,
# retry/ambiguous outcomes and alerts. Do not label this Community restart as failover.
5. Observability and incident runbook
| Symptom | Evidence first | Hypothesis | Bounded action / verification |
|---|---|---|---|
| p99 rises while throughput plateaus | Driver timings, pool wait, Docker CPU/I/O, PROFILE representative query. | Saturation, fan-out, pool queue or store I/O. | Change one factor; rerun same workload; compare p50/p95/p99 and failures. |
| Writes time out | SHOW TRANSACTIONS, logs, hot-node pattern, client retry errors. | Lock contention or overloaded server/client queue. | Remove synthetic contention, reduce concurrency/batch, verify idempotent retries. |
| Process unavailable | Client status codes/exceptions, container state/logs. | DBMS stopped/crashed/network boundary. | Restart isolated process; verify connectivity + semantic canary; escalate if repeated. |
| Restore succeeds but app fails | Counts, constraints/index state, auth config, driver database/URI. | Recovery incomplete at application boundary. | Run full smoke test; restore/RTO ends only when application is useful. |
6. Upgrade gate
Do not “test” upgrades by pointing a new image at the only copy of a store. Record current server/store/Cypher/driver/APOC/GDS versions, read the target upgrade notes, verify Java/platform compatibility, create and restore-test a pre-upgrade backup, run staging application tests, and define rollback. Neo4j 5.26 LTS is the checkpoint for upgrades from older 5.x-era versions; current 2025–2026 upgrades have their own compatibility guide. If a store migration is not backward-compatible, rollback means restoring a validated pre-upgrade artifact into a compatible prior stack—not merely reinstalling an older binary.
7. Deliberately wrong approach: “HA means no backup drill”
Replication can keep serving after a node failure, but it can also faithfully replicate deletion/corruption. Conversely, a backup may be healthy while the application still exceeds RTO because credentials, indexes, routing or smoke tests were omitted.
Repair: test availability/failure handling and backup/restore as independent controls. Measure database recovery and application recovery separately. Keep an evidence bundle with artifact hash, version manifest, restore steps, canary/counts, errors and operator timestamps.
8. Hands-on lab: production-readiness evidence bundle
Setup: start from the Lesson 2/3 source state, capture a version/index/count manifest, run the mixed-load harness, then perform the offline dump/isolated restore and the bounded source-process outage. Enterprise cluster/RBAC/online-backup extensions remain separate optional exercises.
Verification checklist
- The load report records concurrency, operations, p50/p95/p99, throughput and failures together with the machine/container envelope.
- Security evidence shows loopback-only published ports and does not claim Community fine-grained RBAC.
- The dump has a recorded SHA-256 hash and the isolated restore reproduces all Lesson 2 reconciliation counts.
- After the source process restart, the official driver reconnects and semantic canaries pass; the report labels this process recovery, not cluster failover.
- The upgrade checklist includes target-version notes, Java/Cypher/driver/plugin compatibility, a restore-tested pre-upgrade backup and rollback criteria.
Cleanup/reset: remove the isolated restore container/volume after recovery evidence is captured. Keep or restart the source capstone container for Lesson 5. Never delete the only backup artifact until another validated copy exists.
Production judgment
Load, security, recovery and upgrade decisions interact. Larger result sets can dominate network/client time; more heap can starve page-cache/OS headroom; retries can amplify load unless idempotent; a cluster changes failure modes and licensing; observability can add cost; and recovery must include credentials/config/plugins as well as graph data. Capacity headroom is the distance from saturation under the actual mixed workload, not a fixed CPU percentage copied from another system.
Lesson 5 turns all this evidence into an architecture decision package with known limits, costs, exit criteria, and explicit cases where Neo4j should not be the chosen system.
Check your understanding
- Why are p95/p99 required alongside throughput?
- What does the Community outage lab prove?
- Why hash a dump if hashing is not enough?
- When is RTO complete?
- Why must rollback be designed before upgrade?
Review the answers
1. Throughput can stay flat while requests queue and tail latency/failures become unacceptable; percentiles expose service quality under load.
2. It proves client-visible outage/recovery behavior for one process; it does not prove Enterprise cluster failover/quorum.
3. The hash records artifact byte identity/integrity; isolated restore and semantic checks prove recoverability of the graph state.
4. When the required application service is useful again, not merely when database files load.
5. Store/config/plugin changes may not be backward-compatible; a validated pre-upgrade recovery path is needed before mutation.
Summary and next step
Load-Test, Profile, Secure, Back Up, Restore, Fail Over, Monitor, Upgrade, and Execute Incident Runbooks is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Present the Architecture with Evidence, Known Limits, Cost/Capacity Assumptions, and Clear Criteria for When Neo4j Is Not the Right Tool. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Current self-managed operational reference.
- System requirements — Supported Java, OS, memory, storage and filesystem requirements.
- Cypher compatibility and deprecations — Cypher 25 additions, compatibility and release-sensitive syntax.
- Constraints — Integrity constraints and backing-index semantics.
- Indexes — Search-performance and semantic index families.
- LOAD CSV — Transactional CSV import semantics.
- Execution plans — EXPLAIN/PROFILE and evidence-driven query tuning.
- Python driver manual — Official driver sessions, transactions, routing and application integration.
- Python driver performance — Driver-side performance, result handling and database selection.
- Authentication and authorization — Current authentication and edition-aware authorization model.
- Role-based access control — Enterprise/Aura RBAC and least-privilege controls.
- Offline database backup — Community offline dump semantics and backup boundaries.
- Restore a database dump — Community/Enterprise load and restore behavior.
- Online database backup — Enterprise online backup; not available on Aura.
- Neo4j clustering architecture — Enterprise primaries, secondaries, writer election and quorum.
- Neo4j logging — Operational logging surfaces.
- Neo4j metrics — Enterprise metrics surfaces and monitoring reference.
- Full-text indexes — Lexical search, analyzers and full-text query procedures.
- Vector indexes — Current vector-index lifecycle, ANN and score semantics.
- Cypher SEARCH — Cypher 25 vector SEARCH syntax introduced in Neo4j 2026.01.
- Graph Data Science manual — GDS 2026.07 graph projections, algorithms and ML.
- Supported GDS / Neo4j versions — Current GDS-to-Neo4j compatibility matrix.
- APOC installation — Current APOC/Neo4j version pairing and installation boundaries.
- Upgrade to Neo4j 2025–2026 — Supported upgrade paths and 5.26 LTS checkpoint behavior.
- Composite databases — Enterprise-only composite database boundary; unavailable on Aura.
- Built-in CDC procedures — Current db.cdc.* replacements and deprecated cdc.* procedure history.