Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j
Build the Graph Model, Constraints/Indexes, Import Pipeline, Core Cypher Queries, and Application Driver Layer
Build AtlasMart from deterministic source files into a constrained, reconciled property graph and expose it through selective Cypher and a retry-safe official Python driver boundary.
Learning outcomes
Create the disposable AtlasMart capstone DBMS with pinned versions, bounded network exposure, stable IDs, constraints and reproducible source files.
Load customers, products, categories, orders and relationships idempotently with LOAD CSV and reconcile source/graph counts.
Design core Cypher around selective anchors, bounded traversals, deterministic result shaping and indexes/constraints with distinct responsibilities.
Use the Neo4j Python driver with parameterized queries, explicit database selection, managed transactions and retry-safe application semantics.
Prove the build can be repeated without duplicate entities/relationships and preserve an evidence inventory for later search, recovery and performance tests.
Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.
1. AtlasMart problem: a model is production evidence only when it can be rebuilt
Lesson 1 justified a graph for relationship-centric AtlasMart workloads. Now the architecture must be reproducible from an empty environment. A constraint protects an integrity invariant; an index is an access path; an import reconciliation compares source expectations with graph state. These are separate concerns: a fast query can still be wrong, and a unique key does not automatically make every traversal cheap.
| Dimension | Chapter 27 reproducible assumption |
|---|---|
| Neo4j | 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled. |
| Cypher | Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default. |
| Java | Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK. |
| Driver | Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions. |
| Database / auth |
Database neo4j; local user
neo4j; synthetic password
atlasmart-course-2026. Never reuse these lab
credentials in production.
|
| Network / TLS |
Loopback-only HTTP/Bolt for the disposable lab:
127.0.0.1:27474→7474 and
127.0.0.1:27687→7687. No TLS only because
traffic stays on localhost; production/remote connections
require a real TLS policy.
|
| Plugins | No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line. |
| Edition boundary | Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries. |
| Evidence rule | This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally. |
| Continuity | This lesson introduces only capstone-marked entities/relationships and a dedicated container/volume. Re-running the import should not multiply stable-key entities or MERGE-owned relationships. |
# PowerShell — disposable Chapter 27 workspace and Community DBMS.
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Import = Join-Path $Root "import"
$Evidence = Join-Path $Root "evidence"
New-Item -ItemType Directory -Force -Path $Import,$Evidence | Out-Null
docker rm -f atlasmart-neo4j-capstone 2>$null
docker volume rm atlasmart-neo4j-capstone-data 2>$null
docker run -d --name atlasmart-neo4j-capstone `
-p 127.0.0.1:27474:7474 -p 127.0.0.1:27687:7687 `
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
--mount type=bind,source="$Import",target=/var/lib/neo4j/import `
-v atlasmart-neo4j-capstone-data:/data `
neo4j:2026.07.1
docker ps --filter "name=atlasmart-neo4j-capstone"
docker logs --tail 80 atlasmart-neo4j-capstone
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Import = Join-Path $Root "import"
@'
customerId,name,segment
C-1001,Ada,consumer
C-1002,Ben,consumer
C-1003,Cam,business
C-1004,Dara,consumer
C-1005,Eli,business
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'customers.csv')
@'
categoryId,name
CAT-CAMERAS,Cameras
CAT-AUDIO,Audio
CAT-HOME,Smart Home
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'categories.csv')
@'
productId,name,description,categoryId,price,embedding
P-1001,Trail Camera X1,"weatherproof trail camera night vision",CAT-CAMERAS,129.0,"0.98;0.05;0.02;0.01"
P-1002,Mirrorless M20,"compact mirrorless camera travel photography",CAT-CAMERAS,699.0,"0.93;0.08;0.03;0.02"
P-2001,QuietPods,"noise cancelling wireless earbuds",CAT-AUDIO,149.0,"0.04;0.97;0.03;0.01"
P-2002,StudioHead 5,"over ear studio headphones",CAT-AUDIO,199.0,"0.06;0.92;0.05;0.01"
P-3001,HomeCam Mini,"indoor smart home security camera",CAT-HOME,89.0,"0.84;0.08;0.07;0.01"
P-3002,DoorSensor,"wireless door contact sensor smart home",CAT-HOME,39.0,"0.10;0.08;0.80;0.02"
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'products.csv')
@'
orderId,customerId,createdAt,status
O-5001,C-1001,2026-08-01T10:00:00Z,COMPLETE
O-5002,C-1001,2026-08-12T11:00:00Z,COMPLETE
O-5003,C-1002,2026-08-13T09:00:00Z,COMPLETE
O-5004,C-1003,2026-08-14T15:00:00Z,COMPLETE
O-5005,C-1004,2026-08-16T16:00:00Z,RETURNED
O-5006,C-1005,2026-08-18T13:00:00Z,COMPLETE
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'orders.csv')
@'
orderId,productId,quantity
O-5001,P-1001,1
O-5001,P-2001,1
O-5002,P-1002,1
O-5002,P-2001,2
O-5003,P-2001,1
O-5003,P-2002,1
O-5004,P-1001,2
O-5004,P-3001,1
O-5005,P-3001,1
O-5006,P-3002,3
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'contains.csv')
Get-ChildItem $Import | Select-Object Name,Length
// Run against bolt://localhost:27687, database neo4j.
CREATE CONSTRAINT cap_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT cap_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT cap_order_id IF NOT EXISTS
FOR (o:Order) REQUIRE o.orderId IS UNIQUE;
CREATE CONSTRAINT cap_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE INDEX cap_order_created IF NOT EXISTS
FOR (o:Order) ON (o.createdAt);
LOAD CSV WITH HEADERS FROM 'file:///customers.csv' AS row
MERGE (c:Customer {customerId: row.customerId})
SET c.name=row.name, c.segment=row.segment, c.capstone=true;
LOAD CSV WITH HEADERS FROM 'file:///categories.csv' AS row
MERGE (cat:Category {categoryId: row.categoryId})
SET cat.name=row.name, cat.capstone=true;
LOAD CSV WITH HEADERS FROM 'file:///products.csv' AS row
MERGE (p:Product {productId: row.productId})
SET p.name=row.name,
p.description=row.description,
p.price=toFloat(row.price),
p.embedding=[x IN split(row.embedding,';') | toFloat(x)],
p.capstone=true
WITH p,row
MATCH (cat:Category {categoryId:row.categoryId})
MERGE (p)-[:IN_CATEGORY {capstone:true}]->(cat);
LOAD CSV WITH HEADERS FROM 'file:///orders.csv' AS row
MERGE (o:Order {orderId: row.orderId})
SET o.createdAt=datetime(row.createdAt), o.status=row.status, o.capstone=true
WITH o,row
MATCH (c:Customer {customerId:row.customerId})
MERGE (c)-[:PLACED {capstone:true}]->(o);
LOAD CSV WITH HEADERS FROM 'file:///contains.csv' AS row
MATCH (o:Order {orderId:row.orderId})
MATCH (p:Product {productId:row.productId})
MERGE (o)-[r:CONTAINS]->(p)
SET r.quantity=toInteger(row.quantity), r.capstone=true;
2. What Neo4j is doing during the import
Each LOAD CSV statement produces rows from the
file. The uniqueness constraints allow the keyed
MERGE patterns to resolve a single Customer,
Product, Order or Category. Relationship creation is separated
from entity creation so a missing endpoint is observable rather
than silently manufacturing an incomplete entity. For large
sources you would batch transactions and measure memory/lock
behavior; this six-order fixture stays intentionally small so
correctness is hand-checkable.
The list-valued embedding property is deterministic
training data for Lesson 3. It is not claimed to come from a
real embedding model, and it does not make the graph
semantically intelligent by itself.
MATCH (c:Customer) WHERE c.capstone=true WITH count(c) AS customers
MATCH (p:Product) WHERE p.capstone=true WITH customers,count(p) AS products
MATCH (o:Order) WHERE o.capstone=true WITH customers,products,count(o) AS orders
MATCH (cat:Category) WHERE cat.capstone=true WITH customers,products,orders,count(cat) AS categories
MATCH (:Customer)-[pl:PLACED]->(:Order) WHERE pl.capstone=true
WITH customers,products,orders,categories,count(pl) AS placed
MATCH (:Order)-[co:CONTAINS]->(:Product) WHERE co.capstone=true
WITH customers,products,orders,categories,placed,count(co) AS contains
MATCH (:Product)-[ic:IN_CATEGORY]->(:Category) WHERE ic.capstone=true
RETURN customers,products,orders,categories,placed,contains,count(ic) AS inCategory;
// Expected deterministic invariant after one or repeated import:
// customers=5, products=6, orders=6, categories=3,
// placed=6, contains=10, inCategory=6.
3. Core query inventory: selectivity before expansion
A selective anchor starts from a small set of
nodes, ideally one stable key, before expanding relationships. A
two-hop traversal from Customer{customerId} is
predictable on the fixed fixture; a database-wide unbounded path
search is not. The planner may change operator names across
versions/runtimes, so capture actual PROFILE evidence rather
than teaching one plan as eternal.
PROFILE
MATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)-[r:CONTAINS]->(p:Product)
RETURN o.orderId AS orderId,
o.createdAt AS createdAt,
collect({productId:p.productId, quantity:r.quantity}) AS items
ORDER BY createdAt DESC;
// Parameter: $customerId = 'C-1001'
// Expected semantic invariant: orders O-5002 and O-5001 are returned.
// Verify locally that the anchor uses the customer key constraint/index path;
// do not copy numeric DB Hits/Rows from another machine or release.
MATCH (seed:Product {productId:$productId})<-[:CONTAINS]-(o:Order)-[:CONTAINS]->(candidate:Product)
WHERE candidate <> seed
RETURN candidate.productId AS productId,
count(DISTINCT o) AS sharedOrders
ORDER BY sharedOrders DESC, productId
LIMIT $limit;
// $productId='P-2001', $limit=5
// Fixed-fixture candidates include P-1001, P-1002 and P-2002.
// This is association evidence, not causal proof that the products should be recommended.
# pip install neo4j==6.3.0
from neo4j import GraphDatabase
URI = "bolt://localhost:27687"
AUTH = ("neo4j", "atlasmart-course-2026")
DB = "neo4j"
def order_history(tx, customer_id):
result = tx.run("""
MATCH (c:Customer {customerId:$customer_id})-[:PLACED]->(o:Order)-[r:CONTAINS]->(p:Product)
RETURN o.orderId AS orderId, o.createdAt AS createdAt,
collect({productId:p.productId, quantity:r.quantity}) AS items
ORDER BY createdAt DESC
""", customer_id=customer_id)
return [record.data() for record in result]
with GraphDatabase.driver(URI, auth=AUTH) as driver:
driver.verify_connectivity()
with driver.session(database=DB) as session:
rows = session.execute_read(order_history, "C-1001")
print(rows)
# Expected invariant for the fixed fixture: C-1001 has O-5002 and O-5001.
# Exact formatting/timestamps are driver-version dependent; verify the IDs, not a copied screenshot.
4. Driver boundary: retries require idempotency
The driver owns connection pooling and can retry managed transactions for retryable failures. Therefore application transaction functions must tolerate re-execution. A transaction that charges a payment, sends an email, and then writes Neo4j without an idempotency boundary could repeat the external side effect if retried. Keep irreversible external effects outside retryable database functions or coordinate them with an outbox/idempotency key.
5. Deliberately wrong approach: MERGE the whole graph pattern using internal IDs
A practitioner might persist an elementId() in
another service and later MERGE an Order/Product
pattern around it. That couples business identity to a
database-local implementation detail and can create incorrect
relationships after rebuild/migration.
Repair: anchor each entity by a constrained domain key, match endpoints independently, then MERGE only the relationship whose ownership semantics justify idempotency. Re-run the import and verify the deterministic reconciliation counts stay 5/6/6/3/6/10/6.
6. Hands-on lab: empty environment → imported graph → application read
Setup: run the disposable-container block, generate the five CSV files, then execute the import Cypher. The dedicated volume and loopback ports isolate this exercise from the continuity container.
Verification checklist
- Reconciliation returns exactly 5 customers, 6 products, 6 orders, 3 categories, 6 PLACED, 10 CONTAINS and 6 IN_CATEGORY relationships.
- Run the import a second time and confirm those counts do not increase.
-
SHOW CONSTRAINTSshows the four capstone identity constraints andSHOW INDEXESshows the order-date index. - The C-1001 order-history query returns O-5002 and O-5001, and local PROFILE evidence starts from the constrained customer key rather than scanning all customers.
- The Python driver verifies connectivity and returns the same semantic IDs as direct Cypher.
Cleanup/reset: for a standalone rerun, remove
atlasmart-neo4j-capstone and
atlasmart-neo4j-capstone-data, regenerate the CSV
files, and repeat. During the chapter progression, keep the
environment because Lesson 3 consumes it.
Production judgment
Import throughput is secondary to integrity until reconciliation passes. At production scale, measure file size, rows/sec, transaction size, page cache/store I/O, lock contention and transaction-log growth; choose LOAD CSV, Data Importer or offline import based on lifecycle and dataset size. Keep driver timeouts, pool size and retry policy tied to measured concurrency and failure modes rather than universal defaults.
Lesson 3 asks whether search, vector retrieval, GraphRAG or GDS measurably improves a requirement that core Cypher does not already satisfy.
Check your understanding
- Why create constraints before the keyed import?
- What proves the import is idempotent?
- Why separate entity MERGE from relationship MERGE?
- What does PROFILE prove?
- Why must managed-transaction callbacks be retry-safe?
Review the answers
1. They enforce stable identity and give keyed MERGE an integrity/access foundation, especially under repeated or concurrent writes.
2. Re-running it preserves the expected entity/relationship counts and keys rather than multiplying graph state.
3. It makes endpoint ownership and missing references explicit and prevents broad patterns from creating unintended entities.
4. It executes the query and reports runtime evidence for that environment; it does not establish universal operator costs.
5. The driver may re-execute them after retryable failures, so non-idempotent external side effects could duplicate.
Summary and next step
Build the Graph Model, Constraints/Indexes, Import Pipeline, Core Cypher Queries, and Application Driver Layer is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Add Full-Text/Vector/GraphRAG and GDS Analytics Only Where Measured Requirements Justify Them. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Current self-managed operational reference.
- System requirements — Supported Java, OS, memory, storage and filesystem requirements.
- Cypher compatibility and deprecations — Cypher 25 additions, compatibility and release-sensitive syntax.
- Constraints — Integrity constraints and backing-index semantics.
- Indexes — Search-performance and semantic index families.
- LOAD CSV — Transactional CSV import semantics.
- Execution plans — EXPLAIN/PROFILE and evidence-driven query tuning.
- Python driver manual — Official driver sessions, transactions, routing and application integration.
- Python driver performance — Driver-side performance, result handling and database selection.
- Authentication and authorization — Current authentication and edition-aware authorization model.
- Role-based access control — Enterprise/Aura RBAC and least-privilege controls.
- Offline database backup — Community offline dump semantics and backup boundaries.
- Restore a database dump — Community/Enterprise load and restore behavior.
- Online database backup — Enterprise online backup; not available on Aura.
- Neo4j clustering architecture — Enterprise primaries, secondaries, writer election and quorum.
- Neo4j logging — Operational logging surfaces.
- Neo4j metrics — Enterprise metrics surfaces and monitoring reference.
- Full-text indexes — Lexical search, analyzers and full-text query procedures.
- Vector indexes — Current vector-index lifecycle, ANN and score semantics.
- Cypher SEARCH — Cypher 25 vector SEARCH syntax introduced in Neo4j 2026.01.
- Graph Data Science manual — GDS 2026.07 graph projections, algorithms and ML.
- Supported GDS / Neo4j versions — Current GDS-to-Neo4j compatibility matrix.
- APOC installation — Current APOC/Neo4j version pairing and installation boundaries.
- Upgrade to Neo4j 2025–2026 — Supported upgrade paths and 5.26 LTS checkpoint behavior.
- Composite databases — Enterprise-only composite database boundary; unavailable on Aura.
- Built-in CDC procedures — Current db.cdc.* replacements and deprecated cdc.* procedure history.