Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j

Build the Graph Model, Constraints/Indexes, Import Pipeline, Core Cypher Queries, and Application Driver Layer

Build AtlasMart from deterministic source files into a constrained, reconciled property graph and expose it through selective Cypher and a retry-safe official Python driver boundary.

Advanced360–480 minutesModel · import · Cypher · driverNeo4j 2026.07.1 · CommunityCypher 25 · Python driver 6.3.0LOAD CSV · constraints/indexes · reconciliationLoopback 27474/27687 · no APOC/GDS requiredLast reviewed: September 2026

Learning outcomes

01

Create the disposable AtlasMart capstone DBMS with pinned versions, bounded network exposure, stable IDs, constraints and reproducible source files.

02

Load customers, products, categories, orders and relationships idempotently with LOAD CSV and reconcile source/graph counts.

03

Design core Cypher around selective anchors, bounded traversals, deterministic result shaping and indexes/constraints with distinct responsibilities.

04

Use the Neo4j Python driver with parameterized queries, explicit database selection, managed transactions and retry-safe application semantics.

05

Prove the build can be repeated without duplicate entities/relationships and preserve an evidence inventory for later search, recovery and performance tests.

Execution and safety note

Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.

1. AtlasMart problem: a model is production evidence only when it can be rebuilt

Lesson 1 justified a graph for relationship-centric AtlasMart workloads. Now the architecture must be reproducible from an empty environment. A constraint protects an integrity invariant; an index is an access path; an import reconciliation compares source expectations with graph state. These are separate concerns: a fast query can still be wrong, and a unique key does not automatically make every traversal cheap.

Dimension Chapter 27 reproducible assumption
Neo4j 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled.
Cypher Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default.
Java Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK.
Driver Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions.
Database / auth Database neo4j; local user neo4j; synthetic password atlasmart-course-2026. Never reuse these lab credentials in production.
Network / TLS Loopback-only HTTP/Bolt for the disposable lab: 127.0.0.1:27474→7474 and 127.0.0.1:27687→7687. No TLS only because traffic stays on localhost; production/remote connections require a real TLS policy.
Plugins No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line.
Edition boundary Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries.
Evidence rule This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally.
Continuity This lesson introduces only capstone-marked entities/relationships and a dedicated container/volume. Re-running the import should not multiply stable-key entities or MERGE-owned relationships.
Create the disposable capstone DBMS
# PowerShell — disposable Chapter 27 workspace and Community DBMS.
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Import = Join-Path $Root "import"
$Evidence = Join-Path $Root "evidence"
New-Item -ItemType Directory -Force -Path $Import,$Evidence | Out-Null

docker rm -f atlasmart-neo4j-capstone 2>$null
docker volume rm atlasmart-neo4j-capstone-data 2>$null

docker run -d --name atlasmart-neo4j-capstone `
  -p 127.0.0.1:27474:7474 -p 127.0.0.1:27687:7687 `
  -e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
  --mount type=bind,source="$Import",target=/var/lib/neo4j/import `
  -v atlasmart-neo4j-capstone-data:/data `
  neo4j:2026.07.1

docker ps --filter "name=atlasmart-neo4j-capstone"
docker logs --tail 80 atlasmart-neo4j-capstone
Generate deterministic AtlasMart CSV sources
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Import = Join-Path $Root "import"

@'
customerId,name,segment
C-1001,Ada,consumer
C-1002,Ben,consumer
C-1003,Cam,business
C-1004,Dara,consumer
C-1005,Eli,business
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'customers.csv')

@'
categoryId,name
CAT-CAMERAS,Cameras
CAT-AUDIO,Audio
CAT-HOME,Smart Home
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'categories.csv')

@'
productId,name,description,categoryId,price,embedding
P-1001,Trail Camera X1,"weatherproof trail camera night vision",CAT-CAMERAS,129.0,"0.98;0.05;0.02;0.01"
P-1002,Mirrorless M20,"compact mirrorless camera travel photography",CAT-CAMERAS,699.0,"0.93;0.08;0.03;0.02"
P-2001,QuietPods,"noise cancelling wireless earbuds",CAT-AUDIO,149.0,"0.04;0.97;0.03;0.01"
P-2002,StudioHead 5,"over ear studio headphones",CAT-AUDIO,199.0,"0.06;0.92;0.05;0.01"
P-3001,HomeCam Mini,"indoor smart home security camera",CAT-HOME,89.0,"0.84;0.08;0.07;0.01"
P-3002,DoorSensor,"wireless door contact sensor smart home",CAT-HOME,39.0,"0.10;0.08;0.80;0.02"
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'products.csv')

@'
orderId,customerId,createdAt,status
O-5001,C-1001,2026-08-01T10:00:00Z,COMPLETE
O-5002,C-1001,2026-08-12T11:00:00Z,COMPLETE
O-5003,C-1002,2026-08-13T09:00:00Z,COMPLETE
O-5004,C-1003,2026-08-14T15:00:00Z,COMPLETE
O-5005,C-1004,2026-08-16T16:00:00Z,RETURNED
O-5006,C-1005,2026-08-18T13:00:00Z,COMPLETE
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'orders.csv')

@'
orderId,productId,quantity
O-5001,P-1001,1
O-5001,P-2001,1
O-5002,P-1002,1
O-5002,P-2001,2
O-5003,P-2001,1
O-5003,P-2002,1
O-5004,P-1001,2
O-5004,P-3001,1
O-5005,P-3001,1
O-5006,P-3002,3
'@ | Set-Content -Encoding utf8 (Join-Path $Import 'contains.csv')

Get-ChildItem $Import | Select-Object Name,Length
Create integrity/access structures and import the graph
// Run against bolt://localhost:27687, database neo4j.
CREATE CONSTRAINT cap_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT cap_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT cap_order_id IF NOT EXISTS
FOR (o:Order) REQUIRE o.orderId IS UNIQUE;
CREATE CONSTRAINT cap_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE INDEX cap_order_created IF NOT EXISTS
FOR (o:Order) ON (o.createdAt);

LOAD CSV WITH HEADERS FROM 'file:///customers.csv' AS row
MERGE (c:Customer {customerId: row.customerId})
SET c.name=row.name, c.segment=row.segment, c.capstone=true;

LOAD CSV WITH HEADERS FROM 'file:///categories.csv' AS row
MERGE (cat:Category {categoryId: row.categoryId})
SET cat.name=row.name, cat.capstone=true;

LOAD CSV WITH HEADERS FROM 'file:///products.csv' AS row
MERGE (p:Product {productId: row.productId})
SET p.name=row.name,
    p.description=row.description,
    p.price=toFloat(row.price),
    p.embedding=[x IN split(row.embedding,';') | toFloat(x)],
    p.capstone=true
WITH p,row
MATCH (cat:Category {categoryId:row.categoryId})
MERGE (p)-[:IN_CATEGORY {capstone:true}]->(cat);

LOAD CSV WITH HEADERS FROM 'file:///orders.csv' AS row
MERGE (o:Order {orderId: row.orderId})
SET o.createdAt=datetime(row.createdAt), o.status=row.status, o.capstone=true
WITH o,row
MATCH (c:Customer {customerId:row.customerId})
MERGE (c)-[:PLACED {capstone:true}]->(o);

LOAD CSV WITH HEADERS FROM 'file:///contains.csv' AS row
MATCH (o:Order {orderId:row.orderId})
MATCH (p:Product {productId:row.productId})
MERGE (o)-[r:CONTAINS]->(p)
SET r.quantity=toInteger(row.quantity), r.capstone=true;

2. What Neo4j is doing during the import

Each LOAD CSV statement produces rows from the file. The uniqueness constraints allow the keyed MERGE patterns to resolve a single Customer, Product, Order or Category. Relationship creation is separated from entity creation so a missing endpoint is observable rather than silently manufacturing an incomplete entity. For large sources you would batch transactions and measure memory/lock behavior; this six-order fixture stays intentionally small so correctness is hand-checkable.

The list-valued embedding property is deterministic training data for Lesson 3. It is not claimed to come from a real embedding model, and it does not make the graph semantically intelligent by itself.

Reconcile the imported state
MATCH (c:Customer) WHERE c.capstone=true WITH count(c) AS customers
MATCH (p:Product) WHERE p.capstone=true WITH customers,count(p) AS products
MATCH (o:Order) WHERE o.capstone=true WITH customers,products,count(o) AS orders
MATCH (cat:Category) WHERE cat.capstone=true WITH customers,products,orders,count(cat) AS categories
MATCH (:Customer)-[pl:PLACED]->(:Order) WHERE pl.capstone=true
WITH customers,products,orders,categories,count(pl) AS placed
MATCH (:Order)-[co:CONTAINS]->(:Product) WHERE co.capstone=true
WITH customers,products,orders,categories,placed,count(co) AS contains
MATCH (:Product)-[ic:IN_CATEGORY]->(:Category) WHERE ic.capstone=true
RETURN customers,products,orders,categories,placed,contains,count(ic) AS inCategory;

// Expected deterministic invariant after one or repeated import:
// customers=5, products=6, orders=6, categories=3,
// placed=6, contains=10, inCategory=6.

3. Core query inventory: selectivity before expansion

A selective anchor starts from a small set of nodes, ideally one stable key, before expanding relationships. A two-hop traversal from Customer{customerId} is predictable on the fixed fixture; a database-wide unbounded path search is not. The planner may change operator names across versions/runtimes, so capture actual PROFILE evidence rather than teaching one plan as eternal.

EXPLAIN/PROFILE a stable-key order-history query
PROFILE
MATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)-[r:CONTAINS]->(p:Product)
RETURN o.orderId AS orderId,
       o.createdAt AS createdAt,
       collect({productId:p.productId, quantity:r.quantity}) AS items
ORDER BY createdAt DESC;

// Parameter: $customerId = 'C-1001'
// Expected semantic invariant: orders O-5002 and O-5001 are returned.
// Verify locally that the anchor uses the customer key constraint/index path;
// do not copy numeric DB Hits/Rows from another machine or release.
Find co-purchased products with bounded graph expansion
MATCH (seed:Product {productId:$productId})<-[:CONTAINS]-(o:Order)-[:CONTAINS]->(candidate:Product)
WHERE candidate <> seed
RETURN candidate.productId AS productId,
       count(DISTINCT o) AS sharedOrders
ORDER BY sharedOrders DESC, productId
LIMIT $limit;

// $productId='P-2001', $limit=5
// Fixed-fixture candidates include P-1001, P-1002 and P-2002.
// This is association evidence, not causal proof that the products should be recommended.
Connect through the official Python driver
# pip install neo4j==6.3.0
from neo4j import GraphDatabase

URI = "bolt://localhost:27687"
AUTH = ("neo4j", "atlasmart-course-2026")
DB = "neo4j"

def order_history(tx, customer_id):
    result = tx.run("""
        MATCH (c:Customer {customerId:$customer_id})-[:PLACED]->(o:Order)-[r:CONTAINS]->(p:Product)
        RETURN o.orderId AS orderId, o.createdAt AS createdAt,
               collect({productId:p.productId, quantity:r.quantity}) AS items
        ORDER BY createdAt DESC
    """, customer_id=customer_id)
    return [record.data() for record in result]

with GraphDatabase.driver(URI, auth=AUTH) as driver:
    driver.verify_connectivity()
    with driver.session(database=DB) as session:
        rows = session.execute_read(order_history, "C-1001")
        print(rows)

# Expected invariant for the fixed fixture: C-1001 has O-5002 and O-5001.
# Exact formatting/timestamps are driver-version dependent; verify the IDs, not a copied screenshot.

4. Driver boundary: retries require idempotency

The driver owns connection pooling and can retry managed transactions for retryable failures. Therefore application transaction functions must tolerate re-execution. A transaction that charges a payment, sends an email, and then writes Neo4j without an idempotency boundary could repeat the external side effect if retried. Keep irreversible external effects outside retryable database functions or coordinate them with an outbox/idempotency key.

5. Deliberately wrong approach: MERGE the whole graph pattern using internal IDs

A practitioner might persist an elementId() in another service and later MERGE an Order/Product pattern around it. That couples business identity to a database-local implementation detail and can create incorrect relationships after rebuild/migration.

Repair: anchor each entity by a constrained domain key, match endpoints independently, then MERGE only the relationship whose ownership semantics justify idempotency. Re-run the import and verify the deterministic reconciliation counts stay 5/6/6/3/6/10/6.

6. Hands-on lab: empty environment → imported graph → application read

Setup: run the disposable-container block, generate the five CSV files, then execute the import Cypher. The dedicated volume and loopback ports isolate this exercise from the continuity container.

Verification checklist

  • Reconciliation returns exactly 5 customers, 6 products, 6 orders, 3 categories, 6 PLACED, 10 CONTAINS and 6 IN_CATEGORY relationships.
  • Run the import a second time and confirm those counts do not increase.
  • SHOW CONSTRAINTS shows the four capstone identity constraints and SHOW INDEXES shows the order-date index.
  • The C-1001 order-history query returns O-5002 and O-5001, and local PROFILE evidence starts from the constrained customer key rather than scanning all customers.
  • The Python driver verifies connectivity and returns the same semantic IDs as direct Cypher.

Cleanup/reset: for a standalone rerun, remove atlasmart-neo4j-capstone and atlasmart-neo4j-capstone-data, regenerate the CSV files, and repeat. During the chapter progression, keep the environment because Lesson 3 consumes it.

Production judgment

Import throughput is secondary to integrity until reconciliation passes. At production scale, measure file size, rows/sec, transaction size, page cache/store I/O, lock contention and transaction-log growth; choose LOAD CSV, Data Importer or offline import based on lifecycle and dataset size. Keep driver timeouts, pool size and retry policy tied to measured concurrency and failure modes rather than universal defaults.

Lesson 3 asks whether search, vector retrieval, GraphRAG or GDS measurably improves a requirement that core Cypher does not already satisfy.

Check your understanding

  1. Why create constraints before the keyed import?
  2. What proves the import is idempotent?
  3. Why separate entity MERGE from relationship MERGE?
  4. What does PROFILE prove?
  5. Why must managed-transaction callbacks be retry-safe?
Review the answers

1. They enforce stable identity and give keyed MERGE an integrity/access foundation, especially under repeated or concurrent writes.

2. Re-running it preserves the expected entity/relationship counts and keys rather than multiplying graph state.

3. It makes endpoint ownership and missing references explicit and prevents broad patterns from creating unintended entities.

4. It executes the query and reports runtime evidence for that environment; it does not establish universal operator costs.

5. The driver may re-execute them after retryable failures, so non-idempotent external side effects could duplicate.

Summary and next step

Build the Graph Model, Constraints/Indexes, Import Pipeline, Core Cypher Queries, and Application Driver Layer is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Add Full-Text/Vector/GraphRAG and GDS Analytics Only Where Measured Requirements Justify Them. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.