Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j
Present the Architecture with Evidence, Known Limits, Cost/Capacity Assumptions, and Clear Criteria for When Neo4j Is Not the Right Tool
Close the course with an evidence package that states what Neo4j solves, what remains unproven, what the selected edition costs operationally, and when another data platform is the stronger choice.
Learning outcomes
Present the AtlasMart Neo4j decision as a traceable evidence package rather than a feature demo or benchmark anecdote.
Maintain a known-limit register that distinguishes measured limits, design assumptions, edition constraints and untested scenarios.
Build capacity/cost assumptions around workload units, tail latency, failure domains, staffing and licensing rather than hardware folklore.
Define explicit exit/migration criteria and stable export identities so Neo4j can be replaced when requirements change.
Decide when relational, document, search, vector or analytical systems are the better primary platform and state what Neo4j continues to own, if anything.
Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.
1. AtlasMart problem: the final deliverable is a decision, not a screenshot
The capstone now has a reproducible graph, core queries, an application driver boundary, optional retrieval design, load evidence, security boundaries, recovery proof and failure/upgrade runbooks. The final task is to explain what those artifacts support—and what they do not. A production architecture review needs traceability from requirement to mechanism to test to result, a known-limit register, and an architecture decision record (ADR) with alternatives and exit conditions.
| Dimension | Chapter 27 reproducible assumption |
|---|---|
| Neo4j | 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled. |
| Cypher | Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default. |
| Java | Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK. |
| Driver | Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions. |
| Database / auth |
Database neo4j; local user
neo4j; synthetic password
atlasmart-course-2026. Never reuse these lab
credentials in production.
|
| Network / TLS |
Loopback-only HTTP/Bolt for the disposable lab:
127.0.0.1:27474→7474 and
127.0.0.1:27687→7687. No TLS only because
traffic stays on localhost; production/remote connections
require a real TLS policy.
|
| Plugins | No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line. |
| Edition boundary | Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries. |
| Evidence rule | This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally. |
| Continuity | Final review keeps the Community architecture as the minimum reproducible system and treats Enterprise/Aura as separately costed options for requirements Community cannot meet. |
2. Evidence-backed architecture summary
| Architecture area | Chosen baseline | Required proof before production |
|---|---|---|
| Graph model | Customer–PLACED→Order–CONTAINS→Product–IN_CATEGORY→Category with stable constrained keys. | Model/cardinality/degree review; invariant tests; source ownership and temporal semantics. |
| Import | Versioned source files + idempotent LOAD CSV for the capstone fixture. | Source counts/hash, rejected-row policy, reconciliation, transaction/log growth at real scale. |
| Core Cypher | Selective key anchors, bounded traversals, deterministic result shaping. | EXPLAIN/PROFILE on representative cardinality; p50/p95/p99 and result-size evidence. |
| Driver | Python driver 6.3.0, explicit DB, parameterized queries, managed retry-safe transactions. | Pool/timeout/retry tests, idempotency, status-code handling, TLS and secret policy. |
| Search | Full-text and vector only if judged quality requirement justifies them. | Index ONLINE/freshness, recall/ranking metrics, latency, provenance, re-embedding plan. |
| Analytics | No GDS by default for tiny capstone; add only after graph-added-value experiment. | Projection memory estimate, baseline/ablation, reproducibility, refresh/serving boundary. |
| Security | Community native auth + loopback lab only. | Production TLS, secret rotation; Enterprise/Aura RBAC if least privilege/fine-grained access required. |
| Recovery | Community offline dump/load with isolated restore validation. | Measured RPO/RTO, offsite/immutable policy, full app smoke test and operator runbook. |
| HA | Not supplied by Community baseline. | If required: Enterprise cluster design, quorum/failure-domain tests, driver routing/bookmarks and cost. |
| Operations | PROFILE/SHOW/log/OS evidence; Enterprise metrics where licensed. | SLO alerts from baselines, capacity headroom, incident drills, version/upgrade cadence. |
3. Known-limit register
| Limit / uncertainty | Status | Consequence / next action |
|---|---|---|
| Fixture is tiny | Known test limitation | Do not extrapolate latency/QPS. Recreate production-like degree/skew/property/result distributions before sizing. |
| Community lacks true cluster HA | Edition boundary | Choose Enterprise/Aura if availability target requires server failure tolerance; otherwise document downtime/recovery expectation. |
| Community cannot prove fine-grained RBAC | Edition boundary | Do not certify tenant isolation from the Community lab. Evaluate Enterprise/Aura privileges and denial tests. |
| Embedding vectors are synthetic | Experiment boundary | They prove index/retrieval mechanics only. Production semantic quality requires versioned model data and judged evaluation. |
| GDS not justified by current fixture | Intentional omission | Keep it out until a measured algorithm/ML requirement shows graph-added value over baselines. |
| Benchmark results are machine-specific | Evidence boundary | Publish environment + distributions + raw data; use regression gates relative to a controlled baseline. |
| Backup tested locally, not offsite failure domain | Operational gap | Add remote/immutable copy and restore drill appropriate to production threat model. |
4. Cost and capacity model: record variables, not folklore
Capacity planning should expose what drives cost: stored graph/index bytes, growth rate, working-set/page-cache behavior, heap/query/transaction memory, CPU per workload class, disk IOPS/latency, network/result bytes, concurrency, background imports/search/GDS work, backup retention, HA copy count, observability, and operator time. Enterprise/Aura also add licensing/service-tier costs and may reduce some self-managed labor. Do not hard-code a universal “X GB RAM per million nodes” rule.
ADR-027: AtlasMart connected operational graph
Status: Proposed until production-scale evidence passes.
Decision: Use Neo4j for relationship-centric customer/order/product traversals where
graph queries materially reduce complexity and meet SLOs.
Base edition: Community for development/reproducibility.
Production edition: choose only after availability/security/operations requirements are costed.
Core invariants: stable domain IDs + constraints; bounded query inventory; idempotent writes/imports.
Optional features: full-text/vector/GraphRAG/GDS only when judged metrics justify them.
Recovery: independently restorable backups; replication is not backup.
Observability: workload + tail latency + resource + error evidence.
Rejected shortcut: enabling features because they are available.
Exit triggers: graph traversal ceases to add measurable value; cost/staffing/portability or
compliance constraints outweigh benefit; another engine owns dominant workload better.
Migration contract: export stable keys/properties/relationship facts; never depend on internal IDs.
5. When Neo4j is not the right primary tool
| Dominant requirement | Prefer / evaluate first | Why |
|---|---|---|
| Strong relational integrity + mostly tabular keyed joins/transactions | PostgreSQL or another relational DBMS | Mature relational constraints, SQL ecosystem and often simpler operations for non-traversal workloads. |
| Document aggregates with few cross-aggregate traversals | Document database | Natural document ownership and access pattern may avoid graph operational complexity. |
| Large-scale lexical search, faceting, log/search analytics | Elasticsearch/OpenSearch-like search platform | Inverted-index/search-native features may be the primary workload, with graph as optional enrichment. |
| Predominantly ANN/vector retrieval with minimal relationship context | Vector-specialized service/database | Vector lifecycle and scale may dominate while graph traversal adds little. |
| Warehouse/BI scans, large aggregations, columnar analytics | Columnar warehouse/lakehouse | Scan/aggregation economics and SQL BI ecosystem are typically better aligned. |
| Graph traversal is essential but self-managed staffing is unavailable | Managed graph service / Aura after tier review | Operational responsibility may matter more than raw feature parity; verify tier, region, security and cost. |
6. Exit and migration criteria
A responsible architecture includes a way out. Keep business identifiers and source ownership outside Neo4j-specific internal IDs, maintain exportable node/relationship schemas, version Cypher/driver contracts, and preserve reconciliation tests. Trigger a re-evaluation if traversal share falls, search/vector becomes the dominant cost center, compliance requires controls unavailable at the chosen tier, HA/licensing cost exceeds measured value, or production benchmarks show the graph cannot meet SLOs with reasonable headroom.
# Capture before cleanup.
docker exec atlasmart-neo4j-capstone cypher-shell `
-u neo4j -p atlasmart-course-2026 -d neo4j `
"SHOW INDEXES YIELD name,type,state RETURN name,type,state ORDER BY name;"
docker exec atlasmart-neo4j-capstone cypher-shell `
-u neo4j -p atlasmart-course-2026 -d neo4j `
"MATCH (c:Customer) WHERE c.capstone=true RETURN count(c) AS customers;"
# Expected fixed-fixture customer invariant: 5.
# Preserve your evidence files/hashes before deleting the disposable environment.
docker rm -f atlasmart-neo4j-capstone atlasmart-neo4j-capstone-restore 2>$null
docker volume rm atlasmart-neo4j-capstone-data atlasmart-neo4j-capstone-restore-data 2>$null
# Optional after evidence has been archived:
# Remove-Item -Recurse -Force (Join-Path $env:TEMP "atlasmart-neo4j-capstone")
7. Deliberately wrong approach: present the strongest benchmark and hide the limits
A single warm-cache best-case number, a successful screenshot, or a cherry-picked recommendation example is not a production case. It hides skew, failure, security, recovery, cost and alternative-system evidence.
Repair: present the requirement matrix, raw/reproducible test envelope, percentile distributions, recovery drill, security limitations, quality metrics, edition/license manifest, capacity assumptions, known failures and exit criteria. A negative result—such as “GDS adds no signal” or “search engine is cheaper for dominant retrieval”—is a successful architecture finding.
8. Hands-on lab: architecture review and go/no-go decision
Setup: collect the requirement matrix, graph/import reconciliation, query plans, driver tests, search judgments, performance distributions, security boundary, backup/restore evidence, outage/recovery notes, version/license manifest and known-limit register into one review folder.
Verification checklist
- Every production claim links to an artifact or is explicitly labeled an assumption/untested risk.
- Cost/capacity assumptions include workload shape, growth, memory/storage/network, backup retention, HA copies where applicable, observability and operator effort.
- The known-limit register contains edition/tier boundaries and at least one untested failure scenario.
- The ADR names rejected alternatives and explains which future requirement changes would trigger re-evaluation.
- Stable IDs and exportable relationship facts make the exit/migration path technically plausible.
Cleanup/reset: archive evidence first, then run the disposable-container/volume cleanup block. The final proof is the retained evidence and rebuild instructions—not the continued existence of a lab container.
Final production judgment
Neo4j is appropriate when relationship-centric operational questions, path/context reasoning or validated graph analytics create measurable value that exceeds the cost of graph-specific modeling, operations and licensing. It is not automatically the best home for every AtlasMart datum or workload. Keep systems at clear domain boundaries and choose specialized engines when they own a dominant workload more naturally.
This completes the 27-chapter Neo4j path. The durable skill is not memorizing Cypher syntax; it is being able to explain a graph mechanism from model and query through planner/transaction/store/driver behavior, test its limits, recover it, operate it, and defend—or reject—the architecture with evidence.
Check your understanding
- What makes an architecture review evidence-backed?
- Why include exit criteria in an ADR?
- What is a valid negative capstone result?
- Why preserve stable export identities?
- When is Neo4j justified?
Review the answers
1. Each requirement traces to a mechanism, reproducible test, observed result, known limitation and decision consequence.
2. Requirements and economics change; explicit triggers prevent sunk-cost attachment and make migration planning deliberate.
3. Evidence that an optional graph feature adds no value, or that another platform better fits the dominant workload, is useful architecture evidence.
4. They allow reconstruction/migration across databases and prevent lock-in to database-local internal IDs.
5. When measured relationship-centric value, correctness and operability outweigh alternatives and the chosen edition/topology meets security, resilience, capacity and cost constraints.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Current self-managed operational reference.
- System requirements — Supported Java, OS, memory, storage and filesystem requirements.
- Cypher compatibility and deprecations — Cypher 25 additions, compatibility and release-sensitive syntax.
- Constraints — Integrity constraints and backing-index semantics.
- Indexes — Search-performance and semantic index families.
- LOAD CSV — Transactional CSV import semantics.
- Execution plans — EXPLAIN/PROFILE and evidence-driven query tuning.
- Python driver manual — Official driver sessions, transactions, routing and application integration.
- Python driver performance — Driver-side performance, result handling and database selection.
- Authentication and authorization — Current authentication and edition-aware authorization model.
- Role-based access control — Enterprise/Aura RBAC and least-privilege controls.
- Offline database backup — Community offline dump semantics and backup boundaries.
- Restore a database dump — Community/Enterprise load and restore behavior.
- Online database backup — Enterprise online backup; not available on Aura.
- Neo4j clustering architecture — Enterprise primaries, secondaries, writer election and quorum.
- Neo4j logging — Operational logging surfaces.
- Neo4j metrics — Enterprise metrics surfaces and monitoring reference.
- Full-text indexes — Lexical search, analyzers and full-text query procedures.
- Vector indexes — Current vector-index lifecycle, ANN and score semantics.
- Cypher SEARCH — Cypher 25 vector SEARCH syntax introduced in Neo4j 2026.01.
- Graph Data Science manual — GDS 2026.07 graph projections, algorithms and ML.
- Supported GDS / Neo4j versions — Current GDS-to-Neo4j compatibility matrix.
- APOC installation — Current APOC/Neo4j version pairing and installation boundaries.
- Upgrade to Neo4j 2025–2026 — Supported upgrade paths and 5.26 LTS checkpoint behavior.
- Composite databases — Enterprise-only composite database boundary; unavailable on Aura.
- Built-in CDC procedures — Current db.cdc.* replacements and deprecated cdc.* procedure history.