Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j
Define Requirements, Workloads, Availability/RPO/RTO, Data Sources, Security Boundaries, and Graph-Fit Criteria
Turn AtlasMart goals into a falsifiable architecture contract before building the final graph: measurable workloads, SLO/RPO/RTO, source and identity ownership, security boundaries, graph-fit criteria, risks, and exit rules.
Learning outcomes
Translate AtlasMart business goals into measurable graph workloads, SLOs, availability targets, RPO/RTO, and acceptance evidence before choosing topology or features.
Define source-of-truth and stable identity contracts so the graph can be rebuilt, reconciled, restored, or migrated without depending on Neo4j internal IDs.
Draw security and tenant boundaries explicitly, including what Community can prove locally and what requires Enterprise/Aura controls.
Use a graph-fit scorecard that compares connected-query value against relational, document, search, vector, and analytical alternatives.
Create an architecture decision record, known-limit register, version/license manifest, and exit criteria that later lessons can verify.
Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.
1. AtlasMart problem: “use Neo4j” is not yet an architecture
AtlasMart wants one system to answer customer-order-product traversals, recommendations, catalog search, support context, and operational questions. A graph can make relationship-centric questions concise, but visual connectedness is not proof of fit. The capstone therefore begins with an architecture fitness claim: a falsifiable statement that a chosen property-graph model and Neo4j operational envelope satisfy named workloads better enough to justify their complexity and cost.
A service-level objective (SLO) is a measurable reliability/performance target, such as a p95 latency or monthly availability objective. RPO is the maximum acceptable recoverability gap after data loss; RTO is the time allowed to restore useful service. An availability target is not the same as a backup target: Enterprise replication can improve availability, while independently restorable backups protect against corruption, deletion, operator error and other failures that replication can copy.
| Dimension | Chapter 27 reproducible assumption |
|---|---|
| Neo4j | 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled. |
| Cypher | Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default. |
| Java | Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK. |
| Driver | Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions. |
| Database / auth |
Database neo4j; local user
neo4j; synthetic password
atlasmart-course-2026. Never reuse these lab
credentials in production.
|
| Network / TLS |
Loopback-only HTTP/Bolt for the disposable lab:
127.0.0.1:27474→7474 and
127.0.0.1:27687→7687. No TLS only because
traffic stays on localhost; production/remote connections
require a real TLS policy.
|
| Plugins | No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line. |
| Edition boundary | Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries. |
| Evidence rule | This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally. |
| Continuity | The capstone reuses stable IDs such as C-1001, P-1001 and O-5001 and introduces only capstone-tagged data so cleanup is bounded. |
2. Start with a traceability matrix, not a feature checklist
Each requirement needs an owner, a measurable test, a workload shape, and an architectural consequence. This prevents the capstone from adding vector search, GDS, clustering, CDC, or composite databases merely because they appeared earlier in the course. A feature is justified only if its evidence closes a requirement gap more economically or reliably than a simpler mechanism.
| ID | Requirement / workload | Evidence to collect | Design consequence |
|---|---|---|---|
| R1 | Customer order history: keyed customer → orders → products; bounded result. | Correct IDs/items, EXPLAIN/PROFILE access path, p95/p99 under documented concurrency. | Stable keys + uniqueness constraints; parameterized 2-hop Cypher; bounded projection. |
| R2 | Related-product discovery from shared orders/categories. | Candidate relevance on a judged fixture, traversal fan-out, latency distribution. | Graph relationship traversal is a core fit criterion; precompute only if measured. |
| R3 | Catalog lexical/semantic retrieval. | Judged queries, recall@k/MRR, freshness, latency, provenance. | Full-text/vector optional; scores remain retrieval evidence, not business truth. |
| R4 | Recoverability. | Offline Community dump hash, isolated load, canary/count reconciliation, measured recovery steps. | Replication never substitutes for backup; restore testing is mandatory. |
| R5 | HA / node loss. | Community: process outage/restart evidence + deterministic quorum trace. Enterprise option: real cluster writer/quorum test. | True cluster failover is an Enterprise design/cost decision. |
| R6 | Security boundary. | Auth test, network exposure review, secret handling; optional Enterprise deny/allow tests. | Community does not reproduce fine-grained RBAC; do not claim least privilege it cannot enforce. |
CAPSTONE_ID=atlasmart-ch27
SERVER=Neo4j 2026.07.1 Community
LTS_REFERENCE=Neo4j 5.26.30
CYPHER=25 for examples; verify database setting before deployment
JAVA=21 or 25 for Neo4j 2026.07 self-managed; Docker image runtime is pinned by image
DRIVER=neo4j Python 6.3.0
DATABASE=neo4j
BOLT=bolt://localhost:27687 (local disposable lab only)
AUTH=synthetic local native credentials; never production credentials
TLS=off only on loopback lab; production policy must be explicit
APOC=not required; if used pin 2026.07.1
GDS=not required by default; if justified pin 2026.07.0
ENTERPRISE_ONLY=cluster HA, online backup, fine-grained RBAC, composite DB, self-managed CDC
RPO_TARGET=define before backup design
RTO_TARGET=define before recovery design
EXIT_CRITERIA=document conditions that favor relational/search/document/vector/warehouse alternatives
3. Source-of-truth and identity model
The graph must be reconstructible from durable domain
identifiers. AtlasMart therefore treats customerId,
productId, orderId, and
categoryId as cross-system identities and backs
them with constraints. Neo4j element IDs are useful within a
database interaction but are not the contract for rebuilds,
restores, federation or migration. The source system must define
whether an event is authoritative, a correction, a tombstone, or
a derived projection.
For every relationship ask who owns it and how it is refreshed.
PLACED belongs to order ownership;
CONTAINS belongs to order-line facts;
IN_CATEGORY belongs to catalog governance. A
recommendation edge would be derived and should carry
lineage/freshness rather than masquerading as source truth.
4. Security boundary is an architecture input
A trust boundary separates components that may safely trust one another from components that require authentication, authorization, encryption, or validation. The local capstone binds ports to loopback and uses synthetic credentials, so disabling TLS there does not imply that plaintext Bolt is acceptable across a production network. Community supports native users, but current fine-grained role/privilege controls are an Enterprise/Aura capability; therefore the mandatory lab records that limitation instead of inventing a least-privilege role.
5. Graph-fit criteria and explicit alternatives
| Question | Neo4j strengthens the case when… | Another platform strengthens the case when… |
|---|---|---|
| Connected operational reads | Traversal depth and relationship semantics are central and queried online. | Queries are mostly keyed rows/aggregates with shallow joins and relational constraints dominate. |
| Document-shaped state | Connections must be navigated across many entity types with evolving paths. | The unit of access is an independent JSON-like aggregate with few cross-document traversals. |
| Search | Graph context materially improves lexical/semantic candidates. | Primary requirement is large-scale text relevance/faceting/log analytics; a search engine is the natural serving core. |
| Vector retrieval | Neighbors need graph policy/provenance/expansion around candidates. | The workload is predominantly ANN retrieval at vector-centric scale with little graph traversal. |
| Analytics | Graph algorithms/ML add validated signal. | Work is scan/aggregate-heavy BI; a columnar warehouse/lakehouse is usually a better execution engine. |
| Operational footprint | Team can operate Neo4j and benefits exceed licensing/topology cost. | Staffing, portability, compliance or cost requirements favor a simpler existing platform. |
6. Deliberately wrong approach: “capstone = enable every feature”
Adding clustering, CDC, composite databases, vector search, GraphRAG and GDS to look comprehensive increases failure modes, licensing surface, memory, migration work and operational skill requirements. It also makes causality impossible when performance changes because many variables moved together.
Repair: attach every feature to a requirement ID and an acceptance metric. If no requirement fails without it, keep it out of the base architecture. Verify the repaired design by checking that the core order/product workloads, recovery plan and security boundaries are complete even when optional features are disabled.
7. Hands-on lab: freeze the decision contract before touching data
Setup: create a Chapter 27 evidence folder and save the version/license manifest plus the requirement matrix from this lesson. No Neo4j process is required yet, so this lab cannot damage graph state.
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Evidence = Join-Path $Root "evidence"
New-Item -ItemType Directory -Force -Path $Evidence | Out-Null
@"
capstone=atlasmart-ch27
server=Neo4j 2026.07.1 Community
cypher=25 examples
driver=neo4j Python 6.3.0
rpo=DEFINE_ME
rto=DEFINE_ME
"@ | Set-Content -Encoding utf8 (Join-Path $Evidence 'architecture-baseline.txt')
Get-Content (Join-Path $Evidence 'architecture-baseline.txt')
Verification checklist
- Every requirement has an ID, owner, measurable acceptance test, workload shape and evidence artifact.
- RPO and RTO are numbers or explicit “not yet approved” gaps—not vague “high availability” language.
- Stable business IDs and source ownership are named.
- Community/Enterprise/Aura boundaries and optional features are explicit.
- At least one exit criterion names when another platform should replace or own a workload.
Cleanup/reset: if this design iteration is
discarded, remove only the temporary
architecture-baseline.txt; keep approved evidence
under version control or your normal documentation system. No
database cleanup is needed.
Production judgment
Before writing Cypher, quantify cardinality and degree, read/write ratios, concurrency, result size, latency distribution, RPO/RTO, trust boundaries, backup scope, driver timeouts/retries, memory headroom and staffing. Enterprise or Aura can change the availability/security/operability envelope but do not change the need for evidence. A design that cannot state its rollback path, version/license assumptions, known limits, and exit criteria is not production-ready.
Lesson 2 turns this contract into a small but complete, rebuildable AtlasMart graph with constraints, import reconciliation, core queries and a driver boundary.
Check your understanding
- Why is graph visual appeal insufficient evidence of graph fit?
- How do RPO and RTO differ?
- Why avoid Neo4j element IDs as durable business identity?
- Why is replication not a backup?
- When should an optional feature remain absent?
Review the answers
1. Because architecture fit depends on measured workload/query value, correctness, operability, cost and alternatives—not on whether the domain can be drawn as a graph.
2. RPO bounds acceptable recoverability/data-loss gap; RTO bounds time to restore useful service.
3. They are database-local implementation identifiers; stable domain keys are needed across imports, restores and migrations.
4. Replicas can copy corruption/deletion and share failure domains; an independently restorable artifact protects a different failure class.
5. When no measured requirement or acceptance test shows that it improves the system enough to justify its cost and failure surface.
Summary and next step
Define Requirements, Workloads, Availability/RPO/RTO, Data Sources, Security Boundaries, and Graph-Fit Criteria is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Build the Graph Model, Constraints/Indexes, Import Pipeline, Core Cypher Queries, and Application Driver Layer. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Current self-managed operational reference.
- System requirements — Supported Java, OS, memory, storage and filesystem requirements.
- Cypher compatibility and deprecations — Cypher 25 additions, compatibility and release-sensitive syntax.
- Constraints — Integrity constraints and backing-index semantics.
- Indexes — Search-performance and semantic index families.
- LOAD CSV — Transactional CSV import semantics.
- Execution plans — EXPLAIN/PROFILE and evidence-driven query tuning.
- Python driver manual — Official driver sessions, transactions, routing and application integration.
- Python driver performance — Driver-side performance, result handling and database selection.
- Authentication and authorization — Current authentication and edition-aware authorization model.
- Role-based access control — Enterprise/Aura RBAC and least-privilege controls.
- Offline database backup — Community offline dump semantics and backup boundaries.
- Restore a database dump — Community/Enterprise load and restore behavior.
- Online database backup — Enterprise online backup; not available on Aura.
- Neo4j clustering architecture — Enterprise primaries, secondaries, writer election and quorum.
- Neo4j logging — Operational logging surfaces.
- Neo4j metrics — Enterprise metrics surfaces and monitoring reference.
- Full-text indexes — Lexical search, analyzers and full-text query procedures.
- Vector indexes — Current vector-index lifecycle, ANN and score semantics.
- Cypher SEARCH — Cypher 25 vector SEARCH syntax introduced in Neo4j 2026.01.
- Graph Data Science manual — GDS 2026.07 graph projections, algorithms and ML.
- Supported GDS / Neo4j versions — Current GDS-to-Neo4j compatibility matrix.
- APOC installation — Current APOC/Neo4j version pairing and installation boundaries.
- Upgrade to Neo4j 2025–2026 — Supported upgrade paths and 5.26 LTS checkpoint behavior.
- Composite databases — Enterprise-only composite database boundary; unavailable on Aura.
- Built-in CDC procedures — Current db.cdc.* replacements and deprecated cdc.* procedure history.