Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j

Define Requirements, Workloads, Availability/RPO/RTO, Data Sources, Security Boundaries, and Graph-Fit Criteria

Turn AtlasMart goals into a falsifiable architecture contract before building the final graph: measurable workloads, SLO/RPO/RTO, source and identity ownership, security boundaries, graph-fit criteria, risks, and exit rules.

Advanced300–420 minutesRequirements · SLO/RPO/RTO · graph fitNeo4j 2026.07.1 · Community mandatoryCypher 25 · Python driver 6.3.0Architecture evidence · ADR · exit criteriaJava 21/25 · optional features gatedLast reviewed: September 2026

Learning outcomes

01

Translate AtlasMart business goals into measurable graph workloads, SLOs, availability targets, RPO/RTO, and acceptance evidence before choosing topology or features.

02

Define source-of-truth and stable identity contracts so the graph can be rebuilt, reconciled, restored, or migrated without depending on Neo4j internal IDs.

03

Draw security and tenant boundaries explicitly, including what Community can prove locally and what requires Enterprise/Aura controls.

04

Use a graph-fit scorecard that compares connected-query value against relational, document, search, vector, and analytical alternatives.

05

Create an architecture decision record, known-limit register, version/license manifest, and exit criteria that later lessons can verify.

Execution and safety note

Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.

1. AtlasMart problem: “use Neo4j” is not yet an architecture

AtlasMart wants one system to answer customer-order-product traversals, recommendations, catalog search, support context, and operational questions. A graph can make relationship-centric questions concise, but visual connectedness is not proof of fit. The capstone therefore begins with an architecture fitness claim: a falsifiable statement that a chosen property-graph model and Neo4j operational envelope satisfy named workloads better enough to justify their complexity and cost.

A service-level objective (SLO) is a measurable reliability/performance target, such as a p95 latency or monthly availability objective. RPO is the maximum acceptable recoverability gap after data loss; RTO is the time allowed to restore useful service. An availability target is not the same as a backup target: Enterprise replication can improve availability, while independently restorable backups protect against corruption, deletion, operator error and other failures that replication can copy.

Dimension Chapter 27 reproducible assumption
Neo4j 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled.
Cypher Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default.
Java Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK.
Driver Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions.
Database / auth Database neo4j; local user neo4j; synthetic password atlasmart-course-2026. Never reuse these lab credentials in production.
Network / TLS Loopback-only HTTP/Bolt for the disposable lab: 127.0.0.1:27474→7474 and 127.0.0.1:27687→7687. No TLS only because traffic stays on localhost; production/remote connections require a real TLS policy.
Plugins No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line.
Edition boundary Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries.
Evidence rule This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally.
Continuity The capstone reuses stable IDs such as C-1001, P-1001 and O-5001 and introduces only capstone-tagged data so cleanup is bounded.

2. Start with a traceability matrix, not a feature checklist

Each requirement needs an owner, a measurable test, a workload shape, and an architectural consequence. This prevents the capstone from adding vector search, GDS, clustering, CDC, or composite databases merely because they appeared earlier in the course. A feature is justified only if its evidence closes a requirement gap more economically or reliably than a simpler mechanism.

ID Requirement / workload Evidence to collect Design consequence
R1 Customer order history: keyed customer → orders → products; bounded result. Correct IDs/items, EXPLAIN/PROFILE access path, p95/p99 under documented concurrency. Stable keys + uniqueness constraints; parameterized 2-hop Cypher; bounded projection.
R2 Related-product discovery from shared orders/categories. Candidate relevance on a judged fixture, traversal fan-out, latency distribution. Graph relationship traversal is a core fit criterion; precompute only if measured.
R3 Catalog lexical/semantic retrieval. Judged queries, recall@k/MRR, freshness, latency, provenance. Full-text/vector optional; scores remain retrieval evidence, not business truth.
R4 Recoverability. Offline Community dump hash, isolated load, canary/count reconciliation, measured recovery steps. Replication never substitutes for backup; restore testing is mandatory.
R5 HA / node loss. Community: process outage/restart evidence + deterministic quorum trace. Enterprise option: real cluster writer/quorum test. True cluster failover is an Enterprise design/cost decision.
R6 Security boundary. Auth test, network exposure review, secret handling; optional Enterprise deny/allow tests. Community does not reproduce fine-grained RBAC; do not claim least privilege it cannot enforce.
Write the capstone requirements and version manifest
CAPSTONE_ID=atlasmart-ch27
SERVER=Neo4j 2026.07.1 Community
LTS_REFERENCE=Neo4j 5.26.30
CYPHER=25 for examples; verify database setting before deployment
JAVA=21 or 25 for Neo4j 2026.07 self-managed; Docker image runtime is pinned by image
DRIVER=neo4j Python 6.3.0
DATABASE=neo4j
BOLT=bolt://localhost:27687 (local disposable lab only)
AUTH=synthetic local native credentials; never production credentials
TLS=off only on loopback lab; production policy must be explicit
APOC=not required; if used pin 2026.07.1
GDS=not required by default; if justified pin 2026.07.0
ENTERPRISE_ONLY=cluster HA, online backup, fine-grained RBAC, composite DB, self-managed CDC
RPO_TARGET=define before backup design
RTO_TARGET=define before recovery design
EXIT_CRITERIA=document conditions that favor relational/search/document/vector/warehouse alternatives

3. Source-of-truth and identity model

The graph must be reconstructible from durable domain identifiers. AtlasMart therefore treats customerId, productId, orderId, and categoryId as cross-system identities and backs them with constraints. Neo4j element IDs are useful within a database interaction but are not the contract for rebuilds, restores, federation or migration. The source system must define whether an event is authoritative, a correction, a tombstone, or a derived projection.

For every relationship ask who owns it and how it is refreshed. PLACED belongs to order ownership; CONTAINS belongs to order-line facts; IN_CATEGORY belongs to catalog governance. A recommendation edge would be derived and should carry lineage/freshness rather than masquerading as source truth.

4. Security boundary is an architecture input

A trust boundary separates components that may safely trust one another from components that require authentication, authorization, encryption, or validation. The local capstone binds ports to loopback and uses synthetic credentials, so disabling TLS there does not imply that plaintext Bolt is acceptable across a production network. Community supports native users, but current fine-grained role/privilege controls are an Enterprise/Aura capability; therefore the mandatory lab records that limitation instead of inventing a least-privilege role.

5. Graph-fit criteria and explicit alternatives

Question Neo4j strengthens the case when… Another platform strengthens the case when…
Connected operational reads Traversal depth and relationship semantics are central and queried online. Queries are mostly keyed rows/aggregates with shallow joins and relational constraints dominate.
Document-shaped state Connections must be navigated across many entity types with evolving paths. The unit of access is an independent JSON-like aggregate with few cross-document traversals.
Search Graph context materially improves lexical/semantic candidates. Primary requirement is large-scale text relevance/faceting/log analytics; a search engine is the natural serving core.
Vector retrieval Neighbors need graph policy/provenance/expansion around candidates. The workload is predominantly ANN retrieval at vector-centric scale with little graph traversal.
Analytics Graph algorithms/ML add validated signal. Work is scan/aggregate-heavy BI; a columnar warehouse/lakehouse is usually a better execution engine.
Operational footprint Team can operate Neo4j and benefits exceed licensing/topology cost. Staffing, portability, compliance or cost requirements favor a simpler existing platform.

6. Deliberately wrong approach: “capstone = enable every feature”

Adding clustering, CDC, composite databases, vector search, GraphRAG and GDS to look comprehensive increases failure modes, licensing surface, memory, migration work and operational skill requirements. It also makes causality impossible when performance changes because many variables moved together.

Repair: attach every feature to a requirement ID and an acceptance metric. If no requirement fails without it, keep it out of the base architecture. Verify the repaired design by checking that the core order/product workloads, recovery plan and security boundaries are complete even when optional features are disabled.

7. Hands-on lab: freeze the decision contract before touching data

Setup: create a Chapter 27 evidence folder and save the version/license manifest plus the requirement matrix from this lesson. No Neo4j process is required yet, so this lab cannot damage graph state.

Create the capstone evidence workspace
$Root = Join-Path $env:TEMP "atlasmart-neo4j-capstone"
$Evidence = Join-Path $Root "evidence"
New-Item -ItemType Directory -Force -Path $Evidence | Out-Null
@"
capstone=atlasmart-ch27
server=Neo4j 2026.07.1 Community
cypher=25 examples
driver=neo4j Python 6.3.0
rpo=DEFINE_ME
rto=DEFINE_ME
"@ | Set-Content -Encoding utf8 (Join-Path $Evidence 'architecture-baseline.txt')
Get-Content (Join-Path $Evidence 'architecture-baseline.txt')

Verification checklist

  • Every requirement has an ID, owner, measurable acceptance test, workload shape and evidence artifact.
  • RPO and RTO are numbers or explicit “not yet approved” gaps—not vague “high availability” language.
  • Stable business IDs and source ownership are named.
  • Community/Enterprise/Aura boundaries and optional features are explicit.
  • At least one exit criterion names when another platform should replace or own a workload.

Cleanup/reset: if this design iteration is discarded, remove only the temporary architecture-baseline.txt; keep approved evidence under version control or your normal documentation system. No database cleanup is needed.

Production judgment

Before writing Cypher, quantify cardinality and degree, read/write ratios, concurrency, result size, latency distribution, RPO/RTO, trust boundaries, backup scope, driver timeouts/retries, memory headroom and staffing. Enterprise or Aura can change the availability/security/operability envelope but do not change the need for evidence. A design that cannot state its rollback path, version/license assumptions, known limits, and exit criteria is not production-ready.

Lesson 2 turns this contract into a small but complete, rebuildable AtlasMart graph with constraints, import reconciliation, core queries and a driver boundary.

Check your understanding

  1. Why is graph visual appeal insufficient evidence of graph fit?
  2. How do RPO and RTO differ?
  3. Why avoid Neo4j element IDs as durable business identity?
  4. Why is replication not a backup?
  5. When should an optional feature remain absent?
Review the answers

1. Because architecture fit depends on measured workload/query value, correctness, operability, cost and alternatives—not on whether the domain can be drawn as a graph.

2. RPO bounds acceptable recoverability/data-loss gap; RTO bounds time to restore useful service.

3. They are database-local implementation identifiers; stable domain keys are needed across imports, restores and migrations.

4. Replicas can copy corruption/deletion and share failure domains; an independently restorable artifact protects a different failure class.

5. When no measured requirement or acceptance test shows that it improves the system enough to justify its cost and failure surface.

Summary and next step

Define Requirements, Workloads, Availability/RPO/RTO, Data Sources, Security Boundaries, and Graph-Fit Criteria is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Build the Graph Model, Constraints/Indexes, Import Pipeline, Core Cypher Queries, and Application Driver Layer. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.