Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j
Add Full-Text/Vector/GraphRAG and GDS Analytics Only Where Measured Requirements Justify Them
Add search, vector, GraphRAG and GDS only where measured quality or analytical requirements justify their memory, freshness, lifecycle, evaluation and operational cost.
Learning outcomes
Distinguish lexical, vector, graph-expansion, GraphRAG-context and GDS-analytics requirements instead of blending them into one “AI search” feature.
Build Community-reproducible full-text and LIST-backed vector retrieval with current index lifecycle checks and Cypher 25 SEARCH semantics.
Fuse lexical/vector candidates by independent ranks and verify graph/business filters rather than numerically mixing incomparable raw scores.
Assemble provenance-rich GraphRAG context without requiring a paid embedding or LLM service and state what retrieval evidence cannot guarantee.
Use a requirement/quality/cost gate to decide whether a GDS projection or ML pipeline adds enough validated signal to justify its memory and lifecycle.
Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.
1. AtlasMart problem: retrieval features are not architecture trophies
Core Cypher answers exact connected questions, but AtlasMart product discovery has a different requirement: users type words and concepts that may not match exact properties. Lexical retrieval scores token/text matches; vector retrieval searches approximate neighbors of an embedding; graph expansion applies relationship context after candidates exist; GraphRAG assembles graph-grounded retrieval context with provenance for a downstream generation workflow; GDS loads an in-memory projection for graph algorithms/ML. They have different freshness, memory, score and failure semantics.
| Dimension | Chapter 27 reproducible assumption |
|---|---|
| Neo4j | 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled. |
| Cypher | Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default. |
| Java | Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK. |
| Driver | Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions. |
| Database / auth |
Database neo4j; local user
neo4j; synthetic password
atlasmart-course-2026. Never reuse these lab
credentials in production.
|
| Network / TLS |
Loopback-only HTTP/Bolt for the disposable lab:
127.0.0.1:27474→7474 and
127.0.0.1:27687→7687. No TLS only because
traffic stays on localhost; production/remote connections
require a real TLS policy.
|
| Plugins | No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line. |
| Edition boundary | Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries. |
| Evidence rule | This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally. |
| Continuity |
The mandatory search experiment uses only deterministic
product text and precomputed LIST |
2. Requirement gate
| Capability | Add only when… | Evidence / stop condition |
|---|---|---|
| Full-text | Exact terms, phrases, tokenization or typo-tolerant lexical retrieval matters. | Judged query set improves versus exact property filters; index freshness acceptable. |
| Vector | Semantic candidates materially improve judged recall beyond lexical-only retrieval. | Recall@k/MRR and latency improve enough to pay embedding/index lifecycle cost. |
| Graph expansion | Relationships enforce availability, category, ownership, provenance or contextual ranking. | Post-filter quality improves without unacceptable fan-out/tail latency. |
| GraphRAG context | Generated/support answers need explicit graph/document provenance. | Context-retrieval evaluation improves; generation still separately evaluated. |
| GDS | Batch structural analytics/ML answers a decision better than transactional Cypher/baseline. | Ablation demonstrates graph-added value; memory/refresh/edition costs accepted. |
CREATE FULLTEXT INDEX cap_product_fulltext IF NOT EXISTS
FOR (p:Product)
ON EACH [p.name, p.description];
SHOW FULLTEXT INDEXES
YIELD name,state,populationPercent,labelsOrTypes,properties
WHERE name='cap_product_fulltext'
RETURN *;
// Query only after state='ONLINE'.
CALL db.index.fulltext.queryNodes('cap_product_fulltext','noise cancelling')
YIELD node,score
RETURN node.productId AS productId,node.name AS name,score
ORDER BY score DESC
LIMIT 5;
CREATE VECTOR INDEX cap_product_vector IF NOT EXISTS
FOR (p:Product) ON p.embedding
OPTIONS {indexConfig:{
`vector.dimensions`:4,
`vector.similarity_function`:'cosine'
}};
SHOW VECTOR INDEXES
YIELD name,state,populationPercent,labelsOrTypes,properties
WHERE name='cap_product_vector'
RETURN *;
// Neo4j 2026.01+, Cypher 25 preferred query form.
MATCH (p:Product)
SEARCH p IN (
VECTOR INDEX cap_product_vector
FOR [0.05,0.96,0.03,0.01]
LIMIT 4
) SCORE AS vectorScore
RETURN p.productId AS productId,p.name AS name,vectorScore
ORDER BY vectorScore DESC;
3. Score discipline and independent rank fusion
Full-text and vector scores come from different scoring systems. A raw value of 0.9 from one source is not inherently “better” than 0.7 from another. Rank each source independently, then fuse ranks (for example reciprocal-rank fusion) or train a calibrated re-ranker against judgments. Keep source ranks/scores in the evidence record so the result remains explainable.
def reciprocal_rank_fusion(rankings, k=60):
# rankings = list of ordered product-id lists from independent retrieval systems
score = {}
for ranking in rankings:
for rank, product_id in enumerate(ranking, start=1):
score[product_id] = score.get(product_id, 0.0) + 1.0 / (k + rank)
return sorted(score.items(), key=lambda x: (-x[1], x[0]))
lexical = ["P-2001", "P-2002"]
vector = ["P-2001", "P-2002", "P-1002", "P-1001"]
print(reciprocal_rank_fusion([lexical, vector]))
# The input lists are illustrative until you capture them from your local indexes.
# The fusion arithmetic itself is deterministic.
// Assume $candidateIds comes from a lexical/vector retrieval step.
UNWIND $candidateIds AS productId
MATCH (p:Product {productId:productId})-[:IN_CATEGORY]->(cat:Category)
OPTIONAL MATCH (p)<-[:CONTAINS]-(o:Order)<-[:PLACED]-(c:Customer {customerId:$customerId})
WITH p,cat,count(DISTINCT o) AS priorOrders
RETURN p.productId AS productId,
p.name AS name,
cat.categoryId AS categoryId,
priorOrders
ORDER BY priorOrders DESC, productId;
// Graph context is a business/filter signal. It does not convert retrieval similarity into truth.
4. GraphRAG without a paid LLM
The learning objective is retrieval/context design, not API consumption. Build a context record containing the query, retrieved product IDs, category/relationship evidence, source text, timestamp/version, and retrieval method. A downstream LLM may consume it later, but the capstone can evaluate retrieval coverage and provenance without generating text. GraphRAG can reduce unsupported context selection; it cannot guarantee that a language model will never hallucinate.
{
"requestId": "CAP-SEARCH-001",
"query": "quiet wireless audio for commuting",
"retrieval": {
"lexicalIndex": "cap_product_fulltext",
"vectorIndex": "cap_product_vector",
"fusion": "reciprocal-rank",
"candidates": ["P-2001", "P-2002"]
},
"graphContext": {
"categoryIds": ["CAT-AUDIO"],
"relationshipTypes": ["IN_CATEGORY", "CONTAINS", "PLACED"]
},
"provenance": {
"database": "neo4j",
"server": "2026.07.1",
"embeddingFixture": "chapter27-precomputed-v1"
}
}
5. GDS gate: prove graph-added value before loading an in-memory projection
With six products, a Cypher co-purchase baseline is transparent and cheap. Loading GDS only to say “we used graph algorithms” adds memory/lifecycle without evidence. If AtlasMart later has millions of products/interactions and a recommendation-quality target that transactional traversal cannot meet, pin GDS 2026.07.0, estimate projection memory, run a bounded algorithm/ML experiment, compare against a non-graph baseline, and document refresh/serving boundaries. Community GDS is sufficient for algorithm learning but has concurrency/model-catalog limits; Enterprise adds operational capabilities.
6. Deliberately wrong approach: raw-score soup
Summing a Lucene score, vector score and graph count directly creates a number with no common scale or calibration. It can flip rankings after an analyzer/model/index change without any business-quality improvement.
Repair: preserve each signal independently, normalize through ranks or a judged re-ranker, evaluate on a frozen query/judgment set, and version the embedding/analyzer/index configuration. Verify the repair with recall@k/MRR (or task-appropriate metrics), latency distributions, freshness tests and provenance traces.
7. Hands-on lab: prove whether retrieval/analytics earns its complexity
Setup: use the Lesson 2 graph. Create the
full-text and vector indexes, wait for ONLINE,
freeze three to five judged catalog queries, and capture
lexical/vector candidate lists separately before any graph
expansion.
Verification checklist
-
SHOW FULLTEXT INDEXESandSHOW VECTOR INDEXESreport the named indexesONLINE. - The deterministic audio query ranks P-2001/P-2002 among the relevant candidates; record actual local scores without treating them as probabilities.
- Lexical and vector rankings remain separate inputs to the rank-fusion function; no raw-score sum is used.
- Graph expansion records the relationship types and stable IDs that affected the final candidates.
- A short decision note says “add”, “defer”, or “reject” GDS/GraphRAG based on measured quality, latency, memory/lifecycle and provenance requirements—not feature enthusiasm.
Cleanup/reset: if testing the retrieval layer
standalone, use
DROP INDEX cap_product_fulltext IF EXISTS and
DROP INDEX cap_product_vector IF EXISTS; the
underlying Product properties remain. Recreate the indexes
before later tests that require them.
Production judgment
Every retrieval layer adds write amplification, storage/memory, refresh/freshness decisions, security filtering and migration cost. Embeddings also add an external model lifecycle. Keep business rules in explicit graph/policy logic and treat retrieval scores as ranking evidence only. If a dedicated search/vector platform wins the judged quality/latency/cost envelope and the graph adds little context, use that platform rather than forcing all retrieval into Neo4j.
Lesson 4 takes the chosen feature set through load, security, backup/restore, outage recovery, monitoring and upgrade/incident evidence.
Check your understanding
- Why not compare raw full-text and vector scores numerically?
- What must be true before a vector index serves queries?
- Why are the chapter embeddings called deterministic fixtures?
- What does GraphRAG provenance prove?
- When should GDS remain out of the base capstone?
Review the answers
1. They are produced by different scoring systems/scales; fuse independent ranks or calibrate against judgments instead.
2. Its state must be ONLINE; POPULATING indexes are not yet usable.
3. They are precomputed teaching vectors, not outputs from a claimed model/API, so the lab is free and reproducible.
4. It records which sources/graph context were retrieved; it does not prove a downstream generated answer is correct.
5. When a simpler Cypher/tabular baseline meets the requirement and no measured graph-analytics experiment shows enough added value.
Summary and next step
Add Full-Text/Vector/GraphRAG and GDS Analytics Only Where Measured Requirements Justify Them is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Load-Test, Profile, Secure, Back Up, Restore, Fail Over, Monitor, Upgrade, and Execute Incident Runbooks. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Current self-managed operational reference.
- System requirements — Supported Java, OS, memory, storage and filesystem requirements.
- Cypher compatibility and deprecations — Cypher 25 additions, compatibility and release-sensitive syntax.
- Constraints — Integrity constraints and backing-index semantics.
- Indexes — Search-performance and semantic index families.
- LOAD CSV — Transactional CSV import semantics.
- Execution plans — EXPLAIN/PROFILE and evidence-driven query tuning.
- Python driver manual — Official driver sessions, transactions, routing and application integration.
- Python driver performance — Driver-side performance, result handling and database selection.
- Authentication and authorization — Current authentication and edition-aware authorization model.
- Role-based access control — Enterprise/Aura RBAC and least-privilege controls.
- Offline database backup — Community offline dump semantics and backup boundaries.
- Restore a database dump — Community/Enterprise load and restore behavior.
- Online database backup — Enterprise online backup; not available on Aura.
- Neo4j clustering architecture — Enterprise primaries, secondaries, writer election and quorum.
- Neo4j logging — Operational logging surfaces.
- Neo4j metrics — Enterprise metrics surfaces and monitoring reference.
- Full-text indexes — Lexical search, analyzers and full-text query procedures.
- Vector indexes — Current vector-index lifecycle, ANN and score semantics.
- Cypher SEARCH — Cypher 25 vector SEARCH syntax introduced in Neo4j 2026.01.
- Graph Data Science manual — GDS 2026.07 graph projections, algorithms and ML.
- Supported GDS / Neo4j versions — Current GDS-to-Neo4j compatibility matrix.
- APOC installation — Current APOC/Neo4j version pairing and installation boundaries.
- Upgrade to Neo4j 2025–2026 — Supported upgrade paths and 5.26 LTS checkpoint behavior.
- Composite databases — Enterprise-only composite database boundary; unavailable on Aura.
- Built-in CDC procedures — Current db.cdc.* replacements and deprecated cdc.* procedure history.