Chapter 27 · Production Capstone: Model, Import, Query, Search, Analyze, Secure, Fail Over, and Operate Neo4j

Add Full-Text/Vector/GraphRAG and GDS Analytics Only Where Measured Requirements Justify Them

Add search, vector, GraphRAG and GDS only where measured quality or analytical requirements justify their memory, freshness, lifecycle, evaluation and operational cost.

Advanced330–450 minutesFull-text · vector · GraphRAG · GDS gateNeo4j 2026.07.1 · Community search pathCypher 25 SEARCH · deterministic embeddingsFull-text/vector ranks kept independentGDS 2026.07.0 only if justifiedLast reviewed: September 2026

Learning outcomes

01

Distinguish lexical, vector, graph-expansion, GraphRAG-context and GDS-analytics requirements instead of blending them into one “AI search” feature.

02

Build Community-reproducible full-text and LIST-backed vector retrieval with current index lifecycle checks and Cypher 25 SEARCH semantics.

03

Fuse lexical/vector candidates by independent ranks and verify graph/business filters rather than numerically mixing incomparable raw scores.

04

Assemble provenance-rich GraphRAG context without requiring a paid embedding or LLM service and state what retrieval evidence cannot guarantee.

05

Use a requirement/quality/cost gate to decide whether a GDS projection or ML pipeline adds enough validated signal to justify its memory and lifecycle.

Execution and safety note

Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.

1. AtlasMart problem: retrieval features are not architecture trophies

Core Cypher answers exact connected questions, but AtlasMart product discovery has a different requirement: users type words and concepts that may not match exact properties. Lexical retrieval scores token/text matches; vector retrieval searches approximate neighbors of an embedding; graph expansion applies relationship context after candidates exist; GraphRAG assembles graph-grounded retrieval context with provenance for a downstream generation workflow; GDS loads an in-memory projection for graph algorithms/ML. They have different freshness, memory, score and failure semantics.

Dimension Chapter 27 reproducible assumption
Neo4j 2026.07.1 Community for the mandatory capstone. Neo4j 5.26.30 remains the LTS comparison line. Enterprise/Aura-only material is isolated and labeled.
Cypher Cypher 25 examples. Cypher 5 remains a compatibility language; do not assume every existing database has the same default.
Java Neo4j 2026.07 supports Java 21 and Java 25. The official Docker image supplies its runtime; self-managed installs must use a supported JDK.
Driver Neo4j Python driver 6.3.0; Python 3.10–3.14. Use one long-lived driver object and short-lived sessions/managed transactions.
Database / auth Database neo4j; local user neo4j; synthetic password atlasmart-course-2026. Never reuse these lab credentials in production.
Network / TLS Loopback-only HTTP/Bolt for the disposable lab: 127.0.0.1:27474→7474 and 127.0.0.1:27687→7687. No TLS only because traffic stays on localhost; production/remote connections require a real TLS policy.
Plugins No APOC or GDS is required for the mandatory transactional/search/recovery path. If added, pin APOC 2026.07.1 and GDS 2026.07.0 to the 2026.07 server line.
Edition boundary Community provides the free single-instance learning path. Enterprise-only examples include clustering/true failover, online backup, fine-grained RBAC, composite databases and self-managed CDC. Aura has separate managed-tier boundaries.
Evidence rule This generated chapter does not execute your Docker host. Fixed fixture counts and deterministic calculations are expected invariants; latency, plans, DB Hits, resource counters, recovery time and index scores must be captured locally.
Continuity The mandatory search experiment uses only deterministic product text and precomputed LIST embeddings already imported in Lesson 2. No paid model/API is required.

2. Requirement gate

Capability Add only when… Evidence / stop condition
Full-text Exact terms, phrases, tokenization or typo-tolerant lexical retrieval matters. Judged query set improves versus exact property filters; index freshness acceptable.
Vector Semantic candidates materially improve judged recall beyond lexical-only retrieval. Recall@k/MRR and latency improve enough to pay embedding/index lifecycle cost.
Graph expansion Relationships enforce availability, category, ownership, provenance or contextual ranking. Post-filter quality improves without unacceptable fan-out/tail latency.
GraphRAG context Generated/support answers need explicit graph/document provenance. Context-retrieval evaluation improves; generation still separately evaluated.
GDS Batch structural analytics/ML answers a decision better than transactional Cypher/baseline. Ablation demonstrates graph-added value; memory/refresh/edition costs accepted.
Create and inspect a lexical product index
CREATE FULLTEXT INDEX cap_product_fulltext IF NOT EXISTS
FOR (p:Product)
ON EACH [p.name, p.description];

SHOW FULLTEXT INDEXES
YIELD name,state,populationPercent,labelsOrTypes,properties
WHERE name='cap_product_fulltext'
RETURN *;

// Query only after state='ONLINE'.
CALL db.index.fulltext.queryNodes('cap_product_fulltext','noise cancelling')
YIELD node,score
RETURN node.productId AS productId,node.name AS name,score
ORDER BY score DESC
LIMIT 5;
Create a LIST-backed vector index and query with Cypher 25 SEARCH
CREATE VECTOR INDEX cap_product_vector IF NOT EXISTS
FOR (p:Product) ON p.embedding
OPTIONS {indexConfig:{
  `vector.dimensions`:4,
  `vector.similarity_function`:'cosine'
}};

SHOW VECTOR INDEXES
YIELD name,state,populationPercent,labelsOrTypes,properties
WHERE name='cap_product_vector'
RETURN *;

// Neo4j 2026.01+, Cypher 25 preferred query form.
MATCH (p:Product)
SEARCH p IN (
  VECTOR INDEX cap_product_vector
  FOR [0.05,0.96,0.03,0.01]
  LIMIT 4
) SCORE AS vectorScore
RETURN p.productId AS productId,p.name AS name,vectorScore
ORDER BY vectorScore DESC;

3. Score discipline and independent rank fusion

Full-text and vector scores come from different scoring systems. A raw value of 0.9 from one source is not inherently “better” than 0.7 from another. Rank each source independently, then fuse ranks (for example reciprocal-rank fusion) or train a calibrated re-ranker against judgments. Keep source ranks/scores in the evidence record so the result remains explainable.

Fuse ranks without comparing raw score scales
def reciprocal_rank_fusion(rankings, k=60):
    # rankings = list of ordered product-id lists from independent retrieval systems
    score = {}
    for ranking in rankings:
        for rank, product_id in enumerate(ranking, start=1):
            score[product_id] = score.get(product_id, 0.0) + 1.0 / (k + rank)
    return sorted(score.items(), key=lambda x: (-x[1], x[0]))

lexical = ["P-2001", "P-2002"]
vector  = ["P-2001", "P-2002", "P-1002", "P-1001"]
print(reciprocal_rank_fusion([lexical, vector]))

# The input lists are illustrative until you capture them from your local indexes.
# The fusion arithmetic itself is deterministic.
Add graph/business context after candidate retrieval
// Assume $candidateIds comes from a lexical/vector retrieval step.
UNWIND $candidateIds AS productId
MATCH (p:Product {productId:productId})-[:IN_CATEGORY]->(cat:Category)
OPTIONAL MATCH (p)<-[:CONTAINS]-(o:Order)<-[:PLACED]-(c:Customer {customerId:$customerId})
WITH p,cat,count(DISTINCT o) AS priorOrders
RETURN p.productId AS productId,
       p.name AS name,
       cat.categoryId AS categoryId,
       priorOrders
ORDER BY priorOrders DESC, productId;

// Graph context is a business/filter signal. It does not convert retrieval similarity into truth.

4. GraphRAG without a paid LLM

The learning objective is retrieval/context design, not API consumption. Build a context record containing the query, retrieved product IDs, category/relationship evidence, source text, timestamp/version, and retrieval method. A downstream LLM may consume it later, but the capstone can evaluate retrieval coverage and provenance without generating text. GraphRAG can reduce unsupported context selection; it cannot guarantee that a language model will never hallucinate.

Create a provenance-rich context envelope
{
  "requestId": "CAP-SEARCH-001",
  "query": "quiet wireless audio for commuting",
  "retrieval": {
    "lexicalIndex": "cap_product_fulltext",
    "vectorIndex": "cap_product_vector",
    "fusion": "reciprocal-rank",
    "candidates": ["P-2001", "P-2002"]
  },
  "graphContext": {
    "categoryIds": ["CAT-AUDIO"],
    "relationshipTypes": ["IN_CATEGORY", "CONTAINS", "PLACED"]
  },
  "provenance": {
    "database": "neo4j",
    "server": "2026.07.1",
    "embeddingFixture": "chapter27-precomputed-v1"
  }
}

5. GDS gate: prove graph-added value before loading an in-memory projection

With six products, a Cypher co-purchase baseline is transparent and cheap. Loading GDS only to say “we used graph algorithms” adds memory/lifecycle without evidence. If AtlasMart later has millions of products/interactions and a recommendation-quality target that transactional traversal cannot meet, pin GDS 2026.07.0, estimate projection memory, run a bounded algorithm/ML experiment, compare against a non-graph baseline, and document refresh/serving boundaries. Community GDS is sufficient for algorithm learning but has concurrency/model-catalog limits; Enterprise adds operational capabilities.

6. Deliberately wrong approach: raw-score soup

Summing a Lucene score, vector score and graph count directly creates a number with no common scale or calibration. It can flip rankings after an analyzer/model/index change without any business-quality improvement.

Repair: preserve each signal independently, normalize through ranks or a judged re-ranker, evaluate on a frozen query/judgment set, and version the embedding/analyzer/index configuration. Verify the repair with recall@k/MRR (or task-appropriate metrics), latency distributions, freshness tests and provenance traces.

7. Hands-on lab: prove whether retrieval/analytics earns its complexity

Setup: use the Lesson 2 graph. Create the full-text and vector indexes, wait for ONLINE, freeze three to five judged catalog queries, and capture lexical/vector candidate lists separately before any graph expansion.

Verification checklist

  • SHOW FULLTEXT INDEXES and SHOW VECTOR INDEXES report the named indexes ONLINE.
  • The deterministic audio query ranks P-2001/P-2002 among the relevant candidates; record actual local scores without treating them as probabilities.
  • Lexical and vector rankings remain separate inputs to the rank-fusion function; no raw-score sum is used.
  • Graph expansion records the relationship types and stable IDs that affected the final candidates.
  • A short decision note says “add”, “defer”, or “reject” GDS/GraphRAG based on measured quality, latency, memory/lifecycle and provenance requirements—not feature enthusiasm.

Cleanup/reset: if testing the retrieval layer standalone, use DROP INDEX cap_product_fulltext IF EXISTS and DROP INDEX cap_product_vector IF EXISTS; the underlying Product properties remain. Recreate the indexes before later tests that require them.

Production judgment

Every retrieval layer adds write amplification, storage/memory, refresh/freshness decisions, security filtering and migration cost. Embeddings also add an external model lifecycle. Keep business rules in explicit graph/policy logic and treat retrieval scores as ranking evidence only. If a dedicated search/vector platform wins the judged quality/latency/cost envelope and the graph adds little context, use that platform rather than forcing all retrieval into Neo4j.

Lesson 4 takes the chosen feature set through load, security, backup/restore, outage recovery, monitoring and upgrade/incident evidence.

Check your understanding

  1. Why not compare raw full-text and vector scores numerically?
  2. What must be true before a vector index serves queries?
  3. Why are the chapter embeddings called deterministic fixtures?
  4. What does GraphRAG provenance prove?
  5. When should GDS remain out of the base capstone?
Review the answers

1. They are produced by different scoring systems/scales; fuse independent ranks or calibrate against judgments instead.

2. Its state must be ONLINE; POPULATING indexes are not yet usable.

3. They are precomputed teaching vectors, not outputs from a claimed model/API, so the lab is free and reproducible.

4. It records which sources/graph context were retrieved; it does not prove a downstream generated answer is correct.

5. When a simpler Cypher/tabular baseline meets the requirement and no measured graph-analytics experiment shows enough added value.

Summary and next step

Add Full-Text/Vector/GraphRAG and GDS Analytics Only Where Measured Requirements Justify Them is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Load-Test, Profile, Secure, Back Up, Restore, Fail Over, Monitor, Upgrade, and Execute Incident Runbooks. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.