Chapter 22 · Vector Search, Embeddings, Cypher SEARCH, Hybrid Search, and GraphRAG

Design GraphRAG: Chunk/Entity/Relationship Modeling, Retrieval, Context Assembly, Provenance, and Evaluation

Model GraphRAG retrieval with Document/Chunk/entity provenance, retrieve deterministic chunks, expand through graph relationships, assemble bounded evidence-rich context, and evaluate retrieval/grounding without requiring a paid embedding or LLM API.

Advanced270–380 minutesGraphRAG provenance labNeo4j 2026.07.1 · Community mandatoryCypher 25 · Vector SEARCH · HNSW/ANNLIST embeddings mandatory · VECTOR storage optional EE/AuraJava 21/25 · Python driver 6.3 optionalLast reviewed: September 2026

GraphRAG is often described as “vector search plus a graph plus an LLM,” which hides the engineering contract. AtlasMart’s question “Which camera should I use for wildlife on remote trails at night, and why?” requires retrieval, provenance, entity resolution, relationship expansion, context budgeting and evaluation before any text generator is involved. The mandatory lab stops at a deterministic evidence packet so it remains free/local and so retrieval quality can be tested independently from generation quality.

Mental model

Retrieval chooses evidence; the graph connects evidence to entities/facts; provenance says where evidence came from; a generator—if you add one—summarizes that bounded evidence. GraphRAG can improve grounding opportunities, but it does not eliminate hallucination, authorization, citation, or evaluation requirements.

Learning outcomes

01

Model Document/Chunk nodes, NEXT_CHUNK/FROM_DOCUMENT provenance and MENTIONS edges to AtlasMart entities without requiring an LLM extraction pipeline.

02

Retrieve Chunk candidates with current vector SEARCH, preserve vector evidence, and expand to source documents/entities/graph facts.

03

Assemble bounded context packets with chunk/source/entity provenance before any optional generator is called.

04

Define retrieval, provenance, grounding/citation and answer-quality evaluation separately and avoid claiming GraphRAG eliminates hallucination.

05

Design embedding/model/index version upgrades, privacy/security filters, failure injection and rollback for a production GraphRAG service.

Chapter 22 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; the current 5.26 line remains LTS. Version-sensitive examples use explicit CYPHER 25. The mandatory lab uses self-managed Neo4j Community 2026.07.1, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback Bolt 7687 and HTTP 7474, and embeddings stored as LIST<FLOAT>. Neo4j 2026.x supports Java 21/25. Optional client examples pin the official Python driver to neo4j==6.3.0. No APOC, GDS, paid embedding API, paid LLM API, Aura account, or Enterprise license is required.

Community VECTOR/LIST boundary

Vector indexes are available in Community when embeddings are stored as LIST<INTEGER|FLOAT>. The newer fixed-size VECTOR property type requires block-format storage and therefore cannot be persisted as a property in Community; it is an Enterprise/Aura storage capability. The lab deliberately uses LIST embeddings so every mandatory index/search/evaluation step remains free/local. Where VECTOR-specific storage is discussed, it is labeled as an edition-dependent optimization/typing choice rather than a prerequisite.

Current query surface

From Neo4j 2026.01, Cypher 25 SEARCH is the preferred way to query vector indexes and supports in-index filtering when filter properties were declared in the index. db.index.vector.queryNodes() and db.index.vector.queryRelationships() remain useful for older-version compatibility history but are deprecated from Neo4j 2026.04. New course code therefore uses SEARCH.

Lab contract and exact assumptions

Dimension Chapter 22 assumption
server Neo4j Community 2026.07.1, single disposable local database
Cypher Explicit CYPHER 25 for SEARCH and current vector syntax
Java Java 21 or 25 for Neo4j 2026.07
database/auth neo4j / neo4j / atlasmart-course-2026
transport bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; production/remote deployments use verified TLS
plugins none required; APOC/GDS/GenAI are not needed
embedding source deterministic 8-dimensional precomputed AtlasMart vectors; not a paid API and not claimed to be production-quality embeddings
storage LIST so Community can store/index every embedding; VECTOR storage is discussed as Enterprise/Aura-specific
graph 8 Products, 4 Categories, 1 Store, 2 KnowledgeDocuments, 4 Chunks plus provenance/entity edges
indexes full-text product index + 8D product vector index + 8D chunk vector index
measurement learner measures recall@k, runtime latency, index state/options and result IDs; generated lesson never claims that Neo4j was executed here
Term Mechanism-first meaning
embedding Numeric representation produced outside the database by a model or deterministic encoder. Neo4j stores/indexes the values; it does not make semantic truth guarantees about the encoder.
dimension Number of coordinates in an embedding. Index dimension and query-vector dimension must match when dimensions are configured.
LIST embedding Community-compatible numeric property such as [0.95,0.85,...]. Individual elements are list-accessible.
VECTOR value Fixed-length typed vector value introduced in 2025.10; more storage-efficient typing but persisted VECTOR properties require Enterprise/Aura block format.
similarity Function that converts a pair of vectors into an ordering signal. Current vector indexes support cosine and euclidean similarity.
ANN Approximate nearest-neighbor retrieval. It trades guaranteed exactness for scalable search speed/resource behavior.
HNSW Hierarchical Navigable Small World graph used internally by the vector index to navigate candidate neighborhoods rather than compare every stored vector.
recall@k Fraction of the exact top-k neighbors recovered by ANN top-k. It is a retrieval-quality measure, not semantic correctness.
filter property Non-vector property explicitly stored with a 2026.01+ vector index so SEARCH can apply supported predicates inside the ANN search.
quantization Compressed vector representation used inside the index to lower memory/storage and often improve speed, potentially trading accuracy; 2026.07 supports high-fidelity rescoring through search expansion.
GraphRAG Retrieval-augmented generation pattern where graph-structured evidence, provenance, and relationships enrich the context given to a generator. Retrieval quality and generator factuality still require evaluation.

Direct-entry setup

The Chapter 22 fixture already includes two KnowledgeDocument nodes and four Chunk nodes with deterministic embeddings, FROM_DOCUMENT/NEXT_CHUNK/MENTIONS relationships and a chunk vector index. This mirrors the lexical-graph concepts in the official GraphRAG tooling without requiring a paid LLM/entity extractor.

Cypher 25 · setup
CYPHER 25
// Disposable Chapter 22 fixture. Safe to rerun after the cleanup block.
CREATE CONSTRAINT ch22_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch22_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch22_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch22_doc_id IF NOT EXISTS
FOR (d:KnowledgeDocument) REQUIRE d.documentId IS UNIQUE;
CREATE CONSTRAINT ch22_chunk_id IF NOT EXISTS
FOR (c:Chunk) REQUIRE c.chunkId IS UNIQUE;

MERGE (cam:Category {categoryId:'CAT-22-CAM'}) SET cam.name='Cameras', cam.labTag='ch22'
MERGE (out:Category {categoryId:'CAT-22-OUT'}) SET out.name='Outdoor', out.labTag='ch22'
MERGE (sec:Category {categoryId:'CAT-22-SEC'}) SET sec.name='Security', sec.labTag='ch22'
MERGE (acc:Category {categoryId:'CAT-22-ACC'}) SET acc.name='Accessories', acc.labTag='ch22'
MERGE (st:Store {storeId:'ST-22-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch22';

UNWIND [
 {id:'P-2201',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:5,featured:true, emb:[0.95,0.85,0.25,0.05,0.10,0.90,0.05,0.05]},
 {id:'P-2202',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:0,featured:false,emb:[0.90,0.80,0.15,0.05,0.05,0.82,0.05,0.05]},
 {id:'P-2203',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,cat:'CAT-22-SEC',catCode:'SECURITY',region:'CENTRAL',qty:7,featured:false,emb:[0.92,0.10,0.95,0.02,0.05,0.05,0.05,0.05]},
 {id:'P-2204',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:11,featured:false,emb:[0.02,0.88,0.02,0.95,0.10,0.10,0.05,0.02]},
 {id:'P-2205',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:3,featured:true,emb:[0.90,0.45,0.20,0.10,0.95,0.15,0.05,0.03]},
 {id:'P-2206',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:6,featured:false,emb:[0.05,0.55,0.05,0.05,0.05,0.95,0.05,0.02]},
 {id:'P-2207',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,cat:'CAT-22-ACC',catCode:'ACCESSORY',region:'CENTRAL',qty:0,featured:false,emb:[0.08,0.75,0.10,0.05,0.10,0.10,0.95,0.02]},
 {id:'P-2208',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,cat:'CAT-22-CAM',catCode:'CAMERA',region:'ARCHIVE',qty:2,featured:false,emb:[0.88,0.68,0.15,0.05,0.05,0.60,0.05,0.95]}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
    p.active=row.active, p.categoryCode=row.catCode, p.region=row.region,
    p.featured=row.featured, p.embedding=row.emb,
    p.embeddingModel='atlasmart-deterministic-v1', p.embeddingVersion='2026-09-lab',
    p.labTag='ch22'
WITH row,p
MATCH (cat:Category {categoryId:row.cat}), (st:Store {storeId:'ST-22-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch22';

MERGE (d1:KnowledgeDocument {documentId:'DOC-22-TRAILCAM'})
SET d1.title='Trail Camera Pro field manual', d1.uri='atlasmart://manuals/P-2201', d1.version='2026.09', d1.labTag='ch22'
MERGE (d2:KnowledgeDocument {documentId:'DOC-22-SECURITY'})
SET d2.title='Camera selection guide', d2.uri='atlasmart://guides/camera-selection', d2.version='2026.09', d2.labTag='ch22';

UNWIND [
 {id:'CHK-2201',doc:'DOC-22-TRAILCAM',seq:1,text:'Trail Camera Pro is weatherproof and optimized for wildlife monitoring on outdoor trails.',emb:[0.96,0.86,0.12,0.02,0.02,0.94,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
 {id:'CHK-2202',doc:'DOC-22-TRAILCAM',seq:2,text:'Infrared night vision records wildlife without visible illumination and battery life is designed for field deployment.',emb:[0.88,0.72,0.25,0.02,0.02,0.90,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
 {id:'CHK-2203',doc:'DOC-22-SECURITY',seq:1,text:'Indoor Security Camera focuses on Wi-Fi motion alerts and indoor night vision rather than outdoor wildlife use.',emb:[0.84,0.08,0.96,0.02,0.02,0.06,0.02,0.02],products:['P-2203'],cats:['CAT-22-SEC']},
 {id:'CHK-2204',doc:'DOC-22-SECURITY',seq:2,text:'Choose an outdoor trail camera when weather resistance and wildlife observation matter; choose indoor security cameras for home alerting.',emb:[0.90,0.68,0.55,0.02,0.02,0.72,0.02,0.02],products:['P-2201','P-2203'],cats:['CAT-22-CAM','CAT-22-SEC']}
] AS row
MERGE (c:Chunk {chunkId:row.id})
SET c.seq=row.seq, c.text=row.text, c.embedding=row.emb,
    c.embeddingModel='atlasmart-deterministic-v1', c.embeddingVersion='2026-09-lab', c.labTag='ch22'
WITH row,c
MATCH (d:KnowledgeDocument {documentId:row.doc})
MERGE (c)-[:FROM_DOCUMENT]->(d)
WITH row,c
UNWIND row.products AS pid
MATCH (p:Product {productId:pid})
MERGE (c)-[:MENTIONS]->(p)
WITH row,c
UNWIND row.cats AS cid
MATCH (cat:Category {categoryId:cid})
MERGE (c)-[:MENTIONS]->(cat);

MATCH (a:Chunk {chunkId:'CHK-2201'}),(b:Chunk {chunkId:'CHK-2202'}) MERGE (a)-[:NEXT_CHUNK]->(b);
MATCH (a:Chunk {chunkId:'CHK-2203'}),(b:Chunk {chunkId:'CHK-2204'}) MERGE (a)-[:NEXT_CHUNK]->(b);

CREATE FULLTEXT INDEX ch22_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name,p.description,p.tags]
OPTIONS {indexConfig:{`fulltext.analyzer`:'english',`fulltext.eventually_consistent`:false}};

CREATE VECTOR INDEX ch22_product_vector IF NOT EXISTS
FOR (p:Product)
ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
  `vector.dimensions`:8,
  `vector.similarity_function`:'cosine',
  `vector.quantization.type`:'scalar',
  `vector.default_search_expansion_factor`:1.5
}};

CREATE VECTOR INDEX ch22_chunk_vector IF NOT EXISTS
FOR (c:Chunk)
ON c.embedding
WITH [c.embeddingVersion,c.seq]
OPTIONS {indexConfig:{
  `vector.dimensions`:8,
  `vector.similarity_function`:'cosine'
}};

CALL db.awaitIndexes(300);

1. Model retrieval units separately from domain entities

A Product is a domain entity; a Chunk is a retrieval/evidence unit; a KnowledgeDocument is a provenance container. Conflating them creates difficult lifecycle problems: product facts change independently from documentation text, chunk size/embedding changes during retrieval experiments, and source documents need version/URI metadata for audit.

Node/relationship Purpose Lifecycle concern
KnowledgeDocument source identity/title/URI/version document revision, retention, authorization
Chunk bounded retrieval text + embedding + chunk ID/sequence chunking/embedding model version, re-embedding
FROM_DOCUMENT immutable-ish provenance edge must survive context assembly
NEXT_CHUNK adjacent context within document bounded neighbor expansion only
MENTIONS → Product/Category connect retrieved text to graph entities entity resolution and extraction confidence/version
Product/Category/Store operational graph truth transactional business data lifecycle differs from source text

2. Retrieve chunks, then expand evidence—not the other way around

The query first performs vector SEARCH on Chunk with an embedding-version filter. It then follows provenance/entity relationships and returns a structured evidence packet. This keeps the expensive candidate stage bounded and makes every returned statement traceable to a chunk and source.

Cypher 25 · provenance-rich GraphRAG retrieval packet
CYPHER 25
MATCH (chunk:Chunk)
  SEARCH chunk IN (
    VECTOR INDEX ch22_chunk_vector
    FOR $qvec
    WHERE chunk.embeddingVersion = '2026-09-lab'
    LIMIT 3
  ) SCORE AS vectorScore
MATCH (chunk)-[:FROM_DOCUMENT]->(doc:KnowledgeDocument)
OPTIONAL MATCH (chunk)-[:MENTIONS]->(entity)
WITH chunk,doc,vectorScore,
     collect(DISTINCT CASE
       WHEN entity:Product THEN {kind:'Product',id:entity.productId,name:entity.name}
       WHEN entity:Category THEN {kind:'Category',id:entity.categoryId,name:entity.name}
       ELSE NULL END) AS entities
RETURN {
  chunkId:chunk.chunkId,
  text:chunk.text,
  vectorScore:vectorScore,
  source:{documentId:doc.documentId,title:doc.title,uri:doc.uri,version:doc.version},
  entities:[e IN entities WHERE e IS NOT NULL]
} AS evidence;
// This returns a retrieval/context packet, not an LLM answer.
No generator is required to validate retrieval

You can inspect chunk IDs, texts, source URIs/versions and mentioned entities directly. That makes recall, provenance completeness and authorization testable before adding LLM variability/cost.

3. Context assembly is a budgeted data product

A production context assembler should deduplicate chunks, optionally add adjacent NEXT_CHUNK context, fetch authoritative graph facts, enforce tenant/security rules, and cap tokens/bytes. It should not dump an unbounded neighborhood into the model. Every context item should retain source and relationship provenance.

Example context envelope · generator-independent
{
  "queryId":"rag-2026-09-09-001",
  "embedding":{"model":"atlasmart-deterministic-v1","version":"2026-09-lab"},
  "retrieval":{"index":"ch22_chunk_vector","sourceK":3,"strategy":"vector+graph"},
  "evidence":[
    {
      "chunkId":"CHK-2201",
      "source":{"documentId":"DOC-22-TRAILCAM","uri":"atlasmart://manuals/P-2201","version":"2026.09"},
      "entities":[{"kind":"Product","id":"P-2201"},{"kind":"Category","id":"CAT-22-CAM"}],
      "text":"...exact retrieved chunk text..."
    }
  ],
  "graphFacts":[{"productId":"P-2201","storeId":"ST-22-CENTRAL","quantity":5}],
  "policy":{"tenant":"demo","authorizationChecked":true}
}

4. Provenance is not decoration

If a generated answer says “Trail Camera Pro is weatherproof,” the system should be able to point to CHK-2201 and its document URI/version. If it says “five units are available,” that evidence comes from the live STOCKED_AT relationship, not the manual. Keeping provenance classes separate prevents a stale document from being mistaken for transactional inventory truth.

Claim type Preferred evidence Freshness model
manual/specification Chunk → KnowledgeDocument document version/re-ingestion
product identity/category Product/Category graph transactional graph
inventory STOCKED_AT relationship high-frequency transactional fact
semantic similarity vector index rank/score embedding/index version
lexical exact term full-text rank/score text/index freshness

5. Evaluate retrieval and generation separately

GraphRAG can fail even if the final prose sounds plausible. Build evaluation layers. Retrieval metrics ask whether the needed chunks/entities appeared. Context metrics ask whether evidence was authorized, deduplicated and provenance-complete. If a generator is added, answer metrics ask whether claims are supported/cited and whether unsupported statements/hallucinations appear. Latency/cost metrics span all stages.

Layer Example metric/test Failure signal
vector retrieval chunk recall@k, MRR/nDCG on judged chunks needed evidence absent
graph expansion entity/provenance completeness; bounded fan-out wrong entity or missing source
context assembly byte/token cap, duplicate rate, policy pass rate overflow/duplicate/unauthorized evidence
generation optional claim-support precision, citation correctness, abstention/no-answer accuracy unsupported claims despite good retrieval
system p50/p95/p99, fallback/error rate, index readiness tail latency or dependency failure

6. Controlled failure injection

Use reversible failures that exercise the retrieval contract, not destructive production chaos. Examples: remove one chunk embedding, query with wrong dimension, temporarily point the service to a non-existent index name, mark an embeddingVersion filter that matches nothing, duplicate a chunk, or remove a MENTIONS edge. For each, record expected signal, user fallback and reset.

Injected fault Expected evidence Safe repair
Chunk embedding missing chunk absent from vector index / lower retrieval recall backfill embedding and verify index read
wrong embeddingVersion filter zero vector candidates version compatibility check/fallback
duplicate chunk content duplicate context ratio increases dedupe by stable chunk/source identity
MENTIONS edge missing provenance chunk exists but entity expansion incomplete re-run entity linking/reconciliation
index unavailable SEARCH dependency error/readiness failure fallback lexical path or explicit unavailable response

7. Embedding/index migration and rollback

Never overwrite model metadata and hope. Write new embeddings under a new property/version or maintain a migration strategy, populate a new vector index, run the frozen GraphRAG evaluation suite, compare cost/recall/context/answer metrics, switch configuration, and keep the old path until rollback confidence expires. The graph model and provenance IDs should remain stable across encoder changes where possible.

8. Optional official GraphRAG package vs mandatory manual path

Neo4j’s current first-party Python package is neo4j-graphrag; the older neo4j-genai package name is deprecated. The package can orchestrate retrievers, embedders, knowledge-graph building and LLMs, but the course lab intentionally demonstrates the underlying data/query contracts directly. You may later plug in a local or paid model, but do not make the API/provider a prerequisite for understanding GraphRAG.

Paid LLM/embedding services are optional

If you add an external model, document data egress, credentials, regional/privacy constraints, token/cost limits, model/version, retry/idempotency and fallback. The core retrieval/provenance lab remains local and deterministic.

9. Wrong approaches and repairs

Wrong approach Problem Repair
“GraphRAG prevents hallucinations” generator can still make unsupported claims claim/citation evaluation + abstention policy
send top vector text with no source IDs cannot audit/cite/update Document/Chunk provenance envelope
include entire graph neighborhood context explosion/privacy/latency bounded task-specific expansion
mix manual inventory text with live stock as same truth freshness conflict source-class provenance and authority policy
require external OpenAI-style API for course cost/access/privacy dependency deterministic embeddings + generator-independent context
re-embed in place without version metadata mixed neighborhoods and no rollback parallel version/index + evaluation/cutover

10. Cleanup / reset

Run only against the disposable Chapter 22 fixture.

Cypher 25 · cleanup
CYPHER 25
DROP INDEX ch22_product_vector IF EXISTS;
DROP INDEX ch22_product_vector_none IF EXISTS;
DROP INDEX ch22_product_vector_binary IF EXISTS;
DROP INDEX ch22_chunk_vector IF EXISTS;
DROP INDEX ch22_catalog_ft IF EXISTS;
MATCH (n) WHERE n.labTag='ch22' DETACH DELETE n;
DROP CONSTRAINT ch22_product_id IF EXISTS;
DROP CONSTRAINT ch22_category_id IF EXISTS;
DROP CONSTRAINT ch22_store_id IF EXISTS;
DROP CONSTRAINT ch22_doc_id IF EXISTS;
DROP CONSTRAINT ch22_chunk_id IF EXISTS;

11. Final production judgment

Production decision Evidence required
embedding lifecycle Record model/provider/version, dimensions, normalization assumptions, text preprocessing, backfill/re-embedding status, and rollback/index-swap plan.
recall vs latency Measure exact-vs-ANN recall@k on a representative judged set alongside p50/p95/p99; tiny demo recall is not a capacity guarantee.
model/cardinality/degree Separate candidate retrieval count from graph expansion fan-out; bound traversal and final context size.
index memory/storage/write cost Observe vector index size, population/rebuild time, write amplification, page-cache/store pressure, and quantization effects before tuning.
similarity semantics Choose cosine/euclidean from embedding-model semantics; score is a source-specific similarity signal, not factual probability.
filtering/security Declare filter properties intentionally, distinguish in-index from post-filter behavior, and test the real Enterprise service role because semantic-index authorization can suppress candidates.
driver/timeouts/retries Use a long-lived driver, bounded candidate counts, transaction timeouts, idempotent writes and explicit retry/error classification from earlier chapters.
hybrid ranking Fuse independent source ranks (for example RRF/WRRF) or use a trained evaluated re-ranker; never add incomparable raw full-text/vector scores by habit.
GraphRAG provenance Every context unit carries source ID/URI/version/chunk ID and graph entities/relationships so retrieval evidence can be audited.
evaluation Track retrieval recall/precision/MRR/nDCG-style metrics, answer grounding/citation correctness if a generator is added, latency, freshness and zero-result/fallback rates.
privacy/tenant risk Do not embed secrets/PII without policy; authorization must be enforced before context reaches a generator or user.
backup/recovery Rebuild/validate vector indexes and embedding-version metadata in restore drills; restore of graph data is not proof that semantic retrieval is healthy.
Aura/self-managed Aura manages infrastructure and some controls; self-managed exposes server/index/plugin/resource operations. Verify feature/tier availability instead of assuming parity.
licensing/cost Community mandatory lab is free/local. VECTOR storage, Enterprise security/clustering and paid embedding/LLM services are optional and separately costed/licensed.
migration/rollback Run old/new embedding/index versions side by side when possible, freeze evaluation data, cut over by explicit index/query configuration, retain rollback until validation passes.

Check your understanding

  1. Why separate Chunk from Product?
  2. What does FROM_DOCUMENT provide?
  3. Does good retrieval guarantee a factual generated answer?
  4. Why stop the mandatory lab before LLM generation?
  5. How should an embedding-model upgrade be deployed?
Review the answers

1. Chunks are retrieval/source-text units with their own chunking/embedding lifecycle; Product is an operational domain entity.

2. Traceable provenance from retrieved chunk to source document identity/version/URI.

3. No. Generation needs separate grounding/citation/claim evaluation and may need abstention.

4. It keeps the mechanism free/deterministic and lets retrieval/provenance quality be tested independently.

5. Version it, build/reindex in parallel, run frozen evaluation, cut over explicitly, retain rollback until validated.

Summary and next step

Chapter 22 has built a complete retrieval discipline: deterministic embeddings → exact baseline → HNSW ANN → observable index/options → Cypher SEARCH → rank-based hybrid retrieval → graph/business/security context → provenance-rich GraphRAG evidence → layered evaluation. Chapter 23 can now introduce Graph Data Science from the same principle: algorithms operate on explicit projections with measurable assumptions, not as magic graph intelligence.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.