Chapter 22 · Vector Search, Embeddings, Cypher SEARCH, Hybrid Search, and GraphRAG
Hybrid Full-Text + Vector Retrieval, Independent Ranking, Filtering, Re-Ranking, and Graph Expansion
Combine lexical and semantic candidates without comparing raw full-text/vector scores, use rank fusion and graph/business constraints, and make every filtering, re-ranking, security, and latency decision observable.
Vector retrieval solves “similar meaning,” while Chapter 21 full-text retrieval excels at exact terms, model names, acronyms and misspellings. Neither source is universally superior. AtlasMart therefore needs a hybrid pipeline that can preserve each source’s ranking, merge candidates, apply graph/business/security constraints, explain why an item survived, and evaluate whether the combination actually improves relevance.
Hybrid retrieval is not “add the two scores.” It is several independent candidate generators followed by a fusion/ranking policy. Graph traversal then supplies structure and eligibility. Every stage must keep its own evidence so failures can be attributed rather than hidden in one opaque number.
Learning outcomes
Run full-text and vector retrieval as independent bounded sources and retain source-specific ranks/scores for diagnostics.
Use rank fusion such as reciprocal-rank fusion instead of comparing incomparable raw Lucene and vector scores.
Apply graph/business/security filtering after fusion and understand sourceK/finalK survival tradeoffs.
Separate in-index vector filters, graph post-filters, re-ranking features and graph expansion by mechanism.
Evaluate lexical-only, vector-only and hybrid pipelines against the same judged queries and latency/resource budgets.
Current Neo4j Database is 2026.07.1; the current
5.26 line remains LTS. Version-sensitive examples use explicit
CYPHER 25. The mandatory lab uses self-managed
Neo4j Community 2026.07.1, database
neo4j, user neo4j, disposable password
atlasmart-course-2026, loopback Bolt
7687 and HTTP 7474, and embeddings
stored as LIST<FLOAT>. Neo4j 2026.x supports
Java 21/25. Optional client examples pin the official Python
driver to neo4j==6.3.0. No APOC, GDS, paid
embedding API, paid LLM API, Aura account, or Enterprise license
is required.
Vector indexes are available in Community when
embeddings are stored as LIST<INTEGER|FLOAT>.
The newer fixed-size VECTOR property type requires
block-format storage and therefore cannot be persisted as a
property in Community; it is an Enterprise/Aura storage
capability. The lab deliberately uses LIST embeddings so every
mandatory index/search/evaluation step remains free/local. Where
VECTOR-specific storage is discussed, it is labeled as an
edition-dependent optimization/typing choice rather than a
prerequisite.
From Neo4j 2026.01, Cypher 25
SEARCH is the preferred way to query vector indexes
and supports in-index filtering when filter properties were
declared in the index.
db.index.vector.queryNodes() and
db.index.vector.queryRelationships() remain useful
for older-version compatibility history but are deprecated from
Neo4j 2026.04. New course code therefore uses
SEARCH.
Lab contract and exact assumptions
| Dimension | Chapter 22 assumption |
|---|---|
| server | Neo4j Community 2026.07.1, single disposable local database |
| Cypher | Explicit CYPHER 25 for SEARCH and current vector syntax |
| Java | Java 21 or 25 for Neo4j 2026.07 |
| database/auth | neo4j / neo4j / atlasmart-course-2026 |
| transport | bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; production/remote deployments use verified TLS |
| plugins | none required; APOC/GDS/GenAI are not needed |
| embedding source | deterministic 8-dimensional precomputed AtlasMart vectors; not a paid API and not claimed to be production-quality embeddings |
| storage |
LIST |
| graph | 8 Products, 4 Categories, 1 Store, 2 KnowledgeDocuments, 4 Chunks plus provenance/entity edges |
| indexes | full-text product index + 8D product vector index + 8D chunk vector index |
| measurement | learner measures recall@k, runtime latency, index state/options and result IDs; generated lesson never claims that Neo4j was executed here |
| Term | Mechanism-first meaning |
|---|---|
| embedding | Numeric representation produced outside the database by a model or deterministic encoder. Neo4j stores/indexes the values; it does not make semantic truth guarantees about the encoder. |
| dimension | Number of coordinates in an embedding. Index dimension and query-vector dimension must match when dimensions are configured. |
| LIST embedding |
Community-compatible numeric property such as
[0.95,0.85,...]. Individual elements are
list-accessible.
|
| VECTOR value | Fixed-length typed vector value introduced in 2025.10; more storage-efficient typing but persisted VECTOR properties require Enterprise/Aura block format. |
| similarity | Function that converts a pair of vectors into an ordering signal. Current vector indexes support cosine and euclidean similarity. |
| ANN | Approximate nearest-neighbor retrieval. It trades guaranteed exactness for scalable search speed/resource behavior. |
| HNSW | Hierarchical Navigable Small World graph used internally by the vector index to navigate candidate neighborhoods rather than compare every stored vector. |
| recall@k | Fraction of the exact top-k neighbors recovered by ANN top-k. It is a retrieval-quality measure, not semantic correctness. |
| filter property | Non-vector property explicitly stored with a 2026.01+ vector index so SEARCH can apply supported predicates inside the ANN search. |
| quantization | Compressed vector representation used inside the index to lower memory/storage and often improve speed, potentially trading accuracy; 2026.07 supports high-fidelity rescoring through search expansion. |
| GraphRAG | Retrieval-augmented generation pattern where graph-structured evidence, provenance, and relationships enrich the context given to a generator. Retrieval quality and generator factuality still require evaluation. |
Direct-entry setup
The fixture creates both ch22_catalog_ft and
ch22_product_vector. It also creates stock/category
graph facts that intentionally remove some high-relevance
candidates: P-2202 is out of stock and P-2208 is inactive.
CYPHER 25
// Disposable Chapter 22 fixture. Safe to rerun after the cleanup block.
CREATE CONSTRAINT ch22_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch22_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch22_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch22_doc_id IF NOT EXISTS
FOR (d:KnowledgeDocument) REQUIRE d.documentId IS UNIQUE;
CREATE CONSTRAINT ch22_chunk_id IF NOT EXISTS
FOR (c:Chunk) REQUIRE c.chunkId IS UNIQUE;
MERGE (cam:Category {categoryId:'CAT-22-CAM'}) SET cam.name='Cameras', cam.labTag='ch22'
MERGE (out:Category {categoryId:'CAT-22-OUT'}) SET out.name='Outdoor', out.labTag='ch22'
MERGE (sec:Category {categoryId:'CAT-22-SEC'}) SET sec.name='Security', sec.labTag='ch22'
MERGE (acc:Category {categoryId:'CAT-22-ACC'}) SET acc.name='Accessories', acc.labTag='ch22'
MERGE (st:Store {storeId:'ST-22-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch22';
UNWIND [
{id:'P-2201',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:5,featured:true, emb:[0.95,0.85,0.25,0.05,0.10,0.90,0.05,0.05]},
{id:'P-2202',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:0,featured:false,emb:[0.90,0.80,0.15,0.05,0.05,0.82,0.05,0.05]},
{id:'P-2203',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,cat:'CAT-22-SEC',catCode:'SECURITY',region:'CENTRAL',qty:7,featured:false,emb:[0.92,0.10,0.95,0.02,0.05,0.05,0.05,0.05]},
{id:'P-2204',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:11,featured:false,emb:[0.02,0.88,0.02,0.95,0.10,0.10,0.05,0.02]},
{id:'P-2205',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:3,featured:true,emb:[0.90,0.45,0.20,0.10,0.95,0.15,0.05,0.03]},
{id:'P-2206',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:6,featured:false,emb:[0.05,0.55,0.05,0.05,0.05,0.95,0.05,0.02]},
{id:'P-2207',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,cat:'CAT-22-ACC',catCode:'ACCESSORY',region:'CENTRAL',qty:0,featured:false,emb:[0.08,0.75,0.10,0.05,0.10,0.10,0.95,0.02]},
{id:'P-2208',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,cat:'CAT-22-CAM',catCode:'CAMERA',region:'ARCHIVE',qty:2,featured:false,emb:[0.88,0.68,0.15,0.05,0.05,0.60,0.05,0.95]}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
p.active=row.active, p.categoryCode=row.catCode, p.region=row.region,
p.featured=row.featured, p.embedding=row.emb,
p.embeddingModel='atlasmart-deterministic-v1', p.embeddingVersion='2026-09-lab',
p.labTag='ch22'
WITH row,p
MATCH (cat:Category {categoryId:row.cat}), (st:Store {storeId:'ST-22-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch22';
MERGE (d1:KnowledgeDocument {documentId:'DOC-22-TRAILCAM'})
SET d1.title='Trail Camera Pro field manual', d1.uri='atlasmart://manuals/P-2201', d1.version='2026.09', d1.labTag='ch22'
MERGE (d2:KnowledgeDocument {documentId:'DOC-22-SECURITY'})
SET d2.title='Camera selection guide', d2.uri='atlasmart://guides/camera-selection', d2.version='2026.09', d2.labTag='ch22';
UNWIND [
{id:'CHK-2201',doc:'DOC-22-TRAILCAM',seq:1,text:'Trail Camera Pro is weatherproof and optimized for wildlife monitoring on outdoor trails.',emb:[0.96,0.86,0.12,0.02,0.02,0.94,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
{id:'CHK-2202',doc:'DOC-22-TRAILCAM',seq:2,text:'Infrared night vision records wildlife without visible illumination and battery life is designed for field deployment.',emb:[0.88,0.72,0.25,0.02,0.02,0.90,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
{id:'CHK-2203',doc:'DOC-22-SECURITY',seq:1,text:'Indoor Security Camera focuses on Wi-Fi motion alerts and indoor night vision rather than outdoor wildlife use.',emb:[0.84,0.08,0.96,0.02,0.02,0.06,0.02,0.02],products:['P-2203'],cats:['CAT-22-SEC']},
{id:'CHK-2204',doc:'DOC-22-SECURITY',seq:2,text:'Choose an outdoor trail camera when weather resistance and wildlife observation matter; choose indoor security cameras for home alerting.',emb:[0.90,0.68,0.55,0.02,0.02,0.72,0.02,0.02],products:['P-2201','P-2203'],cats:['CAT-22-CAM','CAT-22-SEC']}
] AS row
MERGE (c:Chunk {chunkId:row.id})
SET c.seq=row.seq, c.text=row.text, c.embedding=row.emb,
c.embeddingModel='atlasmart-deterministic-v1', c.embeddingVersion='2026-09-lab', c.labTag='ch22'
WITH row,c
MATCH (d:KnowledgeDocument {documentId:row.doc})
MERGE (c)-[:FROM_DOCUMENT]->(d)
WITH row,c
UNWIND row.products AS pid
MATCH (p:Product {productId:pid})
MERGE (c)-[:MENTIONS]->(p)
WITH row,c
UNWIND row.cats AS cid
MATCH (cat:Category {categoryId:cid})
MERGE (c)-[:MENTIONS]->(cat);
MATCH (a:Chunk {chunkId:'CHK-2201'}),(b:Chunk {chunkId:'CHK-2202'}) MERGE (a)-[:NEXT_CHUNK]->(b);
MATCH (a:Chunk {chunkId:'CHK-2203'}),(b:Chunk {chunkId:'CHK-2204'}) MERGE (a)-[:NEXT_CHUNK]->(b);
CREATE FULLTEXT INDEX ch22_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name,p.description,p.tags]
OPTIONS {indexConfig:{`fulltext.analyzer`:'english',`fulltext.eventually_consistent`:false}};
CREATE VECTOR INDEX ch22_product_vector IF NOT EXISTS
FOR (p:Product)
ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine',
`vector.quantization.type`:'scalar',
`vector.default_search_expansion_factor`:1.5
}};
CREATE VECTOR INDEX ch22_chunk_vector IF NOT EXISTS
FOR (c:Chunk)
ON c.embedding
WITH [c.embeddingVersion,c.seq]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine'
}};
CALL db.awaitIndexes(300);
1. Preserve source identity, rank, and raw diagnostic score
Full-text score and vector similarity score have different generation mechanisms/scales. Current Neo4j guidance says rank each source independently when combining them. Keep raw scores in logs for source diagnostics, but fuse ranks or use an evaluated learned re-ranker. The example uses ordinary RRF because its mechanism is transparent.
| Source | Good at | Raw score policy |
|---|---|---|
| full-text | exact terms, names, lexical intent, phrase/fuzzy syntax | use only within full-text result ordering/diagnostics |
| vector | semantic neighborhood under encoder/metric | use only within vector result ordering/diagnostics |
| graph | eligibility, relationships, inventory, topology, provenance | normally rules/features, not a directly comparable semantic score |
2. Bounded source retrieval + RRF + graph eligibility
The Python harness gets six candidates from each source,
converts each list to rank contributions
1/(k+rank), merges duplicate product IDs, then
fetches graph facts and removes inactive/out-of-stock products.
This is intentionally simple and auditable. Production systems
can weight sources or learn ranking models, but only with judged
data.
from neo4j import GraphDatabase
from collections import defaultdict
URI="bolt://localhost:7687"
AUTH=("neo4j","atlasmart-course-2026")
QTEXT="wildlife trail camera"
QVEC=[1.0,0.8,0.1,0.0,0.0,0.9,0.0,0.0]
SOURCE_K=6
FINAL_K=4
RRF_K=60
LEXICAL="""
CALL db.index.fulltext.queryNodes('ch22_catalog_ft',$q,{limit:$k})
YIELD node,score
RETURN node.productId AS id,score ORDER BY score DESC
"""
VECTOR="""
MATCH (p:Product)
SEARCH p IN (VECTOR INDEX ch22_product_vector FOR $qvec LIMIT $k)
SCORE AS score
RETURN p.productId AS id,score
"""
DETAIL="""
UNWIND $ids AS id
MATCH (p:Product {productId:id})
OPTIONAL MATCH (p)-[:IN_CATEGORY]->(c:Category)
OPTIONAL MATCH (p)-[s:STOCKED_AT]->(st:Store {storeId:'ST-22-CENTRAL'})
RETURN p.productId AS id,p.name AS name,p.active AS active,
p.featured AS featured,c.name AS category,coalesce(s.quantity,0) AS qty
"""
def fuse(rows_by_source):
score=defaultdict(float); evidence=defaultdict(dict)
for source,rows in rows_by_source.items():
for rank,row in enumerate(rows,1):
score[row["id"]]+=1.0/(RRF_K+rank)
evidence[row["id"]][source]={"rank":rank,"rawScore":row["score"]}
return score,evidence
with GraphDatabase.driver(URI,auth=AUTH) as driver:
lex,_ ,_=driver.execute_query(LEXICAL,q=QTEXT,k=SOURCE_K,database_="neo4j")
vec,_ ,_=driver.execute_query(VECTOR,qvec=QVEC,k=SOURCE_K,database_="neo4j")
scores,evidence=fuse({"lexical":lex,"vector":vec})
ids=sorted(scores,key=scores.get,reverse=True)
details,_ ,_=driver.execute_query(DETAIL,ids=ids,database_="neo4j")
by_id={r["id"]:dict(r) for r in details}
survivors=[i for i in ids if by_id[i]["active"] and by_id[i]["qty"]>0]
for rank,i in enumerate(survivors[:FINAL_K],1):
print({"finalRank":rank,"rrf":scores[i],"sources":evidence[i],**by_id[i]})
# Important: raw full-text and vector scores are retained for diagnostics only.
# RRF combines source RANKS; eligibility comes from graph/business data.
The diagnostic payload can show Lucene score and vector score, but the fusion function uses ranks. This prevents accidental interpretation of two unrelated numeric scales as if they were calibrated.
3. sourceK and finalK are separate capacity/recall knobs
If finalK=4 and each source returns only four items, deduplication and business/security filters can leave too few survivors. A larger bounded sourceK gives fusion/filtering room but costs index work, graph lookups, response memory and latency. Measure candidate survival and p95/p99 instead of making sourceK huge.
| Stage | Bound | Failure signal |
|---|---|---|
| full-text source | lexical sourceK | zero/low candidate recall, Lucene query latency |
| vector source | vector sourceK + ANN expansion | recall@k, vector p95/p99, index resource use |
| fusion | union size | duplicate ratio / memory / ranking CPU |
| graph eligibility | degree/pattern bounds | rejection reason counts, DB hits, fan-out |
| response | finalK/context byte cap | payload size and user-perceived latency |
4. Filtering belongs at the lowest correct layer
If “active=true” is stable and selective and declared in the vector index, it may belong inside SEARCH. “Stock at the user’s chosen store” is a relationship traversal and belongs in graph filtering, not vector metadata unless the data model deliberately materializes a filterable property. Tenant/security boundaries must be enforced according to authorization semantics, not merely post-ranked as preferences.
| Rule | Best mechanism when applicable | Reason |
|---|---|---|
| embedding version | vector in-index filter | prevents mixing incompatible vector generations |
| active/category code | vector metadata filter or graph filter depending source consistency | can reduce ANN search space when indexed |
| stock at Store | graph relationship predicate | relationship-local fact varies by store |
| featured promotion | post-rank feature/rule | business policy, not semantic similarity |
| tenant/authorization | security model + query/index design | must not be treated as optional relevance |
5. Graph expansion can enrich explanations without contaminating source scores
For each surviving product, expand to Category, Store inventory, reviews, supplier or knowledge chunks only as bounded application needs. The source rank explains retrieval; graph evidence explains business/context. Avoid unbounded traversal per candidate, especially on hubs. Use PROFILE on representative graph stages and return only fields needed by the client.
6. Hybrid quality must beat baselines, not a screenshot
Create judged queries spanning exact model names, semantic paraphrases, typos, mixed intent and no-answer cases. Compute the same metrics for lexical-only, vector-only and hybrid. A hybrid deployment is justified only if quality gains outweigh extra latency, index/storage cost, complexity and failure modes.
| Experiment | Hold constant | Compare |
|---|---|---|
| lexical vs vector | corpus, judged queries, finalK, business filters | MRR/nDCG/recall, zero-result rate, latency |
| hybrid fusion | same source lists | RRF constant/weights or learned ranker |
| sourceK | same ranking policy | survival/recall vs p95/p99/resource cost |
| graph rule | same retrieval lists | quality slice + rejection reasons + DB hits |
| embedding version | same lexical baseline/judgments | vector/hybrid quality + reindex cost |
7. Security and privacy are before generation
Enterprise fine-grained security can conservatively suppress semantic-index results. This may reduce recall but protects denied data. Do not compensate by broadening privileges. Also do not send unauthorized or sensitive retrieved fields to an LLM. Retrieval traces should record which policy/filter removed candidates without logging secret content unnecessarily.
8. Wrong approaches and repairs
| Wrong approach | Problem | Repair |
|---|---|---|
| 0.5*fulltextScore + 0.5*vectorScore | incomparable scales | rank fusion or evaluated calibration/model |
| sourceK=10000 “for recall” | tail latency/memory/resource explosion | bounded sourceK from survival/quality measurements |
| materialize every graph relationship into vector metadata | stale duplicated state/index bloat | use graph traversal for dynamic relational facts |
| rank authorization as a soft penalty | possible data exposure | hard security enforcement before context/response |
| declare hybrid better from one query | anecdotal overfit | versioned judged suite + latency/resource gates |
9. Production judgment
| Production decision | Evidence required |
|---|---|
| embedding lifecycle | Record model/provider/version, dimensions, normalization assumptions, text preprocessing, backfill/re-embedding status, and rollback/index-swap plan. |
| recall vs latency | Measure exact-vs-ANN recall@k on a representative judged set alongside p50/p95/p99; tiny demo recall is not a capacity guarantee. |
| model/cardinality/degree | Separate candidate retrieval count from graph expansion fan-out; bound traversal and final context size. |
| index memory/storage/write cost | Observe vector index size, population/rebuild time, write amplification, page-cache/store pressure, and quantization effects before tuning. |
| similarity semantics | Choose cosine/euclidean from embedding-model semantics; score is a source-specific similarity signal, not factual probability. |
| filtering/security | Declare filter properties intentionally, distinguish in-index from post-filter behavior, and test the real Enterprise service role because semantic-index authorization can suppress candidates. |
| driver/timeouts/retries | Use a long-lived driver, bounded candidate counts, transaction timeouts, idempotent writes and explicit retry/error classification from earlier chapters. |
| hybrid ranking | Fuse independent source ranks (for example RRF/WRRF) or use a trained evaluated re-ranker; never add incomparable raw full-text/vector scores by habit. |
| GraphRAG provenance | Every context unit carries source ID/URI/version/chunk ID and graph entities/relationships so retrieval evidence can be audited. |
| evaluation | Track retrieval recall/precision/MRR/nDCG-style metrics, answer grounding/citation correctness if a generator is added, latency, freshness and zero-result/fallback rates. |
| privacy/tenant risk | Do not embed secrets/PII without policy; authorization must be enforced before context reaches a generator or user. |
| backup/recovery | Rebuild/validate vector indexes and embedding-version metadata in restore drills; restore of graph data is not proof that semantic retrieval is healthy. |
| Aura/self-managed | Aura manages infrastructure and some controls; self-managed exposes server/index/plugin/resource operations. Verify feature/tier availability instead of assuming parity. |
| licensing/cost | Community mandatory lab is free/local. VECTOR storage, Enterprise security/clustering and paid embedding/LLM services are optional and separately costed/licensed. |
| migration/rollback | Run old/new embedding/index versions side by side when possible, freeze evaluation data, cut over by explicit index/query configuration, retain rollback until validation passes. |
Check your understanding
- Why not add raw full-text and vector scores?
- What does RRF use?
- Why can sourceK exceed finalK?
- Should Store inventory be duplicated into vector filter metadata by default?
- How do you prove hybrid search is better?
Review the answers
1. They arise from different mechanisms/scales and are not calibrated to each other.
2. Rank position in each source list, optionally with source weights, rather than raw score magnitudes.
3. Deduplication, fusion, graph/business/security filters can remove candidates before the final list.
4. No. It is a dynamic relationship fact; keep it in graph filtering unless a measured architecture justifies materialization.
5. Compare lexical-only/vector-only/hybrid against the same judged queries plus latency/resource/freshness/security gates.
Summary and next step
AtlasMart now has a defensible hybrid pipeline: independent lexical/vector candidate lists → rank fusion → bounded graph/security/business filtering → explainable result. Lesson 5 turns the same retrieval discipline into GraphRAG context assembly with explicit Document/Chunk/entity provenance and evaluation.
Authoritative references
- Neo4j Operations Manual — current release — Current server line; the latest release at generation time is Neo4j 2026.07.1.
- Vector indexes — current Cypher Manual — Current CREATE/SHOW syntax, provider capabilities, filter properties, HNSW settings, quantization, population state, and query guidance.
- Cypher 25 SEARCH clause — Preferred Neo4j 2026.01+ vector search syntax, SCORE, in-index filters, post-filters, dimensions, and ANN limitations.
- Vector values and types — VECTOR semantics, coordinate type/dimension, LIST differences, and Enterprise/block-format storage boundary.
- Vector functions — vector.similarity.cosine/euclidean and VECTOR construction/inspection functions.
- Cypher additions/deprecations — SEARCH/vector additions, 2026.06 quantization-setting deprecation, and 2026.07 HFQ status.
- Operations deprecations — db.index.vector.queryNodes/queryRelationships deprecated in 2026.04 in favor of SEARCH.
- Built-in procedures — Legacy vector-query procedure signatures and replacement status; useful for version fallback history.
- Cypher and Neo4j edition differences — Community supports vector indexes over LIST embeddings but cannot persist VECTOR values as properties.
- Semantic-index authorization limitations — Fine-grained authorization can conservatively suppress full-text/vector semantic-index results.
- Hybrid search developer guide — Rank-source independence and weighted reciprocal-rank fusion patterns; do not compare raw heterogeneous scores.
- Hybrid search engineering article — 2026 engineering discussion of words + meaning + graph topology and rank-based fusion.
- Neo4j GraphRAG for Python — Official maintained GraphRAG package, server compatibility, and current package naming.
- GraphRAG RAG guide — Retriever/generation separation and SEARCH-based in-index filtering on Neo4j 2026.01+.
- GraphRAG Knowledge Graph Builder — Document/Chunk lexical graph, NEXT_CHUNK/FROM_DOCUMENT concepts, entity extraction, and provenance structure.
- Python Driver 6.3 — Optional application/evaluation harness; current official Python driver and Bolt compatibility.
- System requirements — Neo4j 2026.07.1 server supports Java 21/25 on documented platforms.