Chapter 22 · Vector Search, Embeddings, Cypher SEARCH, Hybrid Search, and GraphRAG
Embeddings, VECTOR/LIST Values, Dimensionality, Similarity Functions, HNSW Concepts, and ANN Tradeoffs
Build a precise mental model of embeddings and approximate nearest-neighbor retrieval, distinguish Community LIST embeddings from Enterprise/Aura VECTOR storage, and measure exact-vs-HNSW behavior with deterministic AtlasMart vectors.
AtlasMart has a search query that lexical matching only partially solves: “a camera for wildlife on remote trails at night.” The product text contains overlapping words, but the application also wants meaning-level similarity so that semantically related items can be retrieved even when the exact words differ. The engineering risk is to treat an embedding as magical meaning or an ANN index as exact truth. This lesson builds the mechanism from numeric vectors to HNSW candidate search and validates it against an exact baseline.
An embedding is a coordinate supplied by an encoder. Neo4j stores/indexes those coordinates. HNSW is a fast navigation structure for finding nearby coordinates approximately. The database does not know whether “nearby” means useful for AtlasMart until you test the encoder, similarity function, corpus and judged queries.
Learning outcomes
Explain embedding generation as an external/model lifecycle concern and distinguish vector retrieval from graph traversal.
Distinguish Community LIST embeddings from Enterprise/Aura VECTOR property storage, including dimensions and coordinate-type implications.
Choose cosine vs euclidean by embedding semantics and use vector.similarity functions as an exact small-corpus baseline.
Explain HNSW/ANN recall-vs-latency tradeoffs and why approximate results must be measured rather than assumed exact.
Run the deterministic AtlasMart fixture and compare exact top-k with SEARCH top-k without requiring any paid model/API.
Current Neo4j Database is 2026.07.1; the current
5.26 line remains LTS. Version-sensitive examples use explicit
CYPHER 25. The mandatory lab uses self-managed
Neo4j Community 2026.07.1, database
neo4j, user neo4j, disposable password
atlasmart-course-2026, loopback Bolt
7687 and HTTP 7474, and embeddings
stored as LIST<FLOAT>. Neo4j 2026.x supports
Java 21/25. Optional client examples pin the official Python
driver to neo4j==6.3.0. No APOC, GDS, paid
embedding API, paid LLM API, Aura account, or Enterprise license
is required.
Vector indexes are available in Community when
embeddings are stored as LIST<INTEGER|FLOAT>.
The newer fixed-size VECTOR property type requires
block-format storage and therefore cannot be persisted as a
property in Community; it is an Enterprise/Aura storage
capability. The lab deliberately uses LIST embeddings so every
mandatory index/search/evaluation step remains free/local. Where
VECTOR-specific storage is discussed, it is labeled as an
edition-dependent optimization/typing choice rather than a
prerequisite.
From Neo4j 2026.01, Cypher 25
SEARCH is the preferred way to query vector indexes
and supports in-index filtering when filter properties were
declared in the index.
db.index.vector.queryNodes() and
db.index.vector.queryRelationships() remain useful
for older-version compatibility history but are deprecated from
Neo4j 2026.04. New course code therefore uses
SEARCH.
Lab contract and exact assumptions
| Dimension | Chapter 22 assumption |
|---|---|
| server | Neo4j Community 2026.07.1, single disposable local database |
| Cypher | Explicit CYPHER 25 for SEARCH and current vector syntax |
| Java | Java 21 or 25 for Neo4j 2026.07 |
| database/auth | neo4j / neo4j / atlasmart-course-2026 |
| transport | bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; production/remote deployments use verified TLS |
| plugins | none required; APOC/GDS/GenAI are not needed |
| embedding source | deterministic 8-dimensional precomputed AtlasMart vectors; not a paid API and not claimed to be production-quality embeddings |
| storage |
LIST |
| graph | 8 Products, 4 Categories, 1 Store, 2 KnowledgeDocuments, 4 Chunks plus provenance/entity edges |
| indexes | full-text product index + 8D product vector index + 8D chunk vector index |
| measurement | learner measures recall@k, runtime latency, index state/options and result IDs; generated lesson never claims that Neo4j was executed here |
| Term | Mechanism-first meaning |
|---|---|
| embedding | Numeric representation produced outside the database by a model or deterministic encoder. Neo4j stores/indexes the values; it does not make semantic truth guarantees about the encoder. |
| dimension | Number of coordinates in an embedding. Index dimension and query-vector dimension must match when dimensions are configured. |
| LIST embedding |
Community-compatible numeric property such as
[0.95,0.85,...]. Individual elements are
list-accessible.
|
| VECTOR value | Fixed-length typed vector value introduced in 2025.10; more storage-efficient typing but persisted VECTOR properties require Enterprise/Aura block format. |
| similarity | Function that converts a pair of vectors into an ordering signal. Current vector indexes support cosine and euclidean similarity. |
| ANN | Approximate nearest-neighbor retrieval. It trades guaranteed exactness for scalable search speed/resource behavior. |
| HNSW | Hierarchical Navigable Small World graph used internally by the vector index to navigate candidate neighborhoods rather than compare every stored vector. |
| recall@k | Fraction of the exact top-k neighbors recovered by ANN top-k. It is a retrieval-quality measure, not semantic correctness. |
| filter property | Non-vector property explicitly stored with a 2026.01+ vector index so SEARCH can apply supported predicates inside the ANN search. |
| quantization | Compressed vector representation used inside the index to lower memory/storage and often improve speed, potentially trading accuracy; 2026.07 supports high-fidelity rescoring through search expansion. |
| GraphRAG | Retrieval-augmented generation pattern where graph-structured evidence, provenance, and relationships enrich the context given to a generator. Retrieval quality and generator factuality still require evaluation. |
1. Embedding generation is outside the index
The course fixture uses hand-authored deterministic 8D vectors. Each dimension loosely represents a controlled feature family only so the mechanism is inspectable. A production embedding model may have hundreds or thousands of opaque learned dimensions. Neo4j accepts the numeric representation; model choice, text preprocessing, batching, rate limits, version upgrades, drift and privacy are application/MLOps responsibilities.
| Layer | Owns what | Failure if ignored |
|---|---|---|
| encoder | text/image/etc → numeric vector; model/version/dimension/normalization | re-embedding changes neighborhoods; incompatible dimensions break queries |
| Neo4j property | persist LIST or VECTOR values | missing/wrong values are not indexed or cannot be compared |
| vector index | ANN access path + configured similarity/filter metadata | POPULATING/unhealthy/wrong metric changes availability/quality |
| retrieval service | query vector, k, filters, timeouts, fusion | unbounded k, stale embedding version, hidden filters or raw-score misuse |
| evaluation | judged queries + exact baseline + latency | demo anecdotes become false quality claims |
2. LIST vs VECTOR is a storage/type decision, not two meanings
Both forms can represent the same embedding. Community cannot
persist VECTOR properties because VECTOR storage
requires block format, so the mandatory lab uses numeric lists.
Enterprise/Aura can store VECTOR values with fixed dimension and
coordinate type, which improves typing/storage efficiency and
supports vector-specific functions. Do not convert a Community
course into a paid-only lab merely to use the newer property
type.
| Representation | Community 2026.07 | Enterprise/Aura | Operational consequence |
|---|---|---|---|
| LIST<FLOAT> | store + vector-index | store + vector-index | portable course path; dimensions enforced by index if configured |
| VECTOR | cannot persist as property | supported on block format | fixed dimension/coordinate type and potentially more efficient representation |
| query vector | LIST parameter works for lab SEARCH | LIST/VECTOR usable depending API/query | client contract must preserve dimension and numeric validity |
3. Build the deterministic fixture
Run the fixture in a disposable Chapter 22 database. The vectors are deliberately small so you can inspect exact similarities. The labels, stock/category relationships and product text continue the AtlasMart search fixture from Chapter 21, while Document/Chunk nodes create the provenance substrate used in Lesson 5.
CYPHER 25
// Disposable Chapter 22 fixture. Safe to rerun after the cleanup block.
CREATE CONSTRAINT ch22_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch22_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch22_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch22_doc_id IF NOT EXISTS
FOR (d:KnowledgeDocument) REQUIRE d.documentId IS UNIQUE;
CREATE CONSTRAINT ch22_chunk_id IF NOT EXISTS
FOR (c:Chunk) REQUIRE c.chunkId IS UNIQUE;
MERGE (cam:Category {categoryId:'CAT-22-CAM'}) SET cam.name='Cameras', cam.labTag='ch22'
MERGE (out:Category {categoryId:'CAT-22-OUT'}) SET out.name='Outdoor', out.labTag='ch22'
MERGE (sec:Category {categoryId:'CAT-22-SEC'}) SET sec.name='Security', sec.labTag='ch22'
MERGE (acc:Category {categoryId:'CAT-22-ACC'}) SET acc.name='Accessories', acc.labTag='ch22'
MERGE (st:Store {storeId:'ST-22-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch22';
UNWIND [
{id:'P-2201',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:5,featured:true, emb:[0.95,0.85,0.25,0.05,0.10,0.90,0.05,0.05]},
{id:'P-2202',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:0,featured:false,emb:[0.90,0.80,0.15,0.05,0.05,0.82,0.05,0.05]},
{id:'P-2203',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,cat:'CAT-22-SEC',catCode:'SECURITY',region:'CENTRAL',qty:7,featured:false,emb:[0.92,0.10,0.95,0.02,0.05,0.05,0.05,0.05]},
{id:'P-2204',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:11,featured:false,emb:[0.02,0.88,0.02,0.95,0.10,0.10,0.05,0.02]},
{id:'P-2205',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:3,featured:true,emb:[0.90,0.45,0.20,0.10,0.95,0.15,0.05,0.03]},
{id:'P-2206',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:6,featured:false,emb:[0.05,0.55,0.05,0.05,0.05,0.95,0.05,0.02]},
{id:'P-2207',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,cat:'CAT-22-ACC',catCode:'ACCESSORY',region:'CENTRAL',qty:0,featured:false,emb:[0.08,0.75,0.10,0.05,0.10,0.10,0.95,0.02]},
{id:'P-2208',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,cat:'CAT-22-CAM',catCode:'CAMERA',region:'ARCHIVE',qty:2,featured:false,emb:[0.88,0.68,0.15,0.05,0.05,0.60,0.05,0.95]}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
p.active=row.active, p.categoryCode=row.catCode, p.region=row.region,
p.featured=row.featured, p.embedding=row.emb,
p.embeddingModel='atlasmart-deterministic-v1', p.embeddingVersion='2026-09-lab',
p.labTag='ch22'
WITH row,p
MATCH (cat:Category {categoryId:row.cat}), (st:Store {storeId:'ST-22-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch22';
MERGE (d1:KnowledgeDocument {documentId:'DOC-22-TRAILCAM'})
SET d1.title='Trail Camera Pro field manual', d1.uri='atlasmart://manuals/P-2201', d1.version='2026.09', d1.labTag='ch22'
MERGE (d2:KnowledgeDocument {documentId:'DOC-22-SECURITY'})
SET d2.title='Camera selection guide', d2.uri='atlasmart://guides/camera-selection', d2.version='2026.09', d2.labTag='ch22';
UNWIND [
{id:'CHK-2201',doc:'DOC-22-TRAILCAM',seq:1,text:'Trail Camera Pro is weatherproof and optimized for wildlife monitoring on outdoor trails.',emb:[0.96,0.86,0.12,0.02,0.02,0.94,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
{id:'CHK-2202',doc:'DOC-22-TRAILCAM',seq:2,text:'Infrared night vision records wildlife without visible illumination and battery life is designed for field deployment.',emb:[0.88,0.72,0.25,0.02,0.02,0.90,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
{id:'CHK-2203',doc:'DOC-22-SECURITY',seq:1,text:'Indoor Security Camera focuses on Wi-Fi motion alerts and indoor night vision rather than outdoor wildlife use.',emb:[0.84,0.08,0.96,0.02,0.02,0.06,0.02,0.02],products:['P-2203'],cats:['CAT-22-SEC']},
{id:'CHK-2204',doc:'DOC-22-SECURITY',seq:2,text:'Choose an outdoor trail camera when weather resistance and wildlife observation matter; choose indoor security cameras for home alerting.',emb:[0.90,0.68,0.55,0.02,0.02,0.72,0.02,0.02],products:['P-2201','P-2203'],cats:['CAT-22-CAM','CAT-22-SEC']}
] AS row
MERGE (c:Chunk {chunkId:row.id})
SET c.seq=row.seq, c.text=row.text, c.embedding=row.emb,
c.embeddingModel='atlasmart-deterministic-v1', c.embeddingVersion='2026-09-lab', c.labTag='ch22'
WITH row,c
MATCH (d:KnowledgeDocument {documentId:row.doc})
MERGE (c)-[:FROM_DOCUMENT]->(d)
WITH row,c
UNWIND row.products AS pid
MATCH (p:Product {productId:pid})
MERGE (c)-[:MENTIONS]->(p)
WITH row,c
UNWIND row.cats AS cid
MATCH (cat:Category {categoryId:cid})
MERGE (c)-[:MENTIONS]->(cat);
MATCH (a:Chunk {chunkId:'CHK-2201'}),(b:Chunk {chunkId:'CHK-2202'}) MERGE (a)-[:NEXT_CHUNK]->(b);
MATCH (a:Chunk {chunkId:'CHK-2203'}),(b:Chunk {chunkId:'CHK-2204'}) MERGE (a)-[:NEXT_CHUNK]->(b);
CREATE FULLTEXT INDEX ch22_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name,p.description,p.tags]
OPTIONS {indexConfig:{`fulltext.analyzer`:'english',`fulltext.eventually_consistent`:false}};
CREATE VECTOR INDEX ch22_product_vector IF NOT EXISTS
FOR (p:Product)
ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine',
`vector.quantization.type`:'scalar',
`vector.default_search_expansion_factor`:1.5
}};
CREATE VECTOR INDEX ch22_chunk_vector IF NOT EXISTS
FOR (c:Chunk)
ON c.embedding
WITH [c.embeddingVersion,c.seq]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine'
}};
CALL db.awaitIndexes(300);
CYPHER 25
SHOW VECTOR INDEXES
YIELD name,state,populationPercent,entityType,labelsOrTypes,properties,indexProvider,options,failureMessage
WHERE name STARTS WITH 'ch22_'
RETURN * ORDER BY name;
MATCH (p:Product {labTag:'ch22'})
RETURN count(p) AS products,
count { (p)-[:STOCKED_AT]->(:Store {storeId:'ST-22-CENTRAL'}) } AS stockEdges;
MATCH (c:Chunk {labTag:'ch22'}) RETURN count(c) AS chunks;
// Deterministic graph invariants: products=8, stockEdges=8, chunks=4.
// Before vector queries, both vector indexes must be ONLINE and 100% populated.
Vector indexes build in the background.
CALL db.awaitIndexes() makes the lab
deterministic, but production code should observe index
state/startup readiness and handle unavailable dependencies
instead of assuming immediate ONLINE state.
4. Dimensions and similarity functions define valid comparisons
The product index is explicitly 8-dimensional. A configured dimension rejects mismatched query vectors early instead of silently mixing unrelated shapes. Cosine compares direction (angle) while euclidean compares distance. On normalized vectors they often induce the same ordering, but the correct choice should follow the embedding model contract. Do not switch metrics merely because one benchmark looks better on five queries.
:param qvec => [1.0,0.8,0.1,0.0,0.0,0.9,0.0,0.0];
:param k => 4;
CYPHER 25
MATCH (p:Product {labTag:'ch22'})
WITH p, vector.similarity.cosine(p.embedding,$qvec) AS exactScore
ORDER BY exactScore DESC, p.productId
LIMIT $k
RETURN p.productId AS productId,p.name AS name,exactScore;
// For this fixed 8-row fixture, the exact top candidates should begin with P-2202 and P-2201.
// Capture the complete runtime result; do not hard-code ANN equivalence from this tiny graph.
CYPHER 25
:param badVec => [1.0,0.8,0.1,0.0,0.0,0.9,0.0];
MATCH (p:Product)
SEARCH p IN (
VECTOR INDEX ch22_product_vector
FOR $badVec
LIMIT 4
)
RETURN p.productId;
// Expected boundary: query fails because the configured vector dimension is 8 but badVec has 7 values.
5. What HNSW changes
A brute-force exact query computes similarity against every
candidate. HNSW maintains a graph-like navigation structure
inside the index: search enters at sparse upper layers and moves
through promising neighborhoods toward nearby vectors. This
greatly reduces comparisons at scale, but the path can miss an
exact neighbor. Construction parameters such as
vector.hnsw.m and
vector.hnsw.ef_construction influence graph
connectivity/build cost; 2026.07 search expansion and
quantization further alter recall/resource tradeoffs. These are
index-engine controls—not graph-domain relationships and not
universal tuning constants.
| Control | Mechanism | Possible gain | Possible cost/risk |
|---|---|---|---|
| dimensions | validation + index organization | early mismatch detection | requires stable embedding contract |
| HNSW M | maximum graph connectivity per vector | potentially better navigability/recall | larger/slower index population and updates |
| ef_construction | candidate breadth while building HNSW | higher-quality graph | population/update cost |
| quantization | compressed vector representation | lower memory/storage and often faster search | approximation error |
| search expansion | retrieve extra approximate candidates then return requested k; with quantization can rescore unquantized values | higher recall/HFQ | query time |
6. Exact vs ANN is an experiment
Run the exact and SEARCH queries with the same vector and k. For the fixed fixture, exact top candidates begin with P-2202 and P-2201. The ANN result may match exactly on such a tiny corpus; that is not evidence that ANN is exact in production. Compute set overlap as recall@k and repeat on a representative corpus.
CYPHER 25
MATCH (p:Product)
SEARCH p IN (
VECTOR INDEX ch22_product_vector
FOR $qvec
LIMIT $k
) SCORE AS similarityScore
RETURN p.productId AS productId,p.name AS name,similarityScore;
// SEARCH returns approximate nearest neighbors ordered by the index similarity signal.
from neo4j import GraphDatabase
from statistics import median
from time import perf_counter
URI="bolt://localhost:7687"
AUTH=("neo4j","atlasmart-course-2026")
QVEC=[1.0,0.8,0.1,0.0,0.0,0.9,0.0,0.0]
K=4
REPEATS=30
EXACT="""
MATCH (p:Product {labTag:'ch22'})
WITH p,vector.similarity.cosine(p.embedding,$qvec) AS score
ORDER BY score DESC,p.productId LIMIT $k
RETURN p.productId AS id,score
"""
ANN="""
MATCH (p:Product)
SEARCH p IN (VECTOR INDEX ch22_product_vector FOR $qvec LIMIT $k)
SCORE AS score
RETURN p.productId AS id,score
"""
with GraphDatabase.driver(URI,auth=AUTH) as driver:
exact,_,_=driver.execute_query(EXACT,qvec=QVEC,k=K,database_="neo4j")
exact_ids=[r["id"] for r in exact]
times=[]; ann_ids=[]
for _ in range(REPEATS):
t0=perf_counter()
ann,_,_=driver.execute_query(ANN,qvec=QVEC,k=K,database_="neo4j")
times.append((perf_counter()-t0)*1000)
ann_ids=[r["id"] for r in ann]
recall=len(set(exact_ids)&set(ann_ids))/K
print({"exact":exact_ids,"ann":ann_ids,"recall_at_k":recall,
"median_client_ms":median(times),"samples":REPEATS})
# This tiny fixture may produce recall@4 == 1.0. That proves only this fixture/run.
# Repeat on a representative corpus and record p50/p95/p99 plus server/resource evidence.
7. Wrong approaches and repairs
| Wrong approach | Concrete problem | Repair/verification |
|---|---|---|
| require a paid embedding API | lab becomes non-reproducible/cost dependent | precomputed/local deterministic vectors; record model metadata |
| call semantic similarity truth | nearby vectors may be irrelevant/unsafe | judge queries and downstream business/graph rules |
| omit dimensions | mixed dimensions can coexist and fail/confuse later | configure dimensions and test mismatch |
| assume ANN=exact | missed neighbors hidden | exact baseline + recall@k |
| tune HNSW from internet folklore | resource/quality regressions | controlled corpus benchmark with one variable at a time |
| store VECTOR in Community | unsupported property storage boundary | LIST embeddings in Community; label VECTOR as Enterprise/Aura |
8. Production judgment
| Production decision | Evidence required |
|---|---|
| embedding lifecycle | Record model/provider/version, dimensions, normalization assumptions, text preprocessing, backfill/re-embedding status, and rollback/index-swap plan. |
| recall vs latency | Measure exact-vs-ANN recall@k on a representative judged set alongside p50/p95/p99; tiny demo recall is not a capacity guarantee. |
| model/cardinality/degree | Separate candidate retrieval count from graph expansion fan-out; bound traversal and final context size. |
| index memory/storage/write cost | Observe vector index size, population/rebuild time, write amplification, page-cache/store pressure, and quantization effects before tuning. |
| similarity semantics | Choose cosine/euclidean from embedding-model semantics; score is a source-specific similarity signal, not factual probability. |
| filtering/security | Declare filter properties intentionally, distinguish in-index from post-filter behavior, and test the real Enterprise service role because semantic-index authorization can suppress candidates. |
| driver/timeouts/retries | Use a long-lived driver, bounded candidate counts, transaction timeouts, idempotent writes and explicit retry/error classification from earlier chapters. |
| hybrid ranking | Fuse independent source ranks (for example RRF/WRRF) or use a trained evaluated re-ranker; never add incomparable raw full-text/vector scores by habit. |
| GraphRAG provenance | Every context unit carries source ID/URI/version/chunk ID and graph entities/relationships so retrieval evidence can be audited. |
| evaluation | Track retrieval recall/precision/MRR/nDCG-style metrics, answer grounding/citation correctness if a generator is added, latency, freshness and zero-result/fallback rates. |
| privacy/tenant risk | Do not embed secrets/PII without policy; authorization must be enforced before context reaches a generator or user. |
| backup/recovery | Rebuild/validate vector indexes and embedding-version metadata in restore drills; restore of graph data is not proof that semantic retrieval is healthy. |
| Aura/self-managed | Aura manages infrastructure and some controls; self-managed exposes server/index/plugin/resource operations. Verify feature/tier availability instead of assuming parity. |
| licensing/cost | Community mandatory lab is free/local. VECTOR storage, Enterprise security/clustering and paid embedding/LLM services are optional and separately costed/licensed. |
| migration/rollback | Run old/new embedding/index versions side by side when possible, freeze evaluation data, cut over by explicit index/query configuration, retain rollback until validation passes. |
Check your understanding
- Who generates an embedding?
- Why does the Community lab store LIST values?
- What does recall@k compare?
- Does a high vector score prove a product is correct for the user?
- Why specify vector dimensions in the index?
Review the answers
1. An external/local encoder or deterministic application process; Neo4j indexes the supplied numeric values.
2. Community supports vector indexes over numeric LIST properties but cannot persist VECTOR properties because it lacks block-format VECTOR storage.
3. The ANN top-k set against an exact top-k baseline for the same query/corpus.
4. No. It is similarity under one encoder/metric, not business eligibility or factual truth.
5. It enforces the embedding contract and makes mismatched query vectors fail clearly.
Summary and next step
You now have an exact baseline, an ANN vector index, and a measurable definition of retrieval quality. Lesson 2 turns the index itself into an observable production artifact: schema, filter metadata, state, provider, HNSW settings, quantization, expansion, rebuild and rollback.
Authoritative references
- Neo4j Operations Manual — current release — Current server line; the latest release at generation time is Neo4j 2026.07.1.
- Vector indexes — current Cypher Manual — Current CREATE/SHOW syntax, provider capabilities, filter properties, HNSW settings, quantization, population state, and query guidance.
- Cypher 25 SEARCH clause — Preferred Neo4j 2026.01+ vector search syntax, SCORE, in-index filters, post-filters, dimensions, and ANN limitations.
- Vector values and types — VECTOR semantics, coordinate type/dimension, LIST differences, and Enterprise/block-format storage boundary.
- Vector functions — vector.similarity.cosine/euclidean and VECTOR construction/inspection functions.
- Cypher additions/deprecations — SEARCH/vector additions, 2026.06 quantization-setting deprecation, and 2026.07 HFQ status.
- Operations deprecations — db.index.vector.queryNodes/queryRelationships deprecated in 2026.04 in favor of SEARCH.
- Built-in procedures — Legacy vector-query procedure signatures and replacement status; useful for version fallback history.
- Cypher and Neo4j edition differences — Community supports vector indexes over LIST embeddings but cannot persist VECTOR values as properties.
- Semantic-index authorization limitations — Fine-grained authorization can conservatively suppress full-text/vector semantic-index results.
- Hybrid search developer guide — Rank-source independence and weighted reciprocal-rank fusion patterns; do not compare raw heterogeneous scores.
- Hybrid search engineering article — 2026 engineering discussion of words + meaning + graph topology and rank-based fusion.
- Neo4j GraphRAG for Python — Official maintained GraphRAG package, server compatibility, and current package naming.
- GraphRAG RAG guide — Retriever/generation separation and SEARCH-based in-index filtering on Neo4j 2026.01+.
- GraphRAG Knowledge Graph Builder — Document/Chunk lexical graph, NEXT_CHUNK/FROM_DOCUMENT concepts, entity extraction, and provenance structure.
- Python Driver 6.3 — Optional application/evaluation harness; current official Python driver and Bolt compatibility.
- System requirements — Neo4j 2026.07.1 server supports Java 21/25 on documented platforms.