Chapter 22 · Vector Search, Embeddings, Cypher SEARCH, Hybrid Search, and GraphRAG
Create and Inspect Vector Indexes, Additional Filter Properties, Population State, and Quantization Options
Create current 2026.07 vector indexes deliberately: dimensions, similarity, filterable properties, population state, provider/options, HNSW construction settings, scalar/binary quantization, and high-fidelity search expansion without cargo-cult tuning.
AtlasMart can create a vector index with one line, but production reliability depends on what that line commits the organization to: one embedding property, dimensions, similarity, additional filter properties, provider generation, quantization, HNSW build parameters, search expansion, population time, rebuild cost and version behavior. This lesson treats the index as a versioned operational object rather than an invisible “AI feature.”
The index schema is a contract between stored vectors, filter metadata and the SEARCH planner. The options are retrieval-engine policy. Every change may require a new index population and a fresh quality/latency benchmark.
Learning outcomes
Create and inspect current 2026.07 vector indexes and interpret state, populationPercent, provider, properties and options.
Explain 2026.01+ multi-label/type and additional filter-property semantics, including the one-vector-property limit.
Understand scalar/binary/none quantization and 2026.07 high-fidelity search expansion without treating defaults as universal tuning advice.
Observe POPULATING→ONLINE readiness and design rebuild/cutover/rollback rather than querying an unavailable index.
Prove filter-property boundaries with a controlled failure and distinguish Community LIST storage from optional VECTOR storage.
Current Neo4j Database is 2026.07.1; the current
5.26 line remains LTS. Version-sensitive examples use explicit
CYPHER 25. The mandatory lab uses self-managed
Neo4j Community 2026.07.1, database
neo4j, user neo4j, disposable password
atlasmart-course-2026, loopback Bolt
7687 and HTTP 7474, and embeddings
stored as LIST<FLOAT>. Neo4j 2026.x supports
Java 21/25. Optional client examples pin the official Python
driver to neo4j==6.3.0. No APOC, GDS, paid
embedding API, paid LLM API, Aura account, or Enterprise license
is required.
Vector indexes are available in Community when
embeddings are stored as LIST<INTEGER|FLOAT>.
The newer fixed-size VECTOR property type requires
block-format storage and therefore cannot be persisted as a
property in Community; it is an Enterprise/Aura storage
capability. The lab deliberately uses LIST embeddings so every
mandatory index/search/evaluation step remains free/local. Where
VECTOR-specific storage is discussed, it is labeled as an
edition-dependent optimization/typing choice rather than a
prerequisite.
From Neo4j 2026.01, Cypher 25
SEARCH is the preferred way to query vector indexes
and supports in-index filtering when filter properties were
declared in the index.
db.index.vector.queryNodes() and
db.index.vector.queryRelationships() remain useful
for older-version compatibility history but are deprecated from
Neo4j 2026.04. New course code therefore uses
SEARCH.
Lab contract and exact assumptions
| Dimension | Chapter 22 assumption |
|---|---|
| server | Neo4j Community 2026.07.1, single disposable local database |
| Cypher | Explicit CYPHER 25 for SEARCH and current vector syntax |
| Java | Java 21 or 25 for Neo4j 2026.07 |
| database/auth | neo4j / neo4j / atlasmart-course-2026 |
| transport | bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; production/remote deployments use verified TLS |
| plugins | none required; APOC/GDS/GenAI are not needed |
| embedding source | deterministic 8-dimensional precomputed AtlasMart vectors; not a paid API and not claimed to be production-quality embeddings |
| storage |
LIST |
| graph | 8 Products, 4 Categories, 1 Store, 2 KnowledgeDocuments, 4 Chunks plus provenance/entity edges |
| indexes | full-text product index + 8D product vector index + 8D chunk vector index |
| measurement | learner measures recall@k, runtime latency, index state/options and result IDs; generated lesson never claims that Neo4j was executed here |
| Term | Mechanism-first meaning |
|---|---|
| embedding | Numeric representation produced outside the database by a model or deterministic encoder. Neo4j stores/indexes the values; it does not make semantic truth guarantees about the encoder. |
| dimension | Number of coordinates in an embedding. Index dimension and query-vector dimension must match when dimensions are configured. |
| LIST embedding |
Community-compatible numeric property such as
[0.95,0.85,...]. Individual elements are
list-accessible.
|
| VECTOR value | Fixed-length typed vector value introduced in 2025.10; more storage-efficient typing but persisted VECTOR properties require Enterprise/Aura block format. |
| similarity | Function that converts a pair of vectors into an ordering signal. Current vector indexes support cosine and euclidean similarity. |
| ANN | Approximate nearest-neighbor retrieval. It trades guaranteed exactness for scalable search speed/resource behavior. |
| HNSW | Hierarchical Navigable Small World graph used internally by the vector index to navigate candidate neighborhoods rather than compare every stored vector. |
| recall@k | Fraction of the exact top-k neighbors recovered by ANN top-k. It is a retrieval-quality measure, not semantic correctness. |
| filter property | Non-vector property explicitly stored with a 2026.01+ vector index so SEARCH can apply supported predicates inside the ANN search. |
| quantization | Compressed vector representation used inside the index to lower memory/storage and often improve speed, potentially trading accuracy; 2026.07 supports high-fidelity rescoring through search expansion. |
| GraphRAG | Retrieval-augmented generation pattern where graph-structured evidence, provenance, and relationships enrich the context given to a generator. Retrieval quality and generator factuality still require evaluation. |
Direct-entry setup
If Lesson 1 is not already loaded, run the complete fixture. It creates two vector indexes with deterministic LIST embeddings.
CYPHER 25
// Disposable Chapter 22 fixture. Safe to rerun after the cleanup block.
CREATE CONSTRAINT ch22_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch22_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch22_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch22_doc_id IF NOT EXISTS
FOR (d:KnowledgeDocument) REQUIRE d.documentId IS UNIQUE;
CREATE CONSTRAINT ch22_chunk_id IF NOT EXISTS
FOR (c:Chunk) REQUIRE c.chunkId IS UNIQUE;
MERGE (cam:Category {categoryId:'CAT-22-CAM'}) SET cam.name='Cameras', cam.labTag='ch22'
MERGE (out:Category {categoryId:'CAT-22-OUT'}) SET out.name='Outdoor', out.labTag='ch22'
MERGE (sec:Category {categoryId:'CAT-22-SEC'}) SET sec.name='Security', sec.labTag='ch22'
MERGE (acc:Category {categoryId:'CAT-22-ACC'}) SET acc.name='Accessories', acc.labTag='ch22'
MERGE (st:Store {storeId:'ST-22-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch22';
UNWIND [
{id:'P-2201',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:5,featured:true, emb:[0.95,0.85,0.25,0.05,0.10,0.90,0.05,0.05]},
{id:'P-2202',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:0,featured:false,emb:[0.90,0.80,0.15,0.05,0.05,0.82,0.05,0.05]},
{id:'P-2203',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,cat:'CAT-22-SEC',catCode:'SECURITY',region:'CENTRAL',qty:7,featured:false,emb:[0.92,0.10,0.95,0.02,0.05,0.05,0.05,0.05]},
{id:'P-2204',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:11,featured:false,emb:[0.02,0.88,0.02,0.95,0.10,0.10,0.05,0.02]},
{id:'P-2205',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:3,featured:true,emb:[0.90,0.45,0.20,0.10,0.95,0.15,0.05,0.03]},
{id:'P-2206',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:6,featured:false,emb:[0.05,0.55,0.05,0.05,0.05,0.95,0.05,0.02]},
{id:'P-2207',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,cat:'CAT-22-ACC',catCode:'ACCESSORY',region:'CENTRAL',qty:0,featured:false,emb:[0.08,0.75,0.10,0.05,0.10,0.10,0.95,0.02]},
{id:'P-2208',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,cat:'CAT-22-CAM',catCode:'CAMERA',region:'ARCHIVE',qty:2,featured:false,emb:[0.88,0.68,0.15,0.05,0.05,0.60,0.05,0.95]}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
p.active=row.active, p.categoryCode=row.catCode, p.region=row.region,
p.featured=row.featured, p.embedding=row.emb,
p.embeddingModel='atlasmart-deterministic-v1', p.embeddingVersion='2026-09-lab',
p.labTag='ch22'
WITH row,p
MATCH (cat:Category {categoryId:row.cat}), (st:Store {storeId:'ST-22-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch22';
MERGE (d1:KnowledgeDocument {documentId:'DOC-22-TRAILCAM'})
SET d1.title='Trail Camera Pro field manual', d1.uri='atlasmart://manuals/P-2201', d1.version='2026.09', d1.labTag='ch22'
MERGE (d2:KnowledgeDocument {documentId:'DOC-22-SECURITY'})
SET d2.title='Camera selection guide', d2.uri='atlasmart://guides/camera-selection', d2.version='2026.09', d2.labTag='ch22';
UNWIND [
{id:'CHK-2201',doc:'DOC-22-TRAILCAM',seq:1,text:'Trail Camera Pro is weatherproof and optimized for wildlife monitoring on outdoor trails.',emb:[0.96,0.86,0.12,0.02,0.02,0.94,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
{id:'CHK-2202',doc:'DOC-22-TRAILCAM',seq:2,text:'Infrared night vision records wildlife without visible illumination and battery life is designed for field deployment.',emb:[0.88,0.72,0.25,0.02,0.02,0.90,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
{id:'CHK-2203',doc:'DOC-22-SECURITY',seq:1,text:'Indoor Security Camera focuses on Wi-Fi motion alerts and indoor night vision rather than outdoor wildlife use.',emb:[0.84,0.08,0.96,0.02,0.02,0.06,0.02,0.02],products:['P-2203'],cats:['CAT-22-SEC']},
{id:'CHK-2204',doc:'DOC-22-SECURITY',seq:2,text:'Choose an outdoor trail camera when weather resistance and wildlife observation matter; choose indoor security cameras for home alerting.',emb:[0.90,0.68,0.55,0.02,0.02,0.72,0.02,0.02],products:['P-2201','P-2203'],cats:['CAT-22-CAM','CAT-22-SEC']}
] AS row
MERGE (c:Chunk {chunkId:row.id})
SET c.seq=row.seq, c.text=row.text, c.embedding=row.emb,
c.embeddingModel='atlasmart-deterministic-v1', c.embeddingVersion='2026-09-lab', c.labTag='ch22'
WITH row,c
MATCH (d:KnowledgeDocument {documentId:row.doc})
MERGE (c)-[:FROM_DOCUMENT]->(d)
WITH row,c
UNWIND row.products AS pid
MATCH (p:Product {productId:pid})
MERGE (c)-[:MENTIONS]->(p)
WITH row,c
UNWIND row.cats AS cid
MATCH (cat:Category {categoryId:cid})
MERGE (c)-[:MENTIONS]->(cat);
MATCH (a:Chunk {chunkId:'CHK-2201'}),(b:Chunk {chunkId:'CHK-2202'}) MERGE (a)-[:NEXT_CHUNK]->(b);
MATCH (a:Chunk {chunkId:'CHK-2203'}),(b:Chunk {chunkId:'CHK-2204'}) MERGE (a)-[:NEXT_CHUNK]->(b);
CREATE FULLTEXT INDEX ch22_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name,p.description,p.tags]
OPTIONS {indexConfig:{`fulltext.analyzer`:'english',`fulltext.eventually_consistent`:false}};
CREATE VECTOR INDEX ch22_product_vector IF NOT EXISTS
FOR (p:Product)
ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine',
`vector.quantization.type`:'scalar',
`vector.default_search_expansion_factor`:1.5
}};
CREATE VECTOR INDEX ch22_chunk_vector IF NOT EXISTS
FOR (c:Chunk)
ON c.embedding
WITH [c.embeddingVersion,c.seq]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine'
}};
CALL db.awaitIndexes(300);
1. Read the index as schema + provider + state + options
SHOW VECTOR INDEXES is the primary observable
surface. On Neo4j 2026.07, newly created indexes select the
current provider automatically; older providers can continue
functioning after upgrades. Record provider and create statement
in deployment evidence because an index can remain on an older
provider until deliberately rebuilt.
CYPHER 25
SHOW VECTOR INDEXES YIELD *
WHERE name STARTS WITH 'ch22_'
RETURN name,state,populationPercent,indexProvider,entityType,
labelsOrTypes,properties,options,failureMessage,createStatement
ORDER BY name;
| Field | What it proves | What it does not prove |
|---|---|---|
| state=ONLINE | index is available to query | relevance/recall/SLO correctness |
| populationPercent=100 | population finished | every source entity has valid embedding |
| properties | vector property + additional filter properties included | which property is semantically good |
| indexProvider | implementation/provider generation | that old/new provider has equal benchmark behavior |
| options | configured dimensions/similarity/HNSW/quantization/expansion | that defaults suit your workload |
| failureMessage | population/build failure evidence | root cause without logs/data inspection |
2. Additional filter properties are index metadata, not arbitrary Cypher
Since 2026.01, vector indexes can include non-vector properties
with WITH [...]. SEARCH can evaluate supported
predicates on those properties inside ANN retrieval. This
matters for selective filters: in-index filtering keeps
searching until it finds enough qualifying neighbors, whereas a
normal WHERE after SEARCH can simply discard already-selected
candidates.
CYPHER 25
CREATE VECTOR INDEX ch22_product_vector IF NOT EXISTS
FOR (p:Product)
ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine',
`vector.quantization.type`:'scalar',
`vector.default_search_expansion_factor`:1.5
}};
The index may carry multiple labels/types and multiple additional filter properties, but only one property is the indexed vector. Filter metadata must be deliberately chosen because it changes index size/update cost and the predicates SEARCH can push inside the index.
3. Prove the filter-property boundary
The product index declares active,
categoryCode and region as filterable
metadata. featured exists on Product but is not
stored in this vector index. The first query is valid; the
second intentionally fails. That failure is useful evidence that
in-index filtering is tied to index schema, not all graph
properties.
CYPHER 25
MATCH (p:Product)
SEARCH p IN (
VECTOR INDEX ch22_product_vector
FOR $qvec
WHERE p.active = true AND p.categoryCode = 'CAMERA'
LIMIT 4
) SCORE AS similarityScore
RETURN p.productId AS productId,p.name AS name,p.active,p.categoryCode,similarityScore;
// The filter is inside SEARCH because active/categoryCode were declared in WITH [...] at index creation.
CYPHER 25
MATCH (p:Product)
SEARCH p IN (
VECTOR INDEX ch22_product_vector
FOR $qvec
WHERE p.featured = true
LIMIT 4
)
RETURN p.productId;
// Expected boundary: featured was NOT declared as an additional filter property in ch22_product_vector.
4. Population readiness is a dependency, not a sleep timer
Index creation is asynchronous. A fixed sleep such as “wait five
seconds” races with data size/hardware and hides failures. For a
deterministic lab, db.awaitIndexes() blocks to a
timeout; for services/deployments, read state/failure evidence
and refuse vector traffic until required indexes are ONLINE.
Blue/green index rollout can build a new named index, test it,
switch query configuration, then retire the old index after
rollback windows expire.
| Rollout step | Evidence |
|---|---|
| create new named index | create statement committed; old index remains serving |
| population | state POPULATING; populationPercent grows; monitor failureMessage/logs |
| readiness | state ONLINE + evaluation corpus passes |
| cutover | service index name/config updated; p95/p99 + recall checked |
| rollback | switch service back to old index if regression |
| retire | drop old index only after confidence/backup/runbook window |
5. Quantization and high-fidelity search
Neo4j 2026.07 supports none,
scalar and binary quantization types.
Quantization can reduce vector-index memory/storage and improve
speed while introducing approximation error. Search expansion
asks the index for more candidates internally than the requested
k; when quantization is enabled, 2026.07 can rescore returned
neighbors with unquantized values for high-fidelity quantized
search. More expansion can improve accuracy but costs query
time.
| Option | 2026.07 meaning | Benchmark question |
|---|---|---|
| vector.quantization.type=none | unquantized index vectors | Is memory/storage acceptable and recall/latency better? |
| scalar | less aggressive compression; current default | Does compression meet recall@k and latency/resource SLO? |
| binary | 1-bit-per-dimension style aggressive compression | Does higher compression preserve enough recall with expansion? |
| vector.default_search_expansion_factor | internal candidate expansion before final k | How does recall@k/p99/resource use change? |
// Do not run against production blindly. Use a disposable benchmark database.
CREATE VECTOR INDEX ch22_product_vector_none IF NOT EXISTS
FOR (p:Product) ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine',
`vector.quantization.type`:'none',
`vector.default_search_expansion_factor`:1.0
}};
CREATE VECTOR INDEX ch22_product_vector_binary IF NOT EXISTS
FOR (p:Product) ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
`vector.dimensions`:8,
`vector.similarity_function`:'cosine',
`vector.quantization.type`:'binary',
`vector.default_search_expansion_factor`:3.0
}};
// Re-run the same judged queries/recall/latency harness; never infer quality from option names.
6. HNSW construction settings trade build/update resources for graph quality
vector.hnsw.m controls maximum connectivity and
vector.hnsw.ef_construction controls candidate
breadth during insertion. Higher values can improve
navigability/recall with diminishing returns, while increasing
population/update cost. Keep the lab defaults unless a
representative benchmark demonstrates a need; changing them
without a corpus and resource measurements is cargo-cult tuning.
Dimension count, corpus size/distribution, update rate, memory, hardware and required recall all change the optimum. Treat every option as an experimental variable with a rollback path.
7. VECTOR-property path is optional and edition-bound
Enterprise/Aura can use fixed-size VECTOR properties and coordinate types such as FLOAT32. That can improve storage/typing, but it introduces a migration decision: re-write embedding properties, verify block-format/edition compatibility, rebuild indexes if needed, and validate drivers/tooling. The Community LIST path remains semantically valid for indexing/search and is the portable course baseline.
8. Wrong approaches and repairs
| Wrong approach | Problem | Repair |
|---|---|---|
| query immediately after CREATE | POPULATING index is unusable | state/readiness gate |
| use p.featured in SEARCH filter without WITH metadata | query error | declare filter property or post-filter deliberately |
| use deprecated vector.quantization.enabled | stale 2026.06+ configuration | use vector.quantization.type |
| set binary because “fastest” | unknown recall regression | same-corpus recall/latency/resource benchmark |
| rebuild in place with no old index | rollback gap | parallel named index + cutover |
| assume VECTOR is Community storage | unsupported boundary | LIST in Community; VECTOR optional Enterprise/Aura |
9. Production judgment
| Production decision | Evidence required |
|---|---|
| embedding lifecycle | Record model/provider/version, dimensions, normalization assumptions, text preprocessing, backfill/re-embedding status, and rollback/index-swap plan. |
| recall vs latency | Measure exact-vs-ANN recall@k on a representative judged set alongside p50/p95/p99; tiny demo recall is not a capacity guarantee. |
| model/cardinality/degree | Separate candidate retrieval count from graph expansion fan-out; bound traversal and final context size. |
| index memory/storage/write cost | Observe vector index size, population/rebuild time, write amplification, page-cache/store pressure, and quantization effects before tuning. |
| similarity semantics | Choose cosine/euclidean from embedding-model semantics; score is a source-specific similarity signal, not factual probability. |
| filtering/security | Declare filter properties intentionally, distinguish in-index from post-filter behavior, and test the real Enterprise service role because semantic-index authorization can suppress candidates. |
| driver/timeouts/retries | Use a long-lived driver, bounded candidate counts, transaction timeouts, idempotent writes and explicit retry/error classification from earlier chapters. |
| hybrid ranking | Fuse independent source ranks (for example RRF/WRRF) or use a trained evaluated re-ranker; never add incomparable raw full-text/vector scores by habit. |
| GraphRAG provenance | Every context unit carries source ID/URI/version/chunk ID and graph entities/relationships so retrieval evidence can be audited. |
| evaluation | Track retrieval recall/precision/MRR/nDCG-style metrics, answer grounding/citation correctness if a generator is added, latency, freshness and zero-result/fallback rates. |
| privacy/tenant risk | Do not embed secrets/PII without policy; authorization must be enforced before context reaches a generator or user. |
| backup/recovery | Rebuild/validate vector indexes and embedding-version metadata in restore drills; restore of graph data is not proof that semantic retrieval is healthy. |
| Aura/self-managed | Aura manages infrastructure and some controls; self-managed exposes server/index/plugin/resource operations. Verify feature/tier availability instead of assuming parity. |
| licensing/cost | Community mandatory lab is free/local. VECTOR storage, Enterprise security/clustering and paid embedding/LLM services are optional and separately costed/licensed. |
| migration/rollback | Run old/new embedding/index versions side by side when possible, freeze evaluation data, cut over by explicit index/query configuration, retain rollback until validation passes. |
Check your understanding
- What does populationPercent=100 prove?
- Why include a property in WITH [...]?
- Why can binary quantization need more search expansion?
- Why keep an old index during rollout?
- Should you copy HNSW defaults from a tutorial into production tuning policy?
Review the answers
1. Index population completed, not that retrieval quality is acceptable.
2. To make that non-vector property available for supported in-index SEARCH filtering.
3. Aggressive compression can reduce ranking accuracy; more candidates plus high-fidelity rescoring can recover recall at extra query cost.
4. It provides a tested rollback path while the new provider/options are evaluated.
5. No. Start with supported defaults and change only from representative evidence.
Summary and next step
The vector index is now observable and versioned. Lesson 3 moves to the query plane: Cypher 25 SEARCH, SCORE, in-index vs post-filter semantics, PROFILE operators, dimension/null boundaries, deprecated procedure fallback history, and exact-vs-ANN measurement.
Authoritative references
- Neo4j Operations Manual — current release — Current server line; the latest release at generation time is Neo4j 2026.07.1.
- Vector indexes — current Cypher Manual — Current CREATE/SHOW syntax, provider capabilities, filter properties, HNSW settings, quantization, population state, and query guidance.
- Cypher 25 SEARCH clause — Preferred Neo4j 2026.01+ vector search syntax, SCORE, in-index filters, post-filters, dimensions, and ANN limitations.
- Vector values and types — VECTOR semantics, coordinate type/dimension, LIST differences, and Enterprise/block-format storage boundary.
- Vector functions — vector.similarity.cosine/euclidean and VECTOR construction/inspection functions.
- Cypher additions/deprecations — SEARCH/vector additions, 2026.06 quantization-setting deprecation, and 2026.07 HFQ status.
- Operations deprecations — db.index.vector.queryNodes/queryRelationships deprecated in 2026.04 in favor of SEARCH.
- Built-in procedures — Legacy vector-query procedure signatures and replacement status; useful for version fallback history.
- Cypher and Neo4j edition differences — Community supports vector indexes over LIST embeddings but cannot persist VECTOR values as properties.
- Semantic-index authorization limitations — Fine-grained authorization can conservatively suppress full-text/vector semantic-index results.
- Hybrid search developer guide — Rank-source independence and weighted reciprocal-rank fusion patterns; do not compare raw heterogeneous scores.
- Hybrid search engineering article — 2026 engineering discussion of words + meaning + graph topology and rank-based fusion.
- Neo4j GraphRAG for Python — Official maintained GraphRAG package, server compatibility, and current package naming.
- GraphRAG RAG guide — Retriever/generation separation and SEARCH-based in-index filtering on Neo4j 2026.01+.
- GraphRAG Knowledge Graph Builder — Document/Chunk lexical graph, NEXT_CHUNK/FROM_DOCUMENT concepts, entity extraction, and provenance structure.
- Python Driver 6.3 — Optional application/evaluation harness; current official Python driver and Bolt compatibility.
- System requirements — Neo4j 2026.07.1 server supports Java 21/25 on documented platforms.