Chapter 22 · Vector Search, Embeddings, Cypher SEARCH, Hybrid Search, and GraphRAG

Query Vector Indexes with the Modern Cypher SEARCH Clause and Understand Approximate-Neighbor Limits

Use Cypher 25 SEARCH as the preferred current vector-query surface, observe SCORE and query-plan operators, distinguish in-index filtering from post-filtering, and measure recall@k rather than assuming ANN is exact.

Advanced230–330 minutesSEARCH + recall labNeo4j 2026.07.1 · Community mandatoryCypher 25 · Vector SEARCH · HNSW/ANNLIST embeddings mandatory · VECTOR storage optional EE/AuraJava 21/25 · Python driver 6.3 optionalLast reviewed: September 2026

AtlasMart has an ONLINE index, but query semantics still matter. A vector request can fail on dimension mismatch, return fewer rows after a post-filter, return no rows when the query vector is null, or recover a slightly different top-k than brute force because ANN is approximate. The current SEARCH clause makes these mechanics visible in Cypher instead of hiding them inside a procedure call.

Mental model

SEARCH is a planner-visible vector-index access path embedded in MATCH/OPTIONAL MATCH. LIMIT belongs to the ANN search, SCORE is source-specific similarity, and a WHERE placed inside SEARCH is an in-index filter only when its properties were declared in the vector index.

Learning outcomes

01

Use current Cypher 25 SEARCH syntax and SCORE for node vector retrieval over Community LIST embeddings.

02

Distinguish in-index filters from MATCH post-filters and explain why they can return different candidate counts.

03

Inspect PROFILE for vector-index search operators and separate database plan evidence from client/network timing.

04

Handle dimension mismatch, null query vectors, selective-filter/k boundaries and deprecated procedure fallback history.

05

Calculate exact-vs-ANN recall@k and latency distributions instead of describing vector search as exact.

Chapter 22 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; the current 5.26 line remains LTS. Version-sensitive examples use explicit CYPHER 25. The mandatory lab uses self-managed Neo4j Community 2026.07.1, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback Bolt 7687 and HTTP 7474, and embeddings stored as LIST<FLOAT>. Neo4j 2026.x supports Java 21/25. Optional client examples pin the official Python driver to neo4j==6.3.0. No APOC, GDS, paid embedding API, paid LLM API, Aura account, or Enterprise license is required.

Community VECTOR/LIST boundary

Vector indexes are available in Community when embeddings are stored as LIST<INTEGER|FLOAT>. The newer fixed-size VECTOR property type requires block-format storage and therefore cannot be persisted as a property in Community; it is an Enterprise/Aura storage capability. The lab deliberately uses LIST embeddings so every mandatory index/search/evaluation step remains free/local. Where VECTOR-specific storage is discussed, it is labeled as an edition-dependent optimization/typing choice rather than a prerequisite.

Current query surface

From Neo4j 2026.01, Cypher 25 SEARCH is the preferred way to query vector indexes and supports in-index filtering when filter properties were declared in the index. db.index.vector.queryNodes() and db.index.vector.queryRelationships() remain useful for older-version compatibility history but are deprecated from Neo4j 2026.04. New course code therefore uses SEARCH.

Lab contract and exact assumptions

Dimension Chapter 22 assumption
server Neo4j Community 2026.07.1, single disposable local database
Cypher Explicit CYPHER 25 for SEARCH and current vector syntax
Java Java 21 or 25 for Neo4j 2026.07
database/auth neo4j / neo4j / atlasmart-course-2026
transport bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; production/remote deployments use verified TLS
plugins none required; APOC/GDS/GenAI are not needed
embedding source deterministic 8-dimensional precomputed AtlasMart vectors; not a paid API and not claimed to be production-quality embeddings
storage LIST so Community can store/index every embedding; VECTOR storage is discussed as Enterprise/Aura-specific
graph 8 Products, 4 Categories, 1 Store, 2 KnowledgeDocuments, 4 Chunks plus provenance/entity edges
indexes full-text product index + 8D product vector index + 8D chunk vector index
measurement learner measures recall@k, runtime latency, index state/options and result IDs; generated lesson never claims that Neo4j was executed here
Term Mechanism-first meaning
embedding Numeric representation produced outside the database by a model or deterministic encoder. Neo4j stores/indexes the values; it does not make semantic truth guarantees about the encoder.
dimension Number of coordinates in an embedding. Index dimension and query-vector dimension must match when dimensions are configured.
LIST embedding Community-compatible numeric property such as [0.95,0.85,...]. Individual elements are list-accessible.
VECTOR value Fixed-length typed vector value introduced in 2025.10; more storage-efficient typing but persisted VECTOR properties require Enterprise/Aura block format.
similarity Function that converts a pair of vectors into an ordering signal. Current vector indexes support cosine and euclidean similarity.
ANN Approximate nearest-neighbor retrieval. It trades guaranteed exactness for scalable search speed/resource behavior.
HNSW Hierarchical Navigable Small World graph used internally by the vector index to navigate candidate neighborhoods rather than compare every stored vector.
recall@k Fraction of the exact top-k neighbors recovered by ANN top-k. It is a retrieval-quality measure, not semantic correctness.
filter property Non-vector property explicitly stored with a 2026.01+ vector index so SEARCH can apply supported predicates inside the ANN search.
quantization Compressed vector representation used inside the index to lower memory/storage and often improve speed, potentially trading accuracy; 2026.07 supports high-fidelity rescoring through search expansion.
GraphRAG Retrieval-augmented generation pattern where graph-structured evidence, provenance, and relationships enrich the context given to a generator. Retrieval quality and generator factuality still require evaluation.

Direct-entry setup

Run the fixture and parameter block if required, then verify the vector indexes are ONLINE.

Cypher 25 · setup
CYPHER 25
// Disposable Chapter 22 fixture. Safe to rerun after the cleanup block.
CREATE CONSTRAINT ch22_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch22_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch22_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch22_doc_id IF NOT EXISTS
FOR (d:KnowledgeDocument) REQUIRE d.documentId IS UNIQUE;
CREATE CONSTRAINT ch22_chunk_id IF NOT EXISTS
FOR (c:Chunk) REQUIRE c.chunkId IS UNIQUE;

MERGE (cam:Category {categoryId:'CAT-22-CAM'}) SET cam.name='Cameras', cam.labTag='ch22'
MERGE (out:Category {categoryId:'CAT-22-OUT'}) SET out.name='Outdoor', out.labTag='ch22'
MERGE (sec:Category {categoryId:'CAT-22-SEC'}) SET sec.name='Security', sec.labTag='ch22'
MERGE (acc:Category {categoryId:'CAT-22-ACC'}) SET acc.name='Accessories', acc.labTag='ch22'
MERGE (st:Store {storeId:'ST-22-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch22';

UNWIND [
 {id:'P-2201',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:5,featured:true, emb:[0.95,0.85,0.25,0.05,0.10,0.90,0.05,0.05]},
 {id:'P-2202',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:0,featured:false,emb:[0.90,0.80,0.15,0.05,0.05,0.82,0.05,0.05]},
 {id:'P-2203',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,cat:'CAT-22-SEC',catCode:'SECURITY',region:'CENTRAL',qty:7,featured:false,emb:[0.92,0.10,0.95,0.02,0.05,0.05,0.05,0.05]},
 {id:'P-2204',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:11,featured:false,emb:[0.02,0.88,0.02,0.95,0.10,0.10,0.05,0.02]},
 {id:'P-2205',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,cat:'CAT-22-CAM',catCode:'CAMERA',region:'CENTRAL',qty:3,featured:true,emb:[0.90,0.45,0.20,0.10,0.95,0.15,0.05,0.03]},
 {id:'P-2206',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,cat:'CAT-22-OUT',catCode:'OUTDOOR',region:'CENTRAL',qty:6,featured:false,emb:[0.05,0.55,0.05,0.05,0.05,0.95,0.05,0.02]},
 {id:'P-2207',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,cat:'CAT-22-ACC',catCode:'ACCESSORY',region:'CENTRAL',qty:0,featured:false,emb:[0.08,0.75,0.10,0.05,0.10,0.10,0.95,0.02]},
 {id:'P-2208',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,cat:'CAT-22-CAM',catCode:'CAMERA',region:'ARCHIVE',qty:2,featured:false,emb:[0.88,0.68,0.15,0.05,0.05,0.60,0.05,0.95]}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
    p.active=row.active, p.categoryCode=row.catCode, p.region=row.region,
    p.featured=row.featured, p.embedding=row.emb,
    p.embeddingModel='atlasmart-deterministic-v1', p.embeddingVersion='2026-09-lab',
    p.labTag='ch22'
WITH row,p
MATCH (cat:Category {categoryId:row.cat}), (st:Store {storeId:'ST-22-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch22';

MERGE (d1:KnowledgeDocument {documentId:'DOC-22-TRAILCAM'})
SET d1.title='Trail Camera Pro field manual', d1.uri='atlasmart://manuals/P-2201', d1.version='2026.09', d1.labTag='ch22'
MERGE (d2:KnowledgeDocument {documentId:'DOC-22-SECURITY'})
SET d2.title='Camera selection guide', d2.uri='atlasmart://guides/camera-selection', d2.version='2026.09', d2.labTag='ch22';

UNWIND [
 {id:'CHK-2201',doc:'DOC-22-TRAILCAM',seq:1,text:'Trail Camera Pro is weatherproof and optimized for wildlife monitoring on outdoor trails.',emb:[0.96,0.86,0.12,0.02,0.02,0.94,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
 {id:'CHK-2202',doc:'DOC-22-TRAILCAM',seq:2,text:'Infrared night vision records wildlife without visible illumination and battery life is designed for field deployment.',emb:[0.88,0.72,0.25,0.02,0.02,0.90,0.02,0.02],products:['P-2201'],cats:['CAT-22-CAM']},
 {id:'CHK-2203',doc:'DOC-22-SECURITY',seq:1,text:'Indoor Security Camera focuses on Wi-Fi motion alerts and indoor night vision rather than outdoor wildlife use.',emb:[0.84,0.08,0.96,0.02,0.02,0.06,0.02,0.02],products:['P-2203'],cats:['CAT-22-SEC']},
 {id:'CHK-2204',doc:'DOC-22-SECURITY',seq:2,text:'Choose an outdoor trail camera when weather resistance and wildlife observation matter; choose indoor security cameras for home alerting.',emb:[0.90,0.68,0.55,0.02,0.02,0.72,0.02,0.02],products:['P-2201','P-2203'],cats:['CAT-22-CAM','CAT-22-SEC']}
] AS row
MERGE (c:Chunk {chunkId:row.id})
SET c.seq=row.seq, c.text=row.text, c.embedding=row.emb,
    c.embeddingModel='atlasmart-deterministic-v1', c.embeddingVersion='2026-09-lab', c.labTag='ch22'
WITH row,c
MATCH (d:KnowledgeDocument {documentId:row.doc})
MERGE (c)-[:FROM_DOCUMENT]->(d)
WITH row,c
UNWIND row.products AS pid
MATCH (p:Product {productId:pid})
MERGE (c)-[:MENTIONS]->(p)
WITH row,c
UNWIND row.cats AS cid
MATCH (cat:Category {categoryId:cid})
MERGE (c)-[:MENTIONS]->(cat);

MATCH (a:Chunk {chunkId:'CHK-2201'}),(b:Chunk {chunkId:'CHK-2202'}) MERGE (a)-[:NEXT_CHUNK]->(b);
MATCH (a:Chunk {chunkId:'CHK-2203'}),(b:Chunk {chunkId:'CHK-2204'}) MERGE (a)-[:NEXT_CHUNK]->(b);

CREATE FULLTEXT INDEX ch22_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name,p.description,p.tags]
OPTIONS {indexConfig:{`fulltext.analyzer`:'english',`fulltext.eventually_consistent`:false}};

CREATE VECTOR INDEX ch22_product_vector IF NOT EXISTS
FOR (p:Product)
ON p.embedding
WITH [p.active,p.categoryCode,p.region]
OPTIONS {indexConfig:{
  `vector.dimensions`:8,
  `vector.similarity_function`:'cosine',
  `vector.quantization.type`:'scalar',
  `vector.default_search_expansion_factor`:1.5
}};

CREATE VECTOR INDEX ch22_chunk_vector IF NOT EXISTS
FOR (c:Chunk)
ON c.embedding
WITH [c.embeddingVersion,c.seq]
OPTIONS {indexConfig:{
  `vector.dimensions`:8,
  `vector.similarity_function`:'cosine'
}};

CALL db.awaitIndexes(300);
cypher-shell parameters
:param qvec => [1.0,0.8,0.1,0.0,0.0,0.9,0.0,0.0];
:param k => 4;

1. Preferred 2026.01+ SEARCH syntax

The binding variable appears in the MATCH pattern and the SEARCH subclause. FOR supplies the query vector. LIMIT bounds the nearest-neighbor request. Optional SCORE AS binds the vector similarity score for diagnostics/ranking inside this vector result list.

Cypher 25 · vector SEARCH + SCORE
CYPHER 25
MATCH (p:Product)
  SEARCH p IN (
    VECTOR INDEX ch22_product_vector
    FOR $qvec
    LIMIT $k
  ) SCORE AS similarityScore
RETURN p.productId AS productId,p.name AS name,similarityScore;
// SEARCH returns approximate nearest neighbors ordered by the index similarity signal.
Raw score scope

Current vector scores are bounded similarity values, but they are not calibrated probabilities and should not be numerically compared with full-text scores. Lesson 4 preserves source ranks and fuses ranks instead.

2. In-index filtering and post-filtering are not the same mechanism

With a post-filter, ANN first selects up to k neighbors, then Cypher removes rows. With a supported filter inside SEARCH, vector search continues looking for neighbors that satisfy the indexed metadata predicate, which can preserve the requested result count better for selective filters. The property must be declared in the vector index and the predicate must use SEARCH-compatible syntax.

Cypher 25 · in-index filter
CYPHER 25
MATCH (p:Product)
  SEARCH p IN (
    VECTOR INDEX ch22_product_vector
    FOR $qvec
    WHERE p.active = true AND p.categoryCode = 'CAMERA'
    LIMIT 4
  ) SCORE AS similarityScore
RETURN p.productId AS productId,p.name AS name,p.active,p.categoryCode,similarityScore;
// The filter is inside SEARCH because active/categoryCode were declared in WITH [...] at index creation.
Cypher 25 · post-filter for comparison
CYPHER 25
MATCH (p:Product)
  SEARCH p IN (
    VECTOR INDEX ch22_product_vector
    FOR $qvec
    LIMIT 4
  ) SCORE AS similarityScore
WHERE p.active = true AND p.categoryCode = 'CAMERA'
RETURN p.productId AS productId,p.name AS name,similarityScore;
// This WHERE is a post-filter. It can return fewer than four rows because ANN chose candidates first.
Question In-index WHERE MATCH/OPTIONAL MATCH WHERE
when applied during vector search after ANN candidate retrieval
property requirement must be additional filter property in vector index any accessible graph property/pattern under normal Cypher semantics
result-count behavior search attempts to find requested qualifying neighbors can shrink the k candidates after retrieval
flexibility limited supported predicate forms full Cypher post-filter flexibility

3. Approximate LIMIT is a retrieval target, not a promise of exact neighbors

Even with in-index filtering, current documentation notes that ANN can return fewer rows in edge cases when k is close to the qualifying population. And the chosen rows may differ from brute force. The correct acceptance criterion is measured recall@k on judged/representative data plus application-level fallback policy.

Cypher 25 · exact comparator
CYPHER 25
MATCH (p:Product {labTag:'ch22'})
WITH p, vector.similarity.cosine(p.embedding,$qvec) AS exactScore
ORDER BY exactScore DESC, p.productId
LIMIT $k
RETURN p.productId AS productId,p.name AS name,exactScore;
// For this fixed 8-row fixture, the exact top candidates should begin with P-2202 and P-2201.
// Capture the complete runtime result; do not hard-code ANN equivalence from this tiny graph.
Python · repeated ANN measurement
from neo4j import GraphDatabase
from statistics import median
from time import perf_counter

URI="bolt://localhost:7687"
AUTH=("neo4j","atlasmart-course-2026")
QVEC=[1.0,0.8,0.1,0.0,0.0,0.9,0.0,0.0]
K=4
REPEATS=30

EXACT="""
MATCH (p:Product {labTag:'ch22'})
WITH p,vector.similarity.cosine(p.embedding,$qvec) AS score
ORDER BY score DESC,p.productId LIMIT $k
RETURN p.productId AS id,score
"""
ANN="""
MATCH (p:Product)
  SEARCH p IN (VECTOR INDEX ch22_product_vector FOR $qvec LIMIT $k)
  SCORE AS score
RETURN p.productId AS id,score
"""

with GraphDatabase.driver(URI,auth=AUTH) as driver:
    exact,_,_=driver.execute_query(EXACT,qvec=QVEC,k=K,database_="neo4j")
    exact_ids=[r["id"] for r in exact]
    times=[]; ann_ids=[]
    for _ in range(REPEATS):
        t0=perf_counter()
        ann,_,_=driver.execute_query(ANN,qvec=QVEC,k=K,database_="neo4j")
        times.append((perf_counter()-t0)*1000)
        ann_ids=[r["id"] for r in ann]
    recall=len(set(exact_ids)&set(ann_ids))/K
    print({"exact":exact_ids,"ann":ann_ids,"recall_at_k":recall,
           "median_client_ms":median(times),"samples":REPEATS})

# This tiny fixture may produce recall@4 == 1.0. That proves only this fixture/run.
# Repeat on a representative corpus and record p50/p95/p99 plus server/resource evidence.

4. Query-plan evidence

PROFILE executes the query and exposes operator/runtime evidence. Current Cypher 25 plans use vector-index search operators such as NodeVectorIndexSearch. Do not infer server time from client wall-clock alone: result decoding, Bolt/network, queueing and application work can dominate. Pair plan/query evidence with the driver timing discipline from Chapters 13–14.

Cypher 25 · PROFILE SEARCH
CYPHER 25
PROFILE
MATCH (p:Product)
  SEARCH p IN (
    VECTOR INDEX ch22_product_vector
    FOR $qvec
    WHERE p.active = true
    LIMIT 4
  ) SCORE AS similarityScore
RETURN p.productId,p.name,similarityScore;
// On current releases, inspect the plan for the vector-index search operator rather than assuming a label scan.

5. Boundary cases change the result design

Dimension mismatch is a clear error when dimensions are configured. A null query vector inside MATCH produces no match; in OPTIONAL MATCH it follows optional-null semantics. An unsupported in-index property is an error. These are contract conditions that should be validated at the service boundary before issuing expensive requests.

Cypher 25 · wrong dimension
CYPHER 25
:param badVec => [1.0,0.8,0.1,0.0,0.0,0.9,0.0];
MATCH (p:Product)
  SEARCH p IN (
    VECTOR INDEX ch22_product_vector
    FOR $badVec
    LIMIT 4
  )
RETURN p.productId;
// Expected boundary: query fails because the configured vector dimension is 8 but badVec has 7 values.
Cypher 25 · wrong in-index property
CYPHER 25
MATCH (p:Product)
  SEARCH p IN (
    VECTOR INDEX ch22_product_vector
    FOR $qvec
    WHERE p.featured = true
    LIMIT 4
  )
RETURN p.productId;
// Expected boundary: featured was NOT declared as an additional filter property in ch22_product_vector.
Boundary Service response policy
wrong dimension reject as validation/config/model-version mismatch; do not retry blindly
null/missing embedding choose explicit no-result/fallback behavior; record missing-embedding metric
index not ONLINE dependency/readiness failure; fallback or fail according to SLO
filter property absent from index query/schema mismatch; deploy matching index or use deliberate post-filter
too few qualifying neighbors return fewer results or overfetch/fallback according to product contract

6. Procedure history and migration boundary

Older versions query vector indexes with db.index.vector.queryNodes() or queryRelationships(). These are deprecated as of 2026.04, replaced by SEARCH. If your fleet spans older releases, isolate the query adapter by server capability/version and test both paths; do not sprinkle deprecated calls throughout application code.

Compatibility-only legacy form · do not use for new 2026.07 code
// Compatibility history for pre-2026.01 servers / migration adapters:
CALL db.index.vector.queryNodes('ch22_product_vector',4,$qvec)
YIELD node,score
RETURN node.productId,score;
// Current 2026.07 course path uses SEARCH instead.

7. Wrong approaches and repairs

Wrong approach Failure Repair
raw score threshold as semantic truth model/corpus changes invalidate arbitrary threshold judge/calibrate threshold per task if threshold is required
post-filter selective metadata and expect k rows candidate set shrinks in-index filter with declared metadata or increase candidate policy deliberately
retry dimension errors deterministic contract failure validate embedding model/dimension
PROFILE every production request overhead and noisy data sample/diagnostic plan capture; use supported query visibility/metrics
keep deprecated procedure forever loses SEARCH filters/planner integration and future compatibility capability-gated migration to SEARCH

8. Production judgment

Production decision Evidence required
embedding lifecycle Record model/provider/version, dimensions, normalization assumptions, text preprocessing, backfill/re-embedding status, and rollback/index-swap plan.
recall vs latency Measure exact-vs-ANN recall@k on a representative judged set alongside p50/p95/p99; tiny demo recall is not a capacity guarantee.
model/cardinality/degree Separate candidate retrieval count from graph expansion fan-out; bound traversal and final context size.
index memory/storage/write cost Observe vector index size, population/rebuild time, write amplification, page-cache/store pressure, and quantization effects before tuning.
similarity semantics Choose cosine/euclidean from embedding-model semantics; score is a source-specific similarity signal, not factual probability.
filtering/security Declare filter properties intentionally, distinguish in-index from post-filter behavior, and test the real Enterprise service role because semantic-index authorization can suppress candidates.
driver/timeouts/retries Use a long-lived driver, bounded candidate counts, transaction timeouts, idempotent writes and explicit retry/error classification from earlier chapters.
hybrid ranking Fuse independent source ranks (for example RRF/WRRF) or use a trained evaluated re-ranker; never add incomparable raw full-text/vector scores by habit.
GraphRAG provenance Every context unit carries source ID/URI/version/chunk ID and graph entities/relationships so retrieval evidence can be audited.
evaluation Track retrieval recall/precision/MRR/nDCG-style metrics, answer grounding/citation correctness if a generator is added, latency, freshness and zero-result/fallback rates.
privacy/tenant risk Do not embed secrets/PII without policy; authorization must be enforced before context reaches a generator or user.
backup/recovery Rebuild/validate vector indexes and embedding-version metadata in restore drills; restore of graph data is not proof that semantic retrieval is healthy.
Aura/self-managed Aura manages infrastructure and some controls; self-managed exposes server/index/plugin/resource operations. Verify feature/tier availability instead of assuming parity.
licensing/cost Community mandatory lab is free/local. VECTOR storage, Enterprise security/clustering and paid embedding/LLM services are optional and separately costed/licensed.
migration/rollback Run old/new embedding/index versions side by side when possible, freeze evaluation data, cut over by explicit index/query configuration, retain rollback until validation passes.

Check your understanding

  1. Where must WHERE appear for in-index filtering?
  2. Why can post-filtering return fewer than k?
  3. What does SCORE mean?
  4. What is the preferred current replacement for db.index.vector.queryNodes()?
  5. Why inspect exact top-k during evaluation?
Review the answers

1. Inside the SEARCH parentheses, using the same binding variable and properties declared in the vector index.

2. ANN selects k first; the normal WHERE can then discard some of those rows.

3. Similarity under the vector index metric for that source query, not a probability or cross-source comparable unit.

4. Cypher 25 SEARCH on Neo4j 2026.01+; the procedures are deprecated from 2026.04.

5. It provides the ground-truth nearest-neighbor baseline needed to calculate ANN recall@k.

Summary and next step

SEARCH gives AtlasMart a modern, planner-visible vector retrieval path with explicit filters and score evidence. Lesson 4 adds lexical candidates and graph/business context while enforcing the central hybrid rule: preserve independent rankings and fuse ranks—not raw full-text and vector magnitudes.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.