Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context
Full-Text Index Schema, Node vs Relationship Indexes, Analyzers, Tokenization, and Language Handling
Build AtlasMart lexical search from first principles: full-text node and relationship schemas, Lucene-backed tokenization, analyzers, language boundaries, STRING/LIST<STRING> indexing, online state, and Community-safe setup.
AtlasMart product search currently uses
CONTAINS over names and descriptions. It finds
substrings, but it cannot express a search vocabulary such as
“wildlife trail camera,” cannot stem language-specific word
forms, and provides no relevance ordering. The fix is not “turn
on a magic search mode.” A Neo4j full-text index is a separate
Lucene-backed retrieval structure whose schema, analyzer, query
parser, score and freshness contract must be designed
explicitly.
The graph remains the source of entities and relationships. The full-text index is a lexical candidate generator over selected string fields. Search quality depends on what is indexed, how text is analyzed, what query syntax the application allows, and what graph/business filters happen after retrieval.
Learning outcomes
Create current node and relationship full-text indexes and
explain their schema membership rules for STRING and
LIST
Inspect index state/provider/options and list available analyzers before choosing a language/tokenization strategy.
Explain index-time vs query-time analysis, stemming/stop words/case normalization, and why one analyzer is not automatically correct for multilingual content.
Query both Product nodes and REVIEWED relationships without confusing a semantic index with an automatically planner-selected search-performance index.
Build and verify the Community AtlasMart Chapter 21 fixture while preserving stable business identifiers and safe cleanup.
Current Neo4j Database is 2026.07.1; current 5.26
LTS patch is 5.26.30. Version-sensitive examples
use explicit CYPHER 25. The mandatory lab uses
self-managed Neo4j Community 2026.07.1,
database neo4j, user neo4j, disposable
password atlasmart-course-2026, loopback Bolt
7687 and HTTP 7474. Neo4j 2026.x
supports Java 21/25. Optional application examples pin the
official Python driver to neo4j==6.3.0. No APOC,
GDS, embeddings, paid AI service, Aura subscription, or
Enterprise license is required for the chapter.
Full-text indexes and the
db.index.fulltext.* procedures used in the
mandatory lab are part of the normal Neo4j database feature set
and are reproducible in Community. Enterprise adds fine-grained
authorization and built-in metrics surfaces that can change what
full-text queries are allowed to return or what operational
counters are available. Aura also supports database full-text
functionality, but server configuration/filesystem/metrics
controls are platform-managed and must not be presented as
identical to self-managed Neo4j.
Lab contract and exact assumptions
| Dimension | Chapter 21 assumption |
|---|---|
| server | Neo4j Community 2026.07.1, single disposable local database |
| Cypher | Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02 |
| Java | Java 21 or 25 for the 2026.07 server line |
| database/auth | neo4j / neo4j / atlasmart-course-2026 |
| transport | bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS |
| plugins | none required; no APOC/GDS/GenAI plugin |
| graph | 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges |
| full-text | catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer |
| score policy | never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values |
| measurement | learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution |
| Term | Mechanism-first meaning |
|---|---|
| lexical retrieval | Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity. |
| full-text index |
Lucene-backed Neo4j semantic index over one or more STRING
or LIST |
| schema | The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property. |
| tokenization | Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior. |
| analyzer | Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms. |
| query analyzer | Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens. |
| Lucene query string | The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal. |
| score | Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale. |
| candidate set | Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size. |
| eventually consistent index | Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window. |
| freshness SLO | Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured. |
| judged query set | Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly. |
| precision@k | Fraction of the first k returned results judged relevant. |
| recall@k | Fraction of all judged relevant items retrieved in the first k results. |
| MRR | Mean Reciprocal Rank: rewards returning the first relevant result near the top. |
| nDCG | Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier. |
| post-ranking | Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model. |
| rank fusion | Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores. |
1. Full-text is a semantic index, not a normal predicate accelerator
Range/text indexes are planner-selected access paths for
predicates such as equality, prefix, suffix and substring
matches. Full-text indexes are different: they tokenize indexed
text and must be called explicitly through
db.index.fulltext.queryNodes() or
queryRelationships(). The procedure returns the
entity plus a Lucene relevance score already ordered from best
to lower match.
| Question | Search-performance text index | Full-text index |
|---|---|---|
| primary use | exact/prefix/CONTAINS/ENDS WITH predicates | term-oriented lexical retrieval inside text |
| planner use | automatic when compatible | explicit procedure call |
| analysis | trigram/exact string mechanics | Lucene analyzer/tokenizer/stemming/stop words |
| score | none | source-specific relevance score |
| entity type | nodes and relationships | nodes and relationships |
2. Schema membership is OR across labels/types and fields
A node full-text index may name several labels and several
properties. A node can participate if it has at least one
indexed label and at least one indexed property. Relationship
full-text indexes can similarly name multiple relationship types
and properties. Only STRING and
LIST<STRING> values contribute searchable
terms; a tags list is analyzed into the same field namespace for
that property.
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';
UNWIND [
{id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
{id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
{id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
{id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
{id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
{id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
{id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
{id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';
MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
r2.message='Comfortable on long trail runs', r2.labTag='ch21';
CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
`fulltext.analyzer`: 'english',
`fulltext.eventually_consistent`: false
}};
CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};
CALL db.awaitIndexes(300);
CYPHER 25
SHOW FULLTEXT INDEXES
YIELD name, state, populationPercent, entityType, labelsOrTypes,
properties, indexProvider, options, lastRead, readCount
WHERE name STARTS WITH 'ch21_'
RETURN * ORDER BY name;
MATCH (p:Product {labTag:'ch21'})
RETURN count(p) AS products,
count { (p)-[:STOCKED_AT]->(:Store {storeId:'ST-21-CENTRAL'}) } AS stockEdges;
// Deterministic graph invariant: products = 8 and stockEdges = 8.
// Full-text state invariant before search: both indexes are ONLINE and 100% populated.
3. Analyzer choice is part of your data contract
An analyzer determines how both indexed text and, by default,
query text become terms. Current Neo4j defaults new full-text
indexes to standard-no-stop-words. The Chapter 21
catalog intentionally uses english so learners can
observe language-aware stemming/stop-word behavior. Analyzer
names are server capabilities: list them on the exact deployment
rather than copying a blog post.
CYPHER 25
CALL db.index.fulltext.listAvailableAnalyzers()
YIELD analyzer, description, stopwords
WHERE analyzer IN ['standard-no-stop-words','english','whitespace']
RETURN analyzer, description, stopwords;
// Verify the list on your exact server; do not assume every custom/language analyzer exists everywhere.
The optional analyzer entry in
queryNodes/queryRelationships options changes how the query
string is analyzed. It does not retokenize stored index
contents. If index-time and query-time analysis are
incompatible, recall can worsen rather than improve.
4. One analyzer is rarely a complete multilingual strategy
Language analyzers can stem terms and remove language-specific stop words, while the default analyzer is less language-specific. Do not put English, Persian, Arabic, Turkish and Urdu product text into one index and assume one stemming policy is neutral. Common strategies are separate language-specific properties/indexes, a language-neutral analyzer for fields that must mix languages, or a deliberately tested custom analyzer. The correct choice comes from judged queries per locale, not font/script similarity.
| Design | Advantage | Risk / evidence needed |
|---|---|---|
| one neutral index | simple operations; shared IDs | lower recall/precision for morphology-heavy languages |
| per-language property/index | language-specific stemming and stop words | more indexes, routing logic and evaluation sets |
| custom analyzer | domain-specific tokenization possible | plugin/deployment coupling; upgrade/security test required |
5. Node and relationship search solve different questions
The catalog index finds Products from product text. The review
index finds REVIEWED edges from review-message
text; after retrieving a relationship, Cypher can inspect its
endpoints and graph context. This is not equivalent to copying
review text onto a Product merely to make it searchable.
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 10})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;
// Example parameter: :param q => 'wildlife camera'
// Record actual scores; only the relative order in this result set has meaning.
CYPHER 25
CALL db.index.fulltext.queryRelationships('ch21_reviews_ft', 'night vision', {limit: 10})
YIELD relationship, score
RETURN relationship.reviewId AS reviewId,
relationship.message AS message,
score
ORDER BY score DESC;
// Expected invariant from the fixture: REV-2101 is a matching review.
6. What Neo4j is doing
- A transaction creates/updates a STRING or LIST<STRING> value.
- The configured analyzer turns that value into terms.
- Lucene stores terms and postings for the qualifying node/relationship.
- A query procedure analyzes the query string, evaluates Lucene query semantics, obtains matching entries and relevance scores, and returns entities in score order.
- Cypher continues from those entities into normal graph predicates/traversals.
The Cypher planner does not silently decide to use this
full-text index for a normal
MATCH ... WHERE p.description CONTAINS ...
predicate. You explicitly call the semantic index.
7. Deliberately wrong design: “one giant index over everything”
| Wrong move | Concrete failure | Repair |
|---|---|---|
| index every textual label/property together | field meanings and security domains blur; ranking can be dominated by unrelated fields | separate retrieval intents and schemas |
| use one analyzer for all languages without evaluation | stemming/stop-word/tokenization assumptions reduce quality for some locales | route by language and judge each slice |
| treat tags/names/descriptions as equivalent evidence | short exact names and long descriptions can contribute differently to ranking | test field-qualified queries or separate source lists/rank fusion |
| rely on index score as product quality | relevance becomes confused with inventory, margin, policy or popularity | keep lexical score as retrieval evidence; graph/business logic is a separate stage |
8. Verification checklist
- Eight Chapter 21 Product nodes and eight STOCKED_AT edges exist.
-
ch21_catalog_ftandch21_reviews_ftareONLINEbefore testing. - The analyzer list is captured from the actual server.
-
A node query returns product IDs and scores; a relationship
query can return
REV-2101for “night vision.” - No expected numeric score has been copied from this lesson—the learner records current values.
9. Production judgment
| Production decision | Evidence to require before changing the system |
|---|---|
| graph/workload fit | Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient. |
| correctness/non-guarantees | Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability. |
| model/cardinality/degree | Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval. |
| latency | Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call. |
| transactions/freshness | Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore. |
| memory/storage | Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization. |
| CPU/disk/network | Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network. |
| indexes/constraints | Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking. |
| driver/pool/timeouts | Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results. |
| security/tenant risk | Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization. |
| backup/recovery | Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health. |
| observability | Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters. |
| testing/failure injection | Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects. |
| version/tier | Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical. |
| cost/migration | Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval. |
Check your understanding
- Why is a full-text index not a replacement for a text/range index?
- What makes a Product eligible for ch21_catalog_ft?
- What does the analyzer control?
- Why must analyzer names be discovered on the server?
- Why keep relationship text on REVIEWED instead of copying it onto Product automatically?
Review the answers
1. It solves term-oriented lexical retrieval through explicit semantic-index procedures; search-performance indexes solve different predicate-access patterns and are planner-selected.
2. It has the Product label and at least
one indexed STRING or LIST
3. How indexed values and normally the query string are tokenized/normalized/stemmed and which stop words are removed.
4. Available analyzers/custom analyzers are deployment capabilities and may vary; the procedure is the authoritative local inventory.
5. The relationship message is evidence about the review event/edge; keeping its semantics preserves modeling clarity and allows relationship-specific retrieval.
Summary and next step
Full-text retrieval is an explicit Lucene-backed candidate system over carefully chosen graph fields. Lesson 2 now examines the query parser, score interpretation, Boolean/phrase/fuzzy behavior, limits, and the application boundary between safe parameterization and intentionally exposed Lucene syntax.
Authoritative references
- Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
- Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
- Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
- Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
- Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
- Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
- Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
- Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
- Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
- Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
- Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
- Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
- System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
- Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
- Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.