Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context

Full-Text Index Schema, Node vs Relationship Indexes, Analyzers, Tokenization, and Language Handling

Build AtlasMart lexical search from first principles: full-text node and relationship schemas, Lucene-backed tokenization, analyzers, language boundaries, STRING/LIST<STRING> indexing, online state, and Community-safe setup.

Advanced220–300 minutesCatalog + analyzer labNeo4j 2026.07.1 · Community mandatoryCypher 25 · Full-text/Lucene · english analyzerJava 21/25 · Python driver 6.3 optionalLast reviewed: September 2026

AtlasMart product search currently uses CONTAINS over names and descriptions. It finds substrings, but it cannot express a search vocabulary such as “wildlife trail camera,” cannot stem language-specific word forms, and provides no relevance ordering. The fix is not “turn on a magic search mode.” A Neo4j full-text index is a separate Lucene-backed retrieval structure whose schema, analyzer, query parser, score and freshness contract must be designed explicitly.

Mental model

The graph remains the source of entities and relationships. The full-text index is a lexical candidate generator over selected string fields. Search quality depends on what is indexed, how text is analyzed, what query syntax the application allows, and what graph/business filters happen after retrieval.

Learning outcomes

01

Create current node and relationship full-text indexes and explain their schema membership rules for STRING and LIST properties.

02

Inspect index state/provider/options and list available analyzers before choosing a language/tokenization strategy.

03

Explain index-time vs query-time analysis, stemming/stop words/case normalization, and why one analyzer is not automatically correct for multilingual content.

04

Query both Product nodes and REVIEWED relationships without confusing a semantic index with an automatically planner-selected search-performance index.

05

Build and verify the Community AtlasMart Chapter 21 fixture while preserving stable business identifiers and safe cleanup.

Chapter 21 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; current 5.26 LTS patch is 5.26.30. Version-sensitive examples use explicit CYPHER 25. The mandatory lab uses self-managed Neo4j Community 2026.07.1, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback Bolt 7687 and HTTP 7474. Neo4j 2026.x supports Java 21/25. Optional application examples pin the official Python driver to neo4j==6.3.0. No APOC, GDS, embeddings, paid AI service, Aura subscription, or Enterprise license is required for the chapter.

Full-text availability and platform boundary

Full-text indexes and the db.index.fulltext.* procedures used in the mandatory lab are part of the normal Neo4j database feature set and are reproducible in Community. Enterprise adds fine-grained authorization and built-in metrics surfaces that can change what full-text queries are allowed to return or what operational counters are available. Aura also supports database full-text functionality, but server configuration/filesystem/metrics controls are platform-managed and must not be presented as identical to self-managed Neo4j.

Lab contract and exact assumptions

Dimension Chapter 21 assumption
server Neo4j Community 2026.07.1, single disposable local database
Cypher Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02
Java Java 21 or 25 for the 2026.07 server line
database/auth neo4j / neo4j / atlasmart-course-2026
transport bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS
plugins none required; no APOC/GDS/GenAI plugin
graph 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges
full-text catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer
score policy never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values
measurement learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution
Term Mechanism-first meaning
lexical retrieval Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity.
full-text index Lucene-backed Neo4j semantic index over one or more STRING or LIST properties of nodes or relationships, explicitly queried through full-text procedures.
schema The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property.
tokenization Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior.
analyzer Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms.
query analyzer Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens.
Lucene query string The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal.
score Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale.
candidate set Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size.
eventually consistent index Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window.
freshness SLO Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured.
judged query set Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly.
precision@k Fraction of the first k returned results judged relevant.
recall@k Fraction of all judged relevant items retrieved in the first k results.
MRR Mean Reciprocal Rank: rewards returning the first relevant result near the top.
nDCG Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier.
post-ranking Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model.
rank fusion Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores.

1. Full-text is a semantic index, not a normal predicate accelerator

Range/text indexes are planner-selected access paths for predicates such as equality, prefix, suffix and substring matches. Full-text indexes are different: they tokenize indexed text and must be called explicitly through db.index.fulltext.queryNodes() or queryRelationships(). The procedure returns the entity plus a Lucene relevance score already ordered from best to lower match.

Question Search-performance text index Full-text index
primary use exact/prefix/CONTAINS/ENDS WITH predicates term-oriented lexical retrieval inside text
planner use automatic when compatible explicit procedure call
analysis trigram/exact string mechanics Lucene analyzer/tokenizer/stemming/stop words
score none source-specific relevance score
entity type nodes and relationships nodes and relationships

2. Schema membership is OR across labels/types and fields

A node full-text index may name several labels and several properties. A node can participate if it has at least one indexed label and at least one indexed property. Relationship full-text indexes can similarly name multiple relationship types and properties. Only STRING and LIST<STRING> values contribute searchable terms; a tags list is analyzed into the same field namespace for that property.

Cypher 25 · create the complete AtlasMart fixture
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;

MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';

UNWIND [
 {id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
 {id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
 {id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
 {id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
 {id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
 {id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
 {id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
 {id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
    p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';

MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
    r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
    r2.message='Comfortable on long trail runs', r2.labTag='ch21';

CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
  `fulltext.analyzer`: 'english',
  `fulltext.eventually_consistent`: false
}};

CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};

CALL db.awaitIndexes(300);
Cypher 25 · prove graph + index state
CYPHER 25
SHOW FULLTEXT INDEXES
YIELD name, state, populationPercent, entityType, labelsOrTypes,
      properties, indexProvider, options, lastRead, readCount
WHERE name STARTS WITH 'ch21_'
RETURN * ORDER BY name;

MATCH (p:Product {labTag:'ch21'})
RETURN count(p) AS products,
       count { (p)-[:STOCKED_AT]->(:Store {storeId:'ST-21-CENTRAL'}) } AS stockEdges;
// Deterministic graph invariant: products = 8 and stockEdges = 8.
// Full-text state invariant before search: both indexes are ONLINE and 100% populated.

3. Analyzer choice is part of your data contract

An analyzer determines how both indexed text and, by default, query text become terms. Current Neo4j defaults new full-text indexes to standard-no-stop-words. The Chapter 21 catalog intentionally uses english so learners can observe language-aware stemming/stop-word behavior. Analyzer names are server capabilities: list them on the exact deployment rather than copying a blog post.

Cypher 25 · inspect analyzers on this server
CYPHER 25
CALL db.index.fulltext.listAvailableAnalyzers()
YIELD analyzer, description, stopwords
WHERE analyzer IN ['standard-no-stop-words','english','whitespace']
RETURN analyzer, description, stopwords;
// Verify the list on your exact server; do not assume every custom/language analyzer exists everywhere.
Boundary: query analyzer ≠ reindex

The optional analyzer entry in queryNodes/queryRelationships options changes how the query string is analyzed. It does not retokenize stored index contents. If index-time and query-time analysis are incompatible, recall can worsen rather than improve.

4. One analyzer is rarely a complete multilingual strategy

Language analyzers can stem terms and remove language-specific stop words, while the default analyzer is less language-specific. Do not put English, Persian, Arabic, Turkish and Urdu product text into one index and assume one stemming policy is neutral. Common strategies are separate language-specific properties/indexes, a language-neutral analyzer for fields that must mix languages, or a deliberately tested custom analyzer. The correct choice comes from judged queries per locale, not font/script similarity.

Design Advantage Risk / evidence needed
one neutral index simple operations; shared IDs lower recall/precision for morphology-heavy languages
per-language property/index language-specific stemming and stop words more indexes, routing logic and evaluation sets
custom analyzer domain-specific tokenization possible plugin/deployment coupling; upgrade/security test required

5. Node and relationship search solve different questions

The catalog index finds Products from product text. The review index finds REVIEWED edges from review-message text; after retrieving a relationship, Cypher can inspect its endpoints and graph context. This is not equivalent to copying review text onto a Product merely to make it searchable.

Cypher 25 · node full-text query
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 10})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;
// Example parameter: :param q => 'wildlife camera'
// Record actual scores; only the relative order in this result set has meaning.
Cypher 25 · relationship full-text query
CYPHER 25
CALL db.index.fulltext.queryRelationships('ch21_reviews_ft', 'night vision', {limit: 10})
YIELD relationship, score
RETURN relationship.reviewId AS reviewId,
       relationship.message AS message,
       score
ORDER BY score DESC;
// Expected invariant from the fixture: REV-2101 is a matching review.

6. What Neo4j is doing

  1. A transaction creates/updates a STRING or LIST<STRING> value.
  2. The configured analyzer turns that value into terms.
  3. Lucene stores terms and postings for the qualifying node/relationship.
  4. A query procedure analyzes the query string, evaluates Lucene query semantics, obtains matching entries and relevance scores, and returns entities in score order.
  5. Cypher continues from those entities into normal graph predicates/traversals.
Planner boundary

The Cypher planner does not silently decide to use this full-text index for a normal MATCH ... WHERE p.description CONTAINS ... predicate. You explicitly call the semantic index.

7. Deliberately wrong design: “one giant index over everything”

Wrong move Concrete failure Repair
index every textual label/property together field meanings and security domains blur; ranking can be dominated by unrelated fields separate retrieval intents and schemas
use one analyzer for all languages without evaluation stemming/stop-word/tokenization assumptions reduce quality for some locales route by language and judge each slice
treat tags/names/descriptions as equivalent evidence short exact names and long descriptions can contribute differently to ranking test field-qualified queries or separate source lists/rank fusion
rely on index score as product quality relevance becomes confused with inventory, margin, policy or popularity keep lexical score as retrieval evidence; graph/business logic is a separate stage

8. Verification checklist

  • Eight Chapter 21 Product nodes and eight STOCKED_AT edges exist.
  • ch21_catalog_ft and ch21_reviews_ft are ONLINE before testing.
  • The analyzer list is captured from the actual server.
  • A node query returns product IDs and scores; a relationship query can return REV-2101 for “night vision.”
  • No expected numeric score has been copied from this lesson—the learner records current values.

9. Production judgment

Production decision Evidence to require before changing the system
graph/workload fit Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient.
correctness/non-guarantees Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability.
model/cardinality/degree Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval.
latency Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call.
transactions/freshness Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore.
memory/storage Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization.
CPU/disk/network Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network.
indexes/constraints Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking.
driver/pool/timeouts Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results.
security/tenant risk Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization.
backup/recovery Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health.
observability Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters.
testing/failure injection Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects.
version/tier Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical.
cost/migration Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval.

Check your understanding

  1. Why is a full-text index not a replacement for a text/range index?
  2. What makes a Product eligible for ch21_catalog_ft?
  3. What does the analyzer control?
  4. Why must analyzer names be discovered on the server?
  5. Why keep relationship text on REVIEWED instead of copying it onto Product automatically?
Review the answers

1. It solves term-oriented lexical retrieval through explicit semantic-index procedures; search-performance indexes solve different predicate-access patterns and are planner-selected.

2. It has the Product label and at least one indexed STRING or LIST property among name, description, or tags.

3. How indexed values and normally the query string are tokenized/normalized/stemmed and which stop words are removed.

4. Available analyzers/custom analyzers are deployment capabilities and may vary; the procedure is the authoritative local inventory.

5. The relationship message is evidence about the review event/edge; keeping its semantics preserves modeling clarity and allows relationship-specific retrieval.

Summary and next step

Full-text retrieval is an explicit Lucene-backed candidate system over carefully chosen graph fields. Lesson 2 now examines the query parser, score interpretation, Boolean/phrase/fuzzy behavior, limits, and the application boundary between safe parameterization and intentionally exposed Lucene syntax.

Authoritative references

  • Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
  • Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
  • Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
  • Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
  • Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
  • Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
  • Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
  • Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
  • Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
  • Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
  • Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
  • Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
  • System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
  • Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
  • Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.