Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context

Query Syntax, Scoring, Boolean/phrase/fuzzy Patterns, Result Limits, and Score Interpretation

Use current Lucene-style full-text query syntax deliberately, interpret scores as relative ranking evidence rather than probabilities, bound result sets, and test Boolean, phrase, field-qualified, and fuzzy behavior against known AtlasMart documents.

Advanced220–310 minutesQuery/relevance labNeo4j 2026.07.1 · Community mandatoryCypher 25 · Full-text/Lucene · english analyzerJava 21/25 · Python driver 6.3 optionalLast reviewed: September 2026

AtlasMart launches a search box and the first demo looks excellent: “trail camera” returns sensible products. Then users type camra, operators paste Boolean syntax, and a developer adds a global rule “reject every result with score < 0.7.” These are three different problems—query parsing, typo tolerance, and score calibration. A robust search API must define which Lucene features it exposes, how many candidates it requests, and what a score means for that exact query/index/corpus.

Mental model

A full-text query is not just a bag of words. It is a Lucene query string analyzed against a particular full-text index. Parameterization protects Cypher structure, but the query string can still intentionally contain Lucene operators. Relevance scores rank that source result list; they are not probabilities or universal thresholds.

Learning outcomes

01

Use current Boolean, phrase, field-qualified and fuzzy Lucene-style patterns through parameterized full-text procedure calls.

02

Explain why Cypher parameterization does not neutralize Lucene query syntax and design a literal-vs-advanced search API boundary.

03

Interpret full-text scores only as source-local ranking evidence and reject probability/percentage/cross-source comparisons.

04

Choose candidate limits and pagination based on recall, latency and post-filtering requirements rather than arbitrary UI page size.

05

Run a repeatable AtlasMart query matrix that records candidates, ranks and actual scores without fabricating universal score expectations.

Chapter 21 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; current 5.26 LTS patch is 5.26.30. Version-sensitive examples use explicit CYPHER 25. The mandatory lab uses self-managed Neo4j Community 2026.07.1, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback Bolt 7687 and HTTP 7474. Neo4j 2026.x supports Java 21/25. Optional application examples pin the official Python driver to neo4j==6.3.0. No APOC, GDS, embeddings, paid AI service, Aura subscription, or Enterprise license is required for the chapter.

Full-text availability and platform boundary

Full-text indexes and the db.index.fulltext.* procedures used in the mandatory lab are part of the normal Neo4j database feature set and are reproducible in Community. Enterprise adds fine-grained authorization and built-in metrics surfaces that can change what full-text queries are allowed to return or what operational counters are available. Aura also supports database full-text functionality, but server configuration/filesystem/metrics controls are platform-managed and must not be presented as identical to self-managed Neo4j.

Lab contract and exact assumptions

Dimension Chapter 21 assumption
server Neo4j Community 2026.07.1, single disposable local database
Cypher Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02
Java Java 21 or 25 for the 2026.07 server line
database/auth neo4j / neo4j / atlasmart-course-2026
transport bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS
plugins none required; no APOC/GDS/GenAI plugin
graph 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges
full-text catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer
score policy never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values
measurement learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution
Term Mechanism-first meaning
lexical retrieval Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity.
full-text index Lucene-backed Neo4j semantic index over one or more STRING or LIST properties of nodes or relationships, explicitly queried through full-text procedures.
schema The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property.
tokenization Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior.
analyzer Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms.
query analyzer Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens.
Lucene query string The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal.
score Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale.
candidate set Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size.
eventually consistent index Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window.
freshness SLO Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured.
judged query set Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly.
precision@k Fraction of the first k returned results judged relevant.
recall@k Fraction of all judged relevant items retrieved in the first k results.
MRR Mean Reciprocal Rank: rewards returning the first relevant result near the top.
nDCG Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier.
post-ranking Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model.
rank fusion Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores.

Direct-entry setup

If you did not complete Lesson 1 in this disposable database, run the setup now. It is idempotent for the Chapter 21 fixture.

Cypher 25 · fixture and full-text indexes
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;

MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';

UNWIND [
 {id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
 {id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
 {id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
 {id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
 {id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
 {id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
 {id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
 {id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
    p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';

MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
    r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
    r2.message='Comfortable on long trail runs', r2.labTag='ch21';

CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
  `fulltext.analyzer`: 'english',
  `fulltext.eventually_consistent`: false
}};

CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};

CALL db.awaitIndexes(300);

1. Parameterize Cypher; define the Lucene surface separately

Always pass the user search string as a Cypher parameter. This prevents the search text from becoming Cypher syntax. But the full-text procedure deliberately hands that string to Lucene’s query parser, so operators such as AND, quotes, field prefixes and fuzzy markers may still alter retrieval. Decide whether your product exposes “advanced search” or a literal/structured query builder. Escaping is an application-level Lucene concern, not a reason to concatenate Cypher.

Cypher 25 · one stable query API
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 10})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;
// Example parameter: :param q => 'wildlife camera'
// Record actual scores; only the relative order in this result set has meaning.
Two parser layers

$q protects the Cypher layer. It does not promise that the contents of $q are interpreted literally by Lucene. Document or constrain the query grammar your users are allowed to invoke.

2. Boolean, phrase, property and fuzzy patterns

Current Neo4j full-text procedures accept Lucene query language. Neo4j documents Boolean operators, quoted exact/phrase-style matching, and property-qualified queries; Lucene query parsing also supplies fuzzy term syntax. The fuzzy experiment below is deliberately evaluated against the fixture rather than assigned a hard-coded score.

Lucene query-string experiments
// Run each as the $q parameter to db.index.fulltext.queryNodes.
// Boolean: both terms must participate.
'wildlife AND camera'

// Phrase: quoted terms occur as a phrase under the analyzer/query parser.
'"trail camera"'

// Exclusion: camera results except security-oriented terms.
'camera NOT security'

// Property-qualified query against a property included in the index.
'name:"Trail Camera Pro"'

// Fuzzy Lucene term: useful for a misspelling experiment.
'camra~'

// Field + Boolean composition.
'name:camera AND description:night'
Pattern Question it answers Boundary
wildlife AND camera require both analyzed concepts operator semantics are Lucene, not Cypher AND
"trail camera" prefer/require the phrase under the analyzed token stream analysis can alter tokens; punctuation/case may normalize
camera NOT security exclude a lexical term business category != lexical term; still graph-filter later
name:"Trail Camera Pro" limit lexical match to the name field field must be part of the index
camra~ fuzzy typo experiment edit-distance behavior/ranking is query-parser evidence, not a spelling-correction SLA

3. Query-time analyzer is a precision tool, not an index rewrite

The current procedure options support skip, limit, and analyzer. Overriding the query analyzer can be useful when the application has a deliberate query-side normalization policy, but the indexed terms remain those produced at index creation. Compare results under the default and a query-side analyzer only after listing what is available.

Cypher 25 · query analyzer experiment
CYPHER 25
CALL db.index.fulltext.queryNodes(
  'ch21_catalog_ft',
  'trail cameras',
  {limit: 10, analyzer: 'whitespace'}
)
YIELD node, score
RETURN node.productId, node.name, score
ORDER BY score DESC;
// Compare candidate IDs/ranks with the same query using the default index analyzer.

4. Score interpretation: rank evidence, not probability

Neo4j explicitly states that full-text scores are meaningful within the result set from a full-text query. The magnitude depends on terms, document statistics, fields, analyzer and corpus. Adding products can change document-frequency statistics and therefore scores even when one Product did not change. A threshold copied from another index, environment or vector source is not evidence.

Tempting interpretation Why it is wrong Safer alternative
score 0.83 = 83% relevant Lucene score is not calibrated probability judge ranking/thresholds on labeled queries
score > 0.5 always good scale varies by query/corpus/analyzer learn threshold per use case, or prefer rank/candidate limits
full-text 0.7 > vector 0.6 different retrieval sources have incomparable raw scales rank each list independently; fuse ranks
same product should keep same score forever corpus/index statistics can change regression-test ordering/quality, not frozen numbers

5. Candidate limit is part of recall

If AtlasMart ultimately displays five products but graph/security/stock filters remove some lexical candidates, asking Lucene for only five can starve the post-filter stage. Retrieve a bounded sourceK based on measured survival rate and latency. The current full-text procedure supports limit directly; skip exists for pagination, but deep result browsing should still be evaluated for cost and ranking stability.

Cypher 25 · bounded candidate query
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {skip: 0, limit: 20})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;

6. AtlasMart query matrix: record, do not invent

ID Query Expected qualitative evidence
Q1 wildlife camera camera products and wildlife-related text compete; graph context later removes non-product-intent noise
Q2 "trail camera" Trail Camera products should be strong phrase candidates; capture exact rank/score on your server
Q3 night vision camera P-2101 and P-2103 contain strong lexical evidence; relative order is measured, not prescribed
Q4 camra~ fuzzy query should be tested for typo recovery; capture actual candidate set and latency
Q5 camera NOT security compare exclusion effect against plain camera query
Cypher 25 · reusable capture command
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 10})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;
// Example parameter: :param q => 'wildlife camera'
// Record actual scores; only the relative order in this result set has meaning.

7. Wrong approach → failure → repair

Wrong approach Failure Repair
concatenate user input into Cypher query injection / broken syntax risk use parameters; separately define Lucene grammar policy
expose all Lucene syntax accidentally users/operators can produce expensive or surprising queries offer structured search or explicitly documented advanced mode
hard-code one numeric score threshold quality changes when corpus/analyzer/query changes calibrate against judged data and monitor false positives/negatives
return thousands of hits “for recall” latency/network/materialization cost explodes bounded sourceK + measured post-filter survival
compare raw lexical/vector scores ranking has no common scale independent ranks + RRF/WRRF in Chapter 22

8. Verification checklist

  • All query strings are supplied as parameters to Cypher.
  • The allowed Lucene grammar is documented: literal, structured, or advanced.
  • Boolean, phrase, field and fuzzy experiments record IDs/ranks/latency from the current server.
  • No score is described as a probability or compared directly to a vector score.
  • Candidate limit is justified by downstream graph/business filtering and measured latency.

9. Production judgment

Production decision Evidence to require before changing the system
graph/workload fit Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient.
correctness/non-guarantees Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability.
model/cardinality/degree Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval.
latency Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call.
transactions/freshness Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore.
memory/storage Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization.
CPU/disk/network Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network.
indexes/constraints Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking.
driver/pool/timeouts Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results.
security/tenant risk Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization.
backup/recovery Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health.
observability Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters.
testing/failure injection Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects.
version/tier Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical.
cost/migration Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval.

Check your understanding

  1. Why does $q not make a Lucene query literal?
  2. What does a higher full-text score prove?
  3. Why might sourceK exceed the displayed result count?
  4. Why is a fuzzy query not the same as a spelling-correction service?
  5. How should full-text and vector results be combined later?
Review the answers

1. It protects the Cypher parser; the full-text procedure still intentionally parses the parameter value using Lucene query syntax.

2. Within that query result list, the index ranks that entry as a stronger lexical match; it does not prove business relevance or a calibrated probability.

3. Graph/security/business filters can remove candidates, so retrieval needs enough bounded headroom to preserve final recall.

4. It is Lucene fuzzy term matching; behavior/ranking must be evaluated against your domain queries and latency requirements.

5. Rank each source independently and fuse ranks, rather than comparing raw scores.

Summary and next step

Search syntax is an API contract, scores are source-local ranking evidence, and candidate limits are part of recall engineering. Lesson 3 adds the write-path decision: whether Lucene updates must be visible at commit or may enter a background queue with an explicit freshness SLO.

Authoritative references

  • Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
  • Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
  • Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
  • Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
  • Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
  • Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
  • Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
  • Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
  • Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
  • Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
  • Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
  • Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
  • System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
  • Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
  • Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.