Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context
Query Syntax, Scoring, Boolean/phrase/fuzzy Patterns, Result Limits, and Score Interpretation
Use current Lucene-style full-text query syntax deliberately, interpret scores as relative ranking evidence rather than probabilities, bound result sets, and test Boolean, phrase, field-qualified, and fuzzy behavior against known AtlasMart documents.
AtlasMart launches a search box and the first demo looks
excellent: “trail camera” returns sensible products. Then users
type camra, operators paste Boolean syntax, and a
developer adds a global rule “reject every result with score
< 0.7.” These are three different problems—query parsing,
typo tolerance, and score calibration. A robust search API must
define which Lucene features it exposes, how many candidates it
requests, and what a score means for that exact
query/index/corpus.
A full-text query is not just a bag of words. It is a Lucene query string analyzed against a particular full-text index. Parameterization protects Cypher structure, but the query string can still intentionally contain Lucene operators. Relevance scores rank that source result list; they are not probabilities or universal thresholds.
Learning outcomes
Use current Boolean, phrase, field-qualified and fuzzy Lucene-style patterns through parameterized full-text procedure calls.
Explain why Cypher parameterization does not neutralize Lucene query syntax and design a literal-vs-advanced search API boundary.
Interpret full-text scores only as source-local ranking evidence and reject probability/percentage/cross-source comparisons.
Choose candidate limits and pagination based on recall, latency and post-filtering requirements rather than arbitrary UI page size.
Run a repeatable AtlasMart query matrix that records candidates, ranks and actual scores without fabricating universal score expectations.
Current Neo4j Database is 2026.07.1; current 5.26
LTS patch is 5.26.30. Version-sensitive examples
use explicit CYPHER 25. The mandatory lab uses
self-managed Neo4j Community 2026.07.1,
database neo4j, user neo4j, disposable
password atlasmart-course-2026, loopback Bolt
7687 and HTTP 7474. Neo4j 2026.x
supports Java 21/25. Optional application examples pin the
official Python driver to neo4j==6.3.0. No APOC,
GDS, embeddings, paid AI service, Aura subscription, or
Enterprise license is required for the chapter.
Full-text indexes and the
db.index.fulltext.* procedures used in the
mandatory lab are part of the normal Neo4j database feature set
and are reproducible in Community. Enterprise adds fine-grained
authorization and built-in metrics surfaces that can change what
full-text queries are allowed to return or what operational
counters are available. Aura also supports database full-text
functionality, but server configuration/filesystem/metrics
controls are platform-managed and must not be presented as
identical to self-managed Neo4j.
Lab contract and exact assumptions
| Dimension | Chapter 21 assumption |
|---|---|
| server | Neo4j Community 2026.07.1, single disposable local database |
| Cypher | Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02 |
| Java | Java 21 or 25 for the 2026.07 server line |
| database/auth | neo4j / neo4j / atlasmart-course-2026 |
| transport | bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS |
| plugins | none required; no APOC/GDS/GenAI plugin |
| graph | 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges |
| full-text | catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer |
| score policy | never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values |
| measurement | learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution |
| Term | Mechanism-first meaning |
|---|---|
| lexical retrieval | Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity. |
| full-text index |
Lucene-backed Neo4j semantic index over one or more STRING
or LIST |
| schema | The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property. |
| tokenization | Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior. |
| analyzer | Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms. |
| query analyzer | Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens. |
| Lucene query string | The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal. |
| score | Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale. |
| candidate set | Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size. |
| eventually consistent index | Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window. |
| freshness SLO | Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured. |
| judged query set | Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly. |
| precision@k | Fraction of the first k returned results judged relevant. |
| recall@k | Fraction of all judged relevant items retrieved in the first k results. |
| MRR | Mean Reciprocal Rank: rewards returning the first relevant result near the top. |
| nDCG | Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier. |
| post-ranking | Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model. |
| rank fusion | Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores. |
Direct-entry setup
If you did not complete Lesson 1 in this disposable database, run the setup now. It is idempotent for the Chapter 21 fixture.
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';
UNWIND [
{id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
{id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
{id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
{id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
{id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
{id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
{id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
{id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';
MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
r2.message='Comfortable on long trail runs', r2.labTag='ch21';
CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
`fulltext.analyzer`: 'english',
`fulltext.eventually_consistent`: false
}};
CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};
CALL db.awaitIndexes(300);
1. Parameterize Cypher; define the Lucene surface separately
Always pass the user search string as a Cypher parameter. This
prevents the search text from becoming Cypher syntax. But the
full-text procedure deliberately hands that string to Lucene’s
query parser, so operators such as AND, quotes,
field prefixes and fuzzy markers may still alter retrieval.
Decide whether your product exposes “advanced search” or a
literal/structured query builder. Escaping is an
application-level Lucene concern, not a reason to concatenate
Cypher.
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 10})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;
// Example parameter: :param q => 'wildlife camera'
// Record actual scores; only the relative order in this result set has meaning.
$q protects the Cypher layer. It does
not promise that the contents of
$q are interpreted literally by Lucene. Document
or constrain the query grammar your users are allowed to
invoke.
2. Boolean, phrase, property and fuzzy patterns
Current Neo4j full-text procedures accept Lucene query language. Neo4j documents Boolean operators, quoted exact/phrase-style matching, and property-qualified queries; Lucene query parsing also supplies fuzzy term syntax. The fuzzy experiment below is deliberately evaluated against the fixture rather than assigned a hard-coded score.
// Run each as the $q parameter to db.index.fulltext.queryNodes.
// Boolean: both terms must participate.
'wildlife AND camera'
// Phrase: quoted terms occur as a phrase under the analyzer/query parser.
'"trail camera"'
// Exclusion: camera results except security-oriented terms.
'camera NOT security'
// Property-qualified query against a property included in the index.
'name:"Trail Camera Pro"'
// Fuzzy Lucene term: useful for a misspelling experiment.
'camra~'
// Field + Boolean composition.
'name:camera AND description:night'
| Pattern | Question it answers | Boundary |
|---|---|---|
| wildlife AND camera | require both analyzed concepts | operator semantics are Lucene, not Cypher AND |
| "trail camera" | prefer/require the phrase under the analyzed token stream | analysis can alter tokens; punctuation/case may normalize |
| camera NOT security | exclude a lexical term | business category != lexical term; still graph-filter later |
| name:"Trail Camera Pro" | limit lexical match to the name field | field must be part of the index |
| camra~ | fuzzy typo experiment | edit-distance behavior/ranking is query-parser evidence, not a spelling-correction SLA |
3. Query-time analyzer is a precision tool, not an index rewrite
The current procedure options support skip,
limit, and analyzer. Overriding the
query analyzer can be useful when the application has a
deliberate query-side normalization policy, but the indexed
terms remain those produced at index creation. Compare results
under the default and a query-side analyzer only after listing
what is available.
CYPHER 25
CALL db.index.fulltext.queryNodes(
'ch21_catalog_ft',
'trail cameras',
{limit: 10, analyzer: 'whitespace'}
)
YIELD node, score
RETURN node.productId, node.name, score
ORDER BY score DESC;
// Compare candidate IDs/ranks with the same query using the default index analyzer.
4. Score interpretation: rank evidence, not probability
Neo4j explicitly states that full-text scores are meaningful within the result set from a full-text query. The magnitude depends on terms, document statistics, fields, analyzer and corpus. Adding products can change document-frequency statistics and therefore scores even when one Product did not change. A threshold copied from another index, environment or vector source is not evidence.
| Tempting interpretation | Why it is wrong | Safer alternative |
|---|---|---|
| score 0.83 = 83% relevant | Lucene score is not calibrated probability | judge ranking/thresholds on labeled queries |
| score > 0.5 always good | scale varies by query/corpus/analyzer | learn threshold per use case, or prefer rank/candidate limits |
| full-text 0.7 > vector 0.6 | different retrieval sources have incomparable raw scales | rank each list independently; fuse ranks |
| same product should keep same score forever | corpus/index statistics can change | regression-test ordering/quality, not frozen numbers |
5. Candidate limit is part of recall
If AtlasMart ultimately displays five products but
graph/security/stock filters remove some lexical candidates,
asking Lucene for only five can starve the post-filter stage.
Retrieve a bounded sourceK based on measured survival
rate and latency. The current full-text procedure supports
limit directly; skip exists for
pagination, but deep result browsing should still be evaluated
for cost and ranking stability.
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {skip: 0, limit: 20})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;
6. AtlasMart query matrix: record, do not invent
| ID | Query | Expected qualitative evidence |
|---|---|---|
| Q1 | wildlife camera | camera products and wildlife-related text compete; graph context later removes non-product-intent noise |
| Q2 | "trail camera" | Trail Camera products should be strong phrase candidates; capture exact rank/score on your server |
| Q3 | night vision camera | P-2101 and P-2103 contain strong lexical evidence; relative order is measured, not prescribed |
| Q4 | camra~ | fuzzy query should be tested for typo recovery; capture actual candidate set and latency |
| Q5 | camera NOT security | compare exclusion effect against plain camera query |
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 10})
YIELD node, score
RETURN node.productId AS productId, node.name AS name, score
ORDER BY score DESC;
// Example parameter: :param q => 'wildlife camera'
// Record actual scores; only the relative order in this result set has meaning.
7. Wrong approach → failure → repair
| Wrong approach | Failure | Repair |
|---|---|---|
| concatenate user input into Cypher | query injection / broken syntax risk | use parameters; separately define Lucene grammar policy |
| expose all Lucene syntax accidentally | users/operators can produce expensive or surprising queries | offer structured search or explicitly documented advanced mode |
| hard-code one numeric score threshold | quality changes when corpus/analyzer/query changes | calibrate against judged data and monitor false positives/negatives |
| return thousands of hits “for recall” | latency/network/materialization cost explodes | bounded sourceK + measured post-filter survival |
| compare raw lexical/vector scores | ranking has no common scale | independent ranks + RRF/WRRF in Chapter 22 |
8. Verification checklist
- All query strings are supplied as parameters to Cypher.
- The allowed Lucene grammar is documented: literal, structured, or advanced.
- Boolean, phrase, field and fuzzy experiments record IDs/ranks/latency from the current server.
- No score is described as a probability or compared directly to a vector score.
- Candidate limit is justified by downstream graph/business filtering and measured latency.
9. Production judgment
| Production decision | Evidence to require before changing the system |
|---|---|
| graph/workload fit | Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient. |
| correctness/non-guarantees | Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability. |
| model/cardinality/degree | Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval. |
| latency | Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call. |
| transactions/freshness | Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore. |
| memory/storage | Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization. |
| CPU/disk/network | Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network. |
| indexes/constraints | Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking. |
| driver/pool/timeouts | Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results. |
| security/tenant risk | Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization. |
| backup/recovery | Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health. |
| observability | Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters. |
| testing/failure injection | Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects. |
| version/tier | Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical. |
| cost/migration | Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval. |
Check your understanding
- Why does $q not make a Lucene query literal?
- What does a higher full-text score prove?
- Why might sourceK exceed the displayed result count?
- Why is a fuzzy query not the same as a spelling-correction service?
- How should full-text and vector results be combined later?
Review the answers
1. It protects the Cypher parser; the full-text procedure still intentionally parses the parameter value using Lucene query syntax.
2. Within that query result list, the index ranks that entry as a stronger lexical match; it does not prove business relevance or a calibrated probability.
3. Graph/security/business filters can remove candidates, so retrieval needs enough bounded headroom to preserve final recall.
4. It is Lucene fuzzy term matching; behavior/ranking must be evaluated against your domain queries and latency requirements.
5. Rank each source independently and fuse ranks, rather than comparing raw scores.
Summary and next step
Search syntax is an API contract, scores are source-local ranking evidence, and candidate limits are part of recall engineering. Lesson 3 adds the write-path decision: whether Lucene updates must be visible at commit or may enter a background queue with an explicit freshness SLO.
Authoritative references
- Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
- Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
- Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
- Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
- Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
- Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
- Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
- Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
- Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
- Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
- Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
- Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
- System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
- Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
- Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.