Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context
Combine Full-Text Candidates with Graph Traversal, Business Rules, and Post-Ranking
Turn lexical hits into useful AtlasMart results by preserving lexical rank, expanding through graph context, enforcing stock/security/business rules, and post-ranking without confusing graph evidence with the Lucene score.
A search for “wildlife camera” can legitimately retrieve the Wildlife Field Guide because its description contains “wildlife,” or a discontinued/refurbished camera because its text is an excellent lexical match. Lucene is doing its job. AtlasMart’s application has a different job: return purchasable products at a chosen store, respect security/tenant rules, preserve explainable lexical evidence, and possibly promote featured products. The graph is where those constraints and relationships live.
Retrieval first, graph/business decision second. Keep the full-text candidate rank and score visible, then traverse/filter with bounded graph patterns. Post-ranking should explain which signal changed the order rather than hiding everything inside one synthetic score.
Learning outcomes
Preserve lexical rank while expanding candidates through Category and STOCKED_AT graph context.
Apply active/in-stock/business filters after retrieval and explain the recall impact of sourceK vs finalK.
Separate lexical relevance from graph/business ranking instead of adding incomparable signals into an arbitrary numeric soup.
Understand current fine-grained security limitations of Lucene-backed semantic indexes and test the exact service role in Enterprise.
Compare lexical-only and graph-aware result lists with judged queries to determine whether graph context actually improves relevance.
Current Neo4j Database is 2026.07.1; current 5.26
LTS patch is 5.26.30. Version-sensitive examples
use explicit CYPHER 25. The mandatory lab uses
self-managed Neo4j Community 2026.07.1,
database neo4j, user neo4j, disposable
password atlasmart-course-2026, loopback Bolt
7687 and HTTP 7474. Neo4j 2026.x
supports Java 21/25. Optional application examples pin the
official Python driver to neo4j==6.3.0. No APOC,
GDS, embeddings, paid AI service, Aura subscription, or
Enterprise license is required for the chapter.
Full-text indexes and the
db.index.fulltext.* procedures used in the
mandatory lab are part of the normal Neo4j database feature set
and are reproducible in Community. Enterprise adds fine-grained
authorization and built-in metrics surfaces that can change what
full-text queries are allowed to return or what operational
counters are available. Aura also supports database full-text
functionality, but server configuration/filesystem/metrics
controls are platform-managed and must not be presented as
identical to self-managed Neo4j.
Lab contract and exact assumptions
| Dimension | Chapter 21 assumption |
|---|---|
| server | Neo4j Community 2026.07.1, single disposable local database |
| Cypher | Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02 |
| Java | Java 21 or 25 for the 2026.07 server line |
| database/auth | neo4j / neo4j / atlasmart-course-2026 |
| transport | bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS |
| plugins | none required; no APOC/GDS/GenAI plugin |
| graph | 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges |
| full-text | catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer |
| score policy | never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values |
| measurement | learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution |
| Term | Mechanism-first meaning |
|---|---|
| lexical retrieval | Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity. |
| full-text index |
Lucene-backed Neo4j semantic index over one or more STRING
or LIST |
| schema | The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property. |
| tokenization | Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior. |
| analyzer | Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms. |
| query analyzer | Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens. |
| Lucene query string | The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal. |
| score | Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale. |
| candidate set | Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size. |
| eventually consistent index | Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window. |
| freshness SLO | Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured. |
| judged query set | Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly. |
| precision@k | Fraction of the first k returned results judged relevant. |
| recall@k | Fraction of all judged relevant items retrieved in the first k results. |
| MRR | Mean Reciprocal Rank: rewards returning the first relevant result near the top. |
| nDCG | Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier. |
| post-ranking | Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model. |
| rank fusion | Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores. |
Direct-entry setup
Run the AtlasMart fixture if needed. The graph intentionally contains strong lexical candidates that business rules should remove: P-2102 is out of stock and P-2108 is inactive.
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';
UNWIND [
{id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
{id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
{id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
{id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
{id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
{id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
{id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
{id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';
MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
r2.message='Comfortable on long trail runs', r2.labTag='ch21';
CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
`fulltext.analyzer`: 'english',
`fulltext.eventually_consistent`: false
}};
CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};
CALL db.awaitIndexes(300);
1. Keep lexical rank before traversing
Collecting the ordered full-text hits and assigning
lexicalRank=i+1 makes the source order explicit.
This is useful even if the final business ordering changes. It
also creates the same abstraction Chapter 22 will use when
combining full-text and vector sources: preserve source ranks
rather than assuming raw scores share a scale.
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 20})
YIELD node, score
WITH collect({p:node, score:score}) AS hits
UNWIND range(0, size(hits)-1) AS i
WITH hits[i].p AS p, hits[i].score AS lexicalScore, i + 1 AS lexicalRank
MATCH (p)-[stock:STOCKED_AT]->(store:Store {storeId:$storeId})
MATCH (p)-[:IN_CATEGORY]->(category:Category)
WHERE p.active = true AND stock.quantity > 0
WITH p, category, stock, lexicalScore, lexicalRank,
CASE WHEN p.featured THEN 1 ELSE 0 END AS featured
ORDER BY featured DESC, lexicalRank ASC
RETURN p.productId AS productId, p.name AS name,
category.name AS category, stock.quantity AS quantity,
lexicalRank, lexicalScore, featured
LIMIT 5;
// Suggested parameters: q='wildlife camera', storeId='ST-21-CENTRAL'.
// This keeps lexicalRank visible and applies graph/business evidence explicitly.
2. Candidate generation and filtering have separate recall knobs
Suppose finalK=5 and sourceK=5. If two of the top lexical hits are out of stock/inactive, only three may survive. Raising sourceK can recover more valid candidates but increases Lucene, traversal, network and ranking work. Measure the survival rate and tail latency to set a bounded sourceK.
| Stage | Input | Evidence to capture |
|---|---|---|
| lexical retrieval | q, sourceK, analyzer/index | candidate IDs, lexicalRank, lexicalScore, latency |
| graph expansion | candidate nodes | degree/fan-out, category/store matches, DB hits/PROFILE on representative query |
| business/security filter | active/stock/tenant/privilege rules | rejection reason counts |
| post-rank | survivors | final rank + explanation fields |
| UI response | finalK | response bytes, p95/p99, zero-result rate |
3. Do not make one raw score responsible for unrelated policies
Within one full-text result set, lexical score is valid for
ordering lexical evidence. A business rule such as “featured
first” or “must be in stock” is not the same unit. The example
intentionally uses a hard filter for active/stock and a
lexicographic sort of
featured DESC, lexicalRank ASC. That is
explainable: featured status changes priority, then the original
lexical order breaks ties. Other products may choose lexical
rank first and use business signals only for tie-breaking.
Expressions such as
0.73*fulltextScore + 0.19*inventory + 0.08*margin
look scientific but have no meaning until every term is
normalized/calibrated and validated against judgments. Prefer
hard rules, rank features, or a trained/evaluated ranking
model.
4. Graph context can improve or harm relevance
Graph expansion is not automatically beneficial. A category edge can disambiguate product intent; a high-degree popularity relationship can drown niche relevance; an inventory filter can eliminate the best lexical answer even when the UI should instead show “out of stock.” Therefore compare result lists with and without each rule against the same judged queries and record reason codes for removals.
| Graph signal | Possible benefit | Possible harm |
|---|---|---|
| IN_CATEGORY Cameras | removes trail-running/book ambiguity for camera intent | query may intentionally seek accessories or guides |
| STOCKED_AT quantity > 0 | returns purchasable local inventory | hides relevant product the user may want back-in-stock notification for |
| active=true | removes discontinued/internal items | support/help search may need inactive products |
| featured | business promotion | can degrade relevance if allowed to dominate |
| REVIEWED message | social evidence/quality context | review volume/degree bias can favor popular products |
5. Security filtering is not ordinary post-filtering
Enterprise fine-grained authorization changes semantic-index behavior. Because Lucene cannot evaluate Neo4j security rules per entry, Neo4j takes a conservative approach: if a hit might violate denied label/property visibility, the semantic index can return partial or zero results. This protects data but can reduce recall. Test the actual service role and design indexes that do not unnecessarily mix differently restricted fields.
// Enterprise/security review pattern (not a Community RBAC exercise):
// 1. SHOW the full-text index schema/options.
// 2. Identify every indexed label/type/property.
// 3. Compare those fields with role TRAVERSE/READ denials.
// 4. Test the exact service role. Lucene-backed semantic indexes can conservatively
// suppress rows when Neo4j cannot prove that an entry is safe to return.
// 5. Do not compensate by granting wider data access merely to improve recall.
Authorization correctness outranks search recall. If one broad index crosses security domains, split the index/schema or use separate application retrieval paths rather than widening privileges.
6. Explainability response contract
Return enough internal evidence to debug ranking without
exposing sensitive index internals to end users. A service log
might include queryId, productId,
lexicalRank, lexicalScore, filter
reasons, category, quantity, featured flag and final rank.
User-facing explanations can be simpler: “matches trail camera;
in stock at AtlasMart Central.”
{
"queryId": "search-2026-09-09-001",
"productId": "P-2101",
"source": "ch21_catalog_ft",
"lexicalRank": 1,
"lexicalScore": "captured-at-runtime",
"filters": {"active": true, "inStock": true},
"graph": {"category": "Cameras", "quantity": 5},
"business": {"featured": true},
"finalRank": 1
}
7. Wrong approach → failure → repair
| Wrong approach | Failure | Repair |
|---|---|---|
| return lexical hits directly | inactive/out-of-stock/wrong-domain items leak into UX | bounded graph/business stage |
| traverse unbounded neighborhoods for every candidate | fan-out dominates latency | restrict sourceK and graph pattern/degree |
| sort only by featured flag | promotion overwhelms relevance | preserve lexical rank and evaluate rule ordering |
| assume full-text index bypasses authorization cleanly | Enterprise semantic-index security may conservatively suppress data | test role/index schema and accept security-first behavior |
| compare graph bonus directly to vector/full-text score | no shared scale | use ranks/rules/fusion with evaluation |
8. Verification checklist
- P-2102 is demonstrably out of stock and P-2108 inactive in the graph fixture.
- The query returns lexicalRank and lexicalScore before business decisions.
- Final results show category, quantity and featured evidence.
- sourceK and finalK are recorded independently.
- For Enterprise deployments, the exact service role is tested against semantic-index security behavior.
9. Production judgment
| Production decision | Evidence to require before changing the system |
|---|---|
| graph/workload fit | Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient. |
| correctness/non-guarantees | Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability. |
| model/cardinality/degree | Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval. |
| latency | Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call. |
| transactions/freshness | Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore. |
| memory/storage | Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization. |
| CPU/disk/network | Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network. |
| indexes/constraints | Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking. |
| driver/pool/timeouts | Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results. |
| security/tenant risk | Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization. |
| backup/recovery | Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health. |
| observability | Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters. |
| testing/failure injection | Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects. |
| version/tier | Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical. |
| cost/migration | Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval. |
Check your understanding
- Why preserve lexicalRank if the final order changes?
- Why can sourceK need to exceed finalK?
- Why is in-stock usually a filter rather than a raw score addition?
- Can graph context ever reduce search quality?
- What is the security-first response to semantic-index recall loss under denials?
Review the answers
1. It keeps the source retrieval evidence explainable and enables rank-based fusion/comparison later.
2. Graph/security/business filters can remove candidates; bounded headroom preserves final recall.
3. It is a business eligibility rule with a different meaning/unit from lexical relevance.
4. Yes. Overly strict category/stock/popularity rules can remove relevant intent; evaluate each rule against judgments.
5. Redesign index boundaries/retrieval paths or accept conservative filtering; do not broaden privileges just to recover recall.
Summary and next step
AtlasMart now has an explainable retrieval pipeline: lexical candidates → preserved rank → bounded graph expansion → eligibility/security/business rules → final order. Lesson 5 replaces demo-driven tuning with a judged query set and measurable regression gates.
Authoritative references
- Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
- Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
- Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
- Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
- Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
- Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
- Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
- Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
- Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
- Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
- Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
- Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
- System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
- Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
- Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.