Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context

Combine Full-Text Candidates with Graph Traversal, Business Rules, and Post-Ranking

Turn lexical hits into useful AtlasMart results by preserving lexical rank, expanding through graph context, enforcing stock/security/business rules, and post-ranking without confusing graph evidence with the Lucene score.

Advanced230–320 minutesGraph-aware retrieval labNeo4j 2026.07.1 · Community mandatoryCypher 25 · Full-text/Lucene · english analyzerJava 21/25 · Python driver 6.3 optionalLast reviewed: September 2026

A search for “wildlife camera” can legitimately retrieve the Wildlife Field Guide because its description contains “wildlife,” or a discontinued/refurbished camera because its text is an excellent lexical match. Lucene is doing its job. AtlasMart’s application has a different job: return purchasable products at a chosen store, respect security/tenant rules, preserve explainable lexical evidence, and possibly promote featured products. The graph is where those constraints and relationships live.

Mental model

Retrieval first, graph/business decision second. Keep the full-text candidate rank and score visible, then traverse/filter with bounded graph patterns. Post-ranking should explain which signal changed the order rather than hiding everything inside one synthetic score.

Learning outcomes

01

Preserve lexical rank while expanding candidates through Category and STOCKED_AT graph context.

02

Apply active/in-stock/business filters after retrieval and explain the recall impact of sourceK vs finalK.

03

Separate lexical relevance from graph/business ranking instead of adding incomparable signals into an arbitrary numeric soup.

04

Understand current fine-grained security limitations of Lucene-backed semantic indexes and test the exact service role in Enterprise.

05

Compare lexical-only and graph-aware result lists with judged queries to determine whether graph context actually improves relevance.

Chapter 21 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; current 5.26 LTS patch is 5.26.30. Version-sensitive examples use explicit CYPHER 25. The mandatory lab uses self-managed Neo4j Community 2026.07.1, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback Bolt 7687 and HTTP 7474. Neo4j 2026.x supports Java 21/25. Optional application examples pin the official Python driver to neo4j==6.3.0. No APOC, GDS, embeddings, paid AI service, Aura subscription, or Enterprise license is required for the chapter.

Full-text availability and platform boundary

Full-text indexes and the db.index.fulltext.* procedures used in the mandatory lab are part of the normal Neo4j database feature set and are reproducible in Community. Enterprise adds fine-grained authorization and built-in metrics surfaces that can change what full-text queries are allowed to return or what operational counters are available. Aura also supports database full-text functionality, but server configuration/filesystem/metrics controls are platform-managed and must not be presented as identical to self-managed Neo4j.

Lab contract and exact assumptions

Dimension Chapter 21 assumption
server Neo4j Community 2026.07.1, single disposable local database
Cypher Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02
Java Java 21 or 25 for the 2026.07 server line
database/auth neo4j / neo4j / atlasmart-course-2026
transport bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS
plugins none required; no APOC/GDS/GenAI plugin
graph 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges
full-text catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer
score policy never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values
measurement learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution
Term Mechanism-first meaning
lexical retrieval Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity.
full-text index Lucene-backed Neo4j semantic index over one or more STRING or LIST properties of nodes or relationships, explicitly queried through full-text procedures.
schema The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property.
tokenization Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior.
analyzer Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms.
query analyzer Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens.
Lucene query string The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal.
score Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale.
candidate set Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size.
eventually consistent index Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window.
freshness SLO Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured.
judged query set Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly.
precision@k Fraction of the first k returned results judged relevant.
recall@k Fraction of all judged relevant items retrieved in the first k results.
MRR Mean Reciprocal Rank: rewards returning the first relevant result near the top.
nDCG Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier.
post-ranking Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model.
rank fusion Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores.

Direct-entry setup

Run the AtlasMart fixture if needed. The graph intentionally contains strong lexical candidates that business rules should remove: P-2102 is out of stock and P-2108 is inactive.

Cypher 25 · fixture
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;

MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';

UNWIND [
 {id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
 {id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
 {id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
 {id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
 {id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
 {id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
 {id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
 {id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
    p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';

MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
    r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
    r2.message='Comfortable on long trail runs', r2.labTag='ch21';

CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
  `fulltext.analyzer`: 'english',
  `fulltext.eventually_consistent`: false
}};

CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};

CALL db.awaitIndexes(300);

1. Keep lexical rank before traversing

Collecting the ordered full-text hits and assigning lexicalRank=i+1 makes the source order explicit. This is useful even if the final business ordering changes. It also creates the same abstraction Chapter 22 will use when combining full-text and vector sources: preserve source ranks rather than assuming raw scores share a scale.

Cypher 25 · graph-aware retrieval
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_catalog_ft', $q, {limit: 20})
YIELD node, score
WITH collect({p:node, score:score}) AS hits
UNWIND range(0, size(hits)-1) AS i
WITH hits[i].p AS p, hits[i].score AS lexicalScore, i + 1 AS lexicalRank
MATCH (p)-[stock:STOCKED_AT]->(store:Store {storeId:$storeId})
MATCH (p)-[:IN_CATEGORY]->(category:Category)
WHERE p.active = true AND stock.quantity > 0
WITH p, category, stock, lexicalScore, lexicalRank,
     CASE WHEN p.featured THEN 1 ELSE 0 END AS featured
ORDER BY featured DESC, lexicalRank ASC
RETURN p.productId AS productId, p.name AS name,
       category.name AS category, stock.quantity AS quantity,
       lexicalRank, lexicalScore, featured
LIMIT 5;
// Suggested parameters: q='wildlife camera', storeId='ST-21-CENTRAL'.
// This keeps lexicalRank visible and applies graph/business evidence explicitly.

2. Candidate generation and filtering have separate recall knobs

Suppose finalK=5 and sourceK=5. If two of the top lexical hits are out of stock/inactive, only three may survive. Raising sourceK can recover more valid candidates but increases Lucene, traversal, network and ranking work. Measure the survival rate and tail latency to set a bounded sourceK.

Stage Input Evidence to capture
lexical retrieval q, sourceK, analyzer/index candidate IDs, lexicalRank, lexicalScore, latency
graph expansion candidate nodes degree/fan-out, category/store matches, DB hits/PROFILE on representative query
business/security filter active/stock/tenant/privilege rules rejection reason counts
post-rank survivors final rank + explanation fields
UI response finalK response bytes, p95/p99, zero-result rate

3. Do not make one raw score responsible for unrelated policies

Within one full-text result set, lexical score is valid for ordering lexical evidence. A business rule such as “featured first” or “must be in stock” is not the same unit. The example intentionally uses a hard filter for active/stock and a lexicographic sort of featured DESC, lexicalRank ASC. That is explainable: featured status changes priority, then the original lexical order breaks ties. Other products may choose lexical rank first and use business signals only for tie-breaking.

Avoid arbitrary score soup

Expressions such as 0.73*fulltextScore + 0.19*inventory + 0.08*margin look scientific but have no meaning until every term is normalized/calibrated and validated against judgments. Prefer hard rules, rank features, or a trained/evaluated ranking model.

4. Graph context can improve or harm relevance

Graph expansion is not automatically beneficial. A category edge can disambiguate product intent; a high-degree popularity relationship can drown niche relevance; an inventory filter can eliminate the best lexical answer even when the UI should instead show “out of stock.” Therefore compare result lists with and without each rule against the same judged queries and record reason codes for removals.

Graph signal Possible benefit Possible harm
IN_CATEGORY Cameras removes trail-running/book ambiguity for camera intent query may intentionally seek accessories or guides
STOCKED_AT quantity > 0 returns purchasable local inventory hides relevant product the user may want back-in-stock notification for
active=true removes discontinued/internal items support/help search may need inactive products
featured business promotion can degrade relevance if allowed to dominate
REVIEWED message social evidence/quality context review volume/degree bias can favor popular products

5. Security filtering is not ordinary post-filtering

Enterprise fine-grained authorization changes semantic-index behavior. Because Lucene cannot evaluate Neo4j security rules per entry, Neo4j takes a conservative approach: if a hit might violate denied label/property visibility, the semantic index can return partial or zero results. This protects data but can reduce recall. Test the actual service role and design indexes that do not unnecessarily mix differently restricted fields.

Security review checklist · Enterprise concept
// Enterprise/security review pattern (not a Community RBAC exercise):
// 1. SHOW the full-text index schema/options.
// 2. Identify every indexed label/type/property.
// 3. Compare those fields with role TRAVERSE/READ denials.
// 4. Test the exact service role. Lucene-backed semantic indexes can conservatively
//    suppress rows when Neo4j cannot prove that an entry is safe to return.
// 5. Do not compensate by granting wider data access merely to improve recall.
Do not “fix” recall by granting more data

Authorization correctness outranks search recall. If one broad index crosses security domains, split the index/schema or use separate application retrieval paths rather than widening privileges.

6. Explainability response contract

Return enough internal evidence to debug ranking without exposing sensitive index internals to end users. A service log might include queryId, productId, lexicalRank, lexicalScore, filter reasons, category, quantity, featured flag and final rank. User-facing explanations can be simpler: “matches trail camera; in stock at AtlasMart Central.”

Example structured diagnostic row
{
  "queryId": "search-2026-09-09-001",
  "productId": "P-2101",
  "source": "ch21_catalog_ft",
  "lexicalRank": 1,
  "lexicalScore": "captured-at-runtime",
  "filters": {"active": true, "inStock": true},
  "graph": {"category": "Cameras", "quantity": 5},
  "business": {"featured": true},
  "finalRank": 1
}

7. Wrong approach → failure → repair

Wrong approach Failure Repair
return lexical hits directly inactive/out-of-stock/wrong-domain items leak into UX bounded graph/business stage
traverse unbounded neighborhoods for every candidate fan-out dominates latency restrict sourceK and graph pattern/degree
sort only by featured flag promotion overwhelms relevance preserve lexical rank and evaluate rule ordering
assume full-text index bypasses authorization cleanly Enterprise semantic-index security may conservatively suppress data test role/index schema and accept security-first behavior
compare graph bonus directly to vector/full-text score no shared scale use ranks/rules/fusion with evaluation

8. Verification checklist

  • P-2102 is demonstrably out of stock and P-2108 inactive in the graph fixture.
  • The query returns lexicalRank and lexicalScore before business decisions.
  • Final results show category, quantity and featured evidence.
  • sourceK and finalK are recorded independently.
  • For Enterprise deployments, the exact service role is tested against semantic-index security behavior.

9. Production judgment

Production decision Evidence to require before changing the system
graph/workload fit Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient.
correctness/non-guarantees Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability.
model/cardinality/degree Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval.
latency Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call.
transactions/freshness Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore.
memory/storage Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization.
CPU/disk/network Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network.
indexes/constraints Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking.
driver/pool/timeouts Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results.
security/tenant risk Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization.
backup/recovery Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health.
observability Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters.
testing/failure injection Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects.
version/tier Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical.
cost/migration Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval.

Check your understanding

  1. Why preserve lexicalRank if the final order changes?
  2. Why can sourceK need to exceed finalK?
  3. Why is in-stock usually a filter rather than a raw score addition?
  4. Can graph context ever reduce search quality?
  5. What is the security-first response to semantic-index recall loss under denials?
Review the answers

1. It keeps the source retrieval evidence explainable and enables rank-based fusion/comparison later.

2. Graph/security/business filters can remove candidates; bounded headroom preserves final recall.

3. It is a business eligibility rule with a different meaning/unit from lexical relevance.

4. Yes. Overly strict category/stock/popularity rules can remove relevant intent; evaluate each rule against judgments.

5. Redesign index boundaries/retrieval paths or accept conservative filtering; do not broaden privileges just to recover recall.

Summary and next step

AtlasMart now has an explainable retrieval pipeline: lexical candidates → preserved rank → bounded graph expansion → eligibility/security/business rules → final order. Lesson 5 replaces demo-driven tuning with a judged query set and measurable regression gates.

Authoritative references

  • Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
  • Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
  • Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
  • Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
  • Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
  • Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
  • Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
  • Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
  • Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
  • Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
  • Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
  • Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
  • System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
  • Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
  • Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.