Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context
Eventually Consistent Full-Text Indexing, Freshness Expectations, and Operational Monitoring
Decide whether full-text freshness belongs on the transaction commit path or a background queue, observe the visibility window safely, wait for deterministic test refreshes, and monitor index state without turning eventual consistency into an unexplained stale-read bug.
AtlasMart’s merchandising team updates a product title and immediately searches for the new wording. With a synchronous full-text index, commit completion includes the index update. With eventual consistency enabled, the graph commit can succeed before the new text becomes searchable. That is not “random stale data”; it is an explicit queueing trade-off that must be paired with a freshness SLO, measurement, queue/backpressure awareness, and deterministic test barriers.
Synchronous full-text indexing puts index-update work on the commit path. Eventual mode queues full-text updates for a background applier. It can improve write-path behavior, but search visibility now has its own lag distribution and overload mode.
Learning outcomes
Contrast synchronous and eventually consistent full-text update semantics without confusing graph commit visibility with index visibility.
Create an isolated eventual-consistency probe index and measure immediate-vs-after-refresh visibility without asserting fabricated delay numbers.
Use awaitEventuallyConsistentIndexRefresh only as a deterministic test/maintenance barrier, not as a per-request freshness hack.
Inspect SHOW FULLTEXT INDEXES and distinguish Community-observable state from Enterprise-only metrics counters and platform-managed Aura controls.
Define a freshness SLO, alert evidence and overload response for queued full-text updates.
Current Neo4j Database is 2026.07.1; current 5.26
LTS patch is 5.26.30. Version-sensitive examples
use explicit CYPHER 25. The mandatory lab uses
self-managed Neo4j Community 2026.07.1,
database neo4j, user neo4j, disposable
password atlasmart-course-2026, loopback Bolt
7687 and HTTP 7474. Neo4j 2026.x
supports Java 21/25. Optional application examples pin the
official Python driver to neo4j==6.3.0. No APOC,
GDS, embeddings, paid AI service, Aura subscription, or
Enterprise license is required for the chapter.
Full-text indexes and the
db.index.fulltext.* procedures used in the
mandatory lab are part of the normal Neo4j database feature set
and are reproducible in Community. Enterprise adds fine-grained
authorization and built-in metrics surfaces that can change what
full-text queries are allowed to return or what operational
counters are available. Aura also supports database full-text
functionality, but server configuration/filesystem/metrics
controls are platform-managed and must not be presented as
identical to self-managed Neo4j.
Lab contract and exact assumptions
| Dimension | Chapter 21 assumption |
|---|---|
| server | Neo4j Community 2026.07.1, single disposable local database |
| Cypher | Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02 |
| Java | Java 21 or 25 for the 2026.07 server line |
| database/auth | neo4j / neo4j / atlasmart-course-2026 |
| transport | bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS |
| plugins | none required; no APOC/GDS/GenAI plugin |
| graph | 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges |
| full-text | catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer |
| score policy | never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values |
| measurement | learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution |
| Term | Mechanism-first meaning |
|---|---|
| lexical retrieval | Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity. |
| full-text index |
Lucene-backed Neo4j semantic index over one or more STRING
or LIST |
| schema | The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property. |
| tokenization | Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior. |
| analyzer | Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms. |
| query analyzer | Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens. |
| Lucene query string | The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal. |
| score | Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale. |
| candidate set | Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size. |
| eventually consistent index | Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window. |
| freshness SLO | Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured. |
| judged query set | Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly. |
| precision@k | Fraction of the first k returned results judged relevant. |
| recall@k | Fraction of all judged relevant items retrieved in the first k results. |
| MRR | Mean Reciprocal Rank: rewards returning the first relevant result near the top. |
| nDCG | Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier. |
| post-ranking | Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model. |
| rank fusion | Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores. |
Direct-entry setup
Run the core fixture if necessary. The freshness experiment uses
a separate SearchProbe index so it cannot change
the catalog ranking experiment.
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';
UNWIND [
{id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
{id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
{id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
{id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
{id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
{id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
{id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
{id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';
MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
r2.message='Comfortable on long trail runs', r2.labTag='ch21';
CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
`fulltext.analyzer`: 'english',
`fulltext.eventually_consistent`: false
}};
CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};
CALL db.awaitIndexes(300);
CYPHER 25
CREATE FULLTEXT INDEX ch21_freshness_ft IF NOT EXISTS
FOR (n:SearchProbe) ON EACH [n.text]
OPTIONS {indexConfig: {
`fulltext.analyzer`: 'standard-no-stop-words',
`fulltext.eventually_consistent`: true
}};
CALL db.awaitIndexes(300);
MERGE (n:SearchProbe {probeId:'CH21-FRESH-1'})
SET n.text='baseline marker', n.labTag='ch21';
CALL db.index.fulltext.awaitEventuallyConsistentIndexRefresh();
1. Default mode is synchronous
Current Neo4j defaults
db.index.fulltext.eventually_consistent=false. A
new full-text index inherits that default unless its index
configuration says otherwise. In this mode, the index update is
part of commit completion. This reduces search staleness but
means Lucene write work participates in the performance-critical
commit path.
| Mode | Commit path | Search visibility | Primary risk |
|---|---|---|---|
| synchronous / default | store + full-text update before commit returns | commit completion is the normal visibility boundary | write latency/index update cost |
| eventually consistent | store commit queues full-text work | short background-applier window possible | freshness lag / queue saturation |
2. Eventual mode has a queue, not a magic “fast writes” switch
Neo4j queues updates from all eventually consistent full-text indexes and applies them in background threads. Current configuration includes an update-queue maximum; when the queue fills, commits can block waiting for room. That means overload is delayed, not abolished. Do not copy the current default queue length or thread counts as universal tuning advice—measure update volume, heap, write latency and freshness.
As of 2026.07, the documented global defaults include eventual consistency off, update-apply parallelism 1, refresh interval 0s, refresh parallelism 1, and an update-queue maximum of 10,000 entries. These are defaults to inspect, not target values to cargo-cult.
CYPHER 25
SHOW SETTINGS YIELD name, value, defaultValue, description
WHERE name STARTS WITH 'db.index.fulltext.eventually_consistent'
RETURN name, value, defaultValue, description
ORDER BY name;
3. Safe freshness experiment: absence before the barrier is allowed, not guaranteed
After Transaction A commits the sentinel update, the immediate full-text query may already see it because the background applier can run quickly. Therefore the lesson does not fabricate a stale result. The evidence is two timestamps and two candidate sets: “immediate visibility observed?” and “visibility after the supported await barrier?”.
// Transaction A — commit a unique marker.
CYPHER 25
MATCH (n:SearchProbe {probeId:'CH21-FRESH-1'})
SET n.text='baseline marker violet-freshness-sentinel';
// Transaction B — query immediately and record whether it is visible.
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_freshness_ft','violet-freshness-sentinel')
YIELD node, score
RETURN node.probeId, score;
// It MAY already be visible; absence is allowed because the index is eventually consistent.
// Transaction C — test synchronization barrier, then query again.
CYPHER 25
CALL db.index.fulltext.awaitEventuallyConsistentIndexRefresh();
CALL db.index.fulltext.queryNodes('ch21_freshness_ft','violet-freshness-sentinel')
YIELD node, score
RETURN node.probeId, score;
// After the await procedure returns, the fixture expects CH21-FRESH-1 to be visible.
db.index.fulltext.awaitEventuallyConsistentIndexRefresh()
waits for recent queued updates to be applied. It is valuable
for deterministic tests and controlled maintenance checks.
Calling it on every user search would negate the purpose of
eventual consistency and couple read latency to the update
queue.
4. What to monitor in Community vs Enterprise
SHOW FULLTEXT INDEXES is the portable
database-level evidence used here: state, populationPercent,
options, lastRead, readCount and failure information. Neo4j’s
built-in metrics framework—including counters such as full-text
queried/populated—is documented as Enterprise Edition. Community
learners can still measure application latency/freshness, OS
CPU/disk, query counts in their harness and index state; do not
claim Enterprise metrics exist locally when they do not.
CYPHER 25
SHOW FULLTEXT INDEXES
YIELD name, state, populationPercent, lastRead, readCount,
options, failureMessage
WHERE name STARTS WITH 'ch21_'
RETURN name, state, populationPercent, lastRead, readCount,
options, failureMessage
ORDER BY name;
| Signal | Question |
|---|---|
| state/populationPercent | Is the index usable and fully populated? |
| options | Which analyzer and eventual-consistency policy was frozen at creation? |
| lastRead/readCount | Is the index actually serving traffic? |
| application freshness probe | How long from source commit timestamp to first search visibility? |
| write p95/p99 | Did synchronous indexing dominate write tails? Did eventual mode materially help? |
| OS/JVM/Enterprise metrics | Is lag caused by CPU, disk, GC, queue pressure or workload burst? |
5. Freshness SLO design
Do not define freshness as “eventually.” Define a percentile and failure behavior: for example, a product rename should become searchable within a measured budget under normal write load; if not, the search UI may query by durable product ID, show a “recent update pending” state, or temporarily read the graph directly for the edited item. The exact budget comes from product requirements and measured distributions.
| Requirement | Likely choice |
|---|---|
| admin edits must be searchable immediately after save | synchronous index or explicit read-your-own-write UX path |
| high-rate catalog feed tolerates seconds of search lag | eventual mode may be justified after load testing |
| legal/security suppression must disappear immediately | do not rely on a lagging semantic index as the only enforcement boundary |
| nightly bulk enrichment | batch window can tolerate measured catch-up; verify queue/backpressure and completion |
6. Wrong approach → failure → repair
| Wrong approach | Failure | Repair |
|---|---|---|
| enable eventual consistency because “faster” | users see stale search with no requirement or alert | define freshness SLO and benchmark both modes |
| assert immediate query must be stale | background update may already have completed | record observed visibility; only require visibility after the await barrier in test |
| await refresh before every read | read path becomes coupled to update backlog | use it only for tests/controlled workflows |
| watch only index state=ONLINE | ONLINE says usable, not “fresh within SLO” | run sentinel freshness probes and record distributions |
| raise queue/thread settings blindly | heap/CPU contention or commit blocking can move elsewhere | change one setting with workload evidence and rollback |
7. Verification checklist
- The probe full-text index is ONLINE and explicitly configured eventual=true.
- Immediate visibility is recorded as observed yes/no, not predetermined.
-
After
awaitEventuallyConsistentIndexRefresh, the sentinel is retrievable. - The application records commit-to-search visibility latency over repeated probes.
- Community and Enterprise monitoring capabilities are labeled correctly.
8. Production judgment
| Production decision | Evidence to require before changing the system |
|---|---|
| graph/workload fit | Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient. |
| correctness/non-guarantees | Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability. |
| model/cardinality/degree | Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval. |
| latency | Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call. |
| transactions/freshness | Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore. |
| memory/storage | Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization. |
| CPU/disk/network | Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network. |
| indexes/constraints | Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking. |
| driver/pool/timeouts | Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results. |
| security/tenant risk | Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization. |
| backup/recovery | Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health. |
| observability | Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters. |
| testing/failure injection | Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects. |
| version/tier | Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical. |
| cost/migration | Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval. |
Check your understanding
- What becomes eventually consistent when the option is true?
- Why can a commit eventually block even with background indexing?
- What does ONLINE prove?
- Why should a test not require the immediate post-commit query to miss?
- What is the deterministic test barrier?
Review the answers
1. The full-text index update/visibility; the underlying graph transaction still has its normal commit semantics.
2. The shared update queue has a finite maximum; if it fills, commits wait for the applier to create room.
3. The index is usable/populated; it does not prove that eventual updates meet the application freshness SLO.
4. The background updater may apply the change before the query runs; only the possible lag window is guaranteed by the model.
5. db.index.fulltext.awaitEventuallyConsistentIndexRefresh(), followed by a verification query.
Summary and next step
Freshness is now an explicit system property: commit timing, background queue, visibility probe and SLO. Lesson 4 takes the lexical candidate set and adds what the index cannot know—stock, category, active state, permissions and graph relationships—while keeping source ranking evidence observable.
Authoritative references
- Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
- Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
- Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
- Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
- Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
- Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
- Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
- Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
- Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
- Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
- Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
- Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
- System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
- Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
- Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.