Chapter 21 · Full-Text Search, Text Analysis, Relevance, and Hybrid Retrieval with Graph Context

Eventually Consistent Full-Text Indexing, Freshness Expectations, and Operational Monitoring

Decide whether full-text freshness belongs on the transaction commit path or a background queue, observe the visibility window safely, wait for deterministic test refreshes, and monitor index state without turning eventual consistency into an unexplained stale-read bug.

Advanced210–300 minutesFreshness/monitoring labNeo4j 2026.07.1 · Community mandatoryCypher 25 · Full-text/Lucene · english analyzerJava 21/25 · Python driver 6.3 optionalLast reviewed: September 2026

AtlasMart’s merchandising team updates a product title and immediately searches for the new wording. With a synchronous full-text index, commit completion includes the index update. With eventual consistency enabled, the graph commit can succeed before the new text becomes searchable. That is not “random stale data”; it is an explicit queueing trade-off that must be paired with a freshness SLO, measurement, queue/backpressure awareness, and deterministic test barriers.

Mental model

Synchronous full-text indexing puts index-update work on the commit path. Eventual mode queues full-text updates for a background applier. It can improve write-path behavior, but search visibility now has its own lag distribution and overload mode.

Learning outcomes

01

Contrast synchronous and eventually consistent full-text update semantics without confusing graph commit visibility with index visibility.

02

Create an isolated eventual-consistency probe index and measure immediate-vs-after-refresh visibility without asserting fabricated delay numbers.

03

Use awaitEventuallyConsistentIndexRefresh only as a deterministic test/maintenance barrier, not as a per-request freshness hack.

04

Inspect SHOW FULLTEXT INDEXES and distinguish Community-observable state from Enterprise-only metrics counters and platform-managed Aura controls.

05

Define a freshness SLO, alert evidence and overload response for queued full-text updates.

Chapter 21 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; current 5.26 LTS patch is 5.26.30. Version-sensitive examples use explicit CYPHER 25. The mandatory lab uses self-managed Neo4j Community 2026.07.1, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback Bolt 7687 and HTTP 7474. Neo4j 2026.x supports Java 21/25. Optional application examples pin the official Python driver to neo4j==6.3.0. No APOC, GDS, embeddings, paid AI service, Aura subscription, or Enterprise license is required for the chapter.

Full-text availability and platform boundary

Full-text indexes and the db.index.fulltext.* procedures used in the mandatory lab are part of the normal Neo4j database feature set and are reproducible in Community. Enterprise adds fine-grained authorization and built-in metrics surfaces that can change what full-text queries are allowed to return or what operational counters are available. Aura also supports database full-text functionality, but server configuration/filesystem/metrics controls are platform-managed and must not be presented as identical to self-managed Neo4j.

Lab contract and exact assumptions

Dimension Chapter 21 assumption
server Neo4j Community 2026.07.1, single disposable local database
Cypher Explicit CYPHER 25 for version-sensitive examples; current packaged configs default new databases to Cypher 25 from 2026.02
Java Java 21 or 25 for the 2026.07 server line
database/auth neo4j / neo4j / atlasmart-course-2026
transport bolt://localhost:7687 and http://localhost:7474 only for disposable loopback lab; remote/production uses verified TLS
plugins none required; no APOC/GDS/GenAI plugin
graph 8 Product nodes, 4 Category nodes, 1 Store, 2 Customers, 8 STOCKED_AT edges, 2 REVIEWED edges
full-text catalog node index uses english analyzer and synchronous updates; reviews relationship index uses english analyzer
score policy never hard-code expected score numbers; verify candidate IDs/order and collect actual scores because Lucene/corpus changes can alter values
measurement learner captures real query latency, ranking, freshness and index state; generated lesson does not claim server execution
Term Mechanism-first meaning
lexical retrieval Retrieval based on analyzed text terms and Lucene query semantics. It is different from graph traversal and different from embedding/vector similarity.
full-text index Lucene-backed Neo4j semantic index over one or more STRING or LIST properties of nodes or relationships, explicitly queried through full-text procedures.
schema The labels or relationship types plus indexed properties associated with one full-text index. An entity qualifies when it has at least one indexed label/type and at least one indexed property.
tokenization Breaking a character stream into searchable terms. The analyzer decides token boundaries, normalization, stemming and stop-word behavior.
analyzer Index/query text-processing pipeline. The default is standard-no-stop-words; language analyzers can stem/filter language-specific terms.
query analyzer Optional analyzer selected in queryNodes/queryRelationships options. It analyzes the query string only; it does not rebuild or reinterpret already-indexed tokens.
Lucene query string The second argument to a full-text query procedure. It can contain Boolean, phrase, field and fuzzy query syntax; parameterizing it protects Cypher syntax but does not make Lucene operators literal.
score Lucene relevance score returned with each hit. It orders results for that query/index; it is not a probability, confidence percentage or portable score scale.
candidate set Top lexical hits retained before graph/business filtering or downstream rank fusion. Candidate limit is a retrieval-recall decision, not merely a UI page size.
eventually consistent index Full-text mode that removes Lucene update work from the commit path and applies queued updates in the background, introducing a freshness window.
freshness SLO Application requirement for how soon committed text must become searchable. It determines whether eventual consistency is acceptable and how staleness is measured.
judged query set Versioned collection of user queries and relevance labels used to evaluate ranking changes reproducibly.
precision@k Fraction of the first k returned results judged relevant.
recall@k Fraction of all judged relevant items retrieved in the first k results.
MRR Mean Reciprocal Rank: rewards returning the first relevant result near the top.
nDCG Normalized Discounted Cumulative Gain: graded relevance metric that rewards putting highly relevant items earlier.
post-ranking Reordering or filtering a retrieved candidate set using graph context, business rules, permissions or a separate ranking model.
rank fusion Combining multiple retrieval lists by their rank positions instead of directly comparing source-specific raw scores.

Direct-entry setup

Run the core fixture if necessary. The freshness experiment uses a separate SearchProbe index so it cannot change the catalog ranking experiment.

Cypher 25 · core fixture
CYPHER 25
// Safe to rerun in the disposable Chapter 21 lab.
CREATE CONSTRAINT ch21_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch21_category_id IF NOT EXISTS
FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;
CREATE CONSTRAINT ch21_store_id IF NOT EXISTS
FOR (s:Store) REQUIRE s.storeId IS UNIQUE;
CREATE CONSTRAINT ch21_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;

MERGE (cam:Category {categoryId:'CAT-21-CAM'}) SET cam.name='Cameras', cam.labTag='ch21'
MERGE (out:Category {categoryId:'CAT-21-OUT'}) SET out.name='Outdoor', out.labTag='ch21'
MERGE (sec:Category {categoryId:'CAT-21-SEC'}) SET sec.name='Security', sec.labTag='ch21'
MERGE (acc:Category {categoryId:'CAT-21-ACC'}) SET acc.name='Accessories', acc.labTag='ch21'
MERGE (st:Store {storeId:'ST-21-CENTRAL'}) SET st.name='AtlasMart Central', st.labTag='ch21'
MERGE (c1:Customer {customerId:'C-2101'}) SET c1.name='Mina Rahimi', c1.labTag='ch21'
MERGE (c2:Customer {customerId:'C-2102'}) SET c2.name='Omid Karimi', c2.labTag='ch21';

UNWIND [
 {id:'P-2101',name:'Trail Camera Pro',description:'Weatherproof wildlife trail camera with infrared night vision and long battery life',tags:['wildlife','trail','infrared','outdoor'],active:true,category:'CAT-21-CAM',qty:5,featured:true},
 {id:'P-2102',name:'Trail Camera Mini',description:'Compact wildlife camera for trails, gardens, and backyard monitoring',tags:['wildlife','trail','compact'],active:true,category:'CAT-21-CAM',qty:0,featured:false},
 {id:'P-2103',name:'Indoor Security Camera',description:'Wi-Fi home security camera with motion alerts and night vision',tags:['security','indoor','night vision'],active:true,category:'CAT-21-SEC',qty:7,featured:false},
 {id:'P-2104',name:'Trail Running Hydration Vest',description:'Lightweight hydration vest for long trail runs and mountain races',tags:['running','trail','hydration'],active:true,category:'CAT-21-OUT',qty:11,featured:false},
 {id:'P-2105',name:'Action Camera 4K',description:'Water-resistant action sports camera for cycling, hiking, and travel',tags:['action','sports','camera'],active:true,category:'CAT-21-CAM',qty:3,featured:true},
 {id:'P-2106',name:'Wildlife Field Guide',description:'Illustrated guide to birds and mammals for outdoor observation',tags:['wildlife','book','outdoor'],active:true,category:'CAT-21-OUT',qty:6,featured:false},
 {id:'P-2107',name:'Solar Trail Charger',description:'Solar charger for outdoor cameras, sensors, and trail equipment',tags:['solar','trail','charger'],active:true,category:'CAT-21-ACC',qty:0,featured:false},
 {id:'P-2108',name:'Refurbished Trail Camera',description:'Older trail camera unit retained for support reference only',tags:['trail','camera','refurbished'],active:false,category:'CAT-21-CAM',qty:2,featured:false}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.description=row.description, p.tags=row.tags,
    p.active=row.active, p.featured=row.featured, p.labTag='ch21'
WITH row,p
MATCH (cat:Category {categoryId:row.category}), (st:Store {storeId:'ST-21-CENTRAL'})
MERGE (p)-[:IN_CATEGORY]->(cat)
MERGE (p)-[stock:STOCKED_AT]->(st)
SET stock.quantity=row.qty, stock.labTag='ch21';

MATCH (c1:Customer {customerId:'C-2101'}), (p1:Product {productId:'P-2101'})
MERGE (c1)-[r1:REVIEWED]->(p1)
SET r1.reviewId='REV-2101', r1.rating=5,
    r1.message='Excellent night vision for wildlife at the cabin', r1.labTag='ch21';
MATCH (c2:Customer {customerId:'C-2102'}), (p4:Product {productId:'P-2104'})
MERGE (c2)-[r2:REVIEWED]->(p4)
SET r2.reviewId='REV-2102', r2.rating=4,
    r2.message='Comfortable on long trail runs', r2.labTag='ch21';

CREATE FULLTEXT INDEX ch21_catalog_ft IF NOT EXISTS
FOR (p:Product) ON EACH [p.name, p.description, p.tags]
OPTIONS {indexConfig: {
  `fulltext.analyzer`: 'english',
  `fulltext.eventually_consistent`: false
}};

CREATE FULLTEXT INDEX ch21_reviews_ft IF NOT EXISTS
FOR ()-[r:REVIEWED]-() ON EACH [r.message]
OPTIONS {indexConfig: {`fulltext.analyzer`: 'english'}};

CALL db.awaitIndexes(300);
Cypher 25 · eventual-consistency probe index
CYPHER 25
CREATE FULLTEXT INDEX ch21_freshness_ft IF NOT EXISTS
FOR (n:SearchProbe) ON EACH [n.text]
OPTIONS {indexConfig: {
  `fulltext.analyzer`: 'standard-no-stop-words',
  `fulltext.eventually_consistent`: true
}};
CALL db.awaitIndexes(300);
MERGE (n:SearchProbe {probeId:'CH21-FRESH-1'})
SET n.text='baseline marker', n.labTag='ch21';
CALL db.index.fulltext.awaitEventuallyConsistentIndexRefresh();

1. Default mode is synchronous

Current Neo4j defaults db.index.fulltext.eventually_consistent=false. A new full-text index inherits that default unless its index configuration says otherwise. In this mode, the index update is part of commit completion. This reduces search staleness but means Lucene write work participates in the performance-critical commit path.

Mode Commit path Search visibility Primary risk
synchronous / default store + full-text update before commit returns commit completion is the normal visibility boundary write latency/index update cost
eventually consistent store commit queues full-text work short background-applier window possible freshness lag / queue saturation

2. Eventual mode has a queue, not a magic “fast writes” switch

Neo4j queues updates from all eventually consistent full-text indexes and applies them in background threads. Current configuration includes an update-queue maximum; when the queue fills, commits can block waiting for room. That means overload is delayed, not abolished. Do not copy the current default queue length or thread counts as universal tuning advice—measure update volume, heap, write latency and freshness.

Current configuration facts

As of 2026.07, the documented global defaults include eventual consistency off, update-apply parallelism 1, refresh interval 0s, refresh parallelism 1, and an update-queue maximum of 10,000 entries. These are defaults to inspect, not target values to cargo-cult.

Cypher 25 · inspect relevant settings
CYPHER 25
SHOW SETTINGS YIELD name, value, defaultValue, description
WHERE name STARTS WITH 'db.index.fulltext.eventually_consistent'
RETURN name, value, defaultValue, description
ORDER BY name;

3. Safe freshness experiment: absence before the barrier is allowed, not guaranteed

After Transaction A commits the sentinel update, the immediate full-text query may already see it because the background applier can run quickly. Therefore the lesson does not fabricate a stale result. The evidence is two timestamps and two candidate sets: “immediate visibility observed?” and “visibility after the supported await barrier?”.

Cypher 25 · freshness experiment
// Transaction A — commit a unique marker.
CYPHER 25
MATCH (n:SearchProbe {probeId:'CH21-FRESH-1'})
SET n.text='baseline marker violet-freshness-sentinel';

// Transaction B — query immediately and record whether it is visible.
CYPHER 25
CALL db.index.fulltext.queryNodes('ch21_freshness_ft','violet-freshness-sentinel')
YIELD node, score
RETURN node.probeId, score;
// It MAY already be visible; absence is allowed because the index is eventually consistent.

// Transaction C — test synchronization barrier, then query again.
CYPHER 25
CALL db.index.fulltext.awaitEventuallyConsistentIndexRefresh();
CALL db.index.fulltext.queryNodes('ch21_freshness_ft','violet-freshness-sentinel')
YIELD node, score
RETURN node.probeId, score;
// After the await procedure returns, the fixture expects CH21-FRESH-1 to be visible.
Testing boundary

db.index.fulltext.awaitEventuallyConsistentIndexRefresh() waits for recent queued updates to be applied. It is valuable for deterministic tests and controlled maintenance checks. Calling it on every user search would negate the purpose of eventual consistency and couple read latency to the update queue.

4. What to monitor in Community vs Enterprise

SHOW FULLTEXT INDEXES is the portable database-level evidence used here: state, populationPercent, options, lastRead, readCount and failure information. Neo4j’s built-in metrics framework—including counters such as full-text queried/populated—is documented as Enterprise Edition. Community learners can still measure application latency/freshness, OS CPU/disk, query counts in their harness and index state; do not claim Enterprise metrics exist locally when they do not.

Cypher 25 · operational index evidence
CYPHER 25
SHOW FULLTEXT INDEXES
YIELD name, state, populationPercent, lastRead, readCount,
      options, failureMessage
WHERE name STARTS WITH 'ch21_'
RETURN name, state, populationPercent, lastRead, readCount,
       options, failureMessage
ORDER BY name;
Signal Question
state/populationPercent Is the index usable and fully populated?
options Which analyzer and eventual-consistency policy was frozen at creation?
lastRead/readCount Is the index actually serving traffic?
application freshness probe How long from source commit timestamp to first search visibility?
write p95/p99 Did synchronous indexing dominate write tails? Did eventual mode materially help?
OS/JVM/Enterprise metrics Is lag caused by CPU, disk, GC, queue pressure or workload burst?

5. Freshness SLO design

Do not define freshness as “eventually.” Define a percentile and failure behavior: for example, a product rename should become searchable within a measured budget under normal write load; if not, the search UI may query by durable product ID, show a “recent update pending” state, or temporarily read the graph directly for the edited item. The exact budget comes from product requirements and measured distributions.

Requirement Likely choice
admin edits must be searchable immediately after save synchronous index or explicit read-your-own-write UX path
high-rate catalog feed tolerates seconds of search lag eventual mode may be justified after load testing
legal/security suppression must disappear immediately do not rely on a lagging semantic index as the only enforcement boundary
nightly bulk enrichment batch window can tolerate measured catch-up; verify queue/backpressure and completion

6. Wrong approach → failure → repair

Wrong approach Failure Repair
enable eventual consistency because “faster” users see stale search with no requirement or alert define freshness SLO and benchmark both modes
assert immediate query must be stale background update may already have completed record observed visibility; only require visibility after the await barrier in test
await refresh before every read read path becomes coupled to update backlog use it only for tests/controlled workflows
watch only index state=ONLINE ONLINE says usable, not “fresh within SLO” run sentinel freshness probes and record distributions
raise queue/thread settings blindly heap/CPU contention or commit blocking can move elsewhere change one setting with workload evidence and rollback

7. Verification checklist

  • The probe full-text index is ONLINE and explicitly configured eventual=true.
  • Immediate visibility is recorded as observed yes/no, not predetermined.
  • After awaitEventuallyConsistentIndexRefresh, the sentinel is retrievable.
  • The application records commit-to-search visibility latency over repeated probes.
  • Community and Enterprise monitoring capabilities are labeled correctly.

8. Production judgment

Production decision Evidence to require before changing the system
graph/workload fit Search logs and judged queries show a lexical need; traversal-only or exact/text-index predicates are not sufficient.
correctness/non-guarantees Document analyzer, query syntax contract, candidate limit, freshness mode and the fact that relevance score is not a probability.
model/cardinality/degree Measure candidate counts and graph expansion fan-out; cap/bound traversal after retrieval.
latency Track p50/p95/p99 for retrieval plus graph expansion separately; do not optimize only the Lucene call.
transactions/freshness Choose synchronous vs eventual full-text updates from freshness SLO and write-path cost, not folklore.
memory/storage Observe index size, page cache/store pressure, heap impact of eventual-consistency queues and result materialization.
CPU/disk/network Correlate query rate, index update rate, store I/O and response bytes; large candidate sets can shift cost to application/network.
indexes/constraints Keep business-key constraints separate from full-text access paths; wait for ONLINE before querying or benchmarking.
driver/pool/timeouts Use bounded result limits, parameterized Cypher and explicit timeout/retry policy; avoid keeping sessions open while users inspect results.
security/tenant risk Verify graph privileges plus semantic-index conservative filtering; do not assume index membership equals authorization.
backup/recovery Include full-text index recreation/check behavior in recovery drills and verify search after restore rather than assuming index health.
observability Community: SHOW FULLTEXT INDEXES + application logs/OS evidence. Enterprise: add supported metrics such as fulltext queried/populated counters.
testing/failure injection Regression-test analyzer changes, misspellings, empty queries, high-result queries, stale-index windows, denied-data cases and graph-filter effects.
version/tier Record Neo4j/Cypher/analyzer/index provider/platform versions; Aura/self-managed controls and metrics are not identical.
cost/migration Account for reindex time, storage, write amplification, evaluation maintenance and eventual move to vector/hybrid retrieval.

Check your understanding

  1. What becomes eventually consistent when the option is true?
  2. Why can a commit eventually block even with background indexing?
  3. What does ONLINE prove?
  4. Why should a test not require the immediate post-commit query to miss?
  5. What is the deterministic test barrier?
Review the answers

1. The full-text index update/visibility; the underlying graph transaction still has its normal commit semantics.

2. The shared update queue has a finite maximum; if it fills, commits wait for the applier to create room.

3. The index is usable/populated; it does not prove that eventual updates meet the application freshness SLO.

4. The background updater may apply the change before the query runs; only the possible lag window is guaranteed by the model.

5. db.index.fulltext.awaitEventuallyConsistentIndexRefresh(), followed by a verification query.

Summary and next step

Freshness is now an explicit system property: commit timing, background queue, visibility probe and SLO. Lesson 4 takes the lexical candidate set and adds what the index cannot know—stock, category, active state, permissions and graph relationships—while keeping source ranking evidence observable.

Authoritative references

  • Current Neo4j versions — Current database release 2026.07.1 and current 5.26 LTS patch 5.26.30.
  • Full-text indexes — Cypher 25 — Current schema, analyzer, query, score, eventual-consistency, SHOW FULLTEXT INDEXES and procedure semantics.
  • Semantic indexes — Why full-text and vector indexes are explicit semantic retrieval systems and why raw cross-source scores should not be compared.
  • Built-in full-text procedures — Current signatures for queryNodes/queryRelationships/listAvailableAnalyzers/awaitEventuallyConsistentIndexRefresh and limit/skip/query-analyzer options.
  • Index configuration — Default analyzer, eventual-consistency queue model, background update settings and operational implications.
  • Configuration settings — Current db.index.fulltext.* defaults, including standard-no-stop-words and eventual-consistency settings.
  • Index syntax — Current CREATE/SHOW/DROP and semantic-index query syntax; SHOW FULLTEXT INDEXES is the supported filtered SHOW form.
  • Security limitations for semantic indexes — Lucene-backed full-text/vector security filtering can conservatively return partial or zero results under fine-grained restrictions.
  • Hybrid search developer guide — Current rank-fusion guidance for lexical, vector and structural sources; fuse ranks, not incomparable raw scores.
  • Hybrid search engineering article — 2026 worked explanation of combining words, meaning and graph topology with rank-based fusion.
  • Metrics — Enterprise-only built-in metrics surface and monitoring responsibilities.
  • Metrics reference — Current full-text queried/populated counters; edition boundary is Enterprise.
  • System requirements — Neo4j 2026.07 supported Java 21/25 and current OS/runtime boundaries.
  • Python driver 6.3 — Official driver API used by the optional evaluation harness; current 6.3 supports Neo4j 2026.x.
  • Vector SEARCH clause — Bridge to Chapter 22 and explicit current warning to rank vector/full-text sources independently.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.