Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search
Index-Time vs Search-Time Analysis and Why Token Streams Define Search Semantics
Make the token stream visible before writing queries: trace how AtlasMart product text is transformed at index time and search time, then prove why exact-term and full-text queries behave differently.
Learning outcomes
AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.
Explain character filters, one tokenizer, and ordered token filters as the three stages of a custom analyzer.
Separate index-time analysis from search-time analysis and predict how analyzer mismatch changes recall and precision.
Use _analyze to inspect token text, positions,
offsets, and types instead of guessing from the original
string.
Distinguish a full-text match query from a
term-level term query on an analyzed field.
Create a portable AtlasMart v1 analysis contract for product names, descriptions, SKUs, and categories.
Examples target Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep the
course lab contract established earlier: Elasticsearch at
https://localhost:9200 authenticated as the
disposable lab elastic user via
ELASTIC_PASSWORD and verified with the copied
HTTP CA; OpenSearch at
https://localhost:9201 using the disposable
Security-plugin demo admin identity via
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier
OpenSearch -k path is local demo convenience
only, not a production certificate-verification pattern.
Portable examples use core analysis components that both
products document. Product-specific synonym-management and
reload APIs are labeled explicitly instead of pretending the
two platforms expose one interchangeable control plane.
The generation environment does not run the two search
clusters, so expected token/output fragments are documented as
deterministic shapes derived from the configured analyzers,
not falsely presented as captured benchmark output. Run the
requests in the disposable AtlasMart lab from Chapters 01–04
with one local node per product, dedicated
atlasmart-* indices, and the pinned
server/container assumptions from those chapters. Analyzer
experiments use dedicated versioned indices only; cleanup
commands must never target unrelated data, credentials,
certificates, or volumes.
1. The indexed unit is the token stream, not the sentence
A character filter transforms the input character stream before tokenization—for example stripping HTML or applying a configured character mapping. A tokenizer then emits token boundaries. An analyzer has exactly one tokenizer. Ordered token filters may lowercase, stem, remove, expand, or otherwise transform those tokens. Each emitted token has text plus metadata such as a position and character offsets. Lucene builds its inverted index from those analyzed terms.
That model immediately explains why copying the visible product title into a query is insufficient to predict matching. The index analyzer determines what terms exist in the index; the search analyzer determines which terms the user query asks for. Search succeeds only when those two contracts are compatible enough for the query semantics.
POST _analyze
{
"char_filter": ["html_strip"],
"tokenizer": "standard",
"filter": ["lowercase", "asciifolding"],
"text": "<b>Café USB-C Charger</b>"
}
For this fixture, expect terms equivalent to cafe,
usb, c, and charger, with
positions increasing through the stream. The exact response also
reports offsets/types. That evidence proves what this analysis
request emitted; it does not prove how an existing field is
mapped unless you invoke the field's configured analyzer.
2. Index-time and search-time analysis are separate contracts
For a text field, the field's
analyzer is used when text is indexed. By default
the same analyzer is also used for full-text search, but a field
may declare a separate search_analyzer. That
separation is useful when indexing needs extra
terms—autocomplete prefixes are the classic example—while
queries should remain compact. It is also where synonym
operations often belong because changing search-time expansion
can avoid rewriting every indexed document.
| Stage | Input | Output that matters | Operational consequence |
|---|---|---|---|
| Index analysis | Document field value | Terms written to the inverted index | Changing it normally requires reindexing old documents |
| Search analysis | User/query text | Terms/graphs used to construct the query | Some search-only resources can be updated/reloaded product-specifically |
| Normalizer | keyword value |
Exactly one normalized token | No tokenizer; only compatible single-token filters |
3. Build AtlasMart's first explicit analysis contract
Product names need full-text search plus an exact sortable/facetable representation. Descriptions need full-text search but not an exact-value aggregation. SKUs and category codes are identifiers: they should not be split into words. This is query-driven schema design from Chapter 03 carried into analysis.
PUT atlasmart-products-analysis-v1
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0,
"analysis": {
"normalizer": {
"atlas_keyword_lc": {
"type": "custom",
"filter": ["lowercase", "asciifolding"]
}
},
"analyzer": {
"atlas_name_index": {
"type": "custom",
"char_filter": ["html_strip"],
"tokenizer": "standard",
"filter": ["lowercase", "asciifolding"]
},
"atlas_name_search": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "asciifolding"]
}
}
}
},
"mappings": {
"properties": {
"name": {
"type": "text",
"analyzer": "atlas_name_index",
"search_analyzer": "atlas_name_search",
"fields": {"raw": {"type": "keyword", "normalizer": "atlas_keyword_lc"}}
},
"description": {"type": "text", "analyzer": "atlas_name_index"},
"sku": {"type": "keyword", "normalizer": "atlas_keyword_lc"},
"category": {"type": "keyword", "normalizer": "atlas_keyword_lc"}
}
}
}
PUT atlasmart-products-analysis-v1/_doc/p-100?refresh=true
{"name":"Café USB-C Charger","description":"65W fast charger for laptops","sku":"USB-C-65W","category":"POWER"}
PUT atlasmart-products-analysis-v1/_doc/p-101?refresh=true
{"name":"Cafe Travel Mug","description":"Insulated steel mug","sku":"MUG-450","category":"KITCHEN"}
4. The deliberate failure: term query against analyzed user text
A common bug is to send a user's phrase directly to a
term query on name. Term-level queries
do not run the field's full-text analyzer. The indexed field
contains analyzed terms such as cafe and
charger, not the original phrase as one term. Use
match when you want query text analyzed, or query
name.raw when you truly want normalized exact-value
semantics.
GET atlasmart-products-analysis-v1/_search
{"query":{"term":{"name":"Café USB-C Charger"}}}
GET atlasmart-products-analysis-v1/_search
{"query":{"match":{"name":"cafe charger"}}}
GET atlasmart-products-analysis-v1/_search
{"query":{"term":{"name.raw":"cafe usb-c charger"}}}
The first request is intentionally misleading. The second invokes the field search analyzer. The third targets the single-token normalized multi-field. This distinction becomes essential when Chapter 06 adds Query DSL composition.
5. Verification checklist and reset
-
GET atlasmart-products-analysis-v1/_mappingshows the named index and search analyzers plusname.raw. -
POST atlasmart-products-analysis-v1/_analyzewithfield: "name"exposes the field's index analyzer token stream. -
A
matchsearch for unaccentedcafecan match the accented product because folding is part of both relevant analysis paths. -
skuremains one normalized keyword token; it is not split at hyphens. - Deleting this disposable index is sufficient reset for this lesson.
DELETE atlasmart-products-analysis-v1
Check your understanding
-
Why can GET return the original
_sourcewhile search matches terms that do not visibly appear in that exact form? - What does
search_analyzerchange? -
Why is
termusually wrong for free-form user text on atextfield? - When does an index-analyzer change normally require reindexing?
- What does
_analyzeprove?
Review the answers
1. Because _source preserves
the stored JSON representation while the inverted index
stores analyzed terms used for retrieval.
2. It changes how full-text query text is analyzed; it does not retroactively rewrite terms already indexed.
3. It looks for an exact term and does not apply the field full-text analyzer to the supplied value.
4. When already-indexed documents must be represented by a different set of terms; existing Lucene terms do not change merely because settings change.
5. It proves the configured chain’s token output for the supplied input; by itself it does not prove index contents, relevance quality, or production latency.
Production judgment
Analysis is part of the search API contract. Record analyzer names/version, representative token fixtures, language assumptions, and the queries that depend on them. Measure recall, precision, index growth, and query latency on judged traffic before deploying a broader analyzer. More tokens can improve recall while increasing postings, query work, and false positives. Analyzer changes require the same change discipline as schema migrations.
Next, we isolate the tokenizer stage and compare standard, keyword, whitespace, pattern, n-gram, edge-ngram, hierarchy, and language-aware patterns on the same AtlasMart inputs.
Summary and next step
This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.
Authoritative references
- Elastic: standard analyzer — Current reference for the default standard analysis chain.
- Elastic: Analyze API — Token/position/offset inspection for built-in, field-bound, and custom analysis.
- Elastic: tokenizer reference — Official tokenizer families and behavior.
- Elastic: configure synonyms — Search-time synonym-set workflow, validation, and reload behavior.
- Elastic: reload search analyzers API — Reload semantics for file-backed, updateable search analyzers.
- OpenSearch: text analysis — Analyzer, tokenizer, token-filter, and normalizer fundamentals.
- OpenSearch: Analyze API — Index/global analysis requests and detailed token evidence.
- OpenSearch: synonym_graph token filter — Multiword synonym graph behavior.
- OpenSearch: refresh search analyzer — Product-specific analyzer refresh workflow and plugin requirement.
- OpenSearch: normalizers — Single-token keyword normalization constraints.