Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search
Standard, Keyword, Whitespace, Pattern, N-Gram/Edge-N-Gram, Path, and Language Tokenization Patterns
Compare tokenizer families on the same AtlasMart strings so separator rules, partial-word expansion, hierarchy tokens, and language-aware analysis become observable design choices rather than names in a reference table.
Learning outcomes
AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.
Predict the boundaries emitted by standard, keyword, whitespace, pattern, n-gram, edge-ngram, and path-hierarchy tokenizers.
Choose an index-time partial-word strategy without applying the same expansion indiscriminately to query text.
Explain why a tokenizer is not a language analyzer and when language-specific analysis adds stemming or normalization.
Use token counts/positions to reason about index growth and phrase behavior.
Build a repeatable tokenizer comparison harness with AtlasMart names, identifiers, and category paths.
Examples target Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep the
course lab contract established earlier: Elasticsearch at
https://localhost:9200 authenticated as the
disposable lab elastic user via
ELASTIC_PASSWORD and verified with the copied
HTTP CA; OpenSearch at
https://localhost:9201 using the disposable
Security-plugin demo admin identity via
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier
OpenSearch -k path is local demo convenience
only, not a production certificate-verification pattern.
Portable examples use core analysis components that both
products document. Product-specific synonym-management and
reload APIs are labeled explicitly instead of pretending the
two platforms expose one interchangeable control plane.
The generation environment does not run the two search
clusters, so expected token/output fragments are documented as
deterministic shapes derived from the configured analyzers,
not falsely presented as captured benchmark output. Run the
requests in the disposable AtlasMart lab from Chapters 01–04
with one local node per product, dedicated
atlasmart-* indices, and the pinned
server/container assumptions from those chapters. Analyzer
experiments use dedicated versioned indices only; cleanup
commands must never target unrelated data, credentials,
certificates, or volumes.
1. Tokenizers answer one question: where are token boundaries?
The tokenizer receives the character-filter output and turns it into a token stream. The same characters can become one token, many word tokens, overlapping substrings, anchored prefixes, or hierarchy paths depending on the tokenizer. Choosing a tokenizer is therefore choosing the primitive terms available to later token filters and the inverted index.
| Pattern | Typical use | Main risk |
|---|---|---|
standard |
General Unicode word segmentation | Identifiers may split at punctuation |
keyword |
Treat entire input as one token before filters | Not suitable for ordinary word retrieval |
whitespace |
Only whitespace boundaries | Punctuation stays inside tokens |
pattern |
Domain-specific delimiters | A bad regex can create surprising boundaries/cost |
ngram |
Substring/partial matching | Large term expansion and false positives |
edge_ngram |
Prefix autocomplete at index time | Applying it at query time expands user input unnecessarily |
path_hierarchy |
Category/file-style ancestor paths | Wrong delimiter/direction breaks hierarchy semantics |
2. Run the same text through several tokenizer contracts
POST _analyze
{"tokenizer":"standard","text":"USB-C Café Charger/65W"}
POST _analyze
{"tokenizer":"keyword","text":"USB-C Café Charger/65W"}
POST _analyze
{"tokenizer":"whitespace","text":"USB-C Café Charger/65W"}
POST _analyze
{"tokenizer":{"type":"pattern","pattern":"[-_/]+"},"text":"USB-C/65W_POWER"}
Do not memorize only token text. Inspect positions and offsets:
phrase queries and graph-producing filters build on positional
structure. Also note that keyword tokenizer is not
the same as a keyword field. The tokenizer is an
analysis component; the field type has mapping, doc-values,
exact-match, sort, and aggregation semantics.
3. Partial-word retrieval: n-gram vs edge n-gram
An n-gram tokenizer emits sliding substrings; edge n-gram emits prefixes anchored at a token edge. Both can multiply indexed terms. For AtlasMart autocomplete, a controlled edge-ngram index analyzer with a normal search analyzer is usually easier to reason about than applying edge-ngram to every query token. The correct gram range depends on product vocabulary, result quality, and measured resource cost—not folklore.
PUT atlasmart-tokenizers-v1
{
"settings": {"number_of_shards":1,"number_of_replicas":0,
"analysis": {
"tokenizer": {"atlas_edge":{"type":"edge_ngram","min_gram":2,"max_gram":12,"token_chars":["letter","digit"]}},
"analyzer": {
"atlas_autocomplete_index":{"type":"custom","tokenizer":"atlas_edge","filter":["lowercase","asciifolding"]},
"atlas_autocomplete_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","asciifolding"]}
}
}
},
"mappings":{"properties":{"name":{"type":"text","analyzer":"atlas_autocomplete_index","search_analyzer":"atlas_autocomplete_search"}}}
}
POST atlasmart-tokenizers-v1/_analyze
{"analyzer":"atlas_autocomplete_index","text":"Charger"}
POST atlasmart-tokenizers-v1/_analyze
{"analyzer":"atlas_autocomplete_search","text":"Char"}
If the query analyzer emits every prefix of
charger, one user token becomes many query terms.
That changes scoring, clause counts, and resource use. Keep
the expansion where the design intends it and test it with
realistic queries.
4. Paths and language-aware analysis solve different problems
POST _analyze
{"tokenizer":{"type":"path_hierarchy","delimiter":"/"},"text":"electronics/power/chargers"}
A hierarchy tokenizer can emit ancestor-like path terms such as
electronics, electronics/power, and
the full path. A language analyzer, by
contrast, is a packaged analyzer that may combine standard-like
tokenization with language-specific lowercasing, stopwords,
normalization, and stemming. Language behavior must be evaluated
with real catalog/query pairs; stemming can merge forms that
users consider different, and multilingual fields may need
separate language fields or routing rather than one aggressive
analyzer.
POST _analyze
{"analyzer":"standard","text":"Running chargers for laptops"}
POST _analyze
{"analyzer":"english","text":"Running chargers for laptops"}
5. AtlasMart tokenizer regression fixture
Create a small table of inputs and expected token invariants before changing production analyzers: punctuation-heavy SKU, accented brand, hyphenated product phrase, short autocomplete prefix, category path, and at least one language-specific sentence. Store expected tokens and positions in source control and compare both products where portability matters.
Check your understanding
-
Why is the
keywordtokenizer different from akeywordmapping? - Why are edge n-grams commonly asymmetric between index and search analysis?
-
What does
path_hierarchypreserve thatstandarddoes not? - Is a language analyzer merely a tokenizer?
- What should determine gram sizes?
Review the answers
1. The tokenizer is one analysis stage that emits the whole input as one token; the mapping defines exact-value field semantics including indexing/doc-values behavior.
2. Index-time prefixes support prefix lookup, while a normal search analyzer avoids multiplying every user query token into many prefixes.
3. It preserves cumulative path ancestry based on a delimiter, enabling hierarchical category/path matching.
4. No. It is an analyzer chain that can include language-specific normalization, stopwords, stemming and other filters around tokenization.
5. Measured vocabulary/query behavior, relevance requirements, index size, latency, and resource constraints—not a universal constant.
DELETE atlasmart-tokenizers-v1
Production judgment
Tokenizer changes alter the vocabulary of the index and therefore normally require reindexing. Before adopting partial-word or language-specific tokenization, quantify term growth, recall/precision changes, phrase behavior, and p95/p99 search latency under representative load. Keep a simple field for ordinary full-text search when a specialized autocomplete or hierarchy field serves only one interaction.
Next, we keep token boundaries fixed and study filters that normalize, remove, stem, expand, or combine those tokens.
Summary and next step
This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.
Authoritative references
- Elastic: standard analyzer — Current reference for the default standard analysis chain.
- Elastic: Analyze API — Token/position/offset inspection for built-in, field-bound, and custom analysis.
- Elastic: tokenizer reference — Official tokenizer families and behavior.
- Elastic: configure synonyms — Search-time synonym-set workflow, validation, and reload behavior.
- Elastic: reload search analyzers API — Reload semantics for file-backed, updateable search analyzers.
- OpenSearch: text analysis — Analyzer, tokenizer, token-filter, and normalizer fundamentals.
- OpenSearch: Analyze API — Index/global analysis requests and detailed token evidence.
- OpenSearch: synonym_graph token filter — Multiword synonym graph behavior.
- OpenSearch: refresh search analyzer — Product-specific analyzer refresh workflow and plugin requirement.
- OpenSearch: normalizers — Single-token keyword normalization constraints.