Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search
Lowercase, Stop, Stemmer, Synonym, Shingle, ASCII Folding, and Normalization Tradeoffs
Build and inspect token-filter chains for AtlasMart search, then measure the correctness risks introduced by stemming, stopwords, accent folding, phrase shingles, and synonym graphs.
Learning outcomes
AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.
Explain how lowercase, stop, stemmer, ASCII-folding, shingle, and synonym-graph filters transform tokens and positions.
Distinguish character normalization from semantic expansion and understand which operations can collapse meaningful distinctions.
Use synonym_graph for multiword search synonyms
and reason about token graphs instead of flat lists.
Explain why keyword normalizers cannot perform arbitrary stemming or synonym expansion.
Build judged AtlasMart fixtures that reveal recall/precision tradeoffs introduced by each filter.
Examples target Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep the
course lab contract established earlier: Elasticsearch at
https://localhost:9200 authenticated as the
disposable lab elastic user via
ELASTIC_PASSWORD and verified with the copied
HTTP CA; OpenSearch at
https://localhost:9201 using the disposable
Security-plugin demo admin identity via
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier
OpenSearch -k path is local demo convenience
only, not a production certificate-verification pattern.
Portable examples use core analysis components that both
products document. Product-specific synonym-management and
reload APIs are labeled explicitly instead of pretending the
two platforms expose one interchangeable control plane.
The generation environment does not run the two search
clusters, so expected token/output fragments are documented as
deterministic shapes derived from the configured analyzers,
not falsely presented as captured benchmark output. Run the
requests in the disposable AtlasMart lab from Chapters 01–04
with one local node per product, dedicated
atlasmart-* indices, and the pinned
server/container assumptions from those chapters. Analyzer
experiments use dedicated versioned indices only; cleanup
commands must never target unrelated data, credentials,
certificates, or volumes.
1. Token filters are ordered transformations
Filters are applied in the order listed. A later synonym rule sees the tokens produced by earlier filters; a stop filter can remove words that a phrase synonym expected; stemming before or after synonyms can change whether a rule parses or fires. Analysis configuration is therefore an executable pipeline, not an unordered bag of options.
| Filter | Typical effect | Risk to test |
|---|---|---|
| lowercase | Case folding | Case may carry meaning in identifiers |
| stop | Removes configured frequent tokens | Phrase/negation meaning can change |
| stemmer | Reduces inflected forms | Over-stemming can merge distinct concepts |
| asciifolding | Folds many accented forms | Can erase meaningful orthographic distinctions |
| shingle | Adds multi-token word n-grams | Index/query term growth |
| synonym_graph | Represents multiword equivalence/rewrites as a token graph | Phrase semantics and rule lifecycle complexity |
2. Inspect one chain a filter at a time
POST _analyze
{"tokenizer":"standard","filter":["lowercase","asciifolding"],"text":"CAFÉ Résumé"}
POST _analyze
{
"tokenizer":"standard",
"filter":["lowercase",{"type":"stop","stopwords":"_english_"},{"type":"stemmer","language":"light_english"}],
"text":"The chargers are charging laptops"
}
The point is not to prefer these settings globally. The point is to inspect the transformed vocabulary and ask whether the collapsed forms still reflect AtlasMart's product distinctions. For brand names, SKUs, chemical symbols, and multilingual catalog text, aggressive normalization may be incorrect.
3. Multiword synonyms require graph-aware reasoning
Consider users who search power brick,
wall charger, and AC adapter. A multiword
synonym can alter token positions and paths.
synonym_graph is designed for correct graph
representation in search analysis. Phrase queries later consume
that graph. Do not infer graph behavior from a flat list of
displayed token strings alone.
PUT atlasmart-filters-v1
{
"settings":{"number_of_shards":1,"number_of_replicas":0,
"analysis":{"filter":{"atlas_syn":{"type":"synonym_graph","synonyms":[
"power brick, wall charger, ac adapter",
"notebook => laptop"
]}},
"analyzer":{"atlas_search_syn":{"type":"custom","tokenizer":"standard","filter":["lowercase","atlas_syn"]}}}},
"mappings":{"properties":{"name":{"type":"text","analyzer":"standard","search_analyzer":"atlas_search_syn"}}}
}
POST atlasmart-filters-v1/_analyze
{"analyzer":"atlas_search_syn","text":"power brick"}
If a preceding stop filter removes a token required by a synonym rule, rule parsing or matching can change. Invalid synonym rules can also prevent analyzer changes from applying. Keep synonym fixtures that cover equivalent and directional rules, multiword phrases, stopword interactions, and expected phrase matches.
4. Normalizers are deliberately constrained
A normalizer is for keyword values
and emits one token. It has no tokenizer and supports only
filters compatible with single-token normalization. Lowercasing
and ASCII folding are common; stemming and synonym expansion are
not normalizer jobs. If AtlasMart wants SKU matching to ignore
case while preserving the identifier as one value, a normalizer
is appropriate. If it wants free-form semantic expansion, use a
text field/analyzer instead.
PUT atlasmart-normalizer-v1
{
"settings":{"analysis":{"normalizer":{"atlas_keyword":{"type":"custom","filter":["lowercase","asciifolding"]}}}},
"mappings":{"properties":{"sku":{"type":"keyword","normalizer":"atlas_keyword"}}}
}
5. Shingles can help phrases, but they are not free relevance
A shingle filter can emit unigrams plus two-word tokens such as
wireless charger. This can support specific
phrase-oriented designs, but it also creates additional terms.
Test whether the gain is real compared with positional phrase
queries, and account for index size and write cost. Do not add
shingles merely because multiword queries exist.
POST _analyze
{
"tokenizer":"standard",
"filter":["lowercase",{"type":"shingle","min_shingle_size":2,"max_shingle_size":2,"output_unigrams":true}],
"text":"wireless phone charger"
}
Check your understanding
- Why does filter order matter?
-
Why prefer
synonym_graphfor multiword search synonyms? - Can a keyword normalizer stem text?
- What is the principal risk of ASCII folding?
- Why are shingles not automatically better for phrases?
Review the answers
1. Each filter consumes the output of the previous stage, so removed/stemmed/expanded terms change what later filters can see and how token positions/graphs are formed.
2. It represents alternative multi-token paths in a graph so phrase/proximity queries can interpret multiword expansions correctly.
3. No. Normalizers are constrained to single-token-compatible operations and do not support arbitrary stemming or synonym expansion.
4. It can collapse distinctions that are meaningful in some names or languages, so improved recall may come with false positives.
5. They add terms and resource cost; positional phrase queries may already satisfy the requirement, so the benefit must be measured.
DELETE atlasmart-filters-v1
DELETE atlasmart-normalizer-v1
Production judgment
Every normalization choice trades distinctions for matchability. Maintain a judged query set that includes accents, casing, plural/inflected forms, negation/stopwords, multiword synonyms, brand names, SKUs, and multilingual examples. Review false positives as carefully as missed matches. Search analysis can sometimes evolve without rewriting documents; index analysis cannot be assumed to do so.
Next, we turn these components into custom analyzer/multi-field contracts and operate synonym changes without confusing Elastic and OpenSearch lifecycle APIs.
Summary and next step
This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.
Authoritative references
- Elastic: standard analyzer — Current reference for the default standard analysis chain.
- Elastic: Analyze API — Token/position/offset inspection for built-in, field-bound, and custom analysis.
- Elastic: tokenizer reference — Official tokenizer families and behavior.
- Elastic: configure synonyms — Search-time synonym-set workflow, validation, and reload behavior.
- Elastic: reload search analyzers API — Reload semantics for file-backed, updateable search analyzers.
- OpenSearch: text analysis — Analyzer, tokenizer, token-filter, and normalizer fundamentals.
- OpenSearch: Analyze API — Index/global analysis requests and detailed token evidence.
- OpenSearch: synonym_graph token filter — Multiword synonym graph behavior.
- OpenSearch: refresh search analyzer — Product-specific analyzer refresh workflow and plugin requirement.
- OpenSearch: normalizers — Single-token keyword normalization constraints.