Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search
Custom Analyzers, Multi-Fields, Search Analyzers, Synonym Lifecycle, and Reindexing Implications
Turn AtlasMart analysis into an operable schema contract: separate index/search analyzers, preserve exact-value multi-fields, manage synonyms with product-specific workflows, and plan versioned reindex migrations.
Learning outcomes
AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.
Compose custom index/search analyzers and preserve exact-value behavior with multi-fields.
Explain why analysis settings are index-level schema and why many analyzer changes are not in-place mutations of existing terms.
Separate Elastic synonym-set APIs from OpenSearch file/plugin refresh workflows.
Plan a versioned index migration when index-time analysis changes.
Define rollback and relevance-regression gates for analyzer evolution.
Examples target Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep the
course lab contract established earlier: Elasticsearch at
https://localhost:9200 authenticated as the
disposable lab elastic user via
ELASTIC_PASSWORD and verified with the copied
HTTP CA; OpenSearch at
https://localhost:9201 using the disposable
Security-plugin demo admin identity via
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier
OpenSearch -k path is local demo convenience
only, not a production certificate-verification pattern.
Portable examples use core analysis components that both
products document. Product-specific synonym-management and
reload APIs are labeled explicitly instead of pretending the
two platforms expose one interchangeable control plane.
Elastic synonym sets managed through the Synonyms APIs are not
the same control plane as OpenSearch synonym files and the
ISM-plugin refresh-search-analyzer API.
The generation environment does not run the two search
clusters, so expected token/output fragments are documented as
deterministic shapes derived from the configured analyzers,
not falsely presented as captured benchmark output. Run the
requests in the disposable AtlasMart lab from Chapters 01–04
with one local node per product, dedicated
atlasmart-* indices, and the pinned
server/container assumptions from those chapters. Analyzer
experiments use dedicated versioned indices only; cleanup
commands must never target unrelated data, credentials,
certificates, or volumes.
1. One logical value often needs multiple search representations
AtlasMart's product name must support ordinary full-text
retrieval, autocomplete, and exact sorting/faceting. A
multi-field lets the same source value be indexed into separate
field representations, each with its own analysis/mapping
contract. That is different from copying three values into
_source: one JSON value can feed several index
structures.
PUT atlasmart-products-analysis-v2
{
"settings":{"number_of_shards":1,"number_of_replicas":0,
"analysis":{
"tokenizer":{"atlas_edge":{"type":"edge_ngram","min_gram":2,"max_gram":12,"token_chars":["letter","digit"]}},
"normalizer":{"atlas_keyword":{"type":"custom","filter":["lowercase","asciifolding"]}},
"analyzer":{
"atlas_name_index":{"type":"custom","tokenizer":"standard","filter":["lowercase","asciifolding"]},
"atlas_name_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","asciifolding"]},
"atlas_autocomplete_index":{"type":"custom","tokenizer":"atlas_edge","filter":["lowercase","asciifolding"]}
}}},
"mappings":{"properties":{"name":{"type":"text","analyzer":"atlas_name_index","search_analyzer":"atlas_name_search","fields":{
"raw":{"type":"keyword","normalizer":"atlas_keyword"},
"autocomplete":{"type":"text","analyzer":"atlas_autocomplete_index","search_analyzer":"atlas_name_search"}
}}}}
}
2. Analyzer evolution: what can change without reindexing?
Existing indexed terms do not get re-analyzed when you edit a configuration. If the index-time analyzer changes, old and new documents would otherwise represent text under different term contracts, so the safe course is normally a new versioned index plus reindex/dual-write/cutover strategy. Search-time resources are different: some can be reloaded without rewriting stored terms, but the supported mechanism is product and configuration specific.
| Change | Typical migration implication |
|---|---|
| Change index tokenizer/filter | New versioned index and reindex to rebuild terms |
| Add compatible new multi-field to future documents | Old documents do not magically gain terms; backfill/reindex may still be required |
| Change search-only synonym resource | May be reloadable using supported product-specific workflow |
| Change mapping field type | Normally requires new index/reindex rather than in-place incompatible mutation |
3. Elastic synonym lifecycle: synonym sets are a managed API resource
Elasticsearch 9.5 supports managed synonym sets through
/_synonyms. Search analyzers can reference a
synonyms_set using an updateable
synonym_graph filter. Updating the set reloads
associated analyzers. The synonym set must exist before an index
references it, and invalid rules must be treated as deployment
failures rather than silently accepted content.
PUT /_synonyms/atlasmart-products
{
"synonyms_set": [
{"id":"charger","synonyms":"power brick, wall charger, ac adapter"},
{"id":"notebook","synonyms":"notebook => laptop"}
]
}
PUT atlasmart-es-syn-v1
{
"settings":{"analysis":{"filter":{"atlas_syn":{"type":"synonym_graph","synonyms_set":"atlasmart-products","updateable":true}},
"analyzer":{"atlas_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","atlas_syn"]}}}},
"mappings":{"properties":{"name":{"type":"text","analyzer":"standard","search_analyzer":"atlas_search"}}}
}
Do not copy this request into OpenSearch and call it portable. It describes Elastic's current managed synonym-set contract.
4. OpenSearch synonym lifecycle: file-backed resources and refresh API are different
OpenSearch documents synonym_graph with inline or
file-backed rules. Its real-time refresh-search-analyzer API is
under /_plugins/_refresh_search_analyzers/... and
requires the Index State Management plugin; the relevant token
filter must be updateable: true. File-backed
resources must be present where the cluster expects them. That
operational model is not equivalent to Elastic's managed
synonym-set API.
PUT atlasmart-os-syn-v1
{
"settings":{"analysis":{"filter":{"atlas_syn":{"type":"synonym_graph","synonyms_path":"synonyms.txt","updateable":true}},
"analyzer":{"atlas_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","atlas_syn"]}}}},
"mappings":{"properties":{"name":{"type":"text","analyzer":"standard","search_analyzer":"atlas_search"}}}
}
POST /_plugins/_refresh_search_analyzers/atlasmart-os-syn-v1
The chapter's mandatory learning objective does not require modifying real node config files. If your lab image does not mount a controlled synonym file or does not include the required plugin/API path, use inline synonym fixtures to learn token-graph semantics and treat file refresh as an explicitly optional operational exercise.
5. Versioned reindex migration for an index-time analyzer change
Suppose v1 folded accents but the business now needs a language-preserving field plus a folded search subfield. Do not silently change the analyzer name on the old index. Create v2, replay/reindex the source documents, run token and relevance regression tests, compare counts and failure logs, then move a stable read/write alias or application configuration with a rollback window.
POST _reindex
{
"source": {"index":"atlasmart-products-analysis-v1"},
"dest": {"index":"atlasmart-products-analysis-v2","op_type":"create"}
}
GET atlasmart-products-analysis-v2/_count
GET atlasmart-products-analysis-v2/_mapping
POST atlasmart-products-analysis-v2/_analyze
{"field":"name","text":"Café USB-C Charger"}
Check your understanding
- Why can a search-analyzer synonym update avoid reindexing?
- Why does an index-analyzer change usually require a new index?
- Are Elastic synonym-set APIs and OpenSearch refresh-search-analyzer APIs interchangeable?
- Why keep a
name.rawmulti-field? - What is the rollback unit for a risky analyzer migration?
Review the answers
1. Because it changes query-time expansion while leaving the already-indexed term vocabulary untouched, when the product/configuration supports reloadable search analysis.
2. Existing documents retain their old Lucene terms; a new index lets every document be rebuilt consistently under the new analysis contract.
3. No. They are separate product-specific control planes with different resources, endpoints, and operational requirements.
4. It preserves single-token exact-value
semantics for sorting, aggregations and exact filters
while name remains full-text.
5. A versioned index plus controlled alias/application cutover, with the old index retained until acceptance gates pass.
DELETE atlasmart-products-analysis-v2
DELETE atlasmart-es-syn-v1
DELETE atlasmart-os-syn-v1
Production judgment
Analyzer evolution is both a relevance change and a data migration. Gate it with token fixtures, judged queries, count/error reconciliation, latency/resource tests, and a reversible cutover. Synonym administration also has security consequences: the privilege to change search semantics should be limited, audited, and separated from arbitrary application writers.
Next, we make analysis contracts executable as regression tests and use Analyze/term-vector evidence to diagnose unexpected matches.
Summary and next step
This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.
Authoritative references
- Elastic: standard analyzer — Current reference for the default standard analysis chain.
- Elastic: Analyze API — Token/position/offset inspection for built-in, field-bound, and custom analysis.
- Elastic: tokenizer reference — Official tokenizer families and behavior.
- Elastic: configure synonyms — Search-time synonym-set workflow, validation, and reload behavior.
- Elastic: reload search analyzers API — Reload semantics for file-backed, updateable search analyzers.
- OpenSearch: text analysis — Analyzer, tokenizer, token-filter, and normalizer fundamentals.
- OpenSearch: Analyze API — Index/global analysis requests and detailed token evidence.
- OpenSearch: synonym_graph token filter — Multiword synonym graph behavior.
- OpenSearch: refresh search analyzer — Product-specific analyzer refresh workflow and plugin requirement.
- OpenSearch: normalizers — Single-token keyword normalization constraints.