Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search

Custom Analyzers, Multi-Fields, Search Analyzers, Synonym Lifecycle, and Reindexing Implications

Turn AtlasMart analysis into an operable schema contract: separate index/search analyzers, preserve exact-value multi-fields, manage synonyms with product-specific workflows, and plan versioned reindex migrations.

Intermediate → Advanced130–150 minutesAnalyzer evolution labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.

01

Compose custom index/search analyzers and preserve exact-value behavior with multi-fields.

02

Explain why analysis settings are index-level schema and why many analyzer changes are not in-place mutations of existing terms.

03

Separate Elastic synonym-set APIs from OpenSearch file/plugin refresh workflows.

04

Plan a versioned index migration when index-time analysis changes.

05

Define rollback and relevance-regression gates for analyzer evolution.

Chapter baseline reviewed 11 September 2026

Examples target Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course lab contract established earlier: Elasticsearch at https://localhost:9200 authenticated as the disposable lab elastic user via ELASTIC_PASSWORD and verified with the copied HTTP CA; OpenSearch at https://localhost:9201 using the disposable Security-plugin demo admin identity via OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier OpenSearch -k path is local demo convenience only, not a production certificate-verification pattern. Portable examples use core analysis components that both products document. Product-specific synonym-management and reload APIs are labeled explicitly instead of pretending the two platforms expose one interchangeable control plane. Elastic synonym sets managed through the Synonyms APIs are not the same control plane as OpenSearch synonym files and the ISM-plugin refresh-search-analyzer API.

Execution and safety note

The generation environment does not run the two search clusters, so expected token/output fragments are documented as deterministic shapes derived from the configured analyzers, not falsely presented as captured benchmark output. Run the requests in the disposable AtlasMart lab from Chapters 01–04 with one local node per product, dedicated atlasmart-* indices, and the pinned server/container assumptions from those chapters. Analyzer experiments use dedicated versioned indices only; cleanup commands must never target unrelated data, credentials, certificates, or volumes.

1. One logical value often needs multiple search representations

AtlasMart's product name must support ordinary full-text retrieval, autocomplete, and exact sorting/faceting. A multi-field lets the same source value be indexed into separate field representations, each with its own analysis/mapping contract. That is different from copying three values into _source: one JSON value can feed several index structures.

versioned v2 analysis contract
PUT atlasmart-products-analysis-v2
{
  "settings":{"number_of_shards":1,"number_of_replicas":0,
    "analysis":{
      "tokenizer":{"atlas_edge":{"type":"edge_ngram","min_gram":2,"max_gram":12,"token_chars":["letter","digit"]}},
      "normalizer":{"atlas_keyword":{"type":"custom","filter":["lowercase","asciifolding"]}},
      "analyzer":{
        "atlas_name_index":{"type":"custom","tokenizer":"standard","filter":["lowercase","asciifolding"]},
        "atlas_name_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","asciifolding"]},
        "atlas_autocomplete_index":{"type":"custom","tokenizer":"atlas_edge","filter":["lowercase","asciifolding"]}
      }}},
  "mappings":{"properties":{"name":{"type":"text","analyzer":"atlas_name_index","search_analyzer":"atlas_name_search","fields":{
    "raw":{"type":"keyword","normalizer":"atlas_keyword"},
    "autocomplete":{"type":"text","analyzer":"atlas_autocomplete_index","search_analyzer":"atlas_name_search"}
  }}}}
}

2. Analyzer evolution: what can change without reindexing?

Existing indexed terms do not get re-analyzed when you edit a configuration. If the index-time analyzer changes, old and new documents would otherwise represent text under different term contracts, so the safe course is normally a new versioned index plus reindex/dual-write/cutover strategy. Search-time resources are different: some can be reloaded without rewriting stored terms, but the supported mechanism is product and configuration specific.

Change Typical migration implication
Change index tokenizer/filter New versioned index and reindex to rebuild terms
Add compatible new multi-field to future documents Old documents do not magically gain terms; backfill/reindex may still be required
Change search-only synonym resource May be reloadable using supported product-specific workflow
Change mapping field type Normally requires new index/reindex rather than in-place incompatible mutation

3. Elastic synonym lifecycle: synonym sets are a managed API resource

Elasticsearch 9.5 supports managed synonym sets through /_synonyms. Search analyzers can reference a synonyms_set using an updateable synonym_graph filter. Updating the set reloads associated analyzers. The synonym set must exist before an index references it, and invalid rules must be treated as deployment failures rather than silently accepted content.

Elastic only · managed synonym set
PUT /_synonyms/atlasmart-products
{
  "synonyms_set": [
    {"id":"charger","synonyms":"power brick, wall charger, ac adapter"},
    {"id":"notebook","synonyms":"notebook => laptop"}
  ]
}
Elastic only · search analyzer references synonym set
PUT atlasmart-es-syn-v1
{
  "settings":{"analysis":{"filter":{"atlas_syn":{"type":"synonym_graph","synonyms_set":"atlasmart-products","updateable":true}},
    "analyzer":{"atlas_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","atlas_syn"]}}}},
  "mappings":{"properties":{"name":{"type":"text","analyzer":"standard","search_analyzer":"atlas_search"}}}
}

Do not copy this request into OpenSearch and call it portable. It describes Elastic's current managed synonym-set contract.

4. OpenSearch synonym lifecycle: file-backed resources and refresh API are different

OpenSearch documents synonym_graph with inline or file-backed rules. Its real-time refresh-search-analyzer API is under /_plugins/_refresh_search_analyzers/... and requires the Index State Management plugin; the relevant token filter must be updateable: true. File-backed resources must be present where the cluster expects them. That operational model is not equivalent to Elastic's managed synonym-set API.

OpenSearch only · file-backed updateable search synonyms
PUT atlasmart-os-syn-v1
{
  "settings":{"analysis":{"filter":{"atlas_syn":{"type":"synonym_graph","synonyms_path":"synonyms.txt","updateable":true}},
    "analyzer":{"atlas_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","atlas_syn"]}}}},
  "mappings":{"properties":{"name":{"type":"text","analyzer":"standard","search_analyzer":"atlas_search"}}}
}

POST /_plugins/_refresh_search_analyzers/atlasmart-os-syn-v1
Boundary:

The chapter's mandatory learning objective does not require modifying real node config files. If your lab image does not mount a controlled synonym file or does not include the required plugin/API path, use inline synonym fixtures to learn token-graph semantics and treat file refresh as an explicitly optional operational exercise.

5. Versioned reindex migration for an index-time analyzer change

Suppose v1 folded accents but the business now needs a language-preserving field plus a folded search subfield. Do not silently change the analyzer name on the old index. Create v2, replay/reindex the source documents, run token and relevance regression tests, compare counts and failure logs, then move a stable read/write alias or application configuration with a rollback window.

migration skeleton
POST _reindex
{
  "source": {"index":"atlasmart-products-analysis-v1"},
  "dest":   {"index":"atlasmart-products-analysis-v2","op_type":"create"}
}

GET atlasmart-products-analysis-v2/_count
GET atlasmart-products-analysis-v2/_mapping
POST atlasmart-products-analysis-v2/_analyze
{"field":"name","text":"Café USB-C Charger"}

Check your understanding

  1. Why can a search-analyzer synonym update avoid reindexing?
  2. Why does an index-analyzer change usually require a new index?
  3. Are Elastic synonym-set APIs and OpenSearch refresh-search-analyzer APIs interchangeable?
  4. Why keep a name.raw multi-field?
  5. What is the rollback unit for a risky analyzer migration?
Review the answers

1. Because it changes query-time expansion while leaving the already-indexed term vocabulary untouched, when the product/configuration supports reloadable search analysis.

2. Existing documents retain their old Lucene terms; a new index lets every document be rebuilt consistently under the new analysis contract.

3. No. They are separate product-specific control planes with different resources, endpoints, and operational requirements.

4. It preserves single-token exact-value semantics for sorting, aggregations and exact filters while name remains full-text.

5. A versioned index plus controlled alias/application cutover, with the old index retained until acceptance gates pass.

cleanup only lesson-specific experimental indices
DELETE atlasmart-products-analysis-v2
DELETE atlasmart-es-syn-v1
DELETE atlasmart-os-syn-v1

Production judgment

Analyzer evolution is both a relevance change and a data migration. Gate it with token fixtures, judged queries, count/error reconciliation, latency/resource tests, and a reversible cutover. Synonym administration also has security consequences: the privilege to change search semantics should be limited, audited, and separated from arbitrary application writers.

Next, we make analysis contracts executable as regression tests and use Analyze/term-vector evidence to diagnose unexpected matches.

Summary and next step

This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.