Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search

Index-Time vs Search-Time Analysis and Why Token Streams Define Search Semantics

Make the token stream visible before writing queries: trace how AtlasMart product text is transformed at index time and search time, then prove why exact-term and full-text queries behave differently.

Intermediate → Advanced110–130 minutesAnalyzer contract labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.

01

Explain character filters, one tokenizer, and ordered token filters as the three stages of a custom analyzer.

02

Separate index-time analysis from search-time analysis and predict how analyzer mismatch changes recall and precision.

03

Use _analyze to inspect token text, positions, offsets, and types instead of guessing from the original string.

04

Distinguish a full-text match query from a term-level term query on an analyzed field.

05

Create a portable AtlasMart v1 analysis contract for product names, descriptions, SKUs, and categories.

Chapter baseline reviewed 11 September 2026

Examples target Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course lab contract established earlier: Elasticsearch at https://localhost:9200 authenticated as the disposable lab elastic user via ELASTIC_PASSWORD and verified with the copied HTTP CA; OpenSearch at https://localhost:9201 using the disposable Security-plugin demo admin identity via OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier OpenSearch -k path is local demo convenience only, not a production certificate-verification pattern. Portable examples use core analysis components that both products document. Product-specific synonym-management and reload APIs are labeled explicitly instead of pretending the two platforms expose one interchangeable control plane.

Execution and safety note

The generation environment does not run the two search clusters, so expected token/output fragments are documented as deterministic shapes derived from the configured analyzers, not falsely presented as captured benchmark output. Run the requests in the disposable AtlasMart lab from Chapters 01–04 with one local node per product, dedicated atlasmart-* indices, and the pinned server/container assumptions from those chapters. Analyzer experiments use dedicated versioned indices only; cleanup commands must never target unrelated data, credentials, certificates, or volumes.

1. The indexed unit is the token stream, not the sentence

A character filter transforms the input character stream before tokenization—for example stripping HTML or applying a configured character mapping. A tokenizer then emits token boundaries. An analyzer has exactly one tokenizer. Ordered token filters may lowercase, stem, remove, expand, or otherwise transform those tokens. Each emitted token has text plus metadata such as a position and character offsets. Lucene builds its inverted index from those analyzed terms.

That model immediately explains why copying the visible product title into a query is insufficient to predict matching. The index analyzer determines what terms exist in the index; the search analyzer determines which terms the user query asks for. Search succeeds only when those two contracts are compatible enough for the query semantics.

Dev Tools · inspect a portable custom chain
POST _analyze
{
  "char_filter": ["html_strip"],
  "tokenizer": "standard",
  "filter": ["lowercase", "asciifolding"],
  "text": "<b>Café USB-C Charger</b>"
}

For this fixture, expect terms equivalent to cafe, usb, c, and charger, with positions increasing through the stream. The exact response also reports offsets/types. That evidence proves what this analysis request emitted; it does not prove how an existing field is mapped unless you invoke the field's configured analyzer.

2. Index-time and search-time analysis are separate contracts

For a text field, the field's analyzer is used when text is indexed. By default the same analyzer is also used for full-text search, but a field may declare a separate search_analyzer. That separation is useful when indexing needs extra terms—autocomplete prefixes are the classic example—while queries should remain compact. It is also where synonym operations often belong because changing search-time expansion can avoid rewriting every indexed document.

Stage Input Output that matters Operational consequence
Index analysis Document field value Terms written to the inverted index Changing it normally requires reindexing old documents
Search analysis User/query text Terms/graphs used to construct the query Some search-only resources can be updated/reloaded product-specifically
Normalizer keyword value Exactly one normalized token No tokenizer; only compatible single-token filters

3. Build AtlasMart's first explicit analysis contract

Product names need full-text search plus an exact sortable/facetable representation. Descriptions need full-text search but not an exact-value aggregation. SKUs and category codes are identifiers: they should not be split into words. This is query-driven schema design from Chapter 03 carried into analysis.

portable index · AtlasMart analysis v1
PUT atlasmart-products-analysis-v1
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0,
    "analysis": {
      "normalizer": {
        "atlas_keyword_lc": {
          "type": "custom",
          "filter": ["lowercase", "asciifolding"]
        }
      },
      "analyzer": {
        "atlas_name_index": {
          "type": "custom",
          "char_filter": ["html_strip"],
          "tokenizer": "standard",
          "filter": ["lowercase", "asciifolding"]
        },
        "atlas_name_search": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": ["lowercase", "asciifolding"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "name": {
        "type": "text",
        "analyzer": "atlas_name_index",
        "search_analyzer": "atlas_name_search",
        "fields": {"raw": {"type": "keyword", "normalizer": "atlas_keyword_lc"}}
      },
      "description": {"type": "text", "analyzer": "atlas_name_index"},
      "sku": {"type": "keyword", "normalizer": "atlas_keyword_lc"},
      "category": {"type": "keyword", "normalizer": "atlas_keyword_lc"}
    }
  }
}
index deterministic fixtures
PUT atlasmart-products-analysis-v1/_doc/p-100?refresh=true
{"name":"Café USB-C Charger","description":"65W fast charger for laptops","sku":"USB-C-65W","category":"POWER"}

PUT atlasmart-products-analysis-v1/_doc/p-101?refresh=true
{"name":"Cafe Travel Mug","description":"Insulated steel mug","sku":"MUG-450","category":"KITCHEN"}

4. The deliberate failure: term query against analyzed user text

A common bug is to send a user's phrase directly to a term query on name. Term-level queries do not run the field's full-text analyzer. The indexed field contains analyzed terms such as cafe and charger, not the original phrase as one term. Use match when you want query text analyzed, or query name.raw when you truly want normalized exact-value semantics.

wrong then corrected queries
GET atlasmart-products-analysis-v1/_search
{"query":{"term":{"name":"Café USB-C Charger"}}}

GET atlasmart-products-analysis-v1/_search
{"query":{"match":{"name":"cafe charger"}}}

GET atlasmart-products-analysis-v1/_search
{"query":{"term":{"name.raw":"cafe usb-c charger"}}}

The first request is intentionally misleading. The second invokes the field search analyzer. The third targets the single-token normalized multi-field. This distinction becomes essential when Chapter 06 adds Query DSL composition.

5. Verification checklist and reset

  • GET atlasmart-products-analysis-v1/_mapping shows the named index and search analyzers plus name.raw.
  • POST atlasmart-products-analysis-v1/_analyze with field: "name" exposes the field's index analyzer token stream.
  • A match search for unaccented cafe can match the accented product because folding is part of both relevant analysis paths.
  • sku remains one normalized keyword token; it is not split at hyphens.
  • Deleting this disposable index is sufficient reset for this lesson.
cleanup
DELETE atlasmart-products-analysis-v1

Check your understanding

  1. Why can GET return the original _source while search matches terms that do not visibly appear in that exact form?
  2. What does search_analyzer change?
  3. Why is term usually wrong for free-form user text on a text field?
  4. When does an index-analyzer change normally require reindexing?
  5. What does _analyze prove?
Review the answers

1. Because _source preserves the stored JSON representation while the inverted index stores analyzed terms used for retrieval.

2. It changes how full-text query text is analyzed; it does not retroactively rewrite terms already indexed.

3. It looks for an exact term and does not apply the field full-text analyzer to the supplied value.

4. When already-indexed documents must be represented by a different set of terms; existing Lucene terms do not change merely because settings change.

5. It proves the configured chain’s token output for the supplied input; by itself it does not prove index contents, relevance quality, or production latency.

Production judgment

Analysis is part of the search API contract. Record analyzer names/version, representative token fixtures, language assumptions, and the queries that depend on them. Measure recall, precision, index growth, and query latency on judged traffic before deploying a broader analyzer. More tokens can improve recall while increasing postings, query work, and false positives. Analyzer changes require the same change discipline as schema migrations.

Next, we isolate the tokenizer stage and compare standard, keyword, whitespace, pattern, n-gram, edge-ngram, hierarchy, and language-aware patterns on the same AtlasMart inputs.

Summary and next step

This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.