Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search

Standard, Keyword, Whitespace, Pattern, N-Gram/Edge-N-Gram, Path, and Language Tokenization Patterns

Compare tokenizer families on the same AtlasMart strings so separator rules, partial-word expansion, hierarchy tokens, and language-aware analysis become observable design choices rather than names in a reference table.

Intermediate → Advanced120–140 minutesTokenizer comparison labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.

01

Predict the boundaries emitted by standard, keyword, whitespace, pattern, n-gram, edge-ngram, and path-hierarchy tokenizers.

02

Choose an index-time partial-word strategy without applying the same expansion indiscriminately to query text.

03

Explain why a tokenizer is not a language analyzer and when language-specific analysis adds stemming or normalization.

04

Use token counts/positions to reason about index growth and phrase behavior.

05

Build a repeatable tokenizer comparison harness with AtlasMart names, identifiers, and category paths.

Chapter baseline reviewed 11 September 2026

Examples target Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course lab contract established earlier: Elasticsearch at https://localhost:9200 authenticated as the disposable lab elastic user via ELASTIC_PASSWORD and verified with the copied HTTP CA; OpenSearch at https://localhost:9201 using the disposable Security-plugin demo admin identity via OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier OpenSearch -k path is local demo convenience only, not a production certificate-verification pattern. Portable examples use core analysis components that both products document. Product-specific synonym-management and reload APIs are labeled explicitly instead of pretending the two platforms expose one interchangeable control plane.

Execution and safety note

The generation environment does not run the two search clusters, so expected token/output fragments are documented as deterministic shapes derived from the configured analyzers, not falsely presented as captured benchmark output. Run the requests in the disposable AtlasMart lab from Chapters 01–04 with one local node per product, dedicated atlasmart-* indices, and the pinned server/container assumptions from those chapters. Analyzer experiments use dedicated versioned indices only; cleanup commands must never target unrelated data, credentials, certificates, or volumes.

1. Tokenizers answer one question: where are token boundaries?

The tokenizer receives the character-filter output and turns it into a token stream. The same characters can become one token, many word tokens, overlapping substrings, anchored prefixes, or hierarchy paths depending on the tokenizer. Choosing a tokenizer is therefore choosing the primitive terms available to later token filters and the inverted index.

Pattern Typical use Main risk
standard General Unicode word segmentation Identifiers may split at punctuation
keyword Treat entire input as one token before filters Not suitable for ordinary word retrieval
whitespace Only whitespace boundaries Punctuation stays inside tokens
pattern Domain-specific delimiters A bad regex can create surprising boundaries/cost
ngram Substring/partial matching Large term expansion and false positives
edge_ngram Prefix autocomplete at index time Applying it at query time expands user input unnecessarily
path_hierarchy Category/file-style ancestor paths Wrong delimiter/direction breaks hierarchy semantics

2. Run the same text through several tokenizer contracts

portable tokenizer probes
POST _analyze
{"tokenizer":"standard","text":"USB-C Café Charger/65W"}

POST _analyze
{"tokenizer":"keyword","text":"USB-C Café Charger/65W"}

POST _analyze
{"tokenizer":"whitespace","text":"USB-C Café Charger/65W"}

POST _analyze
{"tokenizer":{"type":"pattern","pattern":"[-_/]+"},"text":"USB-C/65W_POWER"}

Do not memorize only token text. Inspect positions and offsets: phrase queries and graph-producing filters build on positional structure. Also note that keyword tokenizer is not the same as a keyword field. The tokenizer is an analysis component; the field type has mapping, doc-values, exact-match, sort, and aggregation semantics.

3. Partial-word retrieval: n-gram vs edge n-gram

An n-gram tokenizer emits sliding substrings; edge n-gram emits prefixes anchored at a token edge. Both can multiply indexed terms. For AtlasMart autocomplete, a controlled edge-ngram index analyzer with a normal search analyzer is usually easier to reason about than applying edge-ngram to every query token. The correct gram range depends on product vocabulary, result quality, and measured resource cost—not folklore.

create a dedicated autocomplete field
PUT atlasmart-tokenizers-v1
{
  "settings": {"number_of_shards":1,"number_of_replicas":0,
    "analysis": {
      "tokenizer": {"atlas_edge":{"type":"edge_ngram","min_gram":2,"max_gram":12,"token_chars":["letter","digit"]}},
      "analyzer": {
        "atlas_autocomplete_index":{"type":"custom","tokenizer":"atlas_edge","filter":["lowercase","asciifolding"]},
        "atlas_autocomplete_search":{"type":"custom","tokenizer":"standard","filter":["lowercase","asciifolding"]}
      }
    }
  },
  "mappings":{"properties":{"name":{"type":"text","analyzer":"atlas_autocomplete_index","search_analyzer":"atlas_autocomplete_search"}}}
}
prove index vs search expansion
POST atlasmart-tokenizers-v1/_analyze
{"analyzer":"atlas_autocomplete_index","text":"Charger"}

POST atlasmart-tokenizers-v1/_analyze
{"analyzer":"atlas_autocomplete_search","text":"Char"}
Deliberate failure: edge n-grams at search time

If the query analyzer emits every prefix of charger, one user token becomes many query terms. That changes scoring, clause counts, and resource use. Keep the expansion where the design intends it and test it with realistic queries.

4. Paths and language-aware analysis solve different problems

category hierarchy tokens
POST _analyze
{"tokenizer":{"type":"path_hierarchy","delimiter":"/"},"text":"electronics/power/chargers"}

A hierarchy tokenizer can emit ancestor-like path terms such as electronics, electronics/power, and the full path. A language analyzer, by contrast, is a packaged analyzer that may combine standard-like tokenization with language-specific lowercasing, stopwords, normalization, and stemming. Language behavior must be evaluated with real catalog/query pairs; stemming can merge forms that users consider different, and multilingual fields may need separate language fields or routing rather than one aggressive analyzer.

compare generic and language analysis
POST _analyze
{"analyzer":"standard","text":"Running chargers for laptops"}

POST _analyze
{"analyzer":"english","text":"Running chargers for laptops"}

5. AtlasMart tokenizer regression fixture

Create a small table of inputs and expected token invariants before changing production analyzers: punctuation-heavy SKU, accented brand, hyphenated product phrase, short autocomplete prefix, category path, and at least one language-specific sentence. Store expected tokens and positions in source control and compare both products where portability matters.

Check your understanding

  1. Why is the keyword tokenizer different from a keyword mapping?
  2. Why are edge n-grams commonly asymmetric between index and search analysis?
  3. What does path_hierarchy preserve that standard does not?
  4. Is a language analyzer merely a tokenizer?
  5. What should determine gram sizes?
Review the answers

1. The tokenizer is one analysis stage that emits the whole input as one token; the mapping defines exact-value field semantics including indexing/doc-values behavior.

2. Index-time prefixes support prefix lookup, while a normal search analyzer avoids multiplying every user query token into many prefixes.

3. It preserves cumulative path ancestry based on a delimiter, enabling hierarchical category/path matching.

4. No. It is an analyzer chain that can include language-specific normalization, stopwords, stemming and other filters around tokenization.

5. Measured vocabulary/query behavior, relevance requirements, index size, latency, and resource constraints—not a universal constant.

cleanup
DELETE atlasmart-tokenizers-v1

Production judgment

Tokenizer changes alter the vocabulary of the index and therefore normally require reindexing. Before adopting partial-word or language-specific tokenization, quantify term growth, recall/precision changes, phrase behavior, and p95/p99 search latency under representative load. Keep a simple field for ordinary full-text search when a specialized autocomplete or hierarchy field serves only one interaction.

Next, we keep token boundaries fixed and study filters that normalize, remove, stem, expand, or combine those tokens.

Summary and next step

This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.