Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search

Use Analyze APIs to Debug Unexpected Matching and Build Tests for Tokenization/Normalization Contracts

Treat analysis as testable application behavior: diagnose unexpected matches with Analyze and term-vector evidence, encode token contracts as regression fixtures, and gate analyzer changes before production reindexing.

Intermediate → Advanced120–145 minutesAnalysis regression-test labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.

01

Use index and global Analyze APIs to reproduce the exact analyzer contract behind an unexpected match or miss.

02

Read token positions, offsets, types, and detail output as debugging evidence rather than only token strings.

03

Use term vectors where useful to compare analyzed terms associated with a stored document.

04

Encode analysis expectations as deterministic regression fixtures for both products.

05

Design a release gate that separates token correctness, relevance quality, and performance evidence.

Chapter baseline reviewed 11 September 2026

Examples target Elasticsearch/Kibana 9.5.3 and OpenSearch/OpenSearch Dashboards 3.8.0. Keep the course lab contract established earlier: Elasticsearch at https://localhost:9200 authenticated as the disposable lab elastic user via ELASTIC_PASSWORD and verified with the copied HTTP CA; OpenSearch at https://localhost:9201 using the disposable Security-plugin demo admin identity via OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier OpenSearch -k path is local demo convenience only, not a production certificate-verification pattern. Portable examples use core analysis components that both products document. Product-specific synonym-management and reload APIs are labeled explicitly instead of pretending the two platforms expose one interchangeable control plane.

Execution and safety note

The generation environment does not run the two search clusters, so expected token/output fragments are documented as deterministic shapes derived from the configured analyzers, not falsely presented as captured benchmark output. Run the requests in the disposable AtlasMart lab from Chapters 01–04 with one local node per product, dedicated atlasmart-* indices, and the pinned server/container assumptions from those chapters. Analyzer experiments use dedicated versioned indices only; cleanup commands must never target unrelated data, credentials, certificates, or volumes.

1. Debug from evidence, not from the visible sentence

When an AtlasMart query unexpectedly matches or misses, capture four things: the field mapping/settings, the index-time token stream for the document text, the search-time token stream for the user text, and the actual query type. This avoids the common error of staring at _source and assuming the engine searches those exact characters.

field-bound Analyze probes
POST atlasmart-products-analysis-v1/_analyze
{"field":"name","text":"Café USB-C Charger"}

POST atlasmart-products-analysis-v1/_analyze
{"analyzer":"atlas_name_search","text":"cafe charger"}

Use the index-scoped API when you need custom analyzer names from that index. Global _analyze is useful for built-in or fully inline chains. With custom components, tying the probe to the index reduces accidental mismatch between the test and deployed settings.

2. Positions and offsets explain phrase and highlighting boundaries

Token position is logical order in the token stream; some filters create gaps or alternatives at the same position. Offsets map tokens back to character ranges in the original/filtered text and can matter for highlighting/debugging. Token type is implementation metadata useful for inspection but should not be treated as a stable business API unless the design explicitly depends on documented behavior.

request detailed analysis stages
POST _analyze
{
  "tokenizer":"standard",
  "filter":["lowercase",{"type":"shingle","min_shingle_size":2,"max_shingle_size":2,"output_unigrams":true}],
  "text":"wireless phone charger",
  "explain":true
}

Detailed/explain-style analysis is diagnostic. It can show how each stage transformed the stream, but it is not a relevance benchmark and should not be used to infer production throughput.

3. Term vectors connect a document to indexed-term evidence

Term-vector APIs can expose terms and statistics for document fields when supported/configured. They are useful for debugging why a particular document contributes certain terms, but the returned statistics depend on index state and options. Treat them as evidence about terms, not as a guarantee of final ranking.

term-vector probe
GET atlasmart-products-analysis-v1/_termvectors/p-100
{
  "fields":["name","description"],
  "term_statistics":false,
  "field_statistics":false,
  "positions":true,
  "offsets":true
}

4. Build deterministic token-contract tests

A lightweight test harness can call _analyze, extract ordered token text/positions, and compare them with version-controlled fixtures. Keep portability fixtures for the common analyzers and separate product-specific fixtures for synonym administration or plugin components. A token test should fail loudly when a release, plugin, configuration, or synonym change alters the contract.

example fixture specification
{
  "case": "accented-product-name",
  "index": "atlasmart-products-analysis-v1",
  "analyzer": "atlas_name_search",
  "input": "Café Charger",
  "expected_tokens": ["cafe", "charger"]
}
{
  "case": "sku-normalization",
  "field": "sku",
  "input": "USB-C-65W",
  "expected_single_token": "usb-c-65w"
}
pseudo-test loop · fail on contract drift
for each fixture:
  response = POST /{index}/_analyze with fixture analyzer/field/input
  actual = ordered response.tokens[].token
  assert actual == expected_tokens
  record server product + exact version + analyzer settings hash

then run judged search queries separately;
do not treat token equality as proof of relevance quality or latency.

5. Failure drills: make the wrong assumption reproducible

Use small fixtures to demonstrate three recurring failures: (1) a term query does not analyze free-form user text; (2) an edge-ngram search analyzer produces more query terms than intended; (3) a synonym/stemmer change alters phrase or recall behavior. For each, capture the before token stream, the failing query/result, the corrected analyzer/query, and the after result. This creates a diagnosis record rather than an anecdote.

Analyzer tests are necessary, not sufficient.

Matching the expected tokens does not prove good ranking, safe latency, acceptable index size, or multilingual correctness. Maintain three gates: token-contract tests, judged relevance tests, and workload/performance tests. Report p95/p99 and resource conditions only from measured runs.

6. Chapter acceptance gate and safe cleanup

  • Every analyzer used by AtlasMart has at least one positive and one boundary token fixture.
  • Product-name punctuation, accents, case, SKUs, category paths, multiword synonyms, and representative language text are covered.
  • Index-time analyzer changes are tied to a versioned reindex plan; search-only synonym changes use the exact supported product workflow.
  • Query tests distinguish match/phrase behavior from term-level exact lookup.
  • Relevance and performance are tested separately from token correctness.

Check your understanding

  1. What is the first evidence to collect for an unexpected match?
  2. Why inspect positions, not only token strings?
  3. What do term vectors prove?
  4. Why version fixtures with the product/server version?
  5. What are the three separate release gates?
Review the answers

1. The field mapping/settings plus the index-time and search-time token streams for the actual document/query text, followed by the exact query type.

2. Phrase/proximity and graph semantics depend on token positions; the same token list with different positions can construct different queries.

3. They expose indexed-term-related evidence for a document/field under requested options; they do not by themselves prove final relevance ranking.

4. Analysis implementations, plugins and defaults can evolve; recording the environment makes drift diagnosable instead of mysterious.

5. Token-contract correctness, judged relevance quality, and workload/performance/resource behavior.

cleanup · only disposable Chapter 05 indices
DELETE atlasmart-products-analysis-v1
DELETE atlasmart-tokenizers-v1
DELETE atlasmart-filters-v1
DELETE atlasmart-normalizer-v1
DELETE atlasmart-products-analysis-v2
DELETE atlasmart-es-syn-v1
DELETE atlasmart-os-syn-v1

Production judgment

Text analysis deserves version control, tests, ownership, and rollback just like application code. Record analyzer settings, synonym sources, plugin/product versions, and judged query fixtures. Watch for relevance regressions, token/index growth, cache and CPU effects, and query tail latency after changes. Keep multilingual behavior explicit; do not assume English stemming/folding strategies generalize to other languages.

Chapter 06 builds on this evidence by composing Query DSL full-text, term-level, boolean, range, nested, geo, and safe user-intent queries. The analysis contract you can now inspect explains why those queries match what they match.

Summary and next step

This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.