Chapter 05 · Text Analysis: Character Filters, Tokenizers, Token Filters, Analyzers, and Language Search
Use Analyze APIs to Debug Unexpected Matching and Build Tests for Tokenization/Normalization Contracts
Treat analysis as testable application behavior: diagnose unexpected matches with Analyze and term-vector evidence, encode token contracts as regression fixtures, and gate analyzer changes before production reindexing.
Learning outcomes
AtlasMart search now has a stable mapping and write pipeline, but users still judge the system by what text does and does not match. A search engine does not index the original sentence as one opaque value: it applies an analysis chain that can rewrite characters, split text into tokens, and transform those tokens before Lucene records searchable terms. This chapter makes that transformation observable so relevance changes are engineered and tested rather than guessed.
Use index and global Analyze APIs to reproduce the exact analyzer contract behind an unexpected match or miss.
Read token positions, offsets, types, and detail output as debugging evidence rather than only token strings.
Use term vectors where useful to compare analyzed terms associated with a stored document.
Encode analysis expectations as deterministic regression fixtures for both products.
Design a release gate that separates token correctness, relevance quality, and performance evidence.
Examples target Elasticsearch/Kibana 9.5.3 and
OpenSearch/OpenSearch Dashboards 3.8.0. Keep the
course lab contract established earlier: Elasticsearch at
https://localhost:9200 authenticated as the
disposable lab elastic user via
ELASTIC_PASSWORD and verified with the copied
HTTP CA; OpenSearch at
https://localhost:9201 using the disposable
Security-plugin demo admin identity via
OPENSEARCH_INITIAL_ADMIN_PASSWORD. The earlier
OpenSearch -k path is local demo convenience
only, not a production certificate-verification pattern.
Portable examples use core analysis components that both
products document. Product-specific synonym-management and
reload APIs are labeled explicitly instead of pretending the
two platforms expose one interchangeable control plane.
The generation environment does not run the two search
clusters, so expected token/output fragments are documented as
deterministic shapes derived from the configured analyzers,
not falsely presented as captured benchmark output. Run the
requests in the disposable AtlasMart lab from Chapters 01–04
with one local node per product, dedicated
atlasmart-* indices, and the pinned
server/container assumptions from those chapters. Analyzer
experiments use dedicated versioned indices only; cleanup
commands must never target unrelated data, credentials,
certificates, or volumes.
1. Debug from evidence, not from the visible sentence
When an AtlasMart query unexpectedly matches or misses, capture
four things: the field mapping/settings, the index-time token
stream for the document text, the search-time token stream for
the user text, and the actual query type. This avoids the common
error of staring at _source and assuming the engine
searches those exact characters.
POST atlasmart-products-analysis-v1/_analyze
{"field":"name","text":"Café USB-C Charger"}
POST atlasmart-products-analysis-v1/_analyze
{"analyzer":"atlas_name_search","text":"cafe charger"}
Use the index-scoped API when you need custom analyzer names
from that index. Global _analyze is useful for
built-in or fully inline chains. With custom components, tying
the probe to the index reduces accidental mismatch between the
test and deployed settings.
2. Positions and offsets explain phrase and highlighting boundaries
Token position is logical order in the token stream; some filters create gaps or alternatives at the same position. Offsets map tokens back to character ranges in the original/filtered text and can matter for highlighting/debugging. Token type is implementation metadata useful for inspection but should not be treated as a stable business API unless the design explicitly depends on documented behavior.
POST _analyze
{
"tokenizer":"standard",
"filter":["lowercase",{"type":"shingle","min_shingle_size":2,"max_shingle_size":2,"output_unigrams":true}],
"text":"wireless phone charger",
"explain":true
}
Detailed/explain-style analysis is diagnostic. It can show how each stage transformed the stream, but it is not a relevance benchmark and should not be used to infer production throughput.
3. Term vectors connect a document to indexed-term evidence
Term-vector APIs can expose terms and statistics for document fields when supported/configured. They are useful for debugging why a particular document contributes certain terms, but the returned statistics depend on index state and options. Treat them as evidence about terms, not as a guarantee of final ranking.
GET atlasmart-products-analysis-v1/_termvectors/p-100
{
"fields":["name","description"],
"term_statistics":false,
"field_statistics":false,
"positions":true,
"offsets":true
}
4. Build deterministic token-contract tests
A lightweight test harness can call _analyze,
extract ordered token text/positions, and compare them with
version-controlled fixtures. Keep portability fixtures for the
common analyzers and separate product-specific fixtures for
synonym administration or plugin components. A token test should
fail loudly when a release, plugin, configuration, or synonym
change alters the contract.
{
"case": "accented-product-name",
"index": "atlasmart-products-analysis-v1",
"analyzer": "atlas_name_search",
"input": "Café Charger",
"expected_tokens": ["cafe", "charger"]
}
{
"case": "sku-normalization",
"field": "sku",
"input": "USB-C-65W",
"expected_single_token": "usb-c-65w"
}
for each fixture:
response = POST /{index}/_analyze with fixture analyzer/field/input
actual = ordered response.tokens[].token
assert actual == expected_tokens
record server product + exact version + analyzer settings hash
then run judged search queries separately;
do not treat token equality as proof of relevance quality or latency.
5. Failure drills: make the wrong assumption reproducible
Use small fixtures to demonstrate three recurring failures: (1)
a term query does not analyze free-form user text;
(2) an edge-ngram search analyzer produces more query terms than
intended; (3) a synonym/stemmer change alters phrase or recall
behavior. For each, capture the before token stream, the failing
query/result, the corrected analyzer/query, and the after
result. This creates a diagnosis record rather than an anecdote.
Matching the expected tokens does not prove good ranking, safe latency, acceptable index size, or multilingual correctness. Maintain three gates: token-contract tests, judged relevance tests, and workload/performance tests. Report p95/p99 and resource conditions only from measured runs.
6. Chapter acceptance gate and safe cleanup
- Every analyzer used by AtlasMart has at least one positive and one boundary token fixture.
- Product-name punctuation, accents, case, SKUs, category paths, multiword synonyms, and representative language text are covered.
- Index-time analyzer changes are tied to a versioned reindex plan; search-only synonym changes use the exact supported product workflow.
-
Query tests distinguish
match/phrase behavior from term-level exact lookup. - Relevance and performance are tested separately from token correctness.
Check your understanding
- What is the first evidence to collect for an unexpected match?
- Why inspect positions, not only token strings?
- What do term vectors prove?
- Why version fixtures with the product/server version?
- What are the three separate release gates?
Review the answers
1. The field mapping/settings plus the index-time and search-time token streams for the actual document/query text, followed by the exact query type.
2. Phrase/proximity and graph semantics depend on token positions; the same token list with different positions can construct different queries.
3. They expose indexed-term-related evidence for a document/field under requested options; they do not by themselves prove final relevance ranking.
4. Analysis implementations, plugins and defaults can evolve; recording the environment makes drift diagnosable instead of mysterious.
5. Token-contract correctness, judged relevance quality, and workload/performance/resource behavior.
DELETE atlasmart-products-analysis-v1
DELETE atlasmart-tokenizers-v1
DELETE atlasmart-filters-v1
DELETE atlasmart-normalizer-v1
DELETE atlasmart-products-analysis-v2
DELETE atlasmart-es-syn-v1
DELETE atlasmart-os-syn-v1
Production judgment
Text analysis deserves version control, tests, ownership, and rollback just like application code. Record analyzer settings, synonym sources, plugin/product versions, and judged query fixtures. Watch for relevance regressions, token/index growth, cache and CPU effects, and query tail latency after changes. Keep multilingual behavior explicit; do not assume English stemming/folding strategies generalize to other languages.
Chapter 06 builds on this evidence by composing Query DSL full-text, term-level, boolean, range, nested, geo, and safe user-intent queries. The analysis contract you can now inspect explains why those queries match what they match.
Summary and next step
This lesson’s analysis contract, observable evidence, failure boundaries, and production implications should now be concrete enough to verify rather than assume.
Authoritative references
- Elastic: standard analyzer — Current reference for the default standard analysis chain.
- Elastic: Analyze API — Token/position/offset inspection for built-in, field-bound, and custom analysis.
- Elastic: tokenizer reference — Official tokenizer families and behavior.
- Elastic: configure synonyms — Search-time synonym-set workflow, validation, and reload behavior.
- Elastic: reload search analyzers API — Reload semantics for file-backed, updateable search analyzers.
- OpenSearch: text analysis — Analyzer, tokenizer, token-filter, and normalizer fundamentals.
- OpenSearch: Analyze API — Index/global analysis requests and detailed token evidence.
- OpenSearch: synonym_graph token filter — Multiword synonym graph behavior.
- OpenSearch: refresh search analyzer — Product-specific analyzer refresh workflow and plugin requirement.
- OpenSearch: normalizers — Single-token keyword normalization constraints.