Chapter 03 · Mappings and Field Types: keyword, text, Numeric, Date, Geo, Object, Nested, and Runtime Fields
text vs keyword: Analysis, Exact Matching, Sorting/Aggregation, Multi-Fields, and Normalizers
Separate full-text semantics from exact-value semantics and prove how analyzers, normalizers, multi-fields, doc values, sorting, and aggregations change observable behavior.
Learning outcomes
AtlasMart wants shoppers to find “USB C charger” with
natural-language matching while merchandisers group results by
brand, sort by SKU, and filter a status value without
tokenization surprises. A single generic “string” type cannot
satisfy all of those requirements correctly. This lesson uses
observable token streams and query results to separate
text from keyword, then uses
multi-fields and normalizers to support multiple access patterns
without duplicating source JSON.
Predict how text analysis changes terms and why
term-level exact matching is a different contract from
full-text matching.
Use keyword for identifiers, statuses and facet
values whose whole value is meaningful, including deliberate
case normalization.
Create a multi-field so one source string supports full-text retrieval plus exact sorting/aggregation semantics.
Prove behavior with mapping inspection, analyze APIs,
match/term queries, sorting and
terms aggregations.
Diagnose the common “aggregate on analyzed text” anti-pattern without blindly enabling memory-expensive fielddata.
Examples are written against
Elasticsearch 9.5.3 and
OpenSearch 3.8.0. Those products share Lucene
ancestry but are not one API surface. Portable examples use
only behavior verified on both platforms; divergent features
are labeled separately. Elasticsearch examples assume the
default self-managed distribution with its bundled JVM.
OpenSearch examples assume the upstream 3.8.0 distribution
with the Security plugin present. Re-check release notes,
support matrices, plugin compatibility, and feature status
before reproducing this chapter later.
The generation environment used to build this chapter does not
run the two search servers. Commands were checked against
current official documentation but were not executed here, so
output blocks describe expected shape and invariant,
not captured benchmark evidence. Use only disposable AtlasMart
indices. Keep Chapter 01/02 endpoints: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD, OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD, and the course
CA/certificate paths established by the earlier labs.
1. text and keyword answer different questions
A text field is intended for full-text search. Its
analyzer converts a string into a token stream at index time and
a compatible analysis process is usually applied to full-text
query input. A keyword field treats the value as
one logical token for exact/term-level operations. That
distinction is semantic, not merely a performance setting.
| Requirement | Preferred mapping | Reason |
|---|---|---|
| Search product names by words/phrases | text |
Analysis creates searchable terms and positions. |
| Filter exact product ID | keyword |
The entire identifier is the value; numeric range semantics are irrelevant. |
| Facet by brand/status | keyword |
Doc values support exact grouping efficiently by default. |
| Sort product name alphabetically | Often a keyword subfield |
Sorting needs whole-value semantics rather than analyzed tokens. |
| Search and facet the same source string | text + keyword multi-field |
One source value is indexed through two search contracts. |
OpenSearch documents keyword as non-analyzed, exact
and case-sensitive by default. Elasticsearch uses the same broad
division, but current Elasticsearch 9.x has additional options
around text/doc-values behavior that should not be projected
onto OpenSearch. A portable course design therefore keeps facets
and sorting on explicit keyword fields/subfields.
2. Observe the token stream
DELETE atlasmart-text-keyword-labPUT atlasmart-text-keyword-lab{ "settings": { "number_of_shards": 1, "number_of_replicas": 0, "analysis": { "normalizer": { "atlas_lower": {"type":"custom","filter":["lowercase","asciifolding"]} } } }, "mappings": { "dynamic":"strict", "properties": { "product_id":{"type":"keyword"}, "name":{"type":"text","fields":{"raw":{"type":"keyword"}}}, "brand":{"type":"keyword","normalizer":"atlas_lower"}, "status":{"type":"keyword"} } }}POST atlasmart-text-keyword-lab/_analyze{"field":"name","text":"USB-C Travel Charger"}POST atlasmart-text-keyword-lab/_analyze{"normalizer":"atlas_lower","text":"ÄTLAS Gear"}
The name analysis should emit multiple terms rather than one literal JSON string. The normalizer should emit exactly one normalized token because normalizers do not tokenize a keyword value into multiple words. Record the actual tokens returned by each product; token details can vary with analyzer configuration and version, so the acceptance criterion is semantic shape, not a memorized token list.
3. One source value, two indexed views
POST atlasmart-text-keyword-lab/_bulk?refresh=true{"index":{"_id":"P-1"}}{"product_id":"P-1","name":"USB-C Travel Charger","brand":"ÄTLAS Gear","status":"ACTIVE"}{"index":{"_id":"P-2"}}{"product_id":"P-2","name":"Compact USB Charger","brand":"Atlas Gear","status":"ACTIVE"}{"index":{"_id":"P-3"}}{"product_id":"P-3","name":"Travel Adapter","brand":"Other","status":"PAUSED"}GET atlasmart-text-keyword-lab/_search{"query":{"match":{"name":"travel charger"}}}GET atlasmart-text-keyword-lab/_search{"query":{"term":{"name.raw":"USB-C Travel Charger"}}}GET atlasmart-text-keyword-lab/_search{"size":0,"aggs":{"brands":{"terms":{"field":"brand"}}}}GET atlasmart-text-keyword-lab/_search{"sort":[{"name.raw":"asc"}],"query":{"match_all":{}}}
name and name.raw are not duplicate
properties in _source; they are two indexed
representations of the same source field. The text view supports
analyzed retrieval. The keyword subfield supports whole-value
operations. The normalized brand values demonstrate that
exact-value semantics can still be normalized for case/diacritic
consistency, but that normalization choice must match business
requirements because it changes what “exact” means.
4. Wrong approach: aggregate on the analyzed text field
A developer may try a terms aggregation directly on
name. On a conventional mapping, this is rejected
or unsupported by default because analyzed text is not the
intended grouping surface. Historically, Elasticsearch users
sometimes enabled fielddata to make such operations possible,
which can consume substantial heap. The safer design for this
requirement is the keyword multi-field you already created.
GET atlasmart-text-keyword-lab/_search{ "size":0, "aggs":{"wrong_product_name_facet":{"terms":{"field":"name"}}}}GET atlasmart-text-keyword-lab/_search{ "size":0, "aggs":{"correct_product_name_facet":{"terms":{"field":"name.raw"}}}}
If the first request fails, that failure is evidence that the mapping does not expose the analyzed text as an aggregation-friendly exact-value surface by default. Do not “fix” it by globally enabling a costly feature without measuring. If the exact behavior differs in a future Elasticsearch release, re-check the mapping reference; the cross-platform contract remains that the dedicated keyword view is the predictable portable choice for facets and sorts.
5. Boundary cases: IDs that look numeric and long keyword values
Field types come from operations, not appearance. A
numeric-looking product ID such as 001234 is
usually better as keyword if AtlasMart performs
only exact lookup and needs to preserve leading zeros.
Conversely, an actual price belongs in a numeric type because
range and numeric sort semantics matter. Dynamic mapping can
infer both incorrectly relative to business intent.
Keyword values also have indexing limits and parameters such as
ignore_above. Dynamic mappings commonly create text
plus keyword subfields with an
ignore_above threshold; a long value may remain in
_source while its keyword subfield is not indexed.
That means “I can see the value in the document” does not prove
“I can facet/filter it as keyword.” Add a boundary test whenever
source length is uncontrolled.
Check your understanding
-
Why does a match query on
textbehave differently from a term query on a keyword subfield? - What is a multi-field?
- What does a normalizer do that an analyzer does not?
- Why is a numeric-looking ID often still keyword?
- Why should fielddata not be the reflexive fix for text aggregations?
Review the answers
1. The text path analyzes input and indexed content into terms, while the keyword path compares whole normalized terms.
2. Multiple indexed representations of
one source field, for example name as text
and name.raw as keyword.
3. It preserves a single token and applies allowed normalization filters instead of tokenizing text into many terms.
4. Its business operation is exact identity, not arithmetic/range search, and lexical form such as leading zeros may matter.
5. It can create significant memory cost; a dedicated keyword field/subfield models the exact-value requirement directly.
Summary and next step
The source string is not the search contract. You have now proved that analysis, exact matching, sorting, and aggregation require distinct indexed representations. The next lesson extends that requirement-driven method to numeric, date, geo, range, vector, object, flattened-like, and computed field families.
Authoritative references
- Elastic mapping overview — Official mapping guidance and schema-evolution constraints.
- Elastic field data types — Current Elasticsearch field-type reference.
- Elastic mapping limits — Mapping-limit settings and mapping-explosion safeguards.
- OpenSearch mapping documentation — Current OpenSearch mapping entry point.
- OpenSearch supported field types — Current OpenSearch field-type inventory.
- OpenSearch mapping explosion — Field growth risks and mapping-limit controls.
- Elastic text field — Text analysis, multi-fields and text-field options.
- Elastic keyword field — Exact-value keyword semantics and parameters.
- OpenSearch keyword field — OpenSearch keyword behavior and multi-field support.
- OpenSearch normalizer — Single-token normalization for keyword fields.