Chapter 03 · Mappings and Field Types: keyword, text, Numeric, Date, Geo, Object, Nested, and Runtime Fields

text vs keyword: Analysis, Exact Matching, Sorting/Aggregation, Multi-Fields, and Normalizers

Separate full-text semantics from exact-value semantics and prove how analyzers, normalizers, multi-fields, doc values, sorting, and aggregations change observable behavior.

Intermediate → Advanced100–120 minutesDual-platform mapping evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart wants shoppers to find “USB C charger” with natural-language matching while merchandisers group results by brand, sort by SKU, and filter a status value without tokenization surprises. A single generic “string” type cannot satisfy all of those requirements correctly. This lesson uses observable token streams and query results to separate text from keyword, then uses multi-fields and normalizers to support multiple access patterns without duplicating source JSON.

01

Predict how text analysis changes terms and why term-level exact matching is a different contract from full-text matching.

02

Use keyword for identifiers, statuses and facet values whose whole value is meaningful, including deliberate case normalization.

03

Create a multi-field so one source string supports full-text retrieval plus exact sorting/aggregation semantics.

04

Prove behavior with mapping inspection, analyze APIs, match/term queries, sorting and terms aggregations.

05

Diagnose the common “aggregate on analyzed text” anti-pattern without blindly enabling memory-expensive fielddata.

Chapter baseline reviewed 10 September 2026

Examples are written against Elasticsearch 9.5.3 and OpenSearch 3.8.0. Those products share Lucene ancestry but are not one API surface. Portable examples use only behavior verified on both platforms; divergent features are labeled separately. Elasticsearch examples assume the default self-managed distribution with its bundled JVM. OpenSearch examples assume the upstream 3.8.0 distribution with the Security plugin present. Re-check release notes, support matrices, plugin compatibility, and feature status before reproducing this chapter later.

Execution and safety note

The generation environment used to build this chapter does not run the two search servers. Commands were checked against current official documentation but were not executed here, so output blocks describe expected shape and invariant, not captured benchmark evidence. Use only disposable AtlasMart indices. Keep Chapter 01/02 endpoints: Elasticsearch at https://localhost:9200 with ELASTIC_PASSWORD, OpenSearch at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD, and the course CA/certificate paths established by the earlier labs.

1. text and keyword answer different questions

A text field is intended for full-text search. Its analyzer converts a string into a token stream at index time and a compatible analysis process is usually applied to full-text query input. A keyword field treats the value as one logical token for exact/term-level operations. That distinction is semantic, not merely a performance setting.

Requirement Preferred mapping Reason
Search product names by words/phrases text Analysis creates searchable terms and positions.
Filter exact product ID keyword The entire identifier is the value; numeric range semantics are irrelevant.
Facet by brand/status keyword Doc values support exact grouping efficiently by default.
Sort product name alphabetically Often a keyword subfield Sorting needs whole-value semantics rather than analyzed tokens.
Search and facet the same source string text + keyword multi-field One source value is indexed through two search contracts.

OpenSearch documents keyword as non-analyzed, exact and case-sensitive by default. Elasticsearch uses the same broad division, but current Elasticsearch 9.x has additional options around text/doc-values behavior that should not be projected onto OpenSearch. A portable course design therefore keeps facets and sorting on explicit keyword fields/subfields.

2. Observe the token stream

portable REST · create a multi-field and normalizer
DELETE atlasmart-text-keyword-labPUT atlasmart-text-keyword-lab{  "settings": {    "number_of_shards": 1,    "number_of_replicas": 0,    "analysis": {      "normalizer": {        "atlas_lower": {"type":"custom","filter":["lowercase","asciifolding"]}      }    }  },  "mappings": {    "dynamic":"strict",    "properties": {      "product_id":{"type":"keyword"},      "name":{"type":"text","fields":{"raw":{"type":"keyword"}}},      "brand":{"type":"keyword","normalizer":"atlas_lower"},      "status":{"type":"keyword"}    }  }}POST atlasmart-text-keyword-lab/_analyze{"field":"name","text":"USB-C Travel Charger"}POST atlasmart-text-keyword-lab/_analyze{"normalizer":"atlas_lower","text":"ÄTLAS Gear"}

The name analysis should emit multiple terms rather than one literal JSON string. The normalizer should emit exactly one normalized token because normalizers do not tokenize a keyword value into multiple words. Record the actual tokens returned by each product; token details can vary with analyzer configuration and version, so the acceptance criterion is semantic shape, not a memorized token list.

3. One source value, two indexed views

portable REST · index and compare query families
POST atlasmart-text-keyword-lab/_bulk?refresh=true{"index":{"_id":"P-1"}}{"product_id":"P-1","name":"USB-C Travel Charger","brand":"ÄTLAS Gear","status":"ACTIVE"}{"index":{"_id":"P-2"}}{"product_id":"P-2","name":"Compact USB Charger","brand":"Atlas Gear","status":"ACTIVE"}{"index":{"_id":"P-3"}}{"product_id":"P-3","name":"Travel Adapter","brand":"Other","status":"PAUSED"}GET atlasmart-text-keyword-lab/_search{"query":{"match":{"name":"travel charger"}}}GET atlasmart-text-keyword-lab/_search{"query":{"term":{"name.raw":"USB-C Travel Charger"}}}GET atlasmart-text-keyword-lab/_search{"size":0,"aggs":{"brands":{"terms":{"field":"brand"}}}}GET atlasmart-text-keyword-lab/_search{"sort":[{"name.raw":"asc"}],"query":{"match_all":{}}}

name and name.raw are not duplicate properties in _source; they are two indexed representations of the same source field. The text view supports analyzed retrieval. The keyword subfield supports whole-value operations. The normalized brand values demonstrate that exact-value semantics can still be normalized for case/diacritic consistency, but that normalization choice must match business requirements because it changes what “exact” means.

4. Wrong approach: aggregate on the analyzed text field

A developer may try a terms aggregation directly on name. On a conventional mapping, this is rejected or unsupported by default because analyzed text is not the intended grouping surface. Historically, Elasticsearch users sometimes enabled fielddata to make such operations possible, which can consume substantial heap. The safer design for this requirement is the keyword multi-field you already created.

controlled failure · wrong field for a facet
GET atlasmart-text-keyword-lab/_search{  "size":0,  "aggs":{"wrong_product_name_facet":{"terms":{"field":"name"}}}}GET atlasmart-text-keyword-lab/_search{  "size":0,  "aggs":{"correct_product_name_facet":{"terms":{"field":"name.raw"}}}}
Diagnosis before tuning

If the first request fails, that failure is evidence that the mapping does not expose the analyzed text as an aggregation-friendly exact-value surface by default. Do not “fix” it by globally enabling a costly feature without measuring. If the exact behavior differs in a future Elasticsearch release, re-check the mapping reference; the cross-platform contract remains that the dedicated keyword view is the predictable portable choice for facets and sorts.

5. Boundary cases: IDs that look numeric and long keyword values

Field types come from operations, not appearance. A numeric-looking product ID such as 001234 is usually better as keyword if AtlasMart performs only exact lookup and needs to preserve leading zeros. Conversely, an actual price belongs in a numeric type because range and numeric sort semantics matter. Dynamic mapping can infer both incorrectly relative to business intent.

Keyword values also have indexing limits and parameters such as ignore_above. Dynamic mappings commonly create text plus keyword subfields with an ignore_above threshold; a long value may remain in _source while its keyword subfield is not indexed. That means “I can see the value in the document” does not prove “I can facet/filter it as keyword.” Add a boundary test whenever source length is uncontrolled.

Check your understanding

  1. Why does a match query on text behave differently from a term query on a keyword subfield?
  2. What is a multi-field?
  3. What does a normalizer do that an analyzer does not?
  4. Why is a numeric-looking ID often still keyword?
  5. Why should fielddata not be the reflexive fix for text aggregations?
Review the answers

1. The text path analyzes input and indexed content into terms, while the keyword path compares whole normalized terms.

2. Multiple indexed representations of one source field, for example name as text and name.raw as keyword.

3. It preserves a single token and applies allowed normalization filters instead of tokenizing text into many terms.

4. Its business operation is exact identity, not arithmetic/range search, and lexical form such as leading zeros may matter.

5. It can create significant memory cost; a dedicated keyword field/subfield models the exact-value requirement directly.

Summary and next step

The source string is not the search contract. You have now proved that analysis, exact matching, sorting, and aggregation require distinct indexed representations. The next lesson extends that requirement-driven method to numeric, date, geo, range, vector, object, flattened-like, and computed field families.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.