Chapter 03 · Mappings and Field Types: keyword, text, Numeric, Date, Geo, Object, Nested, and Runtime Fields
Numeric, Date, Boolean, IP, Geo, Range, Dense Vector, Object, Flattened, and Other Specialized Types
Choose specialized field types from query and aggregation requirements, while keeping Elasticsearch and OpenSearch type-name and capability differences explicit.
Learning outcomes
AtlasMart’s catalog contains prices, timestamps, stock flags, client IPs, warehouse coordinates, promotional validity windows, arbitrary supplier attributes, and later semantic embeddings. Mapping all of them as strings would discard useful semantics; mapping them with platform-specific types without labeling divergence would make the course nonportable. This lesson builds a field-selection matrix and shows where Elasticsearch 9.5.3 and OpenSearch 3.8.0 intentionally use different field families.
Select numeric, date, boolean, IP, geo and range types from operations and validation requirements rather than JSON appearance.
Distinguish object from flattened-like
representations and explain what search/aggregation
precision is traded for schema compactness.
Keep Elasticsearch flattened separate from
OpenSearch flat_object, including their
different subfield capabilities.
Keep Elasticsearch dense_vector separate from
OpenSearch knn_vector, and record
dimensions/model semantics as an external contract.
Compare Elasticsearch runtime fields with OpenSearch derived fields as query-time computation mechanisms without claiming API equivalence.
Examples are written against
Elasticsearch 9.5.3 and
OpenSearch 3.8.0. Those products share Lucene
ancestry but are not one API surface. Portable examples use
only behavior verified on both platforms; divergent features
are labeled separately. Elasticsearch examples assume the
default self-managed distribution with its bundled JVM.
OpenSearch examples assume the upstream 3.8.0 distribution
with the Security plugin present. Re-check release notes,
support matrices, plugin compatibility, and feature status
before reproducing this chapter later.
The generation environment used to build this chapter does not
run the two search servers. Commands were checked against
current official documentation but were not executed here, so
output blocks describe expected shape and invariant,
not captured benchmark evidence. Use only disposable AtlasMart
indices. Keep Chapter 01/02 endpoints: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD, OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD, and the course
CA/certificate paths established by the earlier labs.
1. Choose types from operations
| AtlasMart field | Operation | Portable core type | Key design question |
|---|---|---|---|
| price | Numeric range/sort/metric agg |
double or deliberate scaled representation
|
What precision and unit are contractual? |
| updated_at | Time range/sort | date |
Which formats/time-zone normalization are accepted? |
| in_stock | Exact truth value | boolean |
How are null/unknown states represented? |
| client_ip | Exact/CIDR-style network queries | ip |
Is the field genuinely an IP rather than arbitrary text? |
| warehouse | Distance/bounding-area search | geo_point |
Which coordinate order/input formats are accepted? |
| promo_window | Range overlap/containment | date_range where supported |
Do interval semantics justify a range field? |
| specs | Known structured subfields | object |
Do subfields have stable types and query needs? |
Identifiers deserve special attention. A value containing digits
is not automatically numeric. Product IDs, postal codes and
phone-like values often need exact lexical identity, so
keyword can be more appropriate. Numeric types are
optimized for numeric range semantics.
2. Portable specialized mapping and boundary documents
DELETE atlasmart-special-typesPUT atlasmart-special-types{ "settings":{"number_of_shards":1,"number_of_replicas":0}, "mappings":{ "dynamic":"strict", "properties":{ "product_id":{"type":"keyword"}, "price":{"type":"double"}, "updated_at":{"type":"date"}, "in_stock":{"type":"boolean"}, "client_ip":{"type":"ip"}, "warehouse":{"type":"geo_point"}, "promo_window":{"type":"date_range"}, "specs":{"properties":{"voltage":{"type":"integer"},"connector":{"type":"keyword"}}} } }}POST atlasmart-special-types/_doc/P-300?refresh=true{ "product_id":"P-300","price":49.95, "updated_at":"2026-09-10T12:30:00Z","in_stock":true, "client_ip":"203.0.113.10","warehouse":{"lat":35.6892,"lon":51.3890}, "promo_window":{"gte":"2026-09-01","lt":"2026-10-01"}, "specs":{"voltage":220,"connector":"USB-C"}}
Then test one invalid boundary per type: an unparsable date, invalid IP, malformed coordinate, or string in the integer field. Capture which document is rejected and whether any coercion occurred. Coercion settings are part of the contract; do not infer them from one happy-path document.
3. Flattened-like objects are not the same feature
Both products offer a way to avoid making every arbitrary key a
full mapped field, but the APIs and capabilities diverge.
Elasticsearch’s flattened field indexes leaf values
in a compact single field mapping and supports basic
search/aggregation behavior with documented limitations.
OpenSearch’s flat_object treats the JSON object as
a flat string-oriented structure; current documentation
emphasizes that subfields are not indexed as ordinary typed
fields and that operations such as numeric semantics and
subfield aggregations/filtering have important limitations. They
solve a similar governance problem, not an identical query
problem.
PUT atlasmart-es-flex{ "mappings":{"properties":{"supplier_attributes":{"type":"flattened"}}}}
PUT atlasmart-os-flex{ "mappings":{"properties":{"supplier_attributes":{"type":"flat_object"}}}}
Use flattened-like storage for attributes whose key cardinality is high and whose search requirements are intentionally limited. If AtlasMart needs true numeric range, facet, sort or language analysis semantics on a supplier attribute, promote that attribute into a typed governed field instead of expecting the flat representation to become a universal schema substitute.
4. Dense vectors: shared concept, different field contracts
A dense embedding is a fixed-length numeric vector produced by a
model or feature pipeline. The model name/version, dimension,
normalization and similarity interpretation belong to the data
contract. Elasticsearch uses dense_vector.
OpenSearch’s primary dense-vector mapping is
knn_vector, with its own k-NN settings, engines,
modes and compression options. The course deliberately postpones
serious ANN/HNSW evaluation to Chapter 25; here the goal is to
avoid a schema category error.
"embedding": { "type": "dense_vector", "dims": 384}
"settings": {"index.knn": true},"mappings": {"properties": { "embedding": {"type":"knn_vector","dimension":384}}}
Do not insert fake embeddings and claim semantic quality. A syntactically valid 384-dimensional array only proves the field accepts the shape. Retrieval quality requires a named model, representative corpus and ground-truth evaluation.
5. Query-time fields: Elasticsearch runtime vs OpenSearch derived
Elasticsearch runtime fields and OpenSearch derived fields both
compute field values at query time, but they have different
syntax, supported operations and limitations. Runtime fields can
be declared in Elasticsearch mappings or searches and are
explicitly documented as potentially expensive. OpenSearch
derived fields, introduced earlier in the 2.x line and present
in 3.8, similarly execute scripts over _source or
doc values; current OpenSearch documentation lists restrictions
such as lack of sorting and limitations across some aggregation
families. Treat them as separate product features.
GET atlasmart-special-types/_search{ "runtime_mappings": { "price_with_tax": { "type":"double", "script":"emit(doc['price'].value * 1.10)" } }, "fields":["product_id","price_with_tax"], "query":{"match_all":{}}}
GET atlasmart-special-types/_search{ "derived": { "price_with_tax": { "type":"double", "script":{"source":"emit(doc['price'].value * 1.10)"} } }, "fields":["product_id","price_with_tax"], "query":{"match_all":{}}}
Query-time computation can reduce reindexing pressure for exploratory transformations, but it moves work into reads. Promote heavily used stable computations to indexed fields when measurement shows that repeated script execution harms latency or capacity.
Check your understanding
- Why might a product ID consisting only of digits still be keyword?
-
Are Elasticsearch
flattenedand OpenSearchflat_objectinterchangeable? - What is the OpenSearch dense-vector field called?
- What must accompany a vector dimension in production documentation?
- When should a query-time computed field become indexed?
Review the answers
1. Because the required operation is exact identity rather than numeric range arithmetic, and lexical form may matter.
2. No. They target similar arbitrary-key problems but expose different indexing and query/aggregation capabilities.
3. The main field type is
knn_vector; Elasticsearch uses
dense_vector.
4. At minimum the embedding model/provider/version, dimension, normalization/similarity assumptions, and migration/re-embedding strategy.
5. When the computation and semantics are stable and measured query-time cost or unsupported operations justify paying the indexing/storage cost instead.
Summary and next step
Specialized field types encode operations and validation rules, not cosmetic labels. You now have a portable typed core plus explicit divergence points for flattened, computed and vector fields. The next lesson focuses on the hardest structured-data boundary: arrays of objects, nested tuple semantics, and parent/child joins.
Authoritative references
- Elastic mapping overview — Official mapping guidance and schema-evolution constraints.
- Elastic field data types — Current Elasticsearch field-type reference.
- Elastic mapping limits — Mapping-limit settings and mapping-explosion safeguards.
- OpenSearch mapping documentation — Current OpenSearch mapping entry point.
- OpenSearch supported field types — Current OpenSearch field-type inventory.
- OpenSearch mapping explosion — Field growth risks and mapping-limit controls.
- Elastic flattened field — Elasticsearch compact object mapping and limitations.
- Elastic runtime fields — Elasticsearch query-time field computation.
- OpenSearch flat object — OpenSearch flat-object semantics and limitations.
- OpenSearch derived field — OpenSearch query-time derived field semantics and limitations.
- OpenSearch k-NN vector — OpenSearch dense vector mapping contract.