Chapter 03 · Mappings and Field Types: keyword, text, Numeric, Date, Geo, Object, Nested, and Runtime Fields
Explicit vs Dynamic Mappings, Field Explosion, Mapping Conflicts, and Schema-Governance Strategy
Treat mappings as an executable search contract: control dynamic discovery, prevent field explosion, diagnose type conflicts, and govern schema evolution before ingestion scales.
Learning outcomes
AtlasMart receives catalog feeds from several suppliers. One
feed emits price as a number, another sometimes
sends "unknown", and a marketplace partner invents
a new attribute key for almost every seller. Search still looks
healthy on a tiny sample, but an unconstrained mapping can turn
this ingestion freedom into rejected documents, inconsistent
field semantics, large cluster state, and a costly reindex. The
mapping is therefore not documentation after the fact; it is an
executable contract that decides how incoming JSON becomes
searchable data structures.
Explain explicit and dynamic mapping as different schema-governance choices, including what dynamic discovery can and cannot infer safely.
Detect field explosion and account for objects, leaf fields, aliases, multi-fields, and platform-specific computed fields in mapping growth.
Reproduce a mapping conflict, identify the first incompatible value, and choose reject, transform, rename, or reindex rather than mutating history in place.
Use dynamic: strict, dynamic templates, and
flattened-like representations deliberately instead of
treating arbitrary JSON as free-form schema.
Build a mapping review gate that connects every AtlasMart field to concrete filter, sort, aggregation, full-text, geo, or retrieval requirements.
Examples are written against
Elasticsearch 9.5.3 and
OpenSearch 3.8.0. Those products share Lucene
ancestry but are not one API surface. Portable examples use
only behavior verified on both platforms; divergent features
are labeled separately. Elasticsearch examples assume the
default self-managed distribution with its bundled JVM.
OpenSearch examples assume the upstream 3.8.0 distribution
with the Security plugin present. Re-check release notes,
support matrices, plugin compatibility, and feature status
before reproducing this chapter later.
The generation environment used to build this chapter does not
run the two search servers. Commands were checked against
current official documentation but were not executed here, so
output blocks describe expected shape and invariant,
not captured benchmark evidence. Use only disposable AtlasMart
indices. Keep Chapter 01/02 endpoints: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD, OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD, and the course
CA/certificate paths established by the earlier labs.
1. A mapping is executable schema for search behavior
An Elasticsearch or OpenSearch document arrives as JSON, but the
engine does not query raw JSON directly. A
mapping assigns field types and parameters that
determine parsing, indexing, doc-value creation, analysis, and
which query families are legal. An explicit mapping declares
those choices before documents arrive. Dynamic mapping allows
the engine to add fields when it encounters previously unknown
keys. Dynamic discovery is convenient for exploration, but it
cannot infer business intent: an identifier that consists only
of digits may need exact keyword semantics, a
string that looks like a date may not actually be a date
contract, and a numeric-looking price from one feed can conflict
with a later nonnumeric value.
Both engines support dynamic controls on object
mappings. For a governed catalog,
dynamic: "strict" is a useful boundary: an
unexpected field is rejected rather than silently expanding the
schema. dynamic: false is different—the unknown
value can remain in _source while not becoming a
mapped/searchable field. Elasticsearch also has runtime-oriented
dynamic behavior; OpenSearch has its own derived-field system.
Do not copy one product’s computed-field syntax into the other.
| Policy | What happens to a new key | Useful when | Risk if misunderstood |
|---|---|---|---|
dynamic: true |
Engine adds a mapping inferred from the value | Controlled exploratory data | Input shape can become cluster schema |
dynamic: false |
Unknown fields are not added to the mapping | Payload retention without searchability | Users may assume the field is queryable |
dynamic: strict |
Document with an unknown mapped field is rejected | Contract-driven production ingestion | Requires explicit evolution workflow |
| Dynamic template | Unknown fields matching rules receive predetermined mappings | Namespaced extension fields | Over-broad rules can map the wrong data type |
2. Observe mapping growth instead of guessing
Mapping explosion is runaway growth in mapped field definitions,
often caused by user-controlled keys, telemetry labels, or
arbitrary partner attributes. Both current products document a
default index.mapping.total_fields.limit of
1000. That number is a guardrail, not a sizing
target. Raising it does not fix a bad ownership model; it
increases how much mapping metadata the cluster must coordinate
and how much query/schema state clients may have to process.
GET atlasmart-products-v2/_mappingGET atlasmart-products-v2/_field_caps?fields=*GET atlasmart-products-v2/_settings?filter_path=*.settings.index.mapping.*
Count growth during an ingestion test and ask where each field came from. Elasticsearch counts object mappings, field aliases, multi-fields and mapped runtime fields toward its total-fields limit; OpenSearch documents the same core risk and limit but has different platform-specific field systems. Do not turn a shared default into a promise that every counted construct is identical.
3. Reproduce a conflict safely
Create an intentionally disposable index and let dynamic mapping see a numeric price first. The second document violates the established field contract. The important lesson is not the exact exception wording; it is that a field already mapped one way cannot accept an incompatible representation merely because later JSON asks it to.
DELETE atlasmart-mapping-conflict-demoPUT atlasmart-mapping-conflict-demoPOST atlasmart-mapping-conflict-demo/_doc/1{"sku":"SKU-1","price":19.95}POST atlasmart-mapping-conflict-demo/_doc/2{"sku":"SKU-2","price":"unknown"}GET atlasmart-mapping-conflict-demo/_mapping
The second request should fail with a 4xx mapper/parsing error
while the original price mapping remains numeric.
Do not “repair” this by attempting to change
price from numeric to keyword in
place. Existing indexed structures already encode the old type.
Decide whether "unknown" is invalid input,
transform it to a nullable/missing price, split status into
another field, or create a new versioned index and reindex
deliberately.
4. Wrong approach: every supplier key becomes a top-level field
A tempting AtlasMart design stores partner attributes such as
seller_183_color, seller_481_voltage,
and thousands of ad-hoc names as independent fields. It seems
flexible because each key is individually queryable.
Mechanically, however, each new name expands mapping metadata
and can eventually trip the field limit or burden cluster-state
updates and queries.
The defect is not “the field limit is too small.” The defect
is that unbounded external key cardinality owns the cluster
schema. A safe repair starts by deciding which attributes
truly deserve typed first-class fields. Less important
arbitrary key-value bags can use a product-specific flattened
representation: Elasticsearch flattened and
OpenSearch flat_object are conceptually related
but have different indexing/query capabilities, so they must
be tested separately. Another option is to normalize a bounded
attribute vocabulary upstream.
5. AtlasMart governance lab: strict core plus bounded extensions
Create a portable core mapping that makes business intent explicit. The examples use one primary shard and zero replicas only because the goal is mapping semantics on a disposable single-node lab; Chapter 02 already explained why that topology is not a production availability design.
DELETE atlasmart-products-schema-labPUT atlasmart-products-schema-lab{ "settings": {"number_of_shards": 1, "number_of_replicas": 0}, "mappings": { "dynamic": "strict", "properties": { "product_id": {"type":"keyword"}, "name": {"type":"text", "fields":{"raw":{"type":"keyword"}}}, "brand": {"type":"keyword"}, "price": {"type":"double"}, "in_stock": {"type":"boolean"}, "updated_at": {"type":"date"}, "warehouse": {"type":"geo_point"} } }}
POST atlasmart-products-schema-lab/_doc/P-100?refresh=true{"product_id":"P-100","name":"USB-C Travel Charger","brand":"Atlas","price":29.90,"in_stock":true,"updated_at":"2026-09-10T12:00:00Z","warehouse":{"lat":35.6892,"lon":51.3890}}POST atlasmart-products-schema-lab/_doc/P-101{"product_id":"P-101","name":"Unknown Accessory","supplier_magic":"surprise"}GET atlasmart-products-schema-lab/_mappingGET atlasmart-products-schema-lab/_field_caps?fields=*
The first document should index. The second should be rejected
because supplier_magic is outside the strict
contract. Verification means more than receiving an error:
confirm the field did not appear in the mapping, confirm P-100
is searchable, and record the exact response status/error type
for each platform. Then delete only
atlasmart-products-schema-lab and
atlasmart-mapping-conflict-demo.
Check your understanding
- Why is dynamic mapping not equivalent to schema-free storage?
- What does the total-fields limit protect against?
- Why can’t you simply change an existing numeric field to keyword?
-
How do
dynamic: falseanddynamic: strictdiffer? - What evidence should a schema-governance review retain?
Review the answers
1. Because inferred fields become executable index schema with type, indexing, analysis and resource consequences; business intent is still required.
2. Runaway mapping growth and its cluster/query memory and coordination costs; it is a guardrail rather than a target to maximize.
3. Existing indexed structures were built using the old type. Incompatible type changes normally require a new index and reindex/migration path.
4. False ignores new fields for mapping/indexing while strict rejects documents that introduce unexpected mapped fields.
5. The mapping, field-capabilities output, boundary documents, expected failures, query/aggregation requirements, and the approved evolution/reindex path.
Summary and next step
Mapping governance begins by preventing arbitrary input from
choosing search semantics. You now have an explicit contract, a
field-growth signal, a reproducible conflict, and a safe
evolution decision. The next lesson zooms into the most common
mapping choice—text versus keyword—and
proves why the distinction affects matching, sorting, facets,
and memory.
Authoritative references
- Elastic mapping overview — Official mapping guidance and schema-evolution constraints.
- Elastic field data types — Current Elasticsearch field-type reference.
- Elastic mapping limits — Mapping-limit settings and mapping-explosion safeguards.
- OpenSearch mapping documentation — Current OpenSearch mapping entry point.
- OpenSearch supported field types — Current OpenSearch field-type inventory.
- OpenSearch mapping explosion — Field growth risks and mapping-limit controls.