Chapter 03 · Mappings and Field Types: keyword, text, Numeric, Date, Geo, Object, Nested, and Runtime Fields
Design a Mapping from Query/Aggregation Requirements and Validate It Before Bulk Ingestion
Turn AtlasMart search requirements into a versioned mapping, exercise boundary cases, and establish a repeatable validation gate before high-volume ingestion begins.
Learning outcomes
Before AtlasMart sends millions of catalog documents through the Bulk API, the team needs a mapping acceptance gate. The gate starts from user-visible requirements—search by name, filter brand/status, sort price/date, facet category, query offers by same-object seller/price, find nearby warehouses, preserve arbitrary low-value supplier metadata—and proves that the mapping produces those behaviors on both engines. A mapping review that only checks valid JSON is incomplete.
Translate query, filter, sort, aggregation and retrieval requirements into field types and mapping parameters with explicit rationale.
Build a portable core mapping plus separately labeled Elasticsearch/OpenSearch extensions where APIs diverge.
Validate the candidate with mapping inspection, field capabilities, analyze requests, boundary documents, nested correctness tests, sorting and aggregations.
Treat incompatible changes as versioned-index/reindex migrations instead of in-place schema mutation.
Produce a release checklist that can block bulk ingestion when field behavior, dynamic growth, or cross-platform compatibility is unproven.
Examples are written against
Elasticsearch 9.5.3 and
OpenSearch 3.8.0. Those products share Lucene
ancestry but are not one API surface. Portable examples use
only behavior verified on both platforms; divergent features
are labeled separately. Elasticsearch examples assume the
default self-managed distribution with its bundled JVM.
OpenSearch examples assume the upstream 3.8.0 distribution
with the Security plugin present. Re-check release notes,
support matrices, plugin compatibility, and feature status
before reproducing this chapter later.
The generation environment used to build this chapter does not
run the two search servers. Commands were checked against
current official documentation but were not executed here, so
output blocks describe expected shape and invariant,
not captured benchmark evidence. Use only disposable AtlasMart
indices. Keep Chapter 01/02 endpoints: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD, OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD, and the course
CA/certificate paths established by the earlier labs.
1. Start from acceptance requirements, not field names
| Requirement | Mapping decision | Validation evidence |
|---|---|---|
| Natural-language product-name search | name: text |
Analyze tokens + match query judged fixtures |
| Exact brand facet |
brand: keyword with deliberate normalizer
|
Term query + terms aggregation |
| Price range and numeric sort | price: double |
Range query + numeric sort + invalid value rejection |
| Product identity | product_id: keyword |
Exact lookup preserving lexical ID |
| Offer seller and price must match same offer | offers: nested |
False-positive fixture rejected by nested query |
| Warehouse distance filter | warehouse: geo_point |
Geo query on known coordinates |
| Arbitrary supplier metadata | Product-specific flattened-like extension | No unbounded mapping growth; documented query limits |
Each row has a proof. This prevents mapping review from becoming a style debate. The query requirement owns the field choice.
2. Create the portable AtlasMart candidate mapping
DELETE atlasmart-products-v3-candidatePUT atlasmart-products-v3-candidate{ "settings": { "number_of_shards": 1, "number_of_replicas": 0, "analysis": { "normalizer": { "atlas_lower": {"type":"custom","filter":["lowercase","asciifolding"]} } } }, "mappings": { "dynamic":"strict", "properties": { "product_id":{"type":"keyword"}, "name":{"type":"text","fields":{"raw":{"type":"keyword"}}}, "brand":{"type":"keyword","normalizer":"atlas_lower"}, "category":{"type":"keyword"}, "price":{"type":"double"}, "in_stock":{"type":"boolean"}, "updated_at":{"type":"date"}, "warehouse":{"type":"geo_point"}, "offers":{"type":"nested","properties":{ "seller":{"type":"keyword"}, "price":{"type":"double"}, "available":{"type":"boolean"} }} } }}
One shard/zero replicas is a semantics-lab choice inherited from the disposable single-node setup, not a production prescription. Chapter 13 will size shards from workload/capacity evidence.
3. Run a mapping acceptance test suite
POST atlasmart-products-v3-candidate/_bulk?refresh=true{"index":{"_id":"P-501"}}{"product_id":"P-501","name":"USB-C Travel Charger","brand":"ÄTLAS","category":"chargers","price":39.95,"in_stock":true,"updated_at":"2026-09-10T18:00:00Z","warehouse":{"lat":35.6892,"lon":51.3890},"offers":[{"seller":"Alpha","price":50,"available":true},{"seller":"Beta","price":30,"available":true}]}{"index":{"_id":"P-502"}}{"product_id":"P-502","name":"Travel Adapter Pro","brand":"Atlas","category":"adapters","price":59.00,"in_stock":false,"updated_at":"2026-09-09T12:00:00Z","warehouse":{"lat":40.7128,"lon":-74.0060},"offers":[{"seller":"Alpha","price":55,"available":true}]}GET atlasmart-products-v3-candidate/_mappingGET atlasmart-products-v3-candidate/_field_caps?fields=*POST atlasmart-products-v3-candidate/_analyze{"field":"name","text":"USB-C Travel Charger"}GET atlasmart-products-v3-candidate/_search{"query":{"match":{"name":"travel charger"}}}GET atlasmart-products-v3-candidate/_search{"size":0,"aggs":{"brands":{"terms":{"field":"brand"}}}}GET atlasmart-products-v3-candidate/_search{"sort":[{"price":"asc"}],"query":{"match_all":{}}}GET atlasmart-products-v3-candidate/_search{"query":{"nested":{"path":"offers","query":{"bool":{"filter":[{"term":{"offers.seller":"Alpha"}},{"range":{"offers.price":{"lt":40}}}]}}}}}
The nested query must return no P-501 hit even though P-501 contains Alpha and a separate 30-price Beta offer. That is a correctness test. The brand aggregation should group the normalizer-equivalent spellings according to the configured normalization. The numeric sort must order values numerically. Record actual responses independently for Elasticsearch and OpenSearch.
4. Add product-specific extensions without pretending portability
If AtlasMart must retain arbitrary supplier metadata, add a
separate platform mapping extension after testing its query
needs. Elasticsearch can use flattened; OpenSearch
can use flat_object. If semantic vectors are added
later, Elasticsearch will use dense_vector while
OpenSearch uses knn_vector. If a computed field is
needed before reindex, Elasticsearch runtime and OpenSearch
derived fields need separate definitions. Keep those fragments
in platform-labeled configuration rather than hiding divergence
behind copy/paste.
# Elasticsearch mapping extension"supplier_attributes": {"type":"flattened"}# OpenSearch mapping extension"supplier_attributes": {"type":"flat_object"}# Do not merge these two snippets into one supposedly portable request.
5. Controlled failures are release gates
A robust schema suite proves rejection as well as success. Try
an unknown field under dynamic: strict, a string in
price, a malformed IP/geo value in a dedicated
typed fixture, an aggregation on the analyzed
name instead of name.raw, and an
ordinary-object version of the offers model that produces the
false tuple match. Every failure must have an expected class and
a repair path.
POST atlasmart-products-v3-candidate/_doc/P-599{ "product_id":"P-599", "name":"Bad Feed Record", "brand":"Atlas", "category":"test", "price":"unknown", "in_stock":true, "updated_at":"2026-09-10T18:30:00Z", "warehouse":{"lat":35.6892,"lon":51.3890}, "offers":[], "supplier_unreviewed_key":"must not silently become schema"}
The exact first error can depend on parsing order, so do not assert which invalid field the server reports first. The acceptance criterion is that the document is rejected and neither invalid input is silently accepted as a new governed field. Isolate individual invalid cases when you need deterministic per-field assertions.
6. Evolution means versioned decisions
Suppose AtlasMart later decides category needs both
exact facet and full-text search. Adding a compatible
multi-field can be possible only under specific mapping
constraints, and old documents may not automatically have the
newly indexed representation without reindexing. Changing an
established incompatible field type generally requires a new
index. Use a versioned candidate such as
atlasmart-products-v4, reindex or replay source
data, validate query/relevance behavior, then cut over via the
alias/lifecycle process taught in later chapters.
Approve only when every required field has a business owner and query purpose; dynamic policy is explicit; field-count growth is bounded; exact/full-text behavior is tested; sort/aggregation fields use appropriate doc-value surfaces; nested tuple semantics are proven; invalid inputs fail predictably; platform-specific fields are separated; mapping and analyzer snapshots are version-controlled; and a reindex/rollback path exists.
Check your understanding
- What should drive the mapping before bulk ingestion?
- Why is a valid mapping JSON document not enough evidence?
- What is the purpose of the P-501 nested fixture?
- How should Elasticsearch-only and OpenSearch-only field types be represented in a shared codebase?
- When should bulk ingestion be blocked?
Review the answers
1. Concrete user/application requirements for search, filtering, sorting, aggregation, retrieval, validation and update behavior.
2. It does not prove tokenization, exact-match semantics, sorting, aggregation, nested correctness, invalid-input behavior or cross-platform compatibility.
3. It detects cross-object false positives by placing seller Alpha and the low price on different offer elements.
4. As explicitly separated platform configuration with independent tests, not as supposedly portable syntax.
5. Whenever required field behavior, unexpected-field policy, failure handling, mapping growth, or migration/rollback behavior is unproven.
Production judgment
Mapping choices affect storage, heap and cluster-state size, indexing throughput, refresh/merge cost, query correctness, aggregation memory, reindex duration and migration portability. Measure those effects on representative data before calling a schema production-ready. Do not optimize away doc values, norms, indexing or source retention solely from intuition; each saves one resource while potentially moving cost or removing capability elsewhere. Also remember that mapping is a security surface: user-controlled keys, scripts/computed fields and expensive nested/join queries can become resource-abuse paths unless bounded.
Summary and next step
Chapter 03 ends with a mapping that can be defended field by field and a test suite that proves its semantics before scale. Chapter 04 builds on this contract to teach indexing and CRUD behavior, Bulk API partial failures, refresh/durability distinctions, optimistic concurrency, and retry-safe ingestion.
Authoritative references
- Elastic mapping overview — Official mapping guidance and schema-evolution constraints.
- Elastic field data types — Current Elasticsearch field-type reference.
- Elastic mapping limits — Mapping-limit settings and mapping-explosion safeguards.
- OpenSearch mapping documentation — Current OpenSearch mapping entry point.
- OpenSearch supported field types — Current OpenSearch field-type inventory.
- OpenSearch mapping explosion — Field growth risks and mapping-limit controls.
- Elastic nested field — Nested tuple semantics and limits.
- Elastic flattened field — Elasticsearch arbitrary-key mapping option.
- OpenSearch flat object — OpenSearch arbitrary-key flat-object option.
- OpenSearch derived field — OpenSearch computed derived-field behavior.