Chapter 03 · Mappings and Field Types: keyword, text, Numeric, Date, Geo, Object, Nested, and Runtime Fields
Object vs nested Modeling, Parent/Child/Join Tradeoffs, Denormalization, and Update Cost
Model structured arrays and relationships without false tuple matches, then weigh nested documents, parent/child joins, denormalization, routing, and update amplification.
Learning outcomes
AtlasMart products can have multiple reseller offers and multiple review objects. A plain JSON array looks as if each object remains a tuple, but standard object mapping flattens those values into field-level arrays. A query can therefore combine a seller from one object with a price from another and return a false match. This lesson makes that loss of tuple association observable, then compares nested documents, parent/child joins and deliberate denormalization.
Explain why arrays of ordinary objects do not preserve independent per-object tuple semantics in the inverted index.
Use nested only when the query must keep values
from the same array element together, and measure its
indexing/query cost.
Explain parent/child join as same-index
document relationships that require routing discipline and
add query overhead.
Choose denormalization when read locality and simpler queries justify write amplification and update coordination.
Validate model correctness with paired false-positive/correct nested queries rather than relying on the JSON source shape.
Examples are written against
Elasticsearch 9.5.3 and
OpenSearch 3.8.0. Those products share Lucene
ancestry but are not one API surface. Portable examples use
only behavior verified on both platforms; divergent features
are labeled separately. Elasticsearch examples assume the
default self-managed distribution with its bundled JVM.
OpenSearch examples assume the upstream 3.8.0 distribution
with the Security plugin present. Re-check release notes,
support matrices, plugin compatibility, and feature status
before reproducing this chapter later.
The generation environment used to build this chapter does not
run the two search servers. Commands were checked against
current official documentation but were not executed here, so
output blocks describe expected shape and invariant,
not captured benchmark evidence. Use only disposable AtlasMart
indices. Keep Chapter 01/02 endpoints: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD, OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD, and the course
CA/certificate paths established by the earlier labs.
1. JSON object shape is not tuple-preserving search shape
Consider one product with two reseller offers: seller Alpha
charges 20, seller Beta charges 100. AtlasMart needs “products
where the same offer is from Alpha and below 30.” With
ordinary object mapping, arrays of objects are
flattened into multi-valued fields. The engine can see values
roughly like offers.seller=[Alpha,Beta] and
offers.price=[20,100], but it does not retain the
pair association for an ordinary bool query.
DELETE atlasmart-object-demoPUT atlasmart-object-demo{ "mappings":{"properties":{ "product_id":{"type":"keyword"}, "offers":{"properties":{"seller":{"type":"keyword"},"price":{"type":"double"}}} }}}POST atlasmart-object-demo/_doc/P-1?refresh=true{ "product_id":"P-1", "offers":[ {"seller":"Alpha","price":100}, {"seller":"Beta","price":20} ]}GET atlasmart-object-demo/_search{ "query":{"bool":{"filter":[ {"term":{"offers.seller":"Alpha"}}, {"range":{"offers.price":{"lt":30}}} ]}}}
If P-1 matches, the result is not a search bug; the mapping never promised tuple correlation. The query independently found “Alpha exists” and “some offer below 30 exists.”
2. nested preserves per-object query scope
A nested field indexes each array element as a
separate hidden Lucene document associated with the parent. A
nested query enters that scope, applies predicates to one nested
document, and then returns the parent product. This restores the
semantic requirement at a cost: more Lucene documents, more
complex queries, nested limits, and additional work during
updates because parent/nested block structure is rewritten.
DELETE atlasmart-nested-demoPUT atlasmart-nested-demo{ "mappings":{"properties":{ "product_id":{"type":"keyword"}, "offers":{"type":"nested","properties":{ "seller":{"type":"keyword"},"price":{"type":"double"} }} }}}POST atlasmart-nested-demo/_doc/P-1?refresh=true{ "product_id":"P-1", "offers":[ {"seller":"Alpha","price":100}, {"seller":"Beta","price":20} ]}GET atlasmart-nested-demo/_search{ "query":{"nested":{ "path":"offers", "query":{"bool":{"filter":[ {"term":{"offers.seller":"Alpha"}}, {"range":{"offers.price":{"lt":30}}} ]}} }}}
The repaired query should return zero hits for P-1 because no single nested offer satisfies both predicates. Add an Alpha/20 offer and rerun; only then should the product match. That two-step test proves the tuple contract.
3. Parent/child joins preserve separate document lifecycles—with cost
Both current products expose a join field for
parent/child relationships inside one index. A child must be
routed so it resides with its parent shard; queries such as
has_child or has_parent perform
additional join work. This can help when child entities update
far more frequently than a huge parent and duplicating the
parent on every child would be expensive, but it is not a
relational database hiding inside the search engine.
| Model | Read/query shape | Update shape | Primary risk |
|---|---|---|---|
| Denormalized parent | Simple local document query | Parent rewrite when embedded data changes | Write amplification / stale copies |
nested |
Nested query/aggregation path | Parent block rewritten with nested children | Extra Lucene docs and nested query cost |
Parent/child join |
Join queries across same-shard docs | Child can change independently | Routing discipline and join overhead |
| Separate indices/services | Application-level composition | Independent lifecycles | Cross-system latency/consistency complexity |
For parent/child, the routing key keeps related documents on the same shard. Forgetting routing is not merely a performance regression; it can make the relationship impossible to resolve correctly. Treat routing as a domain invariant and test it during ingestion.
4. Wrong approach: normalize every entity because SQL taught us to
A relational model may keep Product, Brand, Category, Offer and Review in independent tables and join them at read time. Copying that normalization literally into search as many parent/child relations usually fights the engine’s strengths. Search is commonly optimized by indexing the read shape: names, categories, brand facets and other stable display/filter fields can be duplicated into each product document.
Denormalization is not free. If AtlasMart renames a brand, every product copy may need updating. That cost must be compared with search latency, query complexity, index size, refresh/merge work and failure recovery. The right model emerges from read/write ratios and correctness requirements, not ideology.
5. AtlasMart decision lab: offers and reviews
Use three fixtures: ordinary object offers, nested offers, and a small denormalized product document. Record document counts, mapping size, query syntax, correct/incorrect results, and the number of product documents that would need updating when one shared attribute changes. Do not benchmark p99 latency with three documents; this lab establishes semantics.
GET atlasmart-object-demo/_countGET atlasmart-nested-demo/_countGET atlasmart-nested-demo/_mappingGET atlasmart-nested-demo/_search{ "query":{"nested":{ "path":"offers", "query":{"term":{"offers.seller":"Alpha"}}, "inner_hits":{} }}}
Notice that the top-level _count API reports root
documents, while nested fields create additional internal Lucene
documents that affect index size and execution even though they
are not exposed as ordinary top-level hits. Use stats/profile
tools in later chapters to measure that cost under
representative volume.
Check your understanding
- Why can an ordinary object array produce a false positive across two elements?
- What does
nestedchange? - Why is parent/child routing a correctness concern?
- When is denormalization attractive?
- What should you measure before choosing nested or join at scale?
Review the answers
1. Because the subfields are flattened into multi-valued fields and the pairwise association between values from each object is not retained for ordinary queries.
2. Each array element becomes a separate hidden nested document, allowing a nested query to require all predicates to match the same element.
3. Parent and child must be colocated on the same shard for the join relationship to be resolved.
4. When read locality and simple fast queries are more valuable than the write/update amplification needed to maintain duplicated fields.
5. Root/nested/child cardinalities, update frequency, index size, query latency/tail latency, memory, shard distribution, and operational complexity.
Summary and next step
Structured search modeling must preserve exactly the relationships the query needs—no more and no less. You have reproduced an object false positive, repaired it with nested scope, and placed join/denormalization on an explicit cost surface. The final lesson turns all Chapter 03 decisions into a mapping design-and-validation gate for AtlasMart.
Authoritative references
- Elastic mapping overview — Official mapping guidance and schema-evolution constraints.
- Elastic field data types — Current Elasticsearch field-type reference.
- Elastic mapping limits — Mapping-limit settings and mapping-explosion safeguards.
- OpenSearch mapping documentation — Current OpenSearch mapping entry point.
- OpenSearch supported field types — Current OpenSearch field-type inventory.
- OpenSearch mapping explosion — Field growth risks and mapping-limit controls.
- Elastic object field — Object flattening semantics.
- Elastic nested field — Nested document model, querying and limits.
- Elastic join field — Parent/child join semantics and cautions.
- OpenSearch inner hits / parent-child examples — OpenSearch nested and parent/child inner-hit workflows.