Chapter 08 · Aggregations: Metrics, Buckets, Pipeline Aggregations, Cardinality, and Analytical Search
Nested/Reverse Nested, Filter/Global, Sampler, and Modeling Implications for Analytics
Keep analytics aligned with the data model: enter and leave nested-document scope correctly, compare query-scoped and global populations, use samplers deliberately, and understand when modeling choices make analytics expensive or misleading.
Learning outcomes
Analytics follow index modeling. AtlasMart’s order contains a
root order document and nested line-item documents. If you
aggregate items.category without entering nested
scope, the request either fails or answers a different shape
than intended. If you then need a root-level customer count, you
must explicitly return to the parent/root context. Similar scope
transitions occur between query-filtered and global populations,
and sampling deliberately changes the population again.
Enter nested scope for nested item analytics and use reverse_nested to return to the root order scope.
Distinguish query-scoped, filter-scoped and global aggregation populations.
Use sampler/diversified-sampler concepts only when approximate top-document analysis matches the question.
Connect denormalization/nested modeling choices to aggregation complexity, hidden document counts and update cost.
Diagnose analytically plausible but semantically wrong bucket counts caused by scope mistakes.
The reproducible examples target self-managed Elasticsearch
9.5.3 and OpenSearch 3.8.0. The mandatory aggregation
exercises use free/local REST APIs and portable field types.
Aggregation names often look similar across products, but
exact algorithms, defaults, error metadata, scripting
behavior, circuit-breaker accounting and managed-service
limits must be verified per target version. The AtlasMart
items field was explicitly mapped as nested in
Chapter 08’s fixture. Nested objects are indexed as separate
hidden documents; nested aggregation doc_count therefore
counts nested documents, not root orders. That distinction is
central to this lesson.
The generation environment does not provide live Elasticsearch/OpenSearch clusters, so requests are specified as deterministic labs and expected invariants rather than represented as captured output. Run them against the pinned local course clusters, record actual response sizes/timings/error metadata, and remove only the dedicated AtlasMart Chapter 08 indices. Do not use production data, production scripts, or unbounded bucket trees for the failure exercises.
1. Nested arrays preserve tuple semantics by becoming hidden documents
Chapter 03 established why arrays of plain objects can lose
per-object tuple relationships. A nested mapping
fixes that by indexing each nested element as an independent
hidden Lucene document linked to the root. The cost is extra
indexed documents and special query/aggregation scope. Analytics
must respect the same model used for correctness at search time.
DELETE atlasmart-orders-agg-v1
PUT atlasmart-orders-agg-v1
{
"settings": {"number_of_shards": 2, "number_of_replicas": 0},
"mappings": {
"properties": {
"order_id": {"type":"keyword"},
"customer_id": {"type":"keyword"},
"order_date": {"type":"date"},
"region": {"type":"keyword"},
"channel": {"type":"keyword"},
"status": {"type":"keyword"},
"total": {"type":"double"},
"latency_ms": {"type":"long"},
"store": {"properties":{"location":{"type":"geo_point"}}},
"items": {
"type":"nested",
"properties": {
"sku":{"type":"keyword"},
"category":{"type":"keyword"},
"qty":{"type":"integer"},
"line_total":{"type":"double"}
}
}
}
}
}
POST atlasmart-orders-agg-v1/_bulk?refresh=true
{ "index": { "_id": "o1" } }
{ "order_id":"o1","customer_id":"c1","order_date":"2026-09-01T10:00:00Z","region":"north","channel":"web","status":"paid","total":120.0,"latency_ms":80,"store":{"location":{"lat":35.72,"lon":51.41}},"items":[{"sku":"p1","category":"audio","qty":1,"line_total":100.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o2" } }
{ "order_id":"o2","customer_id":"c2","order_date":"2026-09-01T12:00:00Z","region":"south","channel":"mobile","status":"paid","total":80.0,"latency_ms":160,"store":{"location":{"lat":29.61,"lon":52.53}},"items":[{"sku":"p3","category":"audio","qty":1,"line_total":80.0}] }
{ "index": { "_id": "o3" } }
{ "order_id":"o3","customer_id":"c1","order_date":"2026-09-02T08:30:00Z","region":"north","channel":"web","status":"refunded","total":60.0,"latency_ms":250,"store":{"location":{"lat":35.69,"lon":51.39}},"items":[{"sku":"p6","category":"sports","qty":1,"line_total":60.0}] }
{ "index": { "_id": "o4" } }
{ "order_id":"o4","customer_id":"c3","order_date":"2026-09-02T13:20:00Z","region":"west","channel":"store","status":"paid","total":200.0,"latency_ms":110,"store":{"location":{"lat":34.80,"lon":48.51}},"items":[{"sku":"p8","category":"audio","qty":1,"line_total":160.0},{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
{ "index": { "_id": "o5" } }
{ "order_id":"o5","customer_id":"c4","order_date":"2026-09-03T09:15:00Z","region":"north","channel":"mobile","status":"paid","total":150.0,"latency_ms":95,"store":{"location":{"lat":36.26,"lon":59.62}},"items":[{"sku":"p4","category":"audio","qty":1,"line_total":130.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o6" } }
{ "order_id":"o6","customer_id":"c2","order_date":"2026-09-03T18:40:00Z","region":"south","channel":"web","status":"cancelled","total":40.0,"latency_ms":310,"store":{"location":{"lat":31.90,"lon":54.36}},"items":[{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
GET atlasmart-orders-agg-v1/_search
{
"size":0,
"aggs": {
"item_scope": {
"nested":{"path":"items"},
"aggs": {
"categories": {
"terms":{"field":"items.category","size":10},
"aggs": {
"line_revenue":{"sum":{"field":"items.line_total"}},
"back_to_orders": {
"reverse_nested": {},
"aggs": {
"unique_customers":{"cardinality":{"field":"customer_id"}}
}
}
}
}
}
}
}
}
2. Read doc_count at the correct level
The item_scope.doc_count is the number of nested
item documents collected, not the number of orders. A category
bucket counts nested line items in that category. After
reverse_nested, doc_count represents
distinct root documents reached from those nested documents. If
AtlasMart needs line quantity rather than line-document count,
it must aggregate items.qty; the bucket count is
not quantity.
| Level | What doc_count means | Common mistake |
|---|---|---|
| Root search/bucket | Root order documents | Calling it line-item count. |
| nested(items) | Nested line-item documents | Calling it unique orders. |
| terms(items.category) inside nested | Nested line items in category | Calling it total quantity sold. |
| reverse_nested back to root | Root orders connected to current nested bucket | Assuming customer count is identical to order count. |
3. Global changes the population on purpose
GET atlasmart-orders-agg-v1/_search
{
"size":0,
"query":{"term":{"region":"north"}},
"aggs": {
"north_revenue":{"sum":{"field":"total"}},
"all_orders": {
"global": {},
"aggs": {
"all_revenue":{"sum":{"field":"total"}},
"all_regions":{"terms":{"field":"region","size":10}}
}
}
}
}
The ordinary north_revenue aggregation inherits the
query’s north-only document set. The top-level
global aggregation ignores that search query and
re-enters all documents in the target index. This is useful for
“filtered result versus overall catalog” comparisons, but it is
dangerous in multi-tenant applications if the application
expected the tenant filter to apply everywhere. Authorization
filters must not be bypassed accidentally by a global
aggregation.
A global aggregation is an analytical scope primitive, not a permission override. Enforce tenant/document authorization at a layer that the request cannot escape, then test global aggregations under the same authorization model.
4. Filter buckets are explicit subpopulations
GET atlasmart-orders-agg-v1/_search
{
"size":0,
"aggs": {
"paid": {
"filter":{"term":{"status":"paid"}},
"aggs":{"revenue":{"sum":{"field":"total"}}}
},
"slow": {
"filter":{"range":{"latency_ms":{"gte":200}}},
"aggs":{"avg_total":{"avg":{"field":"total"}}}
}
}
}
Named filter buckets improve explainability because each metric
has an explicit population. A filters aggregation
can produce multiple named buckets; remember that those filters
may overlap. Sum their counts only if you proved mutual
exclusivity.
5. Sampler is intentional approximation over top documents per shard
GET atlasmart-orders-agg-v1/_search
{
"size":0,
"query":{"range":{"total":{"gte":40}}},
"aggs": {
"sample": {
"sampler":{"shard_size":2},
"aggs": {
"sample_channels":{"terms":{"field":"channel","size":10}}
}
}
}
}
A sampler limits subaggregation work to top-scoring documents on
each shard. It can be valuable for relevance-oriented
significance analysis, but its bucket counts describe the
sample, not the full matching population. On a
match_all-like query with uninformative equal
scores, a sampler is usually a poor business-analytics
primitive. Its output also depends on shard-local top-document
selection.
Use sampler to speed up the revenue dashboard, then label the sampled sum as exact revenue.
6. Modeling determines analytical cost
| Model | Analytics benefit | Cost / caveat |
|---|---|---|
| Denormalized scalar root fields | Fast straightforward filters/terms/metrics | Duplicated data and update propagation. |
| nested arrays | Correct tuple-preserving item queries/aggregations | Extra hidden docs, nested/reverse_nested syntax and cost. |
| Parent/child | Independent parent/child updates | Join-oriented queries/aggregations are operationally heavier; use only with strong justification. |
| Precomputed summary fields | Cheap dashboards and known KPIs | Freshness/backfill/versioning responsibilities move to ingestion. |
| Separate analytical index/warehouse | Isolation and purpose-built schemas | Pipeline freshness, duplication and operational complexity. |
Search-index modeling is not relational normalization by default. Start from query and analytics requirements, write down update frequency and cardinality, and choose the least complex representation that preserves correctness.
7. Validate scope with hand counts
From the six-order fixture:
root orders = 6
nested item documents = 9
orders containing audio item = 4 (o1,o2,o4,o5)
audio nested line-item docs = 4
unique customers among audio orders = 4 (c1,c2,c3,c4)
Acceptance:
- nested category=audio doc_count must be 4
- reverse_nested doc_count under audio must be 4 root orders
- reverse_nested unique_customers under audio should be 4 on this tiny fixture
- all-orders global sum(total) must be 650
- north query root sum(total) must be 330 (o1+o3+o5)
Never replace these semantic assertions with “the request returned HTTP 200.”
Check your understanding
- What does nested aggregation doc_count count?
- Why use reverse_nested?
- What does global do to the query scope?
- Can sampler counts be labeled exact population counts?
- Why can modeling solve an aggregation performance problem better than tuning?
Review the answers
1. Nested documents at that nested path, not root documents or summed quantity.
2. To move from a nested-document bucket back to the root/parent document context for root-level metrics or buckets.
3. It ignores the search query and aggregates over all documents in the target index, subject to the actual security/access layer.
4. No. They describe the sampled top documents selected per shard.
5. Precomputed/denormalized purpose-built fields or an analytical index can eliminate expensive joins/nested calculations at query time.
Production judgment
Make aggregation scope visible in API names and tests: line-item revenue, order count, and unique customers are different measures. Protect tenant filters from global-scope mistakes, measure nested-document growth, and consider precomputation or analytical isolation when user-facing dashboards repeatedly traverse high-cardinality nested structures.
Summary and next step
Chapter analytics now respect document scope. The next lesson treats bucket outputs as an ordered series and derives rates, moving windows and business KPIs with pipeline aggregations—without confusing post-processing of buckets with a cheaper first-pass query.
Authoritative references
- Elastic aggregations overview — Official aggregation concepts and API entry point.
- Elastic terms aggregation — Shard candidate collection, shard_size and document-count error behavior.
- Elastic cardinality aggregation — HyperLogLog++ approximation and precision_threshold tradeoffs.
- Elastic percentiles aggregation — Approximate percentile calculation and algorithm controls.
- Elastic composite aggregation — Deterministic bucket pagination with after_key and early-termination guidance.
- Elastic pipeline aggregations — Pipeline categories, bucket paths, derivatives and scripts.
- OpenSearch aggregations overview — Metric, bucket and pipeline aggregation structure and resource considerations.
- OpenSearch terms aggregation — Terms size/shard_size behavior and warnings about inaccurate ascending-count ordering.
- OpenSearch cardinality aggregation — Approximate distinct counts, precision_threshold and collector behavior.
- OpenSearch percentile aggregation — Approximate percentiles and TDigest controls.
- OpenSearch composite aggregation — Composite sources and after-key pagination.
- OpenSearch pipeline aggregations — Supported pipeline aggregations including moving_fn, derivative and bucket_script.
- OpenSearch nested aggregation — Nested hidden-document scope and nested metric examples.
- OpenSearch reverse_nested — Moving from nested scope back to parent/root documents.
- OpenSearch global aggregation — Top-level aggregation over all target documents regardless of the search query.
- OpenSearch sampler aggregation — Top-scoring shard-local sample semantics and limitations.