Chapter 08 · Aggregations: Metrics, Buckets, Pipeline Aggregations, Cardinality, and Analytical Search
Bucket Aggregations: terms, range, date histogram, filters, composite, significant terms, and Geo
Build facets and analytical buckets with terms, ranges, date histograms, filters, composite pagination, significance and geo grids while reasoning about shard candidate collection, response size, and approximation.
Learning outcomes
Faceted search is a distributed analytical problem disguised as
a sidebar. AtlasMart wants region/channel facets, order-value
bands, daily revenue, unusual attributes and a way to enumerate
every region/channel combination. These are not interchangeable
bucket types: terms is optimized for top buckets,
composite is the pagination primitive,
range/histogram buckets impose boundaries, significant terms
compare foreground to background, and geo grids trade geographic
precision against bucket count.
Choose terms, range, date_histogram, filters, composite, significant_terms and geo buckets from the analytical question.
Explain how terms candidates are collected per shard and why doc_count can be approximate.
Use composite after_key pagination instead of enormous terms size for exhaustive bucket enumeration.
Distinguish “most common” from “statistically overrepresented” and ordinary categories from geo tiling.
Bound bucket cardinality, response bytes and memory rather than letting UI parameters create unbounded trees.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0. The mandatory aggregation exercises use free/local REST APIs and portable field types. Aggregation names often look similar across products, but exact algorithms, defaults, error metadata, scripting behavior, circuit-breaker accounting and managed-service limits must be verified per target version. For a terms aggregation ordered by descending document count, Elasticsearch exposes document-count error bounds; OpenSearch similarly warns that shard candidate selection can make counts/ordering approximate. Never promise the same error metadata for every order mode or product/version.
The generation environment does not provide live Elasticsearch/OpenSearch clusters, so requests are specified as deterministic labs and expected invariants rather than represented as captured output. Run them against the pinned local course clusters, record actual response sizes/timings/error metadata, and remove only the dedicated AtlasMart Chapter 08 indices. Do not use production data, production scripts, or unbounded bucket trees for the failure exercises.
1. Buckets are partitions of the matched document set
A bucket aggregation assigns documents to groups. Some are
mutually exclusive by construction (carefully designed ranges);
others can overlap (filters); a document with multiple keyword
values can contribute to more than one terms bucket. The bucket
doc_count is a count within that bucket scope, not
necessarily the number of unique business entities.
DELETE atlasmart-orders-agg-v1
PUT atlasmart-orders-agg-v1
{
"settings": {"number_of_shards": 2, "number_of_replicas": 0},
"mappings": {
"properties": {
"order_id": {"type":"keyword"},
"customer_id": {"type":"keyword"},
"order_date": {"type":"date"},
"region": {"type":"keyword"},
"channel": {"type":"keyword"},
"status": {"type":"keyword"},
"total": {"type":"double"},
"latency_ms": {"type":"long"},
"store": {"properties":{"location":{"type":"geo_point"}}},
"items": {
"type":"nested",
"properties": {
"sku":{"type":"keyword"},
"category":{"type":"keyword"},
"qty":{"type":"integer"},
"line_total":{"type":"double"}
}
}
}
}
}
POST atlasmart-orders-agg-v1/_bulk?refresh=true
{ "index": { "_id": "o1" } }
{ "order_id":"o1","customer_id":"c1","order_date":"2026-09-01T10:00:00Z","region":"north","channel":"web","status":"paid","total":120.0,"latency_ms":80,"store":{"location":{"lat":35.72,"lon":51.41}},"items":[{"sku":"p1","category":"audio","qty":1,"line_total":100.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o2" } }
{ "order_id":"o2","customer_id":"c2","order_date":"2026-09-01T12:00:00Z","region":"south","channel":"mobile","status":"paid","total":80.0,"latency_ms":160,"store":{"location":{"lat":29.61,"lon":52.53}},"items":[{"sku":"p3","category":"audio","qty":1,"line_total":80.0}] }
{ "index": { "_id": "o3" } }
{ "order_id":"o3","customer_id":"c1","order_date":"2026-09-02T08:30:00Z","region":"north","channel":"web","status":"refunded","total":60.0,"latency_ms":250,"store":{"location":{"lat":35.69,"lon":51.39}},"items":[{"sku":"p6","category":"sports","qty":1,"line_total":60.0}] }
{ "index": { "_id": "o4" } }
{ "order_id":"o4","customer_id":"c3","order_date":"2026-09-02T13:20:00Z","region":"west","channel":"store","status":"paid","total":200.0,"latency_ms":110,"store":{"location":{"lat":34.80,"lon":48.51}},"items":[{"sku":"p8","category":"audio","qty":1,"line_total":160.0},{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
{ "index": { "_id": "o5" } }
{ "order_id":"o5","customer_id":"c4","order_date":"2026-09-03T09:15:00Z","region":"north","channel":"mobile","status":"paid","total":150.0,"latency_ms":95,"store":{"location":{"lat":36.26,"lon":59.62}},"items":[{"sku":"p4","category":"audio","qty":1,"line_total":130.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o6" } }
{ "order_id":"o6","customer_id":"c2","order_date":"2026-09-03T18:40:00Z","region":"south","channel":"web","status":"cancelled","total":40.0,"latency_ms":310,"store":{"location":{"lat":31.90,"lon":54.36}},"items":[{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
GET atlasmart-orders-agg-v1/_search
{
"size": 0,
"aggs": {
"regions": {
"terms": {
"field": "region",
"size": 10,
"show_term_doc_count_error": true
},
"aggs": {"revenue":{"sum":{"field":"total"}}}
},
"price_bands": {
"range": {
"field":"total",
"ranges":[
{"key":"under_75","to":75},
{"key":"75_to_150","from":75,"to":150},
{"key":"150_plus","from":150}
]
}
},
"daily": {
"date_histogram": {
"field":"order_date",
"calendar_interval":"day",
"min_doc_count":0
},
"aggs":{"revenue":{"sum":{"field":"total"}}}
},
"status_filters": {
"filters": {
"filters": {
"paid":{"term":{"status":"paid"}},
"not_paid":{"bool":{"must_not":{"term":{"status":"paid"}}}}
}
}
}
}
}
2. terms returns top buckets, not a relational GROUP BY oracle
On multiple shards, each shard sends candidate terms to the
coordinating node. If a globally important term falls outside a
shard’s candidate set, the final count can be low. Increasing
shard_size can improve accuracy but consumes
network/memory; increasing size indiscriminately
grows the response and does not turn terms into
safe pagination.
| Need | Use | Do not assume |
|---|---|---|
| Top categories/brands/regions | terms with bounded size and sensible shard_size | Every count is exact under every sort order. |
| Every unique key, page by page | composite + returned after_key | terms size=100000 is pagination. |
| Rare categories | rare_terms where available/appropriate | terms ordered by ascending _count is reliable on shards. |
| Unusual/overrepresented values | significant_terms | The most common term is automatically the most significant. |
Set terms.size to the maximum user-supplied number and return all buckets in one request.
3. Composite is the bucket-pagination primitive
POST atlasmart-orders-agg-v1/_search
{
"size": 0,
"aggs": {
"region_channel": {
"composite": {
"size": 2,
"sources": [
{"region":{"terms":{"field":"region"}}},
{"channel":{"terms":{"field":"channel"}}}
]
},
"aggs":{"revenue":{"sum":{"field":"total"}}}
}
}
}
# For the NEXT page only, copy the returned after_key exactly:
POST atlasmart-orders-agg-v1/_search
{
"size": 0,
"aggs": {
"region_channel": {
"composite": {
"size": 2,
"after": {"region":"<from-after_key>","channel":"<from-after_key>"},
"sources": [
{"region":{"terms":{"field":"region"}}},
{"channel":{"terms":{"field":"channel"}}}
]
}
}
}
}
Use the response’s after_key verbatim for the next
request; do not derive it from the last visible bucket because
implementations can return an after key that is not identical to
that key. Keep the source order stable across pages. Composite
is still expensive on high-cardinality dimensions—page size
bounds response volume, not total work over an export.
4. Range, date histogram and filters encode business boundaries
A price range must define boundary semantics so an order at exactly 75 is not accidentally double-counted or omitted. Date histograms must choose calendar versus fixed intervals deliberately and specify time zone when business days are local rather than UTC. Filters are powerful because categories can overlap, but overlapping buckets should not be summed as though they partitioned the population.
GET atlasmart-orders-agg-v1/_search
{
"size":0,
"aggs": {
"business_day": {
"date_histogram": {
"field":"order_date",
"calendar_interval":"day",
"time_zone":"+03:30"
},
"aggs":{"revenue":{"sum":{"field":"total"}}}
}
}
}
5. Significant terms answers “unusual,” not “largest”
GET atlasmart-orders-agg-v1/_search
{
"size":0,
"query":{"term":{"region":"north"}},
"aggs": {
"unusual_channels": {
"significant_terms": {
"field":"channel",
"background_filter":{"match_all":{}}
}
}
}
}
A significant-term score is meaningful within the
request/heuristic. It compares foreground and background
frequency and can surface small but disproportionately common
values. Tiny foreground sets are unstable; define minimum
evidence and never expose significance as a probability. For
free-text significance, product-specific
significant_text behavior and source re-analysis
cost require separate review.
6. Geo buckets trade map resolution for bucket explosion
GET atlasmart-orders-agg-v1/_search
{
"size":0,
"aggs": {
"stores_by_tile": {
"geotile_grid": {
"field":"store.location",
"precision":5
}
}
}
}
Higher geotile precision creates more, smaller buckets. A UI that lets clients request arbitrary maximum precision over global data can create a resource-abuse path. Bound precision, geographic extent, date range, and maximum returned buckets at the application layer.
7. Measure response complexity as a first-class output
For every facet request record:
- matched document count / time range
- target index and shard count
- bucket type + size + shard_size/precision/interval
- returned bucket count at every tree level
- sum_other_doc_count / doc_count_error metadata where exposed
- response bytes
- took + client end-to-end latency
- _shards.failed and timed_out
- heap/request-breaker signals during representative load
- whether exact enumeration is required (=> composite/export design)
Check your understanding
- Why can terms doc_count be approximate on a multi-shard index?
- What should you use to paginate all unique bucket keys?
- Why is ascending _count terms ordering dangerous?
- What does significant_terms measure?
- Why bound geotile precision?
Review the answers
1. Each shard returns a bounded candidate set; a globally relevant term can be underrepresented if it misses a shard candidate cutoff.
2. Composite aggregation with the returned after_key, keeping source definitions stable.
3. Rare terms are locally/shard dependent, so shard candidate selection can make the global result inaccurate; use purpose-built rare-term semantics where appropriate.
4. Overrepresentation in a foreground set relative to a background set, not raw popularity.
5. Higher precision can explode bucket cardinality, CPU/memory work and response size.
Production judgment
Expose a small, typed facet API rather than arbitrary aggregation JSON. Decide which fields are aggregatable, maximum bucket/page sizes, allowed date ranges and geo precision, and whether counts are contractual or indicative. For exhaustive exports, use composite pagination or an offline analytical path; do not turn the interactive search cluster into an unlimited OLAP endpoint.
Summary and next step
Bucket choice now expresses analytical intent and resource bounds. The next lesson enters nested-document scope and compares query-scoped, global and sampled populations—places where a syntactically valid aggregation can easily answer the wrong business question.
Authoritative references
- Elastic aggregations overview — Official aggregation concepts and API entry point.
- Elastic terms aggregation — Shard candidate collection, shard_size and document-count error behavior.
- Elastic cardinality aggregation — HyperLogLog++ approximation and precision_threshold tradeoffs.
- Elastic percentiles aggregation — Approximate percentile calculation and algorithm controls.
- Elastic composite aggregation — Deterministic bucket pagination with after_key and early-termination guidance.
- Elastic pipeline aggregations — Pipeline categories, bucket paths, derivatives and scripts.
- OpenSearch aggregations overview — Metric, bucket and pipeline aggregation structure and resource considerations.
- OpenSearch terms aggregation — Terms size/shard_size behavior and warnings about inaccurate ascending-count ordering.
- OpenSearch cardinality aggregation — Approximate distinct counts, precision_threshold and collector behavior.
- OpenSearch percentile aggregation — Approximate percentiles and TDigest controls.
- OpenSearch composite aggregation — Composite sources and after-key pagination.
- OpenSearch pipeline aggregations — Supported pipeline aggregations including moving_fn, derivative and bucket_script.
- OpenSearch significant terms — Foreground/background significance semantics and heuristics.
- OpenSearch sampler — Shard-local top-document sampling and limitations.