Chapter 08 · Aggregations: Metrics, Buckets, Pipeline Aggregations, Cardinality, and Analytical Search

Bucket Aggregations: terms, range, date histogram, filters, composite, significant terms, and Geo

Build facets and analytical buckets with terms, ranges, date histograms, filters, composite pagination, significance and geo grids while reasoning about shard candidate collection, response size, and approximation.

Intermediate100–120 minutesDistributed aggregation labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

Faceted search is a distributed analytical problem disguised as a sidebar. AtlasMart wants region/channel facets, order-value bands, daily revenue, unusual attributes and a way to enumerate every region/channel combination. These are not interchangeable bucket types: terms is optimized for top buckets, composite is the pagination primitive, range/histogram buckets impose boundaries, significant terms compare foreground to background, and geo grids trade geographic precision against bucket count.

01

Choose terms, range, date_histogram, filters, composite, significant_terms and geo buckets from the analytical question.

02

Explain how terms candidates are collected per shard and why doc_count can be approximate.

03

Use composite after_key pagination instead of enormous terms size for exhaustive bucket enumeration.

04

Distinguish “most common” from “statistically overrepresented” and ordinary categories from geo tiling.

05

Bound bucket cardinality, response bytes and memory rather than letting UI parameters create unbounded trees.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0. The mandatory aggregation exercises use free/local REST APIs and portable field types. Aggregation names often look similar across products, but exact algorithms, defaults, error metadata, scripting behavior, circuit-breaker accounting and managed-service limits must be verified per target version. For a terms aggregation ordered by descending document count, Elasticsearch exposes document-count error bounds; OpenSearch similarly warns that shard candidate selection can make counts/ordering approximate. Never promise the same error metadata for every order mode or product/version.

Execution and safety note

The generation environment does not provide live Elasticsearch/OpenSearch clusters, so requests are specified as deterministic labs and expected invariants rather than represented as captured output. Run them against the pinned local course clusters, record actual response sizes/timings/error metadata, and remove only the dedicated AtlasMart Chapter 08 indices. Do not use production data, production scripts, or unbounded bucket trees for the failure exercises.

1. Buckets are partitions of the matched document set

A bucket aggregation assigns documents to groups. Some are mutually exclusive by construction (carefully designed ranges); others can overlap (filters); a document with multiple keyword values can contribute to more than one terms bucket. The bucket doc_count is a count within that bucket scope, not necessarily the number of unique business entities.

Dev Tools · create the Chapter 08 analytical fixture
DELETE atlasmart-orders-agg-v1
PUT atlasmart-orders-agg-v1
{
  "settings": {"number_of_shards": 2, "number_of_replicas": 0},
  "mappings": {
    "properties": {
      "order_id":    {"type":"keyword"},
      "customer_id": {"type":"keyword"},
      "order_date":  {"type":"date"},
      "region":      {"type":"keyword"},
      "channel":     {"type":"keyword"},
      "status":      {"type":"keyword"},
      "total":       {"type":"double"},
      "latency_ms":  {"type":"long"},
      "store":       {"properties":{"location":{"type":"geo_point"}}},
      "items": {
        "type":"nested",
        "properties": {
          "sku":{"type":"keyword"},
          "category":{"type":"keyword"},
          "qty":{"type":"integer"},
          "line_total":{"type":"double"}
        }
      }
    }
  }
}
Dev Tools · six hand-computable AtlasMart orders
POST atlasmart-orders-agg-v1/_bulk?refresh=true
{ "index": { "_id": "o1" } }
{ "order_id":"o1","customer_id":"c1","order_date":"2026-09-01T10:00:00Z","region":"north","channel":"web","status":"paid","total":120.0,"latency_ms":80,"store":{"location":{"lat":35.72,"lon":51.41}},"items":[{"sku":"p1","category":"audio","qty":1,"line_total":100.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o2" } }
{ "order_id":"o2","customer_id":"c2","order_date":"2026-09-01T12:00:00Z","region":"south","channel":"mobile","status":"paid","total":80.0,"latency_ms":160,"store":{"location":{"lat":29.61,"lon":52.53}},"items":[{"sku":"p3","category":"audio","qty":1,"line_total":80.0}] }
{ "index": { "_id": "o3" } }
{ "order_id":"o3","customer_id":"c1","order_date":"2026-09-02T08:30:00Z","region":"north","channel":"web","status":"refunded","total":60.0,"latency_ms":250,"store":{"location":{"lat":35.69,"lon":51.39}},"items":[{"sku":"p6","category":"sports","qty":1,"line_total":60.0}] }
{ "index": { "_id": "o4" } }
{ "order_id":"o4","customer_id":"c3","order_date":"2026-09-02T13:20:00Z","region":"west","channel":"store","status":"paid","total":200.0,"latency_ms":110,"store":{"location":{"lat":34.80,"lon":48.51}},"items":[{"sku":"p8","category":"audio","qty":1,"line_total":160.0},{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
{ "index": { "_id": "o5" } }
{ "order_id":"o5","customer_id":"c4","order_date":"2026-09-03T09:15:00Z","region":"north","channel":"mobile","status":"paid","total":150.0,"latency_ms":95,"store":{"location":{"lat":36.26,"lon":59.62}},"items":[{"sku":"p4","category":"audio","qty":1,"line_total":130.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o6" } }
{ "order_id":"o6","customer_id":"c2","order_date":"2026-09-03T18:40:00Z","region":"south","channel":"web","status":"cancelled","total":40.0,"latency_ms":310,"store":{"location":{"lat":31.90,"lon":54.36}},"items":[{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
Dev Tools · bounded facets and time buckets
GET atlasmart-orders-agg-v1/_search
{
  "size": 0,
  "aggs": {
    "regions": {
      "terms": {
        "field": "region",
        "size": 10,
        "show_term_doc_count_error": true
      },
      "aggs": {"revenue":{"sum":{"field":"total"}}}
    },
    "price_bands": {
      "range": {
        "field":"total",
        "ranges":[
          {"key":"under_75","to":75},
          {"key":"75_to_150","from":75,"to":150},
          {"key":"150_plus","from":150}
        ]
      }
    },
    "daily": {
      "date_histogram": {
        "field":"order_date",
        "calendar_interval":"day",
        "min_doc_count":0
      },
      "aggs":{"revenue":{"sum":{"field":"total"}}}
    },
    "status_filters": {
      "filters": {
        "filters": {
          "paid":{"term":{"status":"paid"}},
          "not_paid":{"bool":{"must_not":{"term":{"status":"paid"}}}}
        }
      }
    }
  }
}

2. terms returns top buckets, not a relational GROUP BY oracle

On multiple shards, each shard sends candidate terms to the coordinating node. If a globally important term falls outside a shard’s candidate set, the final count can be low. Increasing shard_size can improve accuracy but consumes network/memory; increasing size indiscriminately grows the response and does not turn terms into safe pagination.

Need Use Do not assume
Top categories/brands/regions terms with bounded size and sensible shard_size Every count is exact under every sort order.
Every unique key, page by page composite + returned after_key terms size=100000 is pagination.
Rare categories rare_terms where available/appropriate terms ordered by ascending _count is reliable on shards.
Unusual/overrepresented values significant_terms The most common term is automatically the most significant.
Deliberately wrong facet

Set terms.size to the maximum user-supplied number and return all buckets in one request.

3. Composite is the bucket-pagination primitive

Dev Tools · page composite buckets
POST atlasmart-orders-agg-v1/_search
{
  "size": 0,
  "aggs": {
    "region_channel": {
      "composite": {
        "size": 2,
        "sources": [
          {"region":{"terms":{"field":"region"}}},
          {"channel":{"terms":{"field":"channel"}}}
        ]
      },
      "aggs":{"revenue":{"sum":{"field":"total"}}}
    }
  }
}

# For the NEXT page only, copy the returned after_key exactly:
POST atlasmart-orders-agg-v1/_search
{
  "size": 0,
  "aggs": {
    "region_channel": {
      "composite": {
        "size": 2,
        "after": {"region":"<from-after_key>","channel":"<from-after_key>"},
        "sources": [
          {"region":{"terms":{"field":"region"}}},
          {"channel":{"terms":{"field":"channel"}}}
        ]
      }
    }
  }
}

Use the response’s after_key verbatim for the next request; do not derive it from the last visible bucket because implementations can return an after key that is not identical to that key. Keep the source order stable across pages. Composite is still expensive on high-cardinality dimensions—page size bounds response volume, not total work over an export.

4. Range, date histogram and filters encode business boundaries

A price range must define boundary semantics so an order at exactly 75 is not accidentally double-counted or omitted. Date histograms must choose calendar versus fixed intervals deliberately and specify time zone when business days are local rather than UTC. Filters are powerful because categories can overlap, but overlapping buckets should not be summed as though they partitioned the population.

Dev Tools · local-time business day example
GET atlasmart-orders-agg-v1/_search
{
  "size":0,
  "aggs": {
    "business_day": {
      "date_histogram": {
        "field":"order_date",
        "calendar_interval":"day",
        "time_zone":"+03:30"
      },
      "aggs":{"revenue":{"sum":{"field":"total"}}}
    }
  }
}

5. Significant terms answers “unusual,” not “largest”

Dev Tools · foreground north versus background all orders
GET atlasmart-orders-agg-v1/_search
{
  "size":0,
  "query":{"term":{"region":"north"}},
  "aggs": {
    "unusual_channels": {
      "significant_terms": {
        "field":"channel",
        "background_filter":{"match_all":{}}
      }
    }
  }
}

A significant-term score is meaningful within the request/heuristic. It compares foreground and background frequency and can surface small but disproportionately common values. Tiny foreground sets are unstable; define minimum evidence and never expose significance as a probability. For free-text significance, product-specific significant_text behavior and source re-analysis cost require separate review.

6. Geo buckets trade map resolution for bucket explosion

Dev Tools · coarse geotile aggregation
GET atlasmart-orders-agg-v1/_search
{
  "size":0,
  "aggs": {
    "stores_by_tile": {
      "geotile_grid": {
        "field":"store.location",
        "precision":5
      }
    }
  }
}

Higher geotile precision creates more, smaller buckets. A UI that lets clients request arbitrary maximum precision over global data can create a resource-abuse path. Bound precision, geographic extent, date range, and maximum returned buckets at the application layer.

7. Measure response complexity as a first-class output

acceptance checklist
For every facet request record:
- matched document count / time range
- target index and shard count
- bucket type + size + shard_size/precision/interval
- returned bucket count at every tree level
- sum_other_doc_count / doc_count_error metadata where exposed
- response bytes
- took + client end-to-end latency
- _shards.failed and timed_out
- heap/request-breaker signals during representative load
- whether exact enumeration is required (=> composite/export design)

Check your understanding

  1. Why can terms doc_count be approximate on a multi-shard index?
  2. What should you use to paginate all unique bucket keys?
  3. Why is ascending _count terms ordering dangerous?
  4. What does significant_terms measure?
  5. Why bound geotile precision?
Review the answers

1. Each shard returns a bounded candidate set; a globally relevant term can be underrepresented if it misses a shard candidate cutoff.

2. Composite aggregation with the returned after_key, keeping source definitions stable.

3. Rare terms are locally/shard dependent, so shard candidate selection can make the global result inaccurate; use purpose-built rare-term semantics where appropriate.

4. Overrepresentation in a foreground set relative to a background set, not raw popularity.

5. Higher precision can explode bucket cardinality, CPU/memory work and response size.

Production judgment

Expose a small, typed facet API rather than arbitrary aggregation JSON. Decide which fields are aggregatable, maximum bucket/page sizes, allowed date ranges and geo precision, and whether counts are contractual or indicative. For exhaustive exports, use composite pagination or an offline analytical path; do not turn the interactive search cluster into an unlimited OLAP endpoint.

Summary and next step

Bucket choice now expresses analytical intent and resource bounds. The next lesson enters nested-document scope and compares query-scoped, global and sampled populations—places where a syntactically valid aggregation can easily answer the wrong business question.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.