Chapter 08 · Aggregations: Metrics, Buckets, Pipeline Aggregations, Cardinality, and Analytical Search

Design a Faceted Analytics Query and Control Accuracy, Memory, Shard Size, and Response Complexity

Assemble a production-style AtlasMart faceted analytics request with bounded bucket trees, measured response complexity, exact-vs-approximate acceptance tests, composite export pagination, and operational guardrails.

Intermediate100–120 minutesDistributed aggregation labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

The final chapter lab turns aggregation knowledge into an API contract. AtlasMart needs fast interactive facets and a separate exhaustive analytical export. Those are different workloads and should not share an “accept arbitrary aggregation JSON” endpoint. The interactive path uses a bounded set of approved facets, explicit approximation labels and response budgets. The export path uses composite pagination and checkpointing rather than a huge terms request.

01

Design a typed faceted analytics request with bounded fields, bucket sizes, ranges and nested depth.

02

State which outputs are exact, approximate, shard-sensitive or sample-dependent in the API contract.

03

Separate interactive top-bucket facets from exhaustive composite-pagination exports.

04

Measure bucket count, response bytes, latency and breaker/resource signals alongside business metrics.

05

Build a deterministic acceptance suite that detects wrong counts, partial shards and resource-boundary regressions.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0. The mandatory aggregation exercises use free/local REST APIs and portable field types. Aggregation names often look similar across products, but exact algorithms, defaults, error metadata, scripting behavior, circuit-breaker accounting and managed-service limits must be verified per target version. This lesson intentionally avoids universal “safe” bucket counts or heap percentages. Limits must be chosen from measured workload/data/shard/cardinality/resource envelopes. The example application policy is a starting structure only.

Execution and safety note

The generation environment does not provide live Elasticsearch/OpenSearch clusters, so requests are specified as deterministic labs and expected invariants rather than represented as captured output. Run them against the pinned local course clusters, record actual response sizes/timings/error metadata, and remove only the dedicated AtlasMart Chapter 08 indices. Do not use production data, production scripts, or unbounded bucket trees for the failure exercises.

1. Build a bounded interactive request

Dev Tools · create the Chapter 08 analytical fixture
DELETE atlasmart-orders-agg-v1
PUT atlasmart-orders-agg-v1
{
  "settings": {"number_of_shards": 2, "number_of_replicas": 0},
  "mappings": {
    "properties": {
      "order_id":    {"type":"keyword"},
      "customer_id": {"type":"keyword"},
      "order_date":  {"type":"date"},
      "region":      {"type":"keyword"},
      "channel":     {"type":"keyword"},
      "status":      {"type":"keyword"},
      "total":       {"type":"double"},
      "latency_ms":  {"type":"long"},
      "store":       {"properties":{"location":{"type":"geo_point"}}},
      "items": {
        "type":"nested",
        "properties": {
          "sku":{"type":"keyword"},
          "category":{"type":"keyword"},
          "qty":{"type":"integer"},
          "line_total":{"type":"double"}
        }
      }
    }
  }
}
Dev Tools · six hand-computable AtlasMart orders
POST atlasmart-orders-agg-v1/_bulk?refresh=true
{ "index": { "_id": "o1" } }
{ "order_id":"o1","customer_id":"c1","order_date":"2026-09-01T10:00:00Z","region":"north","channel":"web","status":"paid","total":120.0,"latency_ms":80,"store":{"location":{"lat":35.72,"lon":51.41}},"items":[{"sku":"p1","category":"audio","qty":1,"line_total":100.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o2" } }
{ "order_id":"o2","customer_id":"c2","order_date":"2026-09-01T12:00:00Z","region":"south","channel":"mobile","status":"paid","total":80.0,"latency_ms":160,"store":{"location":{"lat":29.61,"lon":52.53}},"items":[{"sku":"p3","category":"audio","qty":1,"line_total":80.0}] }
{ "index": { "_id": "o3" } }
{ "order_id":"o3","customer_id":"c1","order_date":"2026-09-02T08:30:00Z","region":"north","channel":"web","status":"refunded","total":60.0,"latency_ms":250,"store":{"location":{"lat":35.69,"lon":51.39}},"items":[{"sku":"p6","category":"sports","qty":1,"line_total":60.0}] }
{ "index": { "_id": "o4" } }
{ "order_id":"o4","customer_id":"c3","order_date":"2026-09-02T13:20:00Z","region":"west","channel":"store","status":"paid","total":200.0,"latency_ms":110,"store":{"location":{"lat":34.80,"lon":48.51}},"items":[{"sku":"p8","category":"audio","qty":1,"line_total":160.0},{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
{ "index": { "_id": "o5" } }
{ "order_id":"o5","customer_id":"c4","order_date":"2026-09-03T09:15:00Z","region":"north","channel":"mobile","status":"paid","total":150.0,"latency_ms":95,"store":{"location":{"lat":36.26,"lon":59.62}},"items":[{"sku":"p4","category":"audio","qty":1,"line_total":130.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o6" } }
{ "order_id":"o6","customer_id":"c2","order_date":"2026-09-03T18:40:00Z","region":"south","channel":"web","status":"cancelled","total":40.0,"latency_ms":310,"store":{"location":{"lat":31.90,"lon":54.36}},"items":[{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
Dev Tools · bounded AtlasMart dashboard/facet request
POST atlasmart-orders-agg-v1/_search
{
  "size": 0,
  "track_total_hits": true,
  "query": {
    "bool": {
      "filter": [
        {"range":{"order_date":{"gte":"2026-09-01","lt":"2026-09-04"}}}
      ]
    }
  },
  "aggs": {
    "regions": {
      "terms": {
        "field":"region",
        "size":10,
        "shard_size":20,
        "show_term_doc_count_error":true
      },
      "aggs": {
        "revenue":{"sum":{"field":"total"}}
      }
    },
    "channels": {"terms":{"field":"channel","size":10}},
    "status": {"terms":{"field":"status","size":10}},
    "daily": {
      "date_histogram":{"field":"order_date","calendar_interval":"day","min_doc_count":0},
      "aggs": {
        "revenue":{"sum":{"field":"total"}},
        "revenue_change":{"derivative":{"buckets_path":"revenue"}}
      }
    },
    "customers": {"cardinality":{"field":"customer_id","precision_threshold":1000}},
    "latency": {"percentiles":{"field":"latency_ms","percents":[50,95,99]}},
    "item_scope": {
      "nested":{"path":"items"},
      "aggs": {
        "categories":{"terms":{"field":"items.category","size":10}}
      }
    }
  }
}

Every aggregation has an owner and a reason: top region/channel/status facets, daily revenue trend, approximate unique customers, approximate latency percentiles, and nested item categories. The request does not expose arbitrary field names, scripts, precision or bucket sizes to the caller. Add a query filter for tenant/authorization in the real application before any analytical scope can be evaluated.

2. Publish semantics next to every metric

Output Contract to document User-facing wording
region/channel/status counts Top terms; counts can be shard-sensitive on large distributed sets Top matching regions/channels/statuses; not an exhaustive export.
daily revenue sum Exact over matched field values if all shards succeed Revenue for indexed/matched orders in this search snapshot.
unique customers Approximate cardinality estimator Estimated unique customers.
latency p95/p99 Approximate percentile estimator Estimated percentile over indexed latency values.
nested item categories Counts nested line-item documents Line-item category occurrences, not unique orders or quantity.
derivative Derived from ordered bucket metric Change from prior returned time bucket under stated gap semantics.

This semantic metadata belongs in API documentation and monitoring dashboards. A number without its population, freshness, approximate/exact status and failure state is not a trustworthy analytical product.

3. Put response and memory complexity under policy

application guardrails · example structure
Interactive analytics budget (EXAMPLE POLICY — calibrate with real measurements):
- max date range: 31 days
- allowed facet fields: region, channel, status, items.category
- max terms size per facet: 50
- max nested aggregation depth: 2
- max geo precision: application-defined by zoom/viewport
- max composite page size: 500 for background export, not browser UI
- reject arbitrary scripts from end users
- reject arbitrary aggregation JSON
- client timeout + server timeout policy explicitly defined
- record response bytes and returned bucket count
- fail closed on _shards.failed > 0 for contractual analytics

These are policy examples, NOT universal Elasticsearch/OpenSearch tuning values.

Validate parameters before generating Query DSL. Enforce allowlists for aggregation fields and bucket types. Track request/response bytes at the application edge and correlate slow queries with search-node heap/circuit-breaker and GC signals. Circuit-breaker trips are safety signals: reducing cardinality/range/tree depth is usually safer than simply raising the breaker.

Deliberately wrong repair

A dashboard trips a request breaker, so raise the breaker until the query succeeds without changing the request.

4. Exhaustive analytics needs a different workflow

Dev Tools · composite export skeleton
# Exhaustive key enumeration belongs in a separate paged workflow.
POST atlasmart-orders-agg-v1/_search
{
  "size":0,
  "aggs": {
    "keys": {
      "composite": {
        "size":100,
        "sources":[
          {"region":{"terms":{"field":"region"}}},
          {"channel":{"terms":{"field":"channel"}}},
          {"day":{"date_histogram":{"field":"order_date","calendar_interval":"day"}}}
        ]
      },
      "aggs":{"revenue":{"sum":{"field":"total"}}}
    }
  }
}

# Repeat with "after": <returned after_key> until no buckets remain.
# Persist checkpoint + product/version + index generation so a restart is explicit.

Composite pages stable bucket keys; it does not promise a transactionally frozen multi-page view by itself. If the underlying index changes during a long export, decide whether that freshness is acceptable or whether the architecture requires a stable index generation/snapshot/point-in-time strategy supported by the exact product/version and workload. Never imply snapshot isolation just because pagination is deterministic.

5. Compare exact fixture truth with distributed approximations

Chapter 08 acceptance oracle
For the six-order fixture, validate at minimum:
- sum(total) = 650
- unique exact customers = 4; cardinality estimate should be compared with 4
- regions: north=3, south=2, west=1
- channels: web=3, mobile=2, store=1
- status: paid=4, refunded=1, cancelled=1
- daily revenue: Sep01=200, Sep02=260, Sep03=190
- nested categories: audio=4, office=2, sports=1, accessories=2 nested line docs
- no shard failures
- response JSON size recorded
- terms error metadata inspected rather than ignored
- composite export visits every expected region/channel/day key exactly once

The tiny fixture lets you prove query semantics independently. Then generate a larger deterministic synthetic dataset with known exact distinct counts and skewed categories to study approximation, shard-size bias, response growth and breaker behavior. Record seed, document count, shard count, mapping, product/version and machine resources so later results are comparable.

6. Failure handling is part of analytical correctness

Inspect _shards.total/successful/skipped/failed, timeout state and per-shard failures. A partial aggregation can be more dangerous than a failed request because it looks like a valid number. For contractual analytics, fail closed when a required shard fails. For best-effort exploratory dashboards, surface a visible partial-results warning and telemetry; do not silently present incomplete numbers as authoritative.

response acceptance pseudocode
if response._shards.failed > 0:
    if request.contract == "authoritative":
        reject_result("partial shard failure")
    else:
        mark_result_partial(response._shards)

record(response.took)
record(serialized_response_bytes)
record(total_returned_bucket_count)
record(index_generation, shard_count, server_version)
record(approximation_settings)

7. Decide when the search cluster is the wrong analytics engine

Signal Search-cluster response Architecture option
Interactive facets over current search result set Strong fit when bounded Keep in search request.
Millions of exhaustive groups Poor interactive fit Composite background export or analytical store.
Exact regulatory/accounting DISTINCT Approximate cardinality mismatch Exact database/warehouse computation.
Long retention + heavy multidimensional scans Can compete with search SLA Separate analytical cluster/store/preaggregation.
Repeated stable dashboard KPIs Query-time recomputation waste Transform/precompute/materialize summaries where appropriate.

Isolation is a correctness and reliability tool, not an admission of failure. Search and analytics can share an engine when workloads fit, but they should not be forced to share the same latency/resource envelope merely because the API supports aggregations.

Check your understanding

  1. Why should interactive facets and exhaustive exports use different request designs?
  2. What must a user be told about cardinality?
  3. Why treat shard failure as an analytical correctness issue?
  4. Why not expose arbitrary aggregation JSON to end users?
  5. When should analytics move off the search cluster?
Review the answers

1. Interactive facets need bounded top buckets and low latency; exhaustive enumeration needs pagination/checkpointing and can consume much more time/resources.

2. It is an estimated unique count under a documented precision/resource setting, not guaranteed exact DISTINCT.

3. Missing shard data can make the final number incomplete while still looking syntactically valid.

4. It enables unbounded bucket/script/geo/range workloads, resource abuse and unstable API semantics.

5. When exactness, exhaustive grouping, retention, scan volume or resource isolation requirements conflict with the search workload/SLO.

Production judgment

A production aggregation API is a contract over population, freshness, exactness, approximation, failure handling and resource bounds. Measure p50/p95/p99 latency with representative cardinalities and concurrency, but keep quality/correctness tests separate from load tests. Version every query template and index generation so regressions can be reproduced and rolled back.

Summary and next step

Chapter 08 now treats aggregations as distributed analytical execution rather than “GROUP BY in JSON”: exact metrics, approximate estimators, shard-sensitive top buckets, nested/global/sample scope, pipeline programs, composite exports and operational limits. Chapter 09 moves those mapping/settings decisions into reusable index/component templates and schema-evolution workflows.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.