Chapter 08 · Aggregations: Metrics, Buckets, Pipeline Aggregations, Cardinality, and Analytical Search
Design a Faceted Analytics Query and Control Accuracy, Memory, Shard Size, and Response Complexity
Assemble a production-style AtlasMart faceted analytics request with bounded bucket trees, measured response complexity, exact-vs-approximate acceptance tests, composite export pagination, and operational guardrails.
Learning outcomes
The final chapter lab turns aggregation knowledge into an API
contract. AtlasMart needs fast interactive facets and a separate
exhaustive analytical export. Those are different workloads and
should not share an “accept arbitrary aggregation JSON”
endpoint. The interactive path uses a bounded set of approved
facets, explicit approximation labels and response budgets. The
export path uses composite pagination and checkpointing rather
than a huge terms request.
Design a typed faceted analytics request with bounded fields, bucket sizes, ranges and nested depth.
State which outputs are exact, approximate, shard-sensitive or sample-dependent in the API contract.
Separate interactive top-bucket facets from exhaustive composite-pagination exports.
Measure bucket count, response bytes, latency and breaker/resource signals alongside business metrics.
Build a deterministic acceptance suite that detects wrong counts, partial shards and resource-boundary regressions.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0. The mandatory aggregation exercises use free/local REST APIs and portable field types. Aggregation names often look similar across products, but exact algorithms, defaults, error metadata, scripting behavior, circuit-breaker accounting and managed-service limits must be verified per target version. This lesson intentionally avoids universal “safe” bucket counts or heap percentages. Limits must be chosen from measured workload/data/shard/cardinality/resource envelopes. The example application policy is a starting structure only.
The generation environment does not provide live Elasticsearch/OpenSearch clusters, so requests are specified as deterministic labs and expected invariants rather than represented as captured output. Run them against the pinned local course clusters, record actual response sizes/timings/error metadata, and remove only the dedicated AtlasMart Chapter 08 indices. Do not use production data, production scripts, or unbounded bucket trees for the failure exercises.
1. Build a bounded interactive request
DELETE atlasmart-orders-agg-v1
PUT atlasmart-orders-agg-v1
{
"settings": {"number_of_shards": 2, "number_of_replicas": 0},
"mappings": {
"properties": {
"order_id": {"type":"keyword"},
"customer_id": {"type":"keyword"},
"order_date": {"type":"date"},
"region": {"type":"keyword"},
"channel": {"type":"keyword"},
"status": {"type":"keyword"},
"total": {"type":"double"},
"latency_ms": {"type":"long"},
"store": {"properties":{"location":{"type":"geo_point"}}},
"items": {
"type":"nested",
"properties": {
"sku":{"type":"keyword"},
"category":{"type":"keyword"},
"qty":{"type":"integer"},
"line_total":{"type":"double"}
}
}
}
}
}
POST atlasmart-orders-agg-v1/_bulk?refresh=true
{ "index": { "_id": "o1" } }
{ "order_id":"o1","customer_id":"c1","order_date":"2026-09-01T10:00:00Z","region":"north","channel":"web","status":"paid","total":120.0,"latency_ms":80,"store":{"location":{"lat":35.72,"lon":51.41}},"items":[{"sku":"p1","category":"audio","qty":1,"line_total":100.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o2" } }
{ "order_id":"o2","customer_id":"c2","order_date":"2026-09-01T12:00:00Z","region":"south","channel":"mobile","status":"paid","total":80.0,"latency_ms":160,"store":{"location":{"lat":29.61,"lon":52.53}},"items":[{"sku":"p3","category":"audio","qty":1,"line_total":80.0}] }
{ "index": { "_id": "o3" } }
{ "order_id":"o3","customer_id":"c1","order_date":"2026-09-02T08:30:00Z","region":"north","channel":"web","status":"refunded","total":60.0,"latency_ms":250,"store":{"location":{"lat":35.69,"lon":51.39}},"items":[{"sku":"p6","category":"sports","qty":1,"line_total":60.0}] }
{ "index": { "_id": "o4" } }
{ "order_id":"o4","customer_id":"c3","order_date":"2026-09-02T13:20:00Z","region":"west","channel":"store","status":"paid","total":200.0,"latency_ms":110,"store":{"location":{"lat":34.80,"lon":48.51}},"items":[{"sku":"p8","category":"audio","qty":1,"line_total":160.0},{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
{ "index": { "_id": "o5" } }
{ "order_id":"o5","customer_id":"c4","order_date":"2026-09-03T09:15:00Z","region":"north","channel":"mobile","status":"paid","total":150.0,"latency_ms":95,"store":{"location":{"lat":36.26,"lon":59.62}},"items":[{"sku":"p4","category":"audio","qty":1,"line_total":130.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o6" } }
{ "order_id":"o6","customer_id":"c2","order_date":"2026-09-03T18:40:00Z","region":"south","channel":"web","status":"cancelled","total":40.0,"latency_ms":310,"store":{"location":{"lat":31.90,"lon":54.36}},"items":[{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
POST atlasmart-orders-agg-v1/_search
{
"size": 0,
"track_total_hits": true,
"query": {
"bool": {
"filter": [
{"range":{"order_date":{"gte":"2026-09-01","lt":"2026-09-04"}}}
]
}
},
"aggs": {
"regions": {
"terms": {
"field":"region",
"size":10,
"shard_size":20,
"show_term_doc_count_error":true
},
"aggs": {
"revenue":{"sum":{"field":"total"}}
}
},
"channels": {"terms":{"field":"channel","size":10}},
"status": {"terms":{"field":"status","size":10}},
"daily": {
"date_histogram":{"field":"order_date","calendar_interval":"day","min_doc_count":0},
"aggs": {
"revenue":{"sum":{"field":"total"}},
"revenue_change":{"derivative":{"buckets_path":"revenue"}}
}
},
"customers": {"cardinality":{"field":"customer_id","precision_threshold":1000}},
"latency": {"percentiles":{"field":"latency_ms","percents":[50,95,99]}},
"item_scope": {
"nested":{"path":"items"},
"aggs": {
"categories":{"terms":{"field":"items.category","size":10}}
}
}
}
}
Every aggregation has an owner and a reason: top region/channel/status facets, daily revenue trend, approximate unique customers, approximate latency percentiles, and nested item categories. The request does not expose arbitrary field names, scripts, precision or bucket sizes to the caller. Add a query filter for tenant/authorization in the real application before any analytical scope can be evaluated.
2. Publish semantics next to every metric
| Output | Contract to document | User-facing wording |
|---|---|---|
| region/channel/status counts | Top terms; counts can be shard-sensitive on large distributed sets | Top matching regions/channels/statuses; not an exhaustive export. |
| daily revenue sum | Exact over matched field values if all shards succeed | Revenue for indexed/matched orders in this search snapshot. |
| unique customers | Approximate cardinality estimator | Estimated unique customers. |
| latency p95/p99 | Approximate percentile estimator | Estimated percentile over indexed latency values. |
| nested item categories | Counts nested line-item documents | Line-item category occurrences, not unique orders or quantity. |
| derivative | Derived from ordered bucket metric | Change from prior returned time bucket under stated gap semantics. |
This semantic metadata belongs in API documentation and monitoring dashboards. A number without its population, freshness, approximate/exact status and failure state is not a trustworthy analytical product.
3. Put response and memory complexity under policy
Interactive analytics budget (EXAMPLE POLICY — calibrate with real measurements):
- max date range: 31 days
- allowed facet fields: region, channel, status, items.category
- max terms size per facet: 50
- max nested aggregation depth: 2
- max geo precision: application-defined by zoom/viewport
- max composite page size: 500 for background export, not browser UI
- reject arbitrary scripts from end users
- reject arbitrary aggregation JSON
- client timeout + server timeout policy explicitly defined
- record response bytes and returned bucket count
- fail closed on _shards.failed > 0 for contractual analytics
These are policy examples, NOT universal Elasticsearch/OpenSearch tuning values.
Validate parameters before generating Query DSL. Enforce allowlists for aggregation fields and bucket types. Track request/response bytes at the application edge and correlate slow queries with search-node heap/circuit-breaker and GC signals. Circuit-breaker trips are safety signals: reducing cardinality/range/tree depth is usually safer than simply raising the breaker.
A dashboard trips a request breaker, so raise the breaker until the query succeeds without changing the request.
4. Exhaustive analytics needs a different workflow
# Exhaustive key enumeration belongs in a separate paged workflow.
POST atlasmart-orders-agg-v1/_search
{
"size":0,
"aggs": {
"keys": {
"composite": {
"size":100,
"sources":[
{"region":{"terms":{"field":"region"}}},
{"channel":{"terms":{"field":"channel"}}},
{"day":{"date_histogram":{"field":"order_date","calendar_interval":"day"}}}
]
},
"aggs":{"revenue":{"sum":{"field":"total"}}}
}
}
}
# Repeat with "after": <returned after_key> until no buckets remain.
# Persist checkpoint + product/version + index generation so a restart is explicit.
Composite pages stable bucket keys; it does not promise a transactionally frozen multi-page view by itself. If the underlying index changes during a long export, decide whether that freshness is acceptable or whether the architecture requires a stable index generation/snapshot/point-in-time strategy supported by the exact product/version and workload. Never imply snapshot isolation just because pagination is deterministic.
5. Compare exact fixture truth with distributed approximations
For the six-order fixture, validate at minimum:
- sum(total) = 650
- unique exact customers = 4; cardinality estimate should be compared with 4
- regions: north=3, south=2, west=1
- channels: web=3, mobile=2, store=1
- status: paid=4, refunded=1, cancelled=1
- daily revenue: Sep01=200, Sep02=260, Sep03=190
- nested categories: audio=4, office=2, sports=1, accessories=2 nested line docs
- no shard failures
- response JSON size recorded
- terms error metadata inspected rather than ignored
- composite export visits every expected region/channel/day key exactly once
The tiny fixture lets you prove query semantics independently. Then generate a larger deterministic synthetic dataset with known exact distinct counts and skewed categories to study approximation, shard-size bias, response growth and breaker behavior. Record seed, document count, shard count, mapping, product/version and machine resources so later results are comparable.
6. Failure handling is part of analytical correctness
Inspect _shards.total/successful/skipped/failed,
timeout state and per-shard failures. A partial aggregation can
be more dangerous than a failed request because it looks like a
valid number. For contractual analytics, fail closed when a
required shard fails. For best-effort exploratory dashboards,
surface a visible partial-results warning and telemetry; do not
silently present incomplete numbers as authoritative.
if response._shards.failed > 0:
if request.contract == "authoritative":
reject_result("partial shard failure")
else:
mark_result_partial(response._shards)
record(response.took)
record(serialized_response_bytes)
record(total_returned_bucket_count)
record(index_generation, shard_count, server_version)
record(approximation_settings)
7. Decide when the search cluster is the wrong analytics engine
| Signal | Search-cluster response | Architecture option |
|---|---|---|
| Interactive facets over current search result set | Strong fit when bounded | Keep in search request. |
| Millions of exhaustive groups | Poor interactive fit | Composite background export or analytical store. |
| Exact regulatory/accounting DISTINCT | Approximate cardinality mismatch | Exact database/warehouse computation. |
| Long retention + heavy multidimensional scans | Can compete with search SLA | Separate analytical cluster/store/preaggregation. |
| Repeated stable dashboard KPIs | Query-time recomputation waste | Transform/precompute/materialize summaries where appropriate. |
Isolation is a correctness and reliability tool, not an admission of failure. Search and analytics can share an engine when workloads fit, but they should not be forced to share the same latency/resource envelope merely because the API supports aggregations.
Check your understanding
- Why should interactive facets and exhaustive exports use different request designs?
- What must a user be told about cardinality?
- Why treat shard failure as an analytical correctness issue?
- Why not expose arbitrary aggregation JSON to end users?
- When should analytics move off the search cluster?
Review the answers
1. Interactive facets need bounded top buckets and low latency; exhaustive enumeration needs pagination/checkpointing and can consume much more time/resources.
2. It is an estimated unique count under a documented precision/resource setting, not guaranteed exact DISTINCT.
3. Missing shard data can make the final number incomplete while still looking syntactically valid.
4. It enables unbounded bucket/script/geo/range workloads, resource abuse and unstable API semantics.
5. When exactness, exhaustive grouping, retention, scan volume or resource isolation requirements conflict with the search workload/SLO.
Production judgment
A production aggregation API is a contract over population, freshness, exactness, approximation, failure handling and resource bounds. Measure p50/p95/p99 latency with representative cardinalities and concurrency, but keep quality/correctness tests separate from load tests. Version every query template and index generation so regressions can be reproduced and rolled back.
Summary and next step
Chapter 08 now treats aggregations as distributed analytical execution rather than “GROUP BY in JSON”: exact metrics, approximate estimators, shard-sensitive top buckets, nested/global/sample scope, pipeline programs, composite exports and operational limits. Chapter 09 moves those mapping/settings decisions into reusable index/component templates and schema-evolution workflows.
Authoritative references
- Elastic aggregations overview — Official aggregation concepts and API entry point.
- Elastic terms aggregation — Shard candidate collection, shard_size and document-count error behavior.
- Elastic cardinality aggregation — HyperLogLog++ approximation and precision_threshold tradeoffs.
- Elastic percentiles aggregation — Approximate percentile calculation and algorithm controls.
- Elastic composite aggregation — Deterministic bucket pagination with after_key and early-termination guidance.
- Elastic pipeline aggregations — Pipeline categories, bucket paths, derivatives and scripts.
- OpenSearch aggregations overview — Metric, bucket and pipeline aggregation structure and resource considerations.
- OpenSearch terms aggregation — Terms size/shard_size behavior and warnings about inaccurate ascending-count ordering.
- OpenSearch cardinality aggregation — Approximate distinct counts, precision_threshold and collector behavior.
- OpenSearch percentile aggregation — Approximate percentiles and TDigest controls.
- OpenSearch composite aggregation — Composite sources and after-key pagination.
- OpenSearch pipeline aggregations — Supported pipeline aggregations including moving_fn, derivative and bucket_script.
- Elastic circuit breaker settings — Request/parent breaker concepts and why breakers are safety mechanisms.
- OpenSearch circuit breaker settings — OpenSearch breaker configuration and resource protection.