Chapter 08 · Aggregations: Metrics, Buckets, Pipeline Aggregations, Cardinality, and Analytical Search

Pipeline Aggregations, Moving Functions, Derivatives, Bucket Scripts, and Time-Series Computation

Compute time-series derivatives, moving functions and business KPIs from bucket outputs while treating bucket paths, gaps, ordering, script cost, and pipeline memory as explicit contracts.

Intermediate100–120 minutesDistributed aggregation labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart wants “daily revenue change,” a two-day moving revenue signal and average revenue per order. Pipeline aggregations compute from aggregation output rather than directly from raw documents. That distinction matters: a pipeline can reduce or transform the final response, but upstream buckets and metrics still had to be collected. A beautiful bucket_sort size:10 at the end does not make an unbounded upstream tree cheap.

01

Classify parent versus sibling pipeline aggregations and resolve buckets_path references correctly.

02

Compute derivatives, moving functions and bucket-script KPIs over ordered date-histogram buckets.

03

Define gap_policy and zero-denominator behavior as business semantics rather than incidental defaults.

04

Explain why pipeline post-processing does not eliminate upstream bucket collection cost.

05

Test pipeline results against a small independent calculation before scaling the time range/cardinality.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0. The mandatory aggregation exercises use free/local REST APIs and portable field types. Aggregation names often look similar across products, but exact algorithms, defaults, error metadata, scripting behavior, circuit-breaker accounting and managed-service limits must be verified per target version. Both target products support core pipeline concepts, but script helpers, deprecated aliases and option details evolve. OpenSearch documents moving_fn as the replacement for deprecated moving_avg; verify the exact supported syntax in each product/version rather than copying one vendor’s script wholesale.

Execution and safety note

The generation environment does not provide live Elasticsearch/OpenSearch clusters, so requests are specified as deterministic labs and expected invariants rather than represented as captured output. Run them against the pinned local course clusters, record actual response sizes/timings/error metadata, and remove only the dedicated AtlasMart Chapter 08 indices. Do not use production data, production scripts, or unbounded bucket trees for the failure exercises.

1. Pipeline aggregations consume bucket outputs

A metric such as daily revenue reads document values. A pipeline aggregation such as a derivative reads the daily revenue values already produced by the parent histogram. Parent pipelines write a result into each bucket; sibling pipelines summarize across a sibling multi-bucket aggregation. The buckets_path is therefore a path through aggregation results, not a field path in _source.

Dev Tools · create the Chapter 08 analytical fixture
DELETE atlasmart-orders-agg-v1
PUT atlasmart-orders-agg-v1
{
  "settings": {"number_of_shards": 2, "number_of_replicas": 0},
  "mappings": {
    "properties": {
      "order_id":    {"type":"keyword"},
      "customer_id": {"type":"keyword"},
      "order_date":  {"type":"date"},
      "region":      {"type":"keyword"},
      "channel":     {"type":"keyword"},
      "status":      {"type":"keyword"},
      "total":       {"type":"double"},
      "latency_ms":  {"type":"long"},
      "store":       {"properties":{"location":{"type":"geo_point"}}},
      "items": {
        "type":"nested",
        "properties": {
          "sku":{"type":"keyword"},
          "category":{"type":"keyword"},
          "qty":{"type":"integer"},
          "line_total":{"type":"double"}
        }
      }
    }
  }
}
Dev Tools · six hand-computable AtlasMart orders
POST atlasmart-orders-agg-v1/_bulk?refresh=true
{ "index": { "_id": "o1" } }
{ "order_id":"o1","customer_id":"c1","order_date":"2026-09-01T10:00:00Z","region":"north","channel":"web","status":"paid","total":120.0,"latency_ms":80,"store":{"location":{"lat":35.72,"lon":51.41}},"items":[{"sku":"p1","category":"audio","qty":1,"line_total":100.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o2" } }
{ "order_id":"o2","customer_id":"c2","order_date":"2026-09-01T12:00:00Z","region":"south","channel":"mobile","status":"paid","total":80.0,"latency_ms":160,"store":{"location":{"lat":29.61,"lon":52.53}},"items":[{"sku":"p3","category":"audio","qty":1,"line_total":80.0}] }
{ "index": { "_id": "o3" } }
{ "order_id":"o3","customer_id":"c1","order_date":"2026-09-02T08:30:00Z","region":"north","channel":"web","status":"refunded","total":60.0,"latency_ms":250,"store":{"location":{"lat":35.69,"lon":51.39}},"items":[{"sku":"p6","category":"sports","qty":1,"line_total":60.0}] }
{ "index": { "_id": "o4" } }
{ "order_id":"o4","customer_id":"c3","order_date":"2026-09-02T13:20:00Z","region":"west","channel":"store","status":"paid","total":200.0,"latency_ms":110,"store":{"location":{"lat":34.80,"lon":48.51}},"items":[{"sku":"p8","category":"audio","qty":1,"line_total":160.0},{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
{ "index": { "_id": "o5" } }
{ "order_id":"o5","customer_id":"c4","order_date":"2026-09-03T09:15:00Z","region":"north","channel":"mobile","status":"paid","total":150.0,"latency_ms":95,"store":{"location":{"lat":36.26,"lon":59.62}},"items":[{"sku":"p4","category":"audio","qty":1,"line_total":130.0},{"sku":"p5","category":"office","qty":1,"line_total":20.0}] }
{ "index": { "_id": "o6" } }
{ "order_id":"o6","customer_id":"c2","order_date":"2026-09-03T18:40:00Z","region":"south","channel":"web","status":"cancelled","total":40.0,"latency_ms":310,"store":{"location":{"lat":31.90,"lon":54.36}},"items":[{"sku":"p7","category":"accessories","qty":2,"line_total":40.0}] }
Pipeline pattern Input Output question
derivative Ordered numeric metric per bucket How much did this metric change from the previous bucket?
moving_fn Sliding window of bucket metric values What is the rolling/smoothed signal?
bucket_script One or more numeric metrics in current bucket What business KPI can be derived from these metrics?
bucket_selector Metrics in current bucket Should this bucket remain in the response?
avg_bucket / sum_bucket / max_bucket Metric across sibling buckets What summary describes all buckets?

2. Build the ordered time-series parent first

Dev Tools · derivative + moving window + business ratio
GET atlasmart-orders-agg-v1/_search
{
  "size":0,
  "aggs": {
    "daily": {
      "date_histogram": {
        "field":"order_date",
        "calendar_interval":"day",
        "min_doc_count":0
      },
      "aggs": {
        "revenue":{"sum":{"field":"total"}},
        "orders":{"value_count":{"field":"order_id"}},
        "revenue_per_order": {
          "bucket_script": {
            "buckets_path":{"r":"revenue","n":"orders"},
            "script":"params.n == 0 ? 0 : params.r / params.n"
          }
        },
        "revenue_change": {
          "derivative":{"buckets_path":"revenue"}
        },
        "moving_revenue": {
          "moving_fn": {
            "buckets_path":"revenue",
            "window":2,
            "script":"MovingFunctions.unweightedAvg(values)"
          }
        }
      }
    }
  }
}

Inspect the raw daily.revenue and daily.orders values before trusting any pipeline output. For this fixture the daily totals are hand-computable. A derivative needs an ordered series; the first bucket has no previous value. A moving window near the beginning contains fewer historical values unless the implementation/window semantics say otherwise.

3. buckets_path is part of the program

Portability review before freezing syntax
Portable concept: parent histogram/date_histogram -> numeric metric -> pipeline over bucket output.

Verify exact syntax on BOTH targets before freezing a lesson/application:
- derivative: buckets_path, gap_policy, optional unit/normalization behavior
- bucket_script: variable map + script language/context
- moving_fn: script functions/window/shift support
- bucket_sort/bucket_selector: post-collection behavior and limits

Do not assume every pipeline option or script helper is identical merely because the aggregation name exists in both products.

Rename an aggregation in a large request and downstream bucket paths can break. Treat aggregation names as stable internal API identifiers, not decorative labels. Unit-test the request body and validate returned paths in integration tests.

Controlled failure

Change buckets_path from revenue to a nonexistent metric in the disposable fixture. Capture the parsing/validation error, restore the correct path, and verify the pipeline values reappear. This is a safe way to learn dependency structure.

4. Gap policy is business semantics

An empty day can mean “zero orders” or “missing telemetry,” and those are not the same. min_doc_count:0 can materialize empty histogram buckets when bounds/range support it; pipeline gap_policy determines how missing metric values are handled. insert_zeros is only correct when zero is the intended meaning. Otherwise it can manufacture a false drop and distort moving averages/derivatives.

thought experiment
Day 1 revenue = 100
Day 2 missing because ingestion failed
Day 3 revenue = 120

If Day 2 is interpreted as zero:
  derivative Day2 = -100
  derivative Day3 = +120
That tells a false business story.

If Day 2 truly had zero orders, those derivatives may be correct.
The pipeline API cannot infer which story is true; ingestion/data-quality semantics must decide.

5. Scripts need bounded inputs and stable ownership

bucket_script is appropriate for simple derived KPIs such as revenue/order when numerator and denominator already exist. It should not become an unrestricted user programming surface. Keep scripts static/versioned, pass data as parameters, handle division by zero, constrain who may submit scripts, and measure CPU/latency. If a KPI is used constantly at large scale, precomputation may be cheaper and easier to govern.

Dev Tools · explicit zero-denominator guard
"revenue_per_order": {
  "bucket_script": {
    "buckets_path": {"r":"revenue", "n":"orders"},
    "script": "params.n == 0 ? null : params.r / params.n"
  }
}

6. Late truncation does not erase upstream cost

resource anti-pattern
ANTI-PATTERN:
1. Create 1-second buckets for a 90-day range.
2. Inside each bucket create 1,000 category buckets.
3. Add several scripted pipeline calculations.
4. Return the entire tree to a browser.

This can create millions of buckets before pipeline post-processing.
A later bucket_sort does NOT refund the work already spent creating upstream buckets.

Pipeline aggregations run after their inputs exist. bucket_sort can reduce what is returned, but parent buckets still had to be built. Control time range, histogram interval, terms cardinality and subaggregation depth before pipeline processing. Monitor the request circuit breaker and search rejection/latency signals; do not raise limits simply to make an unbounded dashboard succeed.

7. Independent oracle for the fixture

small expected series
UTC day        orders    revenue    revenue/order
2026-09-01       2        200           100
2026-09-02       2        260           130
2026-09-03       2        190            95

First derivative of revenue:
2026-09-01: undefined/no previous bucket
2026-09-02: +60
2026-09-03: -70

Two-day trailing unweighted average (check exact window semantics on target):
expected values must be derived from the returned daily revenue series,
not copied from this note without verifying product/version behavior.

Check your understanding

  1. What does buckets_path reference?
  2. Why does bucket_sort not make a huge parent terms aggregation cheap?
  3. When is insert_zeros dangerous?
  4. Why guard a bucket_script denominator?
  5. What should be validated before scaling a moving function?
Review the answers

1. Aggregation output/metric paths, not fields in the original documents.

2. The upstream buckets must already be collected before the pipeline can sort/truncate them.

3. When a missing bucket represents missing/late data rather than a true business zero.

4. Empty/zero-count buckets can otherwise cause invalid or misleading ratios.

5. The parent series, ordering, window/shift/gap semantics and an independent small-fixture result.

Production judgment

Use pipeline aggregations for bounded interactive derivations, not as a substitute for a time-series/warehouse architecture when the request needs millions of buckets or complex analytics. Version scripts and request schemas, measure upstream bucket count and heap pressure, separate missing data from zero, and preserve raw metrics so derived KPIs remain auditable.

Summary and next step

Pipeline aggregations are programs over bucket output. The final lesson combines the chapter into a production-style faceted analytics request with explicit limits, approximation contracts, response-size budgets and an exhaustive composite export path.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.