Chapter 10 · Data Streams, Time-Series Data, Logs, Metrics, and Append-Heavy Workloads

Time-Series Dimensions/Metrics Concepts, Index Modes, Routing, and Compression / Storage Considerations

Model dimensions and metrics for time-series workloads, compare Elastic time-series index mode with portable data-stream patterns, and reason about routing, compression, shard locality, and storage cost.

Intermediate100–120 minutesData stream & telemetry labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart's logs and its metrics are both time stamped, but they have different physical economics. A log event carries free-form text and context; a metric sample repeatedly measures a bounded set of named dimensions. This lesson separates generic data streams from an Elastic-specific time series data stream (TSDS) so you can reason about shard routing and compression without inventing a false cross-platform abstraction.

01

Define time-series dimensions and metrics and distinguish them from arbitrary labels.

02

Explain Elastic index.mode=time_series, time_series_dimension, time_series_metric and routing_path.

03

Avoid claiming OpenSearch has the same TSDS mapping contract when it does not.

04

Reason about dimension-based shard locality, compression opportunities and cardinality cost.

05

Design a portable fallback using an ordinary data stream when product-specific TSDS features are not acceptable.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 with the copied CA certificate, OpenSearch on https://localhost:9201 with the disposable demo certificate explicitly treated as local-only, one-node disposable clusters, pinned versions, and no moving latest tags. Elastic 9.5 supports index.mode=time_series with dimension/metric mappings and time-bound backing indices. OpenSearch 3.8 data streams provide append-oriented backing indices and ISM integration, but this course does not treat Elastic TSDS settings as portable OpenSearch syntax.

Execution and safety note

The generation environment does not run the two search servers, so commands below are deterministic lab specifications and expected invariants, not fabricated captured output. Run them only against disposable AtlasMart course resources. Record your actual responses, timestamps, backing-index names, latency, shard counts, and disk usage before drawing operational conclusions.

1. Dimensions identify a series; metrics vary over time

Field Role Good example Bad example
service Dimension catalog-api Full request URL with unique query string
host_id Dimension node-17 Random request UUID
environment Dimension prod Free-form exception text
request_rate Gauge metric 42.8 requests/s Customer email
request_total Counter metric Monotonic count Arbitrary JSON object

Cardinality matters because each unique dimension combination identifies a distinct time series. A field can be syntactically a keyword yet operationally disastrous as a dimension if almost every event has a unique value.

2. Elastic-specific TSDS mechanics

Elasticsearch 9.5.3 · TSDS template
PUT _index_template/atlasmart-metrics-tsds-v1
{
  "index_patterns":["atlasmart-metrics-tsds"],
  "priority":310,
  "data_stream":{},
  "template":{
    "settings":{
      "index.mode":"time_series",
      "index.routing_path":["service","host_id"]
    },
    "mappings":{
      "properties":{
        "@timestamp":{"type":"date"},
        "service":{"type":"keyword","time_series_dimension":true},
        "host_id":{"type":"keyword","time_series_dimension":true},
        "environment":{"type":"keyword","time_series_dimension":true},
        "request_rate":{"type":"double","time_series_metric":"gauge"},
        "request_total":{"type":"long","time_series_metric":"counter"}
      }
    }
  }
}

Elastic uses dimensions plus timestamp to derive time-series identity and routes documents by dimensions. Current documentation also exposes automatically managed dimension-based routing and a fallback index.routing_path. The relevant point is the mechanism: colocating a series can improve ingest/query locality and storage organization, but a pathological dimension distribution can still create skew.

Timestamp window is part of TSDS semantics

Elastic TSDS backing indices are time bounded. Look-back/look-ahead settings influence which timestamps a current backing index accepts. Late-data policy therefore has to be designed, not guessed.

3. OpenSearch: keep the shared concept, not Elastic-only syntax

OpenSearch 3.8 supports data streams, a required timestamp field (customizable in OpenSearch), backing-index rollover and ISM policies. Do not copy index.mode=time_series, time_series_dimension or time_series_metric into a supposedly portable template unless the target OpenSearch version explicitly documents those exact settings. For the dual-platform AtlasMart lab, ordinary data streams remain the portability baseline.

Portable metrics document
POST atlasmart-telemetry/_doc?refresh=true
{
  "@timestamp":"2026-09-11T06:10:00Z",
  "event_id":"metric-1010",
  "event_kind":"metric",
  "service":"catalog-api",
  "environment":"lab",
  "tenant_id":"tenant-a",
  "host_id":"node-1",
  "level":"INFO",
  "message":"request window",
  "duration_ms":8.7,
  "request_count":150,
  "error_count":3,
  "labels":{}
}

4. Routing can improve locality and still create hot shards

Routing determines which primary shard receives a document. A good routing dimension distributes load while keeping useful query locality. A bad one sends most traffic to one value—for example, routing all events only by environment=prod. Test the actual frequency distribution before making a field part of routing.

Cardinality/distribution reconnaissance
GET atlasmart-telemetry/_search
{
  "size":0,
  "aggs":{
    "services":{"terms":{"field":"service","size":20}},
    "hosts":{"cardinality":{"field":"host_id"}},
    "tenants":{"cardinality":{"field":"tenant_id"}}
  }
}

GET _cat/shards/.ds-atlasmart-telemetry*?v&s=index,shard,prirep

On the tiny one-shard course fixture there is no useful routing benchmark. The lab teaches what to measure; it does not manufacture a throughput win.

5. Compression and storage are emergent, not magic

Repeated dimensions, sorted/time-local data, doc values and segment compression can make metrics compact, but storage depends on mapping, values, source retention, segment state and index codec. Measure store.size, segment counts and representative query latency before/after a design change. Do not extrapolate a production compression ratio from ten course documents.

Inspect physical cost
GET _data_stream/atlasmart-telemetry/_stats
GET _cat/indices/.ds-atlasmart-telemetry*?h=index,docs.count,pri.store.size,store.size
GET .ds-atlasmart-telemetry*/_stats/store,segments,indexing,search

6. High-cardinality budget

Candidate label Expected cardinality Policy
service Tens Safe dimension/facet candidate.
environment Single digits Safe but poor routing key alone.
host_id Hundreds/thousands Useful dimension; monitor churn.
tenant_id Potentially many thousands Security-sensitive; budget query/facet impact and authorization.
request_id Near event count Do not use as metric dimension/facet by default.
raw_url Potentially near event count Normalize route template or store unaggregated text only if required.

7. Portable decision pattern

If you need one schema across Elasticsearch and OpenSearch, use an ordinary data stream with explicit mappings and application-level dimension discipline. If you intentionally specialize for Elasticsearch TSDS, pin the target version and test timestamp windows, duplicate identity, routing and query behavior separately. Portability is a design choice, not an automatic property of similar REST APIs.

Check your understanding

  1. What makes a good time-series dimension?
  2. Is index.mode=time_series portable to OpenSearch?
  3. Why can routing by environment be dangerous?
  4. Can the small lab prove compression gains?
  5. What is the portable fallback?
Review the answers

1. A stable identifying attribute with bounded, meaningful cardinality—not a near-unique event identifier.

2. No. It is taught here as an Elastic-specific capability.

3. If most events have the same value it can concentrate writes on one shard.

4. No. It can show how to inspect storage, but production ratios require representative volume and segment state.

5. An ordinary data stream with explicit mappings and bounded dimension/label governance.

Production judgment

Measure series cardinality, per-key ingest skew, shard utilization, query fan-out, segment/store growth and late-event rejection before changing routing or TSDS settings. Retention reduces long-term storage but does not cure hot-shard design. Keep enough operational headroom for rollover, merge, recovery and snapshots.

Summary and next step

Dimensions and metrics are workload concepts; Elastic TSDS adds a product-specific storage/routing contract. Next we govern the actual event fields and labels so AtlasMart telemetry remains queryable without mapping or cardinality explosion.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.