Chapter 10 · Data Streams, Time-Series Data, Logs, Metrics, and Append-Heavy Workloads
Logs and Metrics Modeling: High Cardinality, Labels, Structured Events, and Field Governance
Govern logs and metrics as bounded structured contracts: control labels, cardinality, field growth, tenant identity, and query requirements instead of indexing arbitrary observability payloads.
Learning outcomes
Observability pipelines often fail gradually: each team adds “just one more label,” JSON keys become fields, tenant/request identifiers become facets, and dashboards later pay the cost. AtlasMart therefore treats telemetry schema as a governed API with explicit high-cardinality budgets and security boundaries.
Separate structured event fields, searchable text, low-cardinality facets, dimensions and opaque identifiers.
Estimate and observe cardinality before promoting a label to an aggregation/routing dimension.
Prevent mapping explosion from arbitrary labels and nested JSON payloads.
Build tenant-aware telemetry queries without relying on client-side filtering for authorization.
Define schema-change and retention evidence for logs and metrics independently.
The reproducible examples target self-managed Elasticsearch
9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab
conventions: Elasticsearch on https://localhost:9200 with the
copied CA certificate, OpenSearch on https://localhost:9201
with the disposable demo certificate explicitly treated as
local-only, one-node disposable clusters, pinned versions, and
no moving latest tags. The shared portable lab deliberately
uses a strict labels object with no undeclared
children. Real deployments may choose bounded explicit label
keys, a flattened-like type, or upstream normalization, but
the exact type/API differs across Elasticsearch and OpenSearch
and should be version tested.
The generation environment does not run the two search servers, so commands below are deterministic lab specifications and expected invariants, not fabricated captured output. Run them only against disposable AtlasMart course resources. Record your actual responses, timestamps, backing-index names, latency, shard counts, and disk usage before drawing operational conclusions.
1. Model by query requirement, not producer convenience
| Field class | Examples | Typical use | Governance |
|---|---|---|---|
| Identity/filter | service, environment, host_id, tenant_id | Filters, group-by, routing candidates | keyword/IP/boolean with bounded contract. |
| Time | @timestamp | Time windows, histograms, retention reasoning | date/date_nanos per product need. |
| Free text | message | Full-text troubleshooting | text; avoid aggregating raw analyzed text. |
| Numeric metric | duration_ms, request_count, error_count | Stats, percentiles, rates | numeric type chosen from range/precision semantics. |
| Opaque event identity | event_id, trace_id, request_id | Exact lookup/correlation | keyword but not automatically a facet/dimension. |
| User-defined labels | merchant/team labels | Selective filtering if approved | Allowlist/budget; never expand arbitrary keys without policy. |
2. Strict mappings expose drift early
PUT _index_template/atlasmart-telemetry-template-v1
{
"index_patterns": ["atlasmart-telemetry"],
"priority": 300,
"data_stream": {},
"template": {
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0
},
"mappings": {
"dynamic": "strict",
"properties": {
"@timestamp": {"type":"date"},
"event_id": {"type":"keyword"},
"event_kind": {"type":"keyword"},
"service": {"type":"keyword"},
"environment": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"host_id": {"type":"keyword"},
"level": {"type":"keyword"},
"message": {"type":"text"},
"duration_ms": {"type":"double"},
"request_count": {"type":"long"},
"error_count": {"type":"long"},
"labels": {"type":"object", "dynamic":"strict"}
}
}
}
}
PUT _data_stream/atlasmart-telemetry
POST atlasmart-telemetry/_doc
{
"@timestamp":"2026-09-11T06:20:00Z",
"event_id":"evt-label-bad",
"event_kind":"log",
"service":"checkout-api",
"environment":"lab",
"tenant_id":"tenant-b",
"host_id":"node-2",
"level":"WARN",
"message":"unexpected label payload",
"duration_ms":21.1,
"request_count":1,
"error_count":0,
"labels":{"request_uuid":"550e8400-e29b-41d4-a716-446655440000"}
}
Because labels is strict and no child is declared,
the write should fail rather than silently creating a new
searchable field. Capture the root cause as schema-drift
evidence. The repair is to update the governed template/version
or normalize/drop the label upstream—not to disable safeguards
during an incident.
3. High cardinality is a budget, not a forbidden number
GET atlasmart-telemetry/_search
{
"size":0,
"aggs":{
"event_ids":{"cardinality":{"field":"event_id"}},
"services":{"cardinality":{"field":"service"}},
"hosts":{"cardinality":{"field":"host_id"}},
"tenants":{"cardinality":{"field":"tenant_id"}}
}
}
cardinality itself is approximate at scale, so use
it as reconnaissance, not a legal/compliance count. More
importantly, cost depends on how a field is used: filtering by a
high-cardinality event ID can be reasonable, while returning a
huge terms aggregation over that same field to every dashboard
is not.
“High cardinality is bad” is too crude. The engineering question is: cardinality of what, used for which query, at what concurrency, over how many shards, with what memory/latency budget?
4. Normalize labels upstream
{
"@timestamp":"2026-09-11T06:22:00Z",
"event_id":"evt-1022",
"event_kind":"log",
"service":"checkout-api",
"environment":"lab",
"tenant_id":"tenant-b",
"host_id":"node-2",
"level":"ERROR",
"message":"payment gateway timeout",
"duration_ms":2503.0,
"request_count":1,
"error_count":1,
"labels":{}
}
Do not put raw customer headers, arbitrary JSON blobs, access tokens, secrets, personal data, or unbounded query strings into labels. Redact/normalize before indexing and use least-privilege ingest credentials. Search authorization must enforce tenant scope server-side; a UI filter is not a security boundary.
5. Structured logs and metrics need different acceptance tests
| Concern | Logs | Metrics |
|---|---|---|
| Primary correctness | Message/context survives normalization; timestamp/service/level are correct | Numeric meaning, unit, temporality and dimensions are correct |
| Typical cardinality hazard | request IDs, URLs, stacktrace-derived keys | container/pod churn, arbitrary labels, user IDs |
| Search need | Full-text + exact filters + time | Time filter + dimension filter + numeric aggregation |
| Retention | Often shorter hot-search window; archive may matter | May downsample/aggregate; raw resolution retention should be explicit |
| Late data | Common from buffers/agents | Can distort windows/rates if silently dropped or double counted |
6. Observe schema and field growth
GET atlasmart-telemetry/_mapping
GET atlasmart-telemetry/_field_caps?fields=*
GET _cluster/stats?filter_path=indices.mappings,indices.fielddata
GET _data_stream/atlasmart-telemetry/_stats
Track field-count growth as a deployment signal. If fields appear without a reviewed schema change, treat that as drift even when indexing still succeeds. For large environments, store this inventory as a CI/CD artifact or periodic control-plane check.
7. Query safely under tenant scope
GET atlasmart-telemetry/_search
{
"size":0,
"query":{
"bool":{
"filter":[
{"term":{"tenant_id":"tenant-b"}},
{"term":{"service":"checkout-api"}},
{"range":{"@timestamp":{"gte":"now-15m"}}}
]
}
},
"aggs":{
"levels":{"terms":{"field":"level","size":10}},
"p95_duration":{"percentiles":{"field":"duration_ms","percents":[95]}}
}
}
Application-side tenant filtering after a broader search is unsafe because unauthorized documents have already crossed the search boundary. Use product security controls and server-side scoped queries appropriate to the deployment.
Check your understanding
- Why is event_id allowed as keyword even if it has very high cardinality?
- What should happen to an undeclared dynamic label in the strict course schema?
- Is cardinality aggregation always exact?
- Why is a UI tenant filter insufficient for security?
- What is a useful drift signal?
Review the answers
1. Exact lookup/correlation can justify it; high cardinality becomes especially costly when used as a broad aggregation/dimension/routing surface.
2. The write should fail so the producer/schema drift is visible.
3. No. At scale it is approximate and should not be used as a compliance count.
4. The search backend must enforce authorization before returning documents.
5. Unexpected mapping/field-count change without a reviewed schema release.
Production judgment
Create budgets for total mapped fields, approved labels, expected cardinality bands, per-tenant query load, response bucket counts and raw event size. Monitor field-growth rate, circuit-breaker/rejection signals, p95/p99 search latency and storage per day. High-cardinality design is sustainable only when the query surface is intentional.
Summary and next step
Telemetry quality depends on bounded structure, not merely valid JSON. Next we handle the events that violate the happy path: late arrivals, corrections and deletes against older backing indices.
Authoritative references
- Elastic data streams — Logical stream, hidden backing indices, write index, @timestamp and rollover behavior.
- Elastic use a data stream — Indexing, searching, rollover, and document updates/deletes through backing indices.
- Elastic time series data streams — Elastic-specific TSDS dimensions, metrics and generated time-series identity.
- Elastic time series index settings — index.mode=time_series, routing path, look-back/look-ahead windows and related settings.
- Elastic time-bound indices and dimension routing — Timestamp acceptance windows and dimension-based shard routing.
- Elastic data stream lifecycle — Built-in lifecycle, rollover, retention, downsampling and storage transitions.
- Elastic data stream retention — Effective retention semantics and lifecycle APIs.
- OpenSearch data streams — Backing indexes, timestamp field, rollover, search and ISM integration.
- OpenSearch rollover API — Manual rollover semantics for data streams and aliases.
- OpenSearch modify data stream API — OpenSearch 3.8 experimental backing-index add/remove operation.
- OpenSearch Index State Management — Policy-driven rollover, retention and index-state automation.
- OpenSearch update document API — Document correction semantics when targeting a concrete backing index.