Chapter 10 · Data Streams, Time-Series Data, Logs, Metrics, and Append-Heavy Workloads

Logs and Metrics Modeling: High Cardinality, Labels, Structured Events, and Field Governance

Govern logs and metrics as bounded structured contracts: control labels, cardinality, field growth, tenant identity, and query requirements instead of indexing arbitrary observability payloads.

Intermediate100–120 minutesData stream & telemetry labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

Observability pipelines often fail gradually: each team adds “just one more label,” JSON keys become fields, tenant/request identifiers become facets, and dashboards later pay the cost. AtlasMart therefore treats telemetry schema as a governed API with explicit high-cardinality budgets and security boundaries.

01

Separate structured event fields, searchable text, low-cardinality facets, dimensions and opaque identifiers.

02

Estimate and observe cardinality before promoting a label to an aggregation/routing dimension.

03

Prevent mapping explosion from arbitrary labels and nested JSON payloads.

04

Build tenant-aware telemetry queries without relying on client-side filtering for authorization.

05

Define schema-change and retention evidence for logs and metrics independently.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 with the copied CA certificate, OpenSearch on https://localhost:9201 with the disposable demo certificate explicitly treated as local-only, one-node disposable clusters, pinned versions, and no moving latest tags. The shared portable lab deliberately uses a strict labels object with no undeclared children. Real deployments may choose bounded explicit label keys, a flattened-like type, or upstream normalization, but the exact type/API differs across Elasticsearch and OpenSearch and should be version tested.

Execution and safety note

The generation environment does not run the two search servers, so commands below are deterministic lab specifications and expected invariants, not fabricated captured output. Run them only against disposable AtlasMart course resources. Record your actual responses, timestamps, backing-index names, latency, shard counts, and disk usage before drawing operational conclusions.

1. Model by query requirement, not producer convenience

Field class Examples Typical use Governance
Identity/filter service, environment, host_id, tenant_id Filters, group-by, routing candidates keyword/IP/boolean with bounded contract.
Time @timestamp Time windows, histograms, retention reasoning date/date_nanos per product need.
Free text message Full-text troubleshooting text; avoid aggregating raw analyzed text.
Numeric metric duration_ms, request_count, error_count Stats, percentiles, rates numeric type chosen from range/precision semantics.
Opaque event identity event_id, trace_id, request_id Exact lookup/correlation keyword but not automatically a facet/dimension.
User-defined labels merchant/team labels Selective filtering if approved Allowlist/budget; never expand arbitrary keys without policy.

2. Strict mappings expose drift early

Dev Tools · recreate governed stream
PUT _index_template/atlasmart-telemetry-template-v1
{
  "index_patterns": ["atlasmart-telemetry"],
  "priority": 300,
  "data_stream": {},
  "template": {
    "settings": {
      "number_of_shards": 1,
      "number_of_replicas": 0
    },
    "mappings": {
      "dynamic": "strict",
      "properties": {
        "@timestamp": {"type":"date"},
        "event_id": {"type":"keyword"},
        "event_kind": {"type":"keyword"},
        "service": {"type":"keyword"},
        "environment": {"type":"keyword"},
        "tenant_id": {"type":"keyword"},
        "host_id": {"type":"keyword"},
        "level": {"type":"keyword"},
        "message": {"type":"text"},
        "duration_ms": {"type":"double"},
        "request_count": {"type":"long"},
        "error_count": {"type":"long"},
        "labels": {"type":"object", "dynamic":"strict"}
      }
    }
  }
}

PUT _data_stream/atlasmart-telemetry
Expected rejection · undeclared label
POST atlasmart-telemetry/_doc
{
  "@timestamp":"2026-09-11T06:20:00Z",
  "event_id":"evt-label-bad",
  "event_kind":"log",
  "service":"checkout-api",
  "environment":"lab",
  "tenant_id":"tenant-b",
  "host_id":"node-2",
  "level":"WARN",
  "message":"unexpected label payload",
  "duration_ms":21.1,
  "request_count":1,
  "error_count":0,
  "labels":{"request_uuid":"550e8400-e29b-41d4-a716-446655440000"}
}

Because labels is strict and no child is declared, the write should fail rather than silently creating a new searchable field. Capture the root cause as schema-drift evidence. The repair is to update the governed template/version or normalize/drop the label upstream—not to disable safeguards during an incident.

3. High cardinality is a budget, not a forbidden number

Measure candidates
GET atlasmart-telemetry/_search
{
  "size":0,
  "aggs":{
    "event_ids":{"cardinality":{"field":"event_id"}},
    "services":{"cardinality":{"field":"service"}},
    "hosts":{"cardinality":{"field":"host_id"}},
    "tenants":{"cardinality":{"field":"tenant_id"}}
  }
}

cardinality itself is approximate at scale, so use it as reconnaissance, not a legal/compliance count. More importantly, cost depends on how a field is used: filtering by a high-cardinality event ID can be reasonable, while returning a huge terms aggregation over that same field to every dashboard is not.

Wrong approach

“High cardinality is bad” is too crude. The engineering question is: cardinality of what, used for which query, at what concurrency, over how many shards, with what memory/latency budget?

4. Normalize labels upstream

Example normalized event
{
  "@timestamp":"2026-09-11T06:22:00Z",
  "event_id":"evt-1022",
  "event_kind":"log",
  "service":"checkout-api",
  "environment":"lab",
  "tenant_id":"tenant-b",
  "host_id":"node-2",
  "level":"ERROR",
  "message":"payment gateway timeout",
  "duration_ms":2503.0,
  "request_count":1,
  "error_count":1,
  "labels":{}
}

Do not put raw customer headers, arbitrary JSON blobs, access tokens, secrets, personal data, or unbounded query strings into labels. Redact/normalize before indexing and use least-privilege ingest credentials. Search authorization must enforce tenant scope server-side; a UI filter is not a security boundary.

5. Structured logs and metrics need different acceptance tests

Concern Logs Metrics
Primary correctness Message/context survives normalization; timestamp/service/level are correct Numeric meaning, unit, temporality and dimensions are correct
Typical cardinality hazard request IDs, URLs, stacktrace-derived keys container/pod churn, arbitrary labels, user IDs
Search need Full-text + exact filters + time Time filter + dimension filter + numeric aggregation
Retention Often shorter hot-search window; archive may matter May downsample/aggregate; raw resolution retention should be explicit
Late data Common from buffers/agents Can distort windows/rates if silently dropped or double counted

6. Observe schema and field growth

Inventory
GET atlasmart-telemetry/_mapping
GET atlasmart-telemetry/_field_caps?fields=*
GET _cluster/stats?filter_path=indices.mappings,indices.fielddata
GET _data_stream/atlasmart-telemetry/_stats

Track field-count growth as a deployment signal. If fields appear without a reviewed schema change, treat that as drift even when indexing still succeeds. For large environments, store this inventory as a CI/CD artifact or periodic control-plane check.

7. Query safely under tenant scope

Server-side query shape
GET atlasmart-telemetry/_search
{
  "size":0,
  "query":{
    "bool":{
      "filter":[
        {"term":{"tenant_id":"tenant-b"}},
        {"term":{"service":"checkout-api"}},
        {"range":{"@timestamp":{"gte":"now-15m"}}}
      ]
    }
  },
  "aggs":{
    "levels":{"terms":{"field":"level","size":10}},
    "p95_duration":{"percentiles":{"field":"duration_ms","percents":[95]}}
  }
}

Application-side tenant filtering after a broader search is unsafe because unauthorized documents have already crossed the search boundary. Use product security controls and server-side scoped queries appropriate to the deployment.

Check your understanding

  1. Why is event_id allowed as keyword even if it has very high cardinality?
  2. What should happen to an undeclared dynamic label in the strict course schema?
  3. Is cardinality aggregation always exact?
  4. Why is a UI tenant filter insufficient for security?
  5. What is a useful drift signal?
Review the answers

1. Exact lookup/correlation can justify it; high cardinality becomes especially costly when used as a broad aggregation/dimension/routing surface.

2. The write should fail so the producer/schema drift is visible.

3. No. At scale it is approximate and should not be used as a compliance count.

4. The search backend must enforce authorization before returning documents.

5. Unexpected mapping/field-count change without a reviewed schema release.

Production judgment

Create budgets for total mapped fields, approved labels, expected cardinality bands, per-tenant query load, response bucket counts and raw event size. Monitor field-growth rate, circuit-breaker/rejection signals, p95/p99 search latency and storage per day. High-cardinality design is sustainable only when the query surface is intentional.

Summary and next step

Telemetry quality depends on bounded structure, not merely valid JSON. Next we handle the events that violate the happy path: late arrivals, corrections and deletes against older backing indices.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.