Normalize append-heavy logs, control cardinality, and make freshness/retention measurable.

Log Analytics: ECS/OpenTelemetry-Like Schema Discipline, Pipelines, Data Streams, Retention, and High Cardinality

Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.

Intermediate → Advanced140–190 minutesStructured logs/data-stream lab · Chapter 28 · Lesson 02Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · ECS 9.5.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Normalize logs with a stable event schema and distinguish ECS conventions from the OpenTelemetry log data model.

02

Use data streams and ingest pipelines for append-heavy log workloads without pretending every event field is safe to index dynamically.

03

Budget high-cardinality dimensions and field growth before dashboards and aggregations make them operationally expensive.

04

Design retention and lifecycle from investigation/business requirements rather than copying the product-search policy.

05

Measure ingest freshness, storage growth, query latency, and cardinality instead of judging log health from document count alone.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned platform baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The current Elastic Common Schema reference is ECS 9.5.0. The mandatory AtlasMart labs use free/local HTTP APIs and deterministic fixtures. The generation environment did not execute live clusters, so numeric latency, throughput, shard growth, and storage values shown as acceptance criteria are measurement instructions—not fabricated captured results.

1. AtlasMart problem: 40 services disagree about what a log event means

One service emits level, another severity, another logLevel. One records requestId, another traceId. Searching all “errors from checkout in the last 15 minutes” becomes a schema problem before it becomes a query problem. A log analytics platform needs normalization: stable field names, types, timestamps, service identity, trace correlation, and controlled custom attributes.

ECS 9.5.0 is Elastic’s current open schema for event data. OpenTelemetry defines a vendor-neutral telemetry data model and semantic conventions. They overlap in concepts but are not identical physical schemas. A pipeline may map OTel TraceId to an Elasticsearch field such as trace.id; that mapping is an ingestion contract, not evidence that ECS and OTel are the same specification.

2. Choose a small, stable log contract

Concept AtlasMart field Reason
Event time @timestamp Required for time-window search/data-stream behavior.
Service identity service.name Stable dimension for ownership and filtering.
Environment service.environment Separates prod/stage without free-form index sprawl.
Severity log.level Human-readable severity; normalize casing.
Message message Searchable text; do not turn arbitrary JSON into thousands of fields.
Trace context trace.id, span.id Correlates logs with distributed traces.
Dataset event.dataset Identifies event family/source.
Tenant tenant.id Authorization/cost partitioning dimension.
Request identity request.id High-cardinality investigation key, not a default dashboard group-by.

3. Portable data-stream template for the lab

All Chapter 28 labs keep the existing local endpoints and security assumptions: Elasticsearch at https://localhost:9200 with ELASTIC_PASSWORD and the copied CA file atlasmart-es-http-ca; OpenSearch at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's demo certificate trust bypass (-k) is acceptable only for this disposable local lab, never production. The shared Docker network remains atlasmart-search. Lab indices use one primary and zero replicas so a single-node workstation can complete the exercises; production redundancy decisions are deliberately separate.

Create a controlled log template
PUT _index_template/atlasmart-logs-v1
{
  "index_patterns": ["atlasmart-logs-*"],
  "data_stream": {},
  "priority": 200,
  "template": {
    "settings": {
      "number_of_shards": 1,
      "number_of_replicas": 0,
      "mapping.total_fields.limit": 300
    },
    "mappings": {
      "dynamic": false,
      "properties": {
        "@timestamp": {"type":"date"},
        "service.name": {"type":"keyword"},
        "service.environment": {"type":"keyword"},
        "log.level": {"type":"keyword"},
        "message": {"type":"text"},
        "trace.id": {"type":"keyword"},
        "span.id": {"type":"keyword"},
        "event.dataset": {"type":"keyword"},
        "tenant.id": {"type":"keyword"},
        "request.id": {"type":"keyword"}
      }
    }
  }
}
Index normalized application logs
POST atlasmart-logs-app/_doc?refresh=wait_for
{
  "@timestamp":"2026-09-12T12:00:00Z",
  "service.name":"checkout",
  "service.environment":"lab",
  "log.level":"ERROR",
  "message":"Payment provider timed out",
  "trace.id":"4bf92f3577b34da6a3ce929d0e0e4736",
  "span.id":"00f067aa0ba902b7",
  "event.dataset":"atlasmart.checkout",
  "tenant.id":"tenant-a",
  "request.id":"req-0001"
}

POST atlasmart-logs-app/_doc?refresh=wait_for
{
  "@timestamp":"2026-09-12T12:00:02Z",
  "service.name":"checkout",
  "service.environment":"lab",
  "log.level":"INFO",
  "message":"Retry succeeded",
  "trace.id":"4bf92f3577b34da6a3ce929d0e0e4736",
  "span.id":"00f067aa0ba902b7",
  "event.dataset":"atlasmart.checkout",
  "tenant.id":"tenant-a",
  "request.id":"req-0001"
}

dynamic:false is deliberate for the shared lab: unknown JSON remains in _source but does not silently become an indexed field. Production designs can use dynamic templates or flattened/flat-object-like patterns where supported, but the portability and query implications must be explicit.

4. High cardinality is a query-and-memory budget, not a banned property

trace.id and request.id are inherently high-cardinality and still useful because investigators search for exact values. The mistake is making them default dimensions for large terms aggregations, dashboards, or per-value alert state. Measure the number of distinct values and query shapes before enabling expensive aggregations.

Observe cardinality without enumerating every ID
POST atlasmart-logs-app/_search
{
  "size": 0,
  "query": {"range":{"@timestamp":{"gte":"now-15m"}}},
  "aggs": {
    "services": {"terms":{"field":"service.name","size":20}},
    "request_ids": {"cardinality":{"field":"request.id","precision_threshold":1000}}
  }
}

Cardinality aggregation can be approximate; the exactness/cost tradeoff is part of the evidence. Do not turn an approximate estimate into a billing or compliance count.

5. Retention is a requirement, not a default copied from another workload

Product catalog data may live for years and be updated in place. Application logs may need seven, thirty, or ninety days depending on debugging and regulatory needs. Security events may require materially longer retention. Observability traces may be sampled and retained less than logs. Express each policy independently, then implement with Elasticsearch ILM/data tiers or OpenSearch ISM according to Chapter 17. ILM and ISM are not JSON-compatible policy engines.

For this lab, use tiny indices and manual cleanup. Do not fake a 30-day lifecycle by waiting. The lesson requires students to document target retention, estimated daily ingest bytes, expected rollover trigger, and restore requirement.

6. Freshness and ingest failure are first-class log SLOs

A log search that is fast but ten minutes behind is operationally broken for incident response. Track source-event time, ingest/observed time, rejected documents, pipeline failures, bulk partial failures, and data-stream growth. A single “documents indexed per second” graph cannot distinguish a quiet service from a broken collector.

Wrong approach. Accept arbitrary nested JSON with unrestricted dynamic mappings because “logs are schemaless.” A new deployment adds unique keys per request and drives mapping growth. Repair: define stable normalized fields, cap field growth, preserve uncontrolled payload as non-indexed/source content or a deliberately chosen flattened representation, and test the queries that justify each indexed field.

7. Mini lab: normalize, ingest, and measure

  1. Index 100 deterministic events for checkout and catalog-api using a script or repeated bulk fixture.
  2. Record source timestamp and client-observed indexing completion to estimate ingest lag.
  3. Query error counts by service.name; then look up one exact request.id.
  4. Attempt to ingest a document with 50 random top-level keys. Verify they do not become searchable fields under the lab mapping.
  5. Record _cat/indices or stats output for backing-index size and document count. These are lab measurements, not production forecasts.

Check your understanding

  1. Are ECS and the OpenTelemetry log data model the same schema?
  2. Why keep request.id indexed if it is high-cardinality?
  3. What does dynamic:false protect?
  4. Why is log freshness an SLO candidate?
  5. What changes in Lesson 3?
Review the answers

1. No. They model overlapping telemetry concepts but have different specifications and physical field conventions; pipelines may map between them.

2. Exact lookup is valuable; high cardinality is dangerous mainly when query/aggregation state scales with distinct values.

3. It prevents unexpected fields from silently becoming indexed mappings while preserving the original document in _source.

4. An incident query over stale data can be operationally useless even if query latency is low.

5. Security analytics keeps structured events but adds detection logic, alert volume, stricter credentials, investigative search, and often longer retention.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.