Correlate logs, metrics, and traces without collapsing distinct telemetry signals into one schema.

Observability: Logs/Metrics/Traces Correlation, Service Maps, APM-Like Workflows, and Cost Controls

Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.

Intermediate → Advanced145–195 minutesTelemetry correlation/APM lab · Chapter 28 · Lesson 04Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · ECS 9.5.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Model logs, metrics, and traces as related but distinct signals with explicit correlation identifiers.

02

Use trace/service identity to move from a user symptom to logs and spans without forcing every signal into one document.

03

Explain service maps/APM-like workflows as derived views whose correctness depends on telemetry and sampling.

04

Control telemetry cost through sampling, cardinality budgets, retention, and signal-specific storage choices.

05

Compare Elastic and OpenSearch observability surfaces without claiming their APM/OTel integrations are interchangeable.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned platform baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The current Elastic Common Schema reference is ECS 9.5.0. The mandatory AtlasMart labs use free/local HTTP APIs and deterministic fixtures. The generation environment did not execute live clusters, so numeric latency, throughput, shard growth, and storage values shown as acceptance criteria are measurement instructions—not fabricated captured results.

1. AtlasMart problem: checkout is slow, but which signal proves why?

A shopper reports a 2.4-second checkout. A log line says “payment timeout.” A trace shows the payment span dominates the request. A metric shows provider error rate rising. These are different observations. The observability architecture must preserve a common service and trace identity while letting each signal use a schema and retention model appropriate to its shape.

OpenTelemetry defines stable concepts such as service.name, trace ID, span ID, log severity, resource attributes, and metric instruments. Elastic commonly maps telemetry into ECS-oriented fields such as trace.id/span.id; OpenSearch observability workflows can ingest OTel through collectors/Data Prepper and build trace/service-map views. Treat those mappings as product/pipeline contracts, not as proof that every backend stores OTel identically.

2. Separate stores, shared correlation keys

Signal Typical document shape Primary questions High-risk cardinality
Logs event/message + attributes What happened? What context/error text exists? request/trace IDs, arbitrary labels
Traces span timing + parent/child + attributes Where did the request spend time? span IDs, URLs if not templated
Metrics time series + dimensions + numeric samples How is behavior changing over time? unbounded labels such as user/request ID

Do not place user ID or request ID into metric labels merely because those fields exist in logs. Metrics work because dimensions stay bounded enough for efficient time-series aggregation.

3. Deterministic correlation fixture

All Chapter 28 labs keep the existing local endpoints and security assumptions: Elasticsearch at https://localhost:9200 with ELASTIC_PASSWORD and the copied CA file atlasmart-es-http-ca; OpenSearch at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's demo certificate trust bypass (-k) is acceptable only for this disposable local lab, never production. The shared Docker network remains atlasmart-search. Lab indices use one primary and zero replicas so a single-node workstation can complete the exercises; production redundancy decisions are deliberately separate.

Create three small signal indices
PUT atlasmart-otel-logs-v1
{"settings":{"number_of_shards":1,"number_of_replicas":0},"mappings":{"dynamic":"strict","properties":{"@timestamp":{"type":"date"},"service.name":{"type":"keyword"},"trace.id":{"type":"keyword"},"span.id":{"type":"keyword"},"log.level":{"type":"keyword"},"message":{"type":"text"},"tenant.id":{"type":"keyword"}}}}

PUT atlasmart-otel-spans-v1
{"settings":{"number_of_shards":1,"number_of_replicas":0},"mappings":{"dynamic":"strict","properties":{"@timestamp":{"type":"date"},"service.name":{"type":"keyword"},"trace.id":{"type":"keyword"},"span.id":{"type":"keyword"},"parent.id":{"type":"keyword"},"span.name":{"type":"keyword"},"duration_ms":{"type":"float"},"status":{"type":"keyword"},"tenant.id":{"type":"keyword"}}}}

PUT atlasmart-otel-metrics-v1
{"settings":{"number_of_shards":1,"number_of_replicas":0},"mappings":{"dynamic":"strict","properties":{"@timestamp":{"type":"date"},"service.name":{"type":"keyword"},"metric.name":{"type":"keyword"},"metric.value":{"type":"double"},"tenant.id":{"type":"keyword"}}}}
Index one correlated incident
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-otel-logs-v1","_id":"L-1"}}
{"@timestamp":"2026-09-12T12:30:01Z","service.name":"checkout","trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa","span.id":"1111111111111111","log.level":"ERROR","message":"Payment provider timed out","tenant.id":"tenant-a"}
{"index":{"_index":"atlasmart-otel-spans-v1","_id":"T-1"}}
{"@timestamp":"2026-09-12T12:30:00Z","service.name":"checkout","trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa","span.id":"1111111111111111","parent.id":"","span.name":"POST /checkout","duration_ms":2400,"status":"ERROR","tenant.id":"tenant-a"}
{"index":{"_index":"atlasmart-otel-spans-v1","_id":"T-2"}}
{"@timestamp":"2026-09-12T12:30:00.100Z","service.name":"payment-client","trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa","span.id":"2222222222222222","parent.id":"1111111111111111","span.name":"POST payment-provider","duration_ms":2200,"status":"ERROR","tenant.id":"tenant-a"}
{"index":{"_index":"atlasmart-otel-metrics-v1","_id":"M-1"}}
{"@timestamp":"2026-09-12T12:30:00Z","service.name":"payment-client","metric.name":"http.client.error_ratio","metric.value":0.18,"tenant.id":"tenant-a"}

4. Move from user symptom to evidence

Trace-first correlation query
POST atlasmart-otel-spans-v1/_search
{
  "size":20,
  "query":{"term":{"trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"}},
  "sort":[{"@timestamp":"asc"}]
}

POST atlasmart-otel-logs-v1/_search
{
  "query":{"term":{"trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"}}
}

POST atlasmart-otel-metrics-v1/_search
{
  "size":0,
  "query":{"bool":{"filter":[
    {"term":{"service.name":"payment-client"}},
    {"range":{"@timestamp":{"gte":"2026-09-12T12:25:00Z","lte":"2026-09-12T12:35:00Z"}}}
  ]}},
  "aggs":{"avg_error_ratio":{"avg":{"field":"metric.value"}}}
}

Expected evidence: the checkout trace contains a long failing payment-client child span and a correlated error log. The single metric sample is not enough to prove a fleet-wide regression; a real incident requires a time series and baseline. Correlation narrows hypotheses—it does not automatically establish causality.

5. Service maps and APM-like workflows

Service maps are derived from trace relationships. Missing spans, sampling, broken context propagation, or inconsistent service.name can make the map incomplete. Current OpenSearch APM documentation combines trace datasets, a service-map index pattern, and Prometheus RED metrics; Elastic APM/Observability offers its own integrated telemetry model and UI. These are operationally similar goals with different storage/UI/plugin contracts.

Use the platform UI as a view over evidence, not as the only copy of the reasoning. Runbooks should preserve raw trace IDs, query time ranges, service names, and screenshots/exports where needed so an incident can be reconstructed.

6. Cost controls are part of observability correctness

Control Benefit Risk if overused
Trace sampling Reduces span volume and network/storage cost Rare failures may be underrepresented.
Log level/filtering Cuts low-value event volume Missing detail during incident.
Metric dimension budget Prevents time-series explosion Too few dimensions can hide localized failures.
Signal-specific retention Aligns cost with investigation value Cross-signal lookback windows no longer overlap.
Aggregation/downsampling Cheaper long-term trends Cannot reconstruct fine-grained events.
Wrong approach. Put every user ID, request ID, URL, exception string, and dynamic label on every metric and retain every trace forever. Repair: classify dimensions, sample traces intentionally, keep high-cardinality IDs in logs/traces where justified, set per-signal retention, and validate that incident questions remain answerable.

7. Mini lab: prove correlation and a missing-context failure

  1. Query the trace and identify the longest child span.
  2. Use the trace ID to retrieve the log event.
  3. Query payment-client metrics for the same time window.
  4. Insert a second log event without trace.id. Demonstrate that service/time filtering can find it but trace correlation cannot.
  5. Write the alert/runbook evidence you would need before declaring the payment provider the root cause.

Check your understanding

  1. Why not store logs, spans, and metrics as one document type?
  2. What does a service map prove?
  3. Why is trace.id useful in logs but dangerous as a metric label?
  4. Does a long child span prove root cause?
  5. What does Lesson 5 decide?
Review the answers

1. They have different shapes, cardinality, query patterns, retention, and update/aggregation semantics even though they share correlation fields.

2. It shows relationships derived from observed/sampled traces; missing telemetry can make it incomplete.

3. Exact correlation benefits logs, while per-request metric dimensions can explode time-series cardinality.

4. It is strong evidence about where time was spent, but causality still needs supporting metrics/logs and comparison to baseline.

5. Whether these four workloads should share a cluster, share only infrastructure, or be physically isolated based on SLO, security, retention, and noisy-neighbor evidence.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.