Correlate logs, metrics, and traces without collapsing distinct telemetry signals into one schema.
Observability: Logs/Metrics/Traces Correlation, Service Maps, APM-Like Workflows, and Cost Controls
Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.
Learning outcomes
Model logs, metrics, and traces as related but distinct signals with explicit correlation identifiers.
Use trace/service identity to move from a user symptom to logs and spans without forcing every signal into one document.
Explain service maps/APM-like workflows as derived views whose correctness depends on telemetry and sampling.
Control telemetry cost through sampling, cardinality budgets, retention, and signal-specific storage choices.
Compare Elastic and OpenSearch observability surfaces without claiming their APM/OTel integrations are interchangeable.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: checkout is slow, but which signal proves why?
A shopper reports a 2.4-second checkout. A log line says “payment timeout.” A trace shows the payment span dominates the request. A metric shows provider error rate rising. These are different observations. The observability architecture must preserve a common service and trace identity while letting each signal use a schema and retention model appropriate to its shape.
OpenTelemetry defines stable concepts such as
service.name, trace ID, span ID, log severity,
resource attributes, and metric instruments. Elastic commonly
maps telemetry into ECS-oriented fields such as
trace.id/span.id; OpenSearch
observability workflows can ingest OTel through collectors/Data
Prepper and build trace/service-map views. Treat those mappings
as product/pipeline contracts, not as proof that every backend
stores OTel identically.
2. Separate stores, shared correlation keys
| Signal | Typical document shape | Primary questions | High-risk cardinality |
|---|---|---|---|
| Logs | event/message + attributes | What happened? What context/error text exists? | request/trace IDs, arbitrary labels |
| Traces | span timing + parent/child + attributes | Where did the request spend time? | span IDs, URLs if not templated |
| Metrics | time series + dimensions + numeric samples | How is behavior changing over time? | unbounded labels such as user/request ID |
Do not place user ID or request ID into metric labels merely because those fields exist in logs. Metrics work because dimensions stay bounded enough for efficient time-series aggregation.
3. Deterministic correlation fixture
All Chapter 28 labs keep the existing local endpoints and
security assumptions: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD and the copied CA file
atlasmart-es-http-ca; OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's
demo certificate trust bypass (-k) is acceptable
only for this disposable local lab, never production. The shared
Docker network remains atlasmart-search. Lab
indices use one primary and zero replicas so a single-node
workstation can complete the exercises; production redundancy
decisions are deliberately separate.
PUT atlasmart-otel-logs-v1
{"settings":{"number_of_shards":1,"number_of_replicas":0},"mappings":{"dynamic":"strict","properties":{"@timestamp":{"type":"date"},"service.name":{"type":"keyword"},"trace.id":{"type":"keyword"},"span.id":{"type":"keyword"},"log.level":{"type":"keyword"},"message":{"type":"text"},"tenant.id":{"type":"keyword"}}}}
PUT atlasmart-otel-spans-v1
{"settings":{"number_of_shards":1,"number_of_replicas":0},"mappings":{"dynamic":"strict","properties":{"@timestamp":{"type":"date"},"service.name":{"type":"keyword"},"trace.id":{"type":"keyword"},"span.id":{"type":"keyword"},"parent.id":{"type":"keyword"},"span.name":{"type":"keyword"},"duration_ms":{"type":"float"},"status":{"type":"keyword"},"tenant.id":{"type":"keyword"}}}}
PUT atlasmart-otel-metrics-v1
{"settings":{"number_of_shards":1,"number_of_replicas":0},"mappings":{"dynamic":"strict","properties":{"@timestamp":{"type":"date"},"service.name":{"type":"keyword"},"metric.name":{"type":"keyword"},"metric.value":{"type":"double"},"tenant.id":{"type":"keyword"}}}}
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-otel-logs-v1","_id":"L-1"}}
{"@timestamp":"2026-09-12T12:30:01Z","service.name":"checkout","trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa","span.id":"1111111111111111","log.level":"ERROR","message":"Payment provider timed out","tenant.id":"tenant-a"}
{"index":{"_index":"atlasmart-otel-spans-v1","_id":"T-1"}}
{"@timestamp":"2026-09-12T12:30:00Z","service.name":"checkout","trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa","span.id":"1111111111111111","parent.id":"","span.name":"POST /checkout","duration_ms":2400,"status":"ERROR","tenant.id":"tenant-a"}
{"index":{"_index":"atlasmart-otel-spans-v1","_id":"T-2"}}
{"@timestamp":"2026-09-12T12:30:00.100Z","service.name":"payment-client","trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa","span.id":"2222222222222222","parent.id":"1111111111111111","span.name":"POST payment-provider","duration_ms":2200,"status":"ERROR","tenant.id":"tenant-a"}
{"index":{"_index":"atlasmart-otel-metrics-v1","_id":"M-1"}}
{"@timestamp":"2026-09-12T12:30:00Z","service.name":"payment-client","metric.name":"http.client.error_ratio","metric.value":0.18,"tenant.id":"tenant-a"}
4. Move from user symptom to evidence
POST atlasmart-otel-spans-v1/_search
{
"size":20,
"query":{"term":{"trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"}},
"sort":[{"@timestamp":"asc"}]
}
POST atlasmart-otel-logs-v1/_search
{
"query":{"term":{"trace.id":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"}}
}
POST atlasmart-otel-metrics-v1/_search
{
"size":0,
"query":{"bool":{"filter":[
{"term":{"service.name":"payment-client"}},
{"range":{"@timestamp":{"gte":"2026-09-12T12:25:00Z","lte":"2026-09-12T12:35:00Z"}}}
]}},
"aggs":{"avg_error_ratio":{"avg":{"field":"metric.value"}}}
}
Expected evidence: the checkout trace contains a long failing payment-client child span and a correlated error log. The single metric sample is not enough to prove a fleet-wide regression; a real incident requires a time series and baseline. Correlation narrows hypotheses—it does not automatically establish causality.
5. Service maps and APM-like workflows
Service maps are derived from trace relationships. Missing
spans, sampling, broken context propagation, or inconsistent
service.name can make the map incomplete. Current
OpenSearch APM documentation combines trace datasets, a
service-map index pattern, and Prometheus RED metrics; Elastic
APM/Observability offers its own integrated telemetry model and
UI. These are operationally similar goals with different
storage/UI/plugin contracts.
Use the platform UI as a view over evidence, not as the only copy of the reasoning. Runbooks should preserve raw trace IDs, query time ranges, service names, and screenshots/exports where needed so an incident can be reconstructed.
6. Cost controls are part of observability correctness
| Control | Benefit | Risk if overused |
|---|---|---|
| Trace sampling | Reduces span volume and network/storage cost | Rare failures may be underrepresented. |
| Log level/filtering | Cuts low-value event volume | Missing detail during incident. |
| Metric dimension budget | Prevents time-series explosion | Too few dimensions can hide localized failures. |
| Signal-specific retention | Aligns cost with investigation value | Cross-signal lookback windows no longer overlap. |
| Aggregation/downsampling | Cheaper long-term trends | Cannot reconstruct fine-grained events. |
7. Mini lab: prove correlation and a missing-context failure
- Query the trace and identify the longest child span.
- Use the trace ID to retrieve the log event.
- Query payment-client metrics for the same time window.
-
Insert a second log event without
trace.id. Demonstrate that service/time filtering can find it but trace correlation cannot. - Write the alert/runbook evidence you would need before declaring the payment provider the root cause.
Check your understanding
- Why not store logs, spans, and metrics as one document type?
- What does a service map prove?
- Why is trace.id useful in logs but dangerous as a metric label?
- Does a long child span prove root cause?
- What does Lesson 5 decide?
Review the answers
1. They have different shapes, cardinality, query patterns, retention, and update/aggregation semantics even though they share correlation fields.
2. It shows relationships derived from observed/sampled traces; missing telemetry can make it incomplete.
3. Exact correlation benefits logs, while per-request metric dimensions can explode time-series cardinality.
4. It is strong evidence about where time was spent, but causality still needs supporting metrics/logs and comparison to baseline.
5. Whether these four workloads should share a cluster, share only infrastructure, or be physically isolated based on SLO, security, retention, and noisy-neighbor evidence.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Elasticsearch 9.5.3 release notes
- Elastic Common Schema 9.5 reference
- ECS getting started and normalization
- ECS log fields
- Elastic Observability fields and object schemas
- Elastic ECS-formatted application logs
- OpenSearch 3.8 version history
- OpenSearch Security Analytics overview
- OpenSearch Security Analytics detectors
- OpenSearch Security Analytics access control
- OpenSearch APM configuration
- OpenSearch Trace Analytics
- OpenTelemetry logs data model
- OpenTelemetry service semantic conventions