Chapter 10 · Data Streams, Time-Series Data, Logs, Metrics, and Append-Heavy Workloads
Data Streams, Backing Indices, Write Index, @timestamp, Rollover, and Append-Only Design
Build an append-heavy AtlasMart telemetry stream, trace writes into hidden backing indices, roll the stream, and distinguish logical stream identity from the current physical write index.
Learning outcomes
AtlasMart emits application logs and operational metrics continuously. A single permanent index would grow without a clean write-generation boundary, while date-stamped manual indices would force every producer and consumer to know physical names. A data stream solves the naming problem by giving clients one logical target while the cluster owns a sequence of hidden backing indices.
Distinguish a data stream from its hidden backing indices and identify the current write index.
Explain the timestamp contract, rollover generation, and why appends target only the write index.
Observe a manual rollover and prove searches span old and new backing indices.
Explain why append-oriented design is not the same as “data can never be corrected.”
Choose a data stream only when the write/update pattern and retention model actually fit.
The reproducible examples target self-managed Elasticsearch
9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab
conventions: Elasticsearch on https://localhost:9200 with the
copied CA certificate, OpenSearch on https://localhost:9201
with the disposable demo certificate explicitly treated as
local-only, one-node disposable clusters, pinned versions, and
no moving latest tags. Elastic requires
@timestamp for ordinary data streams. OpenSearch
defaults to @timestamp but also supports a custom
timestamp field declared inside the template's
data_stream.timestamp_field; that custom-field
capability is therefore not presented as a portable
Elastic/OpenSearch contract.
The generation environment does not run the two search servers, so commands below are deterministic lab specifications and expected invariants, not fabricated captured output. Run them only against disposable AtlasMart course resources. Record your actual responses, timestamps, backing-index names, latency, shard counts, and disk usage before drawing operational conclusions.
1. Logical stream, physical generations
A data stream is a logical name backed by one or more hidden indices. The newest backing index is the write index. Indexing to the stream routes a new document to that write index; searching the stream expands to all backing indices. A rollover creates a new backing index, increments the stream generation, and moves the write role forward. The old backing index remains searchable and can later be managed by lifecycle policy.
| Construct | Meaning | AtlasMart consequence |
|---|---|---|
| Data stream | Stable logical target for append-heavy time-series data | Producers write to atlasmart-telemetry; dashboards search the same name. |
| Backing index | Hidden physical index belonging to the stream | Shard/storage state is inspected here; clients normally do not hard-code its name. |
| Write index | Newest backing index receiving ordinary stream writes | Rollover changes it without changing producer configuration. |
| Timestamp field | Field required to place an event on the time axis | AtlasMart standardizes on @timestamp even though OpenSearch can configure another field. |
| Generation | Monotonic stream generation associated with backing-index progression | Generation order is operational metadata, not proof of event-time ordering. |
2. Create the portable AtlasMart stream
PUT _index_template/atlasmart-telemetry-template-v1
{
"index_patterns": ["atlasmart-telemetry"],
"priority": 300,
"data_stream": {},
"template": {
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0
},
"mappings": {
"dynamic": "strict",
"properties": {
"@timestamp": {"type":"date"},
"event_id": {"type":"keyword"},
"event_kind": {"type":"keyword"},
"service": {"type":"keyword"},
"environment": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"host_id": {"type":"keyword"},
"level": {"type":"keyword"},
"message": {"type":"text"},
"duration_ms": {"type":"double"},
"request_count": {"type":"long"},
"error_count": {"type":"long"},
"labels": {"type":"object", "dynamic":"strict"}
}
}
}
}
DELETE _data_stream/atlasmart-telemetry
PUT _data_stream/atlasmart-telemetry
GET _data_stream/atlasmart-telemetry
GET _data_stream/atlasmart-telemetry/_stats
The exact generated backing-index name differs by product and version. Treat the API response as the source of truth. Verify exactly one backing index initially, confirm its health, and record store bytes before ingesting data.
3. Append deterministic events
POST atlasmart-telemetry/_doc?refresh=true
{
"@timestamp":"2026-09-11T06:00:00Z",
"event_id":"evt-1001",
"event_kind":"log",
"service":"catalog-api",
"environment":"lab",
"tenant_id":"tenant-a",
"host_id":"node-1",
"level":"INFO",
"message":"product P-1001 indexed",
"duration_ms":12.4,
"request_count":1,
"error_count":0,
"labels":{}
}
POST atlasmart-telemetry/_doc?refresh=true
{
"@timestamp":"2026-09-11T05:58:30Z",
"event_id":"evt-0999-late",
"event_kind":"log",
"service":"catalog-api",
"environment":"lab",
"tenant_id":"tenant-a",
"host_id":"node-1",
"level":"WARN",
"message":"late arrival from retry buffer",
"duration_ms":31.0,
"request_count":1,
"error_count":0,
"labels":{}
}
GET atlasmart-telemetry/_search
{
"sort":[{"@timestamp":"asc"}],
"query":{"match_all":{}}
}
The second event arrives later in wall-clock ingestion order but has an earlier event timestamp. A data stream is append-oriented; it does not promise ingest order equals event-time order. Search and aggregation logic must use the timestamp field explicitly when temporal ordering matters.
4. Rollover changes the write generation, not the stream name
GET _data_stream/atlasmart-telemetry
POST atlasmart-telemetry/_rollover
GET _data_stream/atlasmart-telemetry
POST atlasmart-telemetry/_doc?refresh=true
{
"@timestamp":"2026-09-11T06:05:00Z",
"event_id":"evt-1002",
"event_kind":"metric",
"service":"catalog-api",
"environment":"lab",
"tenant_id":"tenant-a",
"host_id":"node-1",
"level":"INFO",
"message":"five minute request rollup",
"duration_ms":9.8,
"request_count":120,
"error_count":2,
"labels":{}
}
GET atlasmart-telemetry/_search
{
"size":0,
"aggs":{"by_physical_index":{"terms":{"field":"_index","size":10}}}
}
Expected invariant: after rollover there are at least two backing indices; the post-rollover event lands in the newest one; a search against the stream sees documents from both generations. Do not assert a specific hidden name without reading the actual stream metadata.
5. Wrong model: “append-only means immutable forever”
Append-heavy systems optimize for mostly-new events, but observability data can contain corrections: a parser bug, a redaction requirement, a duplicate, or a compliance deletion. The safe lesson is not “never mutate”; it is “mutations are exceptional operations that must identify the correct backing index, preserve auditability, and respect lifecycle/retention constraints.” Elastic explicitly documents updating/deleting documents through the backing index. OpenSearch documents data streams as primarily append-only; direct correction should likewise use a concrete backing index after resolving the hit.
Rollover partitions storage generations. It does not deduplicate event IDs, guarantee ordering, make a write durable forever, or turn retention into backup.
6. Data stream or alias-with-write-index?
| Workload | Prefer | Reason |
|---|---|---|
| Continuously generated logs/metrics/events, mostly appends | Data stream | Timestamp-aware append model and lifecycle integration match the workload. |
| Frequent last-write-wins updates by stable ID | Write alias / ordinary index family | Data streams deliberately optimize a different write pattern. |
| Business records requiring arbitrary in-place mutation | Ordinary index or alias-managed generations | Avoid fighting backing-index targeting rules. |
| Telemetry with explicit event corrections only | Data stream with controlled correction path | Keep normal ingest simple while treating correction as privileged maintenance. |
7. Verification and cleanup
GET _data_stream/atlasmart-telemetry
GET _data_stream/atlasmart-telemetry/_stats
GET _cat/indices/.ds-atlasmart-telemetry*?v
GET atlasmart-telemetry/_count
# Cleanup only this disposable stream/template:
DELETE _data_stream/atlasmart-telemetry
DELETE _index_template/atlasmart-telemetry-template-v1
Record backing-index count, document count, maximum timestamp, store bytes, shard health and the exact generation before cleanup. Those observations are the foundation for later rollover and retention acceptance criteria.
Check your understanding
- Where does a normal write to a data stream go?
- Does rollover change the producer target name?
- Does generation order imply event-time order?
- Why standardize AtlasMart on @timestamp even though OpenSearch supports a custom field?
- Are backing indices backups?
Review the answers
1. To the current write backing index.
2. No. The stream name remains stable while the write backing index changes.
3. No. Late events can have earlier timestamps and still be written to the current generation.
4. It keeps the course contract portable and avoids depending on an OpenSearch-specific capability.
5. No. They are the physical storage generations of the live stream and share the same failure/retention domain unless separately protected.
Production judgment
Track ingest rate, rejected writes, indexing lag, backing-index count, shard growth, storage bytes, rollover frequency and search fan-out. A permanent giant index concentrates operational risk, but overly frequent rollover creates too many shards and metadata. Set rollover from measured workload/SLOs, not folklore. Snapshots and restore drills remain separate from retention.
Summary and next step
You can now predict where an append lands, what rollover changes, and why the stream name hides physical generations. Next we distinguish ordinary data streams from Elastic time-series index mode and reason about dimensions, metrics, shard routing and storage layout.
Authoritative references
- Elastic data streams — Logical stream, hidden backing indices, write index, @timestamp and rollover behavior.
- Elastic use a data stream — Indexing, searching, rollover, and document updates/deletes through backing indices.
- Elastic time series data streams — Elastic-specific TSDS dimensions, metrics and generated time-series identity.
- Elastic time series index settings — index.mode=time_series, routing path, look-back/look-ahead windows and related settings.
- Elastic time-bound indices and dimension routing — Timestamp acceptance windows and dimension-based shard routing.
- Elastic data stream lifecycle — Built-in lifecycle, rollover, retention, downsampling and storage transitions.
- Elastic data stream retention — Effective retention semantics and lifecycle APIs.
- OpenSearch data streams — Backing indexes, timestamp field, rollover, search and ISM integration.
- OpenSearch rollover API — Manual rollover semantics for data streams and aliases.
- OpenSearch modify data stream API — OpenSearch 3.8 experimental backing-index add/remove operation.
- OpenSearch Index State Management — Policy-driven rollover, retention and index-state automation.
- OpenSearch update document API — Document correction semantics when targeting a concrete backing index.