Chapter 10 · Data Streams, Time-Series Data, Logs, Metrics, and Append-Heavy Workloads

Data Streams, Backing Indices, Write Index, @timestamp, Rollover, and Append-Only Design

Build an append-heavy AtlasMart telemetry stream, trace writes into hidden backing indices, roll the stream, and distinguish logical stream identity from the current physical write index.

Intermediate100–120 minutesData stream & telemetry labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart emits application logs and operational metrics continuously. A single permanent index would grow without a clean write-generation boundary, while date-stamped manual indices would force every producer and consumer to know physical names. A data stream solves the naming problem by giving clients one logical target while the cluster owns a sequence of hidden backing indices.

01

Distinguish a data stream from its hidden backing indices and identify the current write index.

02

Explain the timestamp contract, rollover generation, and why appends target only the write index.

03

Observe a manual rollover and prove searches span old and new backing indices.

04

Explain why append-oriented design is not the same as “data can never be corrected.”

05

Choose a data stream only when the write/update pattern and retention model actually fit.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 with the copied CA certificate, OpenSearch on https://localhost:9201 with the disposable demo certificate explicitly treated as local-only, one-node disposable clusters, pinned versions, and no moving latest tags. Elastic requires @timestamp for ordinary data streams. OpenSearch defaults to @timestamp but also supports a custom timestamp field declared inside the template's data_stream.timestamp_field; that custom-field capability is therefore not presented as a portable Elastic/OpenSearch contract.

Execution and safety note

The generation environment does not run the two search servers, so commands below are deterministic lab specifications and expected invariants, not fabricated captured output. Run them only against disposable AtlasMart course resources. Record your actual responses, timestamps, backing-index names, latency, shard counts, and disk usage before drawing operational conclusions.

1. Logical stream, physical generations

A data stream is a logical name backed by one or more hidden indices. The newest backing index is the write index. Indexing to the stream routes a new document to that write index; searching the stream expands to all backing indices. A rollover creates a new backing index, increments the stream generation, and moves the write role forward. The old backing index remains searchable and can later be managed by lifecycle policy.

Construct Meaning AtlasMart consequence
Data stream Stable logical target for append-heavy time-series data Producers write to atlasmart-telemetry; dashboards search the same name.
Backing index Hidden physical index belonging to the stream Shard/storage state is inspected here; clients normally do not hard-code its name.
Write index Newest backing index receiving ordinary stream writes Rollover changes it without changing producer configuration.
Timestamp field Field required to place an event on the time axis AtlasMart standardizes on @timestamp even though OpenSearch can configure another field.
Generation Monotonic stream generation associated with backing-index progression Generation order is operational metadata, not proof of event-time ordering.

2. Create the portable AtlasMart stream

Dev Tools · portable template
PUT _index_template/atlasmart-telemetry-template-v1
{
  "index_patterns": ["atlasmart-telemetry"],
  "priority": 300,
  "data_stream": {},
  "template": {
    "settings": {
      "number_of_shards": 1,
      "number_of_replicas": 0
    },
    "mappings": {
      "dynamic": "strict",
      "properties": {
        "@timestamp": {"type":"date"},
        "event_id": {"type":"keyword"},
        "event_kind": {"type":"keyword"},
        "service": {"type":"keyword"},
        "environment": {"type":"keyword"},
        "tenant_id": {"type":"keyword"},
        "host_id": {"type":"keyword"},
        "level": {"type":"keyword"},
        "message": {"type":"text"},
        "duration_ms": {"type":"double"},
        "request_count": {"type":"long"},
        "error_count": {"type":"long"},
        "labels": {"type":"object", "dynamic":"strict"}
      }
    }
  }
}
Dev Tools · create and inspect
DELETE _data_stream/atlasmart-telemetry
PUT _data_stream/atlasmart-telemetry
GET _data_stream/atlasmart-telemetry
GET _data_stream/atlasmart-telemetry/_stats

The exact generated backing-index name differs by product and version. Treat the API response as the source of truth. Verify exactly one backing index initially, confirm its health, and record store bytes before ingesting data.

3. Append deterministic events

Dev Tools · normal and late event
POST atlasmart-telemetry/_doc?refresh=true
{
  "@timestamp":"2026-09-11T06:00:00Z",
  "event_id":"evt-1001",
  "event_kind":"log",
  "service":"catalog-api",
  "environment":"lab",
  "tenant_id":"tenant-a",
  "host_id":"node-1",
  "level":"INFO",
  "message":"product P-1001 indexed",
  "duration_ms":12.4,
  "request_count":1,
  "error_count":0,
  "labels":{}
}

POST atlasmart-telemetry/_doc?refresh=true
{
  "@timestamp":"2026-09-11T05:58:30Z",
  "event_id":"evt-0999-late",
  "event_kind":"log",
  "service":"catalog-api",
  "environment":"lab",
  "tenant_id":"tenant-a",
  "host_id":"node-1",
  "level":"WARN",
  "message":"late arrival from retry buffer",
  "duration_ms":31.0,
  "request_count":1,
  "error_count":0,
  "labels":{}
}

GET atlasmart-telemetry/_search
{
  "sort":[{"@timestamp":"asc"}],
  "query":{"match_all":{}}
}

The second event arrives later in wall-clock ingestion order but has an earlier event timestamp. A data stream is append-oriented; it does not promise ingest order equals event-time order. Search and aggregation logic must use the timestamp field explicitly when temporal ordering matters.

4. Rollover changes the write generation, not the stream name

Dev Tools · rollover and inspect
GET _data_stream/atlasmart-telemetry
POST atlasmart-telemetry/_rollover
GET _data_stream/atlasmart-telemetry

POST atlasmart-telemetry/_doc?refresh=true
{
  "@timestamp":"2026-09-11T06:05:00Z",
  "event_id":"evt-1002",
  "event_kind":"metric",
  "service":"catalog-api",
  "environment":"lab",
  "tenant_id":"tenant-a",
  "host_id":"node-1",
  "level":"INFO",
  "message":"five minute request rollup",
  "duration_ms":9.8,
  "request_count":120,
  "error_count":2,
  "labels":{}
}

GET atlasmart-telemetry/_search
{
  "size":0,
  "aggs":{"by_physical_index":{"terms":{"field":"_index","size":10}}}
}

Expected invariant: after rollover there are at least two backing indices; the post-rollover event lands in the newest one; a search against the stream sees documents from both generations. Do not assert a specific hidden name without reading the actual stream metadata.

5. Wrong model: “append-only means immutable forever”

Append-heavy systems optimize for mostly-new events, but observability data can contain corrections: a parser bug, a redaction requirement, a duplicate, or a compliance deletion. The safe lesson is not “never mutate”; it is “mutations are exceptional operations that must identify the correct backing index, preserve auditability, and respect lifecycle/retention constraints.” Elastic explicitly documents updating/deleting documents through the backing index. OpenSearch documents data streams as primarily append-only; direct correction should likewise use a concrete backing index after resolving the hit.

Do not use rollover as a correctness mechanism

Rollover partitions storage generations. It does not deduplicate event IDs, guarantee ordering, make a write durable forever, or turn retention into backup.

6. Data stream or alias-with-write-index?

Workload Prefer Reason
Continuously generated logs/metrics/events, mostly appends Data stream Timestamp-aware append model and lifecycle integration match the workload.
Frequent last-write-wins updates by stable ID Write alias / ordinary index family Data streams deliberately optimize a different write pattern.
Business records requiring arbitrary in-place mutation Ordinary index or alias-managed generations Avoid fighting backing-index targeting rules.
Telemetry with explicit event corrections only Data stream with controlled correction path Keep normal ingest simple while treating correction as privileged maintenance.

7. Verification and cleanup

Verification checklist
GET _data_stream/atlasmart-telemetry
GET _data_stream/atlasmart-telemetry/_stats
GET _cat/indices/.ds-atlasmart-telemetry*?v
GET atlasmart-telemetry/_count

# Cleanup only this disposable stream/template:
DELETE _data_stream/atlasmart-telemetry
DELETE _index_template/atlasmart-telemetry-template-v1

Record backing-index count, document count, maximum timestamp, store bytes, shard health and the exact generation before cleanup. Those observations are the foundation for later rollover and retention acceptance criteria.

Check your understanding

  1. Where does a normal write to a data stream go?
  2. Does rollover change the producer target name?
  3. Does generation order imply event-time order?
  4. Why standardize AtlasMart on @timestamp even though OpenSearch supports a custom field?
  5. Are backing indices backups?
Review the answers

1. To the current write backing index.

2. No. The stream name remains stable while the write backing index changes.

3. No. Late events can have earlier timestamps and still be written to the current generation.

4. It keeps the course contract portable and avoids depending on an OpenSearch-specific capability.

5. No. They are the physical storage generations of the live stream and share the same failure/retention domain unless separately protected.

Production judgment

Track ingest rate, rejected writes, indexing lag, backing-index count, shard growth, storage bytes, rollover frequency and search fan-out. A permanent giant index concentrates operational risk, but overly frequent rollover creates too many shards and metadata. Set rollover from measured workload/SLOs, not folklore. Snapshots and restore drills remain separate from retention.

Summary and next step

You can now predict where an append lands, what rollover changes, and why the stream name hides physical generations. Next we distinguish ordinary data streams from Elastic time-series index mode and reason about dimensions, metrics, shard routing and storage layout.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.