Chapter 10 · Data Streams, Time-Series Data, Logs, Metrics, and Append-Heavy Workloads
Late/Out-of-Order Events, Updates/Deletes Against Backing Indices, and Operational Constraints
Handle late and out-of-order events without confusing append-only design with immutable truth, then target corrections safely while respecting backing-index, timestamp, lifecycle, and retention constraints.
Learning outcomes
AtlasMart discovers that a retry buffer delivered an event ten minutes late and that an older log contains a value that must be redacted. Data streams are optimized for append-heavy ingestion, but real systems still need a deliberate correction path. This lesson separates event-time lateness from document mutation and shows why the physical backing index matters.
Distinguish late arrival, out-of-order arrival, duplicate event and correction of an existing event.
Locate the concrete backing index that owns a hit before an exceptional update/delete.
Explain Elastic TSDS timestamp-window constraints separately from ordinary data-stream behavior.
Apply optimistic concurrency and audit controls to privileged corrections.
Design retention so deletion, restore and compliance requirements do not contradict each other.
The reproducible examples target self-managed Elasticsearch
9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab
conventions: Elasticsearch on https://localhost:9200 with the
copied CA certificate, OpenSearch on https://localhost:9201
with the disposable demo certificate explicitly treated as
local-only, one-node disposable clusters, pinned versions, and
no moving latest tags. Elastic documents update/delete
operations against a data stream by targeting the backing
index for the document. OpenSearch data streams are documented
as primarily append-only; the safe cross-product operational
pattern is to resolve the hit's concrete
_index and perform any exceptional document
mutation only against that physical index after verifying your
product/version behavior and permissions.
The generation environment does not run the two search servers, so commands below are deterministic lab specifications and expected invariants, not fabricated captured output. Run them only against disposable AtlasMart course resources. Record your actual responses, timestamps, backing-index names, latency, shard counts, and disk usage before drawing operational conclusions.
1. Four different “late data” problems
| Case | Meaning | Correct response |
|---|---|---|
| Late arrival | Event created earlier but ingested now | Usually append with its original timestamp if accepted by the stream/index mode. |
| Out-of-order arrival | Ingest order differs from event-time order | Sort/window by timestamp; do not assume append order is temporal order. |
| Duplicate delivery | Same event delivered again after timeout/retry | Use stable event identity/dedup policy; TSDS may have product-specific duplicate identity semantics. |
| Correction | Previously stored event is wrong or must be redacted | Find physical document, authorize correction, update/delete with audit and concurrency controls. |
2. Prove late arrival is not automatically a correction
POST atlasmart-telemetry/_doc?refresh=true
{
"@timestamp":"2026-09-11T05:50:00Z",
"event_id":"evt-late-1050",
"event_kind":"log",
"service":"catalog-api",
"environment":"lab",
"tenant_id":"tenant-a",
"host_id":"node-1",
"level":"WARN",
"message":"agent buffer replay",
"duration_ms":17.0,
"request_count":1,
"error_count":0,
"labels":{}
}
GET atlasmart-telemetry/_search
{
"query":{"term":{"event_id":"evt-late-1050"}},
"sort":[{"@timestamp":"asc"}]
}
Ordinary data streams can accept an older timestamp so long as
the mapping and product rules allow it. Elastic TSDS is more
constrained because backing indices have accepted time ranges. A
TSDS late-data policy must account for
index.time_series.start_time/end_time,
look-back/look-ahead behavior, and reindex/import procedures
rather than assuming every historical timestamp can be appended
to the current write index.
3. Resolve the backing index before correction
GET atlasmart-telemetry/_search
{
"seq_no_primary_term":true,
"query":{"term":{"event_id":"evt-late-1050"}},
"_source":["event_id","message","@timestamp"]
}
Record the returned _index, _id,
_seq_no and _primary_term. The backing
index is part of the mutation address. Never derive it from an
assumed date/generation convention; rollover, restore, shrink or
product behavior can invalidate that guess.
POST /.ds-atlasmart-telemetry-OBSERVED-BACKING/_update/OBSERVED_ID?if_seq_no=OBSERVED_SEQ&if_primary_term=OBSERVED_TERM&refresh=true
{
"doc":{"message":"redacted operational event"}
}
The placeholders are deliberate because the values must come from your actual search response. A 409 conflict means somebody changed the document after you read it; investigate rather than blindly retrying a stale correction.
4. Deletion requires the same physical discipline
DELETE /.ds-atlasmart-telemetry-OBSERVED-BACKING/_doc/OBSERVED_ID?if_seq_no=OBSERVED_SEQ&if_primary_term=OBSERVED_TERM&refresh=true
A delete from the live index does not erase copies in snapshots, downstream exports, caches or replicated systems. Compliance deletion therefore needs a system-level retention/erasure procedure, not one REST call.
Lifecycle deletion controls how long live backing indices remain. Snapshots may intentionally preserve data longer. Write the restore/retention/legal policy together so operators know which copy is authoritative and when each copy expires.
5. Rollover during correction: why IDs are not enough
If an event has a stable event_id but the stream
has rolled over, an update addressed only to the logical stream
and ID can be ambiguous or unsupported depending on the
API/product. Resolve the document through search first. Store
the physical index in your correction job and re-check
concurrency metadata immediately before mutation.
hit = search_stream(event_id)
assert hit.count == 1
authorize_correction(hit.source.tenant_id)
log_audit("before", hit.index, hit.id, hit.seq_no, hit.primary_term)
update_backing_index(
index=hit.index,
id=hit.id,
if_seq_no=hit.seq_no,
if_primary_term=hit.primary_term,
patch=approved_patch,
)
verify_stream_search(event_id, expected_patch)
log_audit("after", ...)
6. OpenSearch 3.8 backing-index metadata operations are not document edits
OpenSearch 3.8 introduced an experimental
_data_stream/_modify API for attaching or detaching
backing indexes. That is a metadata operation; it does not
rewrite documents or move shards. Do not confuse “remove backing
index from stream” with “delete the data in the index,” and do
not use an experimental feature as a substitute for a correction
workflow.
7. Retention and late-data acceptance criteria
| Requirement | Question to answer before production |
|---|---|
| Maximum lateness | How old may an event be and still belong in the searchable hot stream? |
| Duplicate policy | What stable identity detects replay, and what happens on ambiguity? |
| Correction SLA | Who may change old telemetry, through which audited service? |
| Retention | How long must raw data remain searchable, and what is only archived? |
| Restore | Can a snapshot restore recreate required stream/template/lifecycle state? |
| Compliance | Which snapshots/exports also require expiry or legal hold? |
Check your understanding
- Is a late event necessarily an update?
- Why search before correcting a stream document?
- What does a 409 OCC conflict mean?
- Does deleting the live document remove snapshot copies?
- What does OpenSearch _data_stream/_modify change?
Review the answers
1. No. It is usually a new append carrying an older event timestamp.
2. To discover the actual backing index, ID and concurrency metadata rather than guessing physical location.
3. The document changed after the correction job read it; re-read and decide intentionally.
4. No. Backup/export retention must be handled separately.
5. Data-stream membership metadata for backing indices, not document contents.
Production judgment
Measure late-arrival rate and distribution, correction volume, duplicate rate, update/delete failures, OCC conflicts and time between event timestamp and searchable timestamp. If corrections are common rather than exceptional, reconsider whether a data stream is the right storage contract. Failure injection should stay in disposable streams: roll over during a correction test, intentionally use stale concurrency metadata, and verify that the system fails closed without losing audit evidence.
Summary and next step
Append-heavy does not mean “pretend historical facts never change.” You now have a controlled model for lateness, duplicate delivery and physical backing-index corrections. The final lesson converts these mechanisms into a complete telemetry architecture and acceptance test.
Authoritative references
- Elastic data streams — Logical stream, hidden backing indices, write index, @timestamp and rollover behavior.
- Elastic use a data stream — Indexing, searching, rollover, and document updates/deletes through backing indices.
- Elastic time series data streams — Elastic-specific TSDS dimensions, metrics and generated time-series identity.
- Elastic time series index settings — index.mode=time_series, routing path, look-back/look-ahead windows and related settings.
- Elastic time-bound indices and dimension routing — Timestamp acceptance windows and dimension-based shard routing.
- Elastic data stream lifecycle — Built-in lifecycle, rollover, retention, downsampling and storage transitions.
- Elastic data stream retention — Effective retention semantics and lifecycle APIs.
- OpenSearch data streams — Backing indexes, timestamp field, rollover, search and ISM integration.
- OpenSearch rollover API — Manual rollover semantics for data streams and aliases.
- OpenSearch modify data stream API — OpenSearch 3.8 experimental backing-index add/remove operation.
- OpenSearch Index State Management — Policy-driven rollover, retention and index-state automation.
- OpenSearch update document API — Document correction semantics when targeting a concrete backing index.