Chapter 10 · Data Streams, Time-Series Data, Logs, Metrics, and Append-Heavy Workloads

Late/Out-of-Order Events, Updates/Deletes Against Backing Indices, and Operational Constraints

Handle late and out-of-order events without confusing append-only design with immutable truth, then target corrections safely while respecting backing-index, timestamp, lifecycle, and retention constraints.

Intermediate100–120 minutesData stream & telemetry labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart discovers that a retry buffer delivered an event ten minutes late and that an older log contains a value that must be redacted. Data streams are optimized for append-heavy ingestion, but real systems still need a deliberate correction path. This lesson separates event-time lateness from document mutation and shows why the physical backing index matters.

01

Distinguish late arrival, out-of-order arrival, duplicate event and correction of an existing event.

02

Locate the concrete backing index that owns a hit before an exceptional update/delete.

03

Explain Elastic TSDS timestamp-window constraints separately from ordinary data-stream behavior.

04

Apply optimistic concurrency and audit controls to privileged corrections.

05

Design retention so deletion, restore and compliance requirements do not contradict each other.

Chapter baseline reviewed 11 September 2026

The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 with the copied CA certificate, OpenSearch on https://localhost:9201 with the disposable demo certificate explicitly treated as local-only, one-node disposable clusters, pinned versions, and no moving latest tags. Elastic documents update/delete operations against a data stream by targeting the backing index for the document. OpenSearch data streams are documented as primarily append-only; the safe cross-product operational pattern is to resolve the hit's concrete _index and perform any exceptional document mutation only against that physical index after verifying your product/version behavior and permissions.

Execution and safety note

The generation environment does not run the two search servers, so commands below are deterministic lab specifications and expected invariants, not fabricated captured output. Run them only against disposable AtlasMart course resources. Record your actual responses, timestamps, backing-index names, latency, shard counts, and disk usage before drawing operational conclusions.

1. Four different “late data” problems

Case Meaning Correct response
Late arrival Event created earlier but ingested now Usually append with its original timestamp if accepted by the stream/index mode.
Out-of-order arrival Ingest order differs from event-time order Sort/window by timestamp; do not assume append order is temporal order.
Duplicate delivery Same event delivered again after timeout/retry Use stable event identity/dedup policy; TSDS may have product-specific duplicate identity semantics.
Correction Previously stored event is wrong or must be redacted Find physical document, authorize correction, update/delete with audit and concurrency controls.

2. Prove late arrival is not automatically a correction

Append a late event
POST atlasmart-telemetry/_doc?refresh=true
{
  "@timestamp":"2026-09-11T05:50:00Z",
  "event_id":"evt-late-1050",
  "event_kind":"log",
  "service":"catalog-api",
  "environment":"lab",
  "tenant_id":"tenant-a",
  "host_id":"node-1",
  "level":"WARN",
  "message":"agent buffer replay",
  "duration_ms":17.0,
  "request_count":1,
  "error_count":0,
  "labels":{}
}

GET atlasmart-telemetry/_search
{
  "query":{"term":{"event_id":"evt-late-1050"}},
  "sort":[{"@timestamp":"asc"}]
}

Ordinary data streams can accept an older timestamp so long as the mapping and product rules allow it. Elastic TSDS is more constrained because backing indices have accepted time ranges. A TSDS late-data policy must account for index.time_series.start_time/end_time, look-back/look-ahead behavior, and reindex/import procedures rather than assuming every historical timestamp can be appended to the current write index.

3. Resolve the backing index before correction

Search for the physical owner
GET atlasmart-telemetry/_search
{
  "seq_no_primary_term":true,
  "query":{"term":{"event_id":"evt-late-1050"}},
  "_source":["event_id","message","@timestamp"]
}

Record the returned _index, _id, _seq_no and _primary_term. The backing index is part of the mutation address. Never derive it from an assumed date/generation convention; rollover, restore, shrink or product behavior can invalidate that guess.

Exceptional correction · substitute observed values
POST /.ds-atlasmart-telemetry-OBSERVED-BACKING/_update/OBSERVED_ID?if_seq_no=OBSERVED_SEQ&if_primary_term=OBSERVED_TERM&refresh=true
{
  "doc":{"message":"redacted operational event"}
}

The placeholders are deliberate because the values must come from your actual search response. A 409 conflict means somebody changed the document after you read it; investigate rather than blindly retrying a stale correction.

4. Deletion requires the same physical discipline

Exceptional delete · only after approval
DELETE /.ds-atlasmart-telemetry-OBSERVED-BACKING/_doc/OBSERVED_ID?if_seq_no=OBSERVED_SEQ&if_primary_term=OBSERVED_TERM&refresh=true

A delete from the live index does not erase copies in snapshots, downstream exports, caches or replicated systems. Compliance deletion therefore needs a system-level retention/erasure procedure, not one REST call.

Retention is not backup or erasure proof

Lifecycle deletion controls how long live backing indices remain. Snapshots may intentionally preserve data longer. Write the restore/retention/legal policy together so operators know which copy is authoritative and when each copy expires.

5. Rollover during correction: why IDs are not enough

If an event has a stable event_id but the stream has rolled over, an update addressed only to the logical stream and ID can be ambiguous or unsupported depending on the API/product. Resolve the document through search first. Store the physical index in your correction job and re-check concurrency metadata immediately before mutation.

Correction job pseudo-code
hit = search_stream(event_id)
assert hit.count == 1
authorize_correction(hit.source.tenant_id)
log_audit("before", hit.index, hit.id, hit.seq_no, hit.primary_term)
update_backing_index(
    index=hit.index,
    id=hit.id,
    if_seq_no=hit.seq_no,
    if_primary_term=hit.primary_term,
    patch=approved_patch,
)
verify_stream_search(event_id, expected_patch)
log_audit("after", ...)

6. OpenSearch 3.8 backing-index metadata operations are not document edits

OpenSearch 3.8 introduced an experimental _data_stream/_modify API for attaching or detaching backing indexes. That is a metadata operation; it does not rewrite documents or move shards. Do not confuse “remove backing index from stream” with “delete the data in the index,” and do not use an experimental feature as a substitute for a correction workflow.

7. Retention and late-data acceptance criteria

Requirement Question to answer before production
Maximum lateness How old may an event be and still belong in the searchable hot stream?
Duplicate policy What stable identity detects replay, and what happens on ambiguity?
Correction SLA Who may change old telemetry, through which audited service?
Retention How long must raw data remain searchable, and what is only archived?
Restore Can a snapshot restore recreate required stream/template/lifecycle state?
Compliance Which snapshots/exports also require expiry or legal hold?

Check your understanding

  1. Is a late event necessarily an update?
  2. Why search before correcting a stream document?
  3. What does a 409 OCC conflict mean?
  4. Does deleting the live document remove snapshot copies?
  5. What does OpenSearch _data_stream/_modify change?
Review the answers

1. No. It is usually a new append carrying an older event timestamp.

2. To discover the actual backing index, ID and concurrency metadata rather than guessing physical location.

3. The document changed after the correction job read it; re-read and decide intentionally.

4. No. Backup/export retention must be handled separately.

5. Data-stream membership metadata for backing indices, not document contents.

Production judgment

Measure late-arrival rate and distribution, correction volume, duplicate rate, update/delete failures, OCC conflicts and time between event timestamp and searchable timestamp. If corrections are common rather than exceptional, reconsider whether a data stream is the right storage contract. Failure injection should stay in disposable streams: roll over during a correction test, intentionally use stale concurrency metadata, and verify that the system fails closed without losing audit evidence.

Summary and next step

Append-heavy does not mean “pretend historical facts never change.” You now have a controlled model for lateness, duplicate delivery and physical backing-index corrections. The final lesson converts these mechanisms into a complete telemetry architecture and acceptance test.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.