Triage AtlasMart incidents by mechanism—source failure, pipeline bug, warehouse performance, or semantic error—using evidence before repair, and trace blast radius through lineage.

Incident Triage: Source Failure vs Pipeline Bug vs Warehouse Performance vs Semantic Error

Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.

Intermediate → Advanced150–190 minutesIncident triage labClassification + blast radius + repairLast reviewed: September 2026

Learning outcomes

01

Classify incidents using evidence rather than symptoms.

02

Distinguish source failure, pipeline bug, warehouse performance, and semantic error.

03

Freeze unsafe publication while preserving raw evidence for diagnosis/replay.

04

Use lineage to prioritize consumers and backfill scope.

05

Verify repair through reconciliation rather than task success.

Continuity: observability watches the governed system; it does not redefine it

Chapter 25 begins from the accepted AtlasMart state established through Chapters 01–24: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current accepted paid order line, source progress remains committed through sequence 208, Chapter 20 metric contracts remain authoritative, Chapter 21 certified marts remain dependent on conformed assets, Chapter 22 security/privacy controls remain in force, Chapter 23 lineage/ownership metadata supplies blast-radius context, and Chapter 24 tests remain the correctness gates. Observability adds continuous evidence and incident handling; it must not silently reinterpret business rules merely to make a dashboard look healthy.

Executed local reliability fixture

Runtime: Python 3.13.5 + SQLite 3.46.1. Storage: local in-memory structures plus SQLite semantics where transactions matter. Clock: UTC with explicit timestamps. Security: synthetic identifiers only. Cost: free/local. Healthy run: 10 records in/out, 4,820 bytes, 7.0 minutes, zero rejects/retries, 860 ms fixture CPU, 34 MB fixture peak memory, and certification 12 minutes after source readiness. Important limitation: these CPU/memory numbers are fixture measurements used to teach signal relationships; they are not performance recommendations for any warehouse engine or cloud service.

1. The realistic problem: “dashboard wrong” does not identify the failure layer

A consumer reports stale or incorrect revenue. Possible causes include an upstream extract that never arrived, transform code that dropped rows, a warehouse query that queued/spilled and missed a deadline, or a semantic layer/dashboard filter that changed metric population. Treating every report as “rerun the ETL” can worsen the incident or destroy evidence.

2. Triage by mechanism

Class Primary evidence Typical first containment
Source failure manifest missing/late; source counts absent; downstream idle mark product stale; contact source owner; do not fabricate data
Pipeline bug source ready, transform/reject/test failures; mismatch introduced inside job stop certification; preserve raw; rollback/fix job; replay affected partition
Warehouse performance inputs/semantics correct; queue/scan/spill/latency abnormal protect workload, tune/scale/isolate based on measured bottleneck
Semantic error pipeline/schema healthy; governed metric vs consumer query differs invalidate affected semantic/cache/report outputs; restore/version metric logic

3. A compact decision procedure

Evidence-first classifier (didactic)
def classify(e):    if not e["source_ready"]:        return "source_failure"    if not e["contract_ok"] or not e["pipeline_tests_ok"]:        return "pipeline_or_schema_bug"    if not e["metric_reconciles"]:        return "semantic_or_data_bug"    if e["certification_late"] and e["resource_or_queue_abnormal"]:        return "warehouse_performance"    return "needs_more_evidence"

This is not an automated universal root-cause engine. It is a teaching ordering: verify upstream availability, structural/pipeline correctness, semantic reconciliation, then performance evidence. Real incidents can have multiple causes.

4. Triage the three injected incidents

Incident Key evidence Classification Repair
09:05 source manifest source-ready SLO missed; later controls correct Source failure / upstream lateness wait for valid manifest, run/backfill partition, communicate freshness miss
line_amount_usd renamed contract diff before transform; zero rows published Breaking schema/source-contract incident resolve semantics, compatibility mapping + contract v2, replay raw partition
665 USD dashboard pipeline success + schema/volume normal; governed total 820 Semantic error restore governed metric/filter, invalidate/rebuild affected products, reconcile

5. Warehouse performance is a separate diagnosis

A slow query or queue can delay certification while data itself is correct. Confirm workload identity, scan bytes, pruning, join/skew, spill, queue time, cache state, concurrency, and engine version before changing physical design. Do not “fix performance” by dropping rows, using stale aggregates outside their freshness contract, or weakening correctness tests. Chapter 26 will treat performance engineering in depth.

6. Containment comes before speculative repair

  1. Record incident start, dataset/partition, observed SLO/test evidence, and run/source IDs.
  2. Mark affected certified products stale/unsafe when correctness is uncertain.
  3. Preserve raw inputs, failed outputs, logs, schema snapshots, and test evidence.
  4. Use lineage to notify named consumers and prevent stale cache/report reuse.
  5. Only then apply a repair whose mechanism matches the evidence.

7. Controlled failure: repair silently, notify nobody

Wrong: correct the bad dashboard query, refresh it, and close the issue because the number is now 820. Consumers may already have exported 665 USD, made decisions, or copied it to another report. Repair: communicate detection time, affected period/products, known/unknown impact, last-known-good state, repair/backfill status, and final reconciliation. Communication is part of data reliability.

8. Blast radius determines backfill scope

Incident Observed impacted assets Safe scope principle
Breaking source-field rename 15 rebuild/revalidate every dependent path from changed contract edge
Revenue semantic bug 9 invalidate/recompute products consuming the wrong metric definition/result
Source lateness all consumers of that partition until certification backfill only missed partition unless evidence shows broader gaps

Do not default to “rebuild all history.” Partition-aware backfills reduce resource and correctness risk, but only when lineage/history evidence proves the defect is scoped.

9. Repair verification

Acceptance controls after each isolated repair
paid lines       10paid orders       8units             12gross revenue    820.00 USDcost             495.00 USDgross profit     325.00 USDsource sequence  208source_late repair:   partition 2026-09-21 reconciledschema repair:        partition 2026-09-21 reconciledsemantic repair:      semantic cache/product reconciled

A successful retry is not sufficient. Re-run the relevant Chapter 24 contract/history/reconciliation tests, compare lineage/certification state, and only then restore the certified status.

10. Incident timeline and diagnostic latency

Fixture timeline excerpt
08:04 healthy source manifest available08:05 healthy orchestrator start08:12 tasks complete08:16 certified product published09:05 late-source manifest finally arrives09:18 late scenario certified (48 min past SLO)

Mean time to diagnose/repair can be useful operational measures, but only if event boundaries are defined consistently. A team that starts the clock at ticket creation cannot compare directly with one that starts at first SLO breach.

11. Production judgment and bridge

Diagnosis: symptoms are not causes. Recoverability: preserve raw and lineage evidence before mutation. Idempotency: backfills/replays must converge under repeats. Security: incident artifacts may contain PII, SQL, credentials, or internal topology; scope access. Operator ergonomics: a runbook should point to the next evidence query, not merely say “investigate.” Next: Lesson 5 packages detection, communication, repair, backfill, and learning into one repeatable runbook.

12. Verification checklist

  1. Classify the 09:05 incident as source lateness, not a transform bug.
  2. Confirm schema drift publishes zero rows before contract resolution.
  3. Confirm the 665 USD result is a semantic error despite green pipeline status.
  4. Confirm performance tuning is not used to bypass correctness gates.
  5. Confirm every repair ends with canonical reconciliation and certification review.

Knowledge check

Acceptance questions

  1. What evidence separates source lateness from a slow warehouse?
  2. Why freeze certification when correctness is uncertain?
  3. Why can a full-history backfill be dangerous?
  4. What closes an incident: task success or reconciled consumer-safe data?
Review the answers

1. Source readiness/manifest timing versus source-ready data plus abnormal queue/scan/runtime evidence.

2. It prevents downstream consumers from treating potentially wrong outputs as certified truth.

3. It expands blast radius, resource contention, and opportunity for unintended history changes.

4. Reconciled, tested, re-certified data plus communication of affected consumers and resolution.

Authoritative references

13. Lab cleanup/reset

Incident scenarios are intentionally isolated. Discard the in-memory state after each scenario so classification/repair evidence does not leak into the next one.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.