Triage AtlasMart incidents by mechanism—source failure, pipeline bug, warehouse performance, or semantic error—using evidence before repair, and trace blast radius through lineage.
Incident Triage: Source Failure vs Pipeline Bug vs Warehouse Performance vs Semantic Error
Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.
Learning outcomes
Classify incidents using evidence rather than symptoms.
Distinguish source failure, pipeline bug, warehouse performance, and semantic error.
Freeze unsafe publication while preserving raw evidence for diagnosis/replay.
Use lineage to prioritize consumers and backfill scope.
Verify repair through reconciliation rather than task success.
Chapter 25 begins from the accepted AtlasMart state established through Chapters 01–24: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current accepted paid order line, source progress remains committed through sequence 208, Chapter 20 metric contracts remain authoritative, Chapter 21 certified marts remain dependent on conformed assets, Chapter 22 security/privacy controls remain in force, Chapter 23 lineage/ownership metadata supplies blast-radius context, and Chapter 24 tests remain the correctness gates. Observability adds continuous evidence and incident handling; it must not silently reinterpret business rules merely to make a dashboard look healthy.
Runtime: Python 3.13.5 + SQLite 3.46.1. Storage: local in-memory structures plus SQLite semantics where transactions matter. Clock: UTC with explicit timestamps. Security: synthetic identifiers only. Cost: free/local. Healthy run: 10 records in/out, 4,820 bytes, 7.0 minutes, zero rejects/retries, 860 ms fixture CPU, 34 MB fixture peak memory, and certification 12 minutes after source readiness. Important limitation: these CPU/memory numbers are fixture measurements used to teach signal relationships; they are not performance recommendations for any warehouse engine or cloud service.
1. The realistic problem: “dashboard wrong” does not identify the failure layer
A consumer reports stale or incorrect revenue. Possible causes include an upstream extract that never arrived, transform code that dropped rows, a warehouse query that queued/spilled and missed a deadline, or a semantic layer/dashboard filter that changed metric population. Treating every report as “rerun the ETL” can worsen the incident or destroy evidence.
2. Triage by mechanism
| Class | Primary evidence | Typical first containment |
|---|---|---|
| Source failure | manifest missing/late; source counts absent; downstream idle | mark product stale; contact source owner; do not fabricate data |
| Pipeline bug | source ready, transform/reject/test failures; mismatch introduced inside job | stop certification; preserve raw; rollback/fix job; replay affected partition |
| Warehouse performance | inputs/semantics correct; queue/scan/spill/latency abnormal | protect workload, tune/scale/isolate based on measured bottleneck |
| Semantic error | pipeline/schema healthy; governed metric vs consumer query differs | invalidate affected semantic/cache/report outputs; restore/version metric logic |
3. A compact decision procedure
def classify(e): if not e["source_ready"]: return "source_failure" if not e["contract_ok"] or not e["pipeline_tests_ok"]: return "pipeline_or_schema_bug" if not e["metric_reconciles"]: return "semantic_or_data_bug" if e["certification_late"] and e["resource_or_queue_abnormal"]: return "warehouse_performance" return "needs_more_evidence"
This is not an automated universal root-cause engine. It is a teaching ordering: verify upstream availability, structural/pipeline correctness, semantic reconciliation, then performance evidence. Real incidents can have multiple causes.
4. Triage the three injected incidents
| Incident | Key evidence | Classification | Repair |
|---|---|---|---|
| 09:05 source manifest | source-ready SLO missed; later controls correct | Source failure / upstream lateness | wait for valid manifest, run/backfill partition, communicate freshness miss |
| line_amount_usd renamed | contract diff before transform; zero rows published | Breaking schema/source-contract incident | resolve semantics, compatibility mapping + contract v2, replay raw partition |
| 665 USD dashboard | pipeline success + schema/volume normal; governed total 820 | Semantic error | restore governed metric/filter, invalidate/rebuild affected products, reconcile |
5. Warehouse performance is a separate diagnosis
A slow query or queue can delay certification while data itself is correct. Confirm workload identity, scan bytes, pruning, join/skew, spill, queue time, cache state, concurrency, and engine version before changing physical design. Do not “fix performance” by dropping rows, using stale aggregates outside their freshness contract, or weakening correctness tests. Chapter 26 will treat performance engineering in depth.
6. Containment comes before speculative repair
- Record incident start, dataset/partition, observed SLO/test evidence, and run/source IDs.
- Mark affected certified products stale/unsafe when correctness is uncertain.
- Preserve raw inputs, failed outputs, logs, schema snapshots, and test evidence.
- Use lineage to notify named consumers and prevent stale cache/report reuse.
- Only then apply a repair whose mechanism matches the evidence.
7. Controlled failure: repair silently, notify nobody
Wrong: correct the bad dashboard query, refresh it, and close the issue because the number is now 820. Consumers may already have exported 665 USD, made decisions, or copied it to another report. Repair: communicate detection time, affected period/products, known/unknown impact, last-known-good state, repair/backfill status, and final reconciliation. Communication is part of data reliability.
8. Blast radius determines backfill scope
| Incident | Observed impacted assets | Safe scope principle |
|---|---|---|
| Breaking source-field rename | 15 | rebuild/revalidate every dependent path from changed contract edge |
| Revenue semantic bug | 9 | invalidate/recompute products consuming the wrong metric definition/result |
| Source lateness | all consumers of that partition until certification | backfill only missed partition unless evidence shows broader gaps |
Do not default to “rebuild all history.” Partition-aware backfills reduce resource and correctness risk, but only when lineage/history evidence proves the defect is scoped.
9. Repair verification
paid lines 10paid orders 8units 12gross revenue 820.00 USDcost 495.00 USDgross profit 325.00 USDsource sequence 208source_late repair: partition 2026-09-21 reconciledschema repair: partition 2026-09-21 reconciledsemantic repair: semantic cache/product reconciled
A successful retry is not sufficient. Re-run the relevant Chapter 24 contract/history/reconciliation tests, compare lineage/certification state, and only then restore the certified status.
10. Incident timeline and diagnostic latency
08:04 healthy source manifest available08:05 healthy orchestrator start08:12 tasks complete08:16 certified product published09:05 late-source manifest finally arrives09:18 late scenario certified (48 min past SLO)
Mean time to diagnose/repair can be useful operational measures, but only if event boundaries are defined consistently. A team that starts the clock at ticket creation cannot compare directly with one that starts at first SLO breach.
11. Production judgment and bridge
Diagnosis: symptoms are not causes. Recoverability: preserve raw and lineage evidence before mutation. Idempotency: backfills/replays must converge under repeats. Security: incident artifacts may contain PII, SQL, credentials, or internal topology; scope access. Operator ergonomics: a runbook should point to the next evidence query, not merely say “investigate.” Next: Lesson 5 packages detection, communication, repair, backfill, and learning into one repeatable runbook.
12. Verification checklist
- Classify the 09:05 incident as source lateness, not a transform bug.
- Confirm schema drift publishes zero rows before contract resolution.
- Confirm the 665 USD result is a semantic error despite green pipeline status.
- Confirm performance tuning is not used to bypass correctness gates.
- Confirm every repair ends with canonical reconciliation and certification review.
Knowledge check
Acceptance questions
- What evidence separates source lateness from a slow warehouse?
- Why freeze certification when correctness is uncertain?
- Why can a full-history backfill be dangerous?
- What closes an incident: task success or reconciled consumer-safe data?
Review the answers
1. Source readiness/manifest timing versus source-ready data plus abnormal queue/scan/runtime evidence.
2. It prevents downstream consumers from treating potentially wrong outputs as certified truth.
3. It expands blast radius, resource contention, and opportunity for unintended history changes.
4. Reconciled, tested, re-certified data plus communication of affected consumers and resolution.
Authoritative references
- OpenTelemetry — Metrics specificationAuthoritative metric concepts and separation of instrumentation from collection/export behavior.
- OpenTelemetry — Metrics data modelUseful for reasoning about time-series measurements, aggregation, temporality, and preserving metric semantics.
- Google SRE Book — Service Level ObjectivesDefines SLIs/SLOs and emphasizes selecting indicators from user-relevant service behavior rather than everything that is easy to measure.
- Google SRE Book — Monitoring Distributed SystemsBackground on actionable monitoring signals and avoiding monitoring that produces noise without operator decisions.
- SQLite — TransactionsSupports the local demonstration of atomic repair/backfill state and safe commit boundaries.
- Python — statisticsUsed for small deterministic baseline examples; production anomaly detection needs workload-specific statistical validation.
- Python — hashlibUsed to fingerprint deterministic incident evidence and repaired outputs.
13. Lab cleanup/reset
Incident scenarios are intentionally isolated. Discard the in-memory state after each scenario so classification/repair evidence does not leak into the next one.