Operate an AtlasMart data-incident runbook that detects, communicates, repairs, backfills, reconciles, and learns from failures with owned postmortem actions.
Build a Data Incident Runbook with Detection, Blast Radius, Consumer Communication, Repair, Backfill, and Postmortem
Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.
Learning outcomes
Build a runbook with detection, containment, blast-radius analysis, communication, repair, backfill, and verification.
Record incident timelines and last-known-good state.
Produce concise postmortems with concrete owners and due dates.
Demonstrate safe partition-scoped repair and re-certification.
Turn incident learning into test/contract/observability improvements.
Chapter 25 begins from the accepted AtlasMart state established through Chapters 01–24: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current accepted paid order line, source progress remains committed through sequence 208, Chapter 20 metric contracts remain authoritative, Chapter 21 certified marts remain dependent on conformed assets, Chapter 22 security/privacy controls remain in force, Chapter 23 lineage/ownership metadata supplies blast-radius context, and Chapter 24 tests remain the correctness gates. Observability adds continuous evidence and incident handling; it must not silently reinterpret business rules merely to make a dashboard look healthy.
Runtime: Python 3.13.5 + SQLite 3.46.1. Storage: local in-memory structures plus SQLite semantics where transactions matter. Clock: UTC with explicit timestamps. Security: synthetic identifiers only. Cost: free/local. Healthy run: 10 records in/out, 4,820 bytes, 7.0 minutes, zero rejects/retries, 860 ms fixture CPU, 34 MB fixture peak memory, and certification 12 minutes after source readiness. Important limitation: these CPU/memory numbers are fixture measurements used to teach signal relationships; they are not performance recommendations for any warehouse engine or cloud service.
1. The realistic problem: reliability depends on what happens after the alert
AtlasMart can detect a freshness miss, schema drift, or revenue discrepancy. Reliability still fails if nobody knows who owns the response, consumers are not warned, raw evidence is overwritten, a backfill corrupts current partitions, or the postmortem produces no corrective action. A runbook is an operational decision path, not a list of vendor buttons.
2. AtlasMart runbook: the invariant sequence
- Detect and timestamp: capture the SLO/test/metric that fired plus run/source/partition identifiers.
- Contain: mark affected products stale/unsafe; stop further certification if correctness is uncertain.
- Assess blast radius: traverse lineage to models, metrics, marts, reports, and consumers.
- Communicate: state observed impact, affected period/products, last-known-good state, and next update trigger.
- Diagnose: classify source, pipeline/schema, warehouse performance, or semantic mechanism from evidence.
- Repair: change the smallest justified component; preserve raw evidence and audit trail.
- Backfill/replay: process only affected partitions/assets when evidence supports scoping.
- Verify/reconcile: rerun critical Chapter 24 tests and canonical controls.
- Re-certify and communicate resolution.
- Postmortem: record cause, contributing conditions, detection gaps, and owned prevention actions.
3. Incident 1 runbook: source lateness
| Step | Evidence/action |
|---|---|
| Detect | source manifest absent at 08:10; stale-product SLO opens |
| Contain | keep previous certified partition labeled with its as-of time; do not fabricate current rows |
| Communicate | finance/operations told 2026-09-21 partition is delayed |
| Repair | source manifest arrives 09:05; normal validated load begins |
| Backfill | process only 2026-09-21 partition |
| Verify | 10/8/12/820/495/325 controls pass |
| Resolve | certified 09:18; record 48-minute certified SLO miss |
4. Incident 2 runbook: schema drift
| Step | Evidence/action |
|---|---|
| Detect |
line_amount_usd missing;
gross_line_amount_usd added
|
| Contain | contract gate blocks transform/publication; retain raw payload |
| Blast radius | 15 dependent catalog assets from source field to consumer groups |
| Diagnose | producer schema/contract change; semantics not assumed |
| Repair | owner-approved compatibility mapping and contract v2 |
| Backfill | replay 2026-09-21 from preserved raw |
| Verify | schema + lineage + metric controls pass before certification |
5. Incident 3 runbook: semantic metric bug
| Step | Evidence/action |
|---|---|
| Detect | cross-tool reconciliation: dashboard 665 USD vs governed 820 USD |
| Contain | invalidate affected semantic cache/report certification |
| Blast radius | 9 metric/mart/report/consumer assets |
| Diagnose | dashboard-local channel filter excluded 155 USD sales population |
| Repair | compile/consume governed metric definition |
| Backfill | recompute semantic cache/product for affected partition/window |
| Verify | 820 USD plus lineage/version checks; notify consumers of corrected result |
6. Consumer communication is evidence-bearing, not vague
incident: INC-SEM-20260921status: identified -> repairingaffected: gross_revenue_usd.v1 consumers for 2026-09-21observed: dashboard 665 USD; governed control 820 USDlast known good: prior certified metric contract/resultcontainment: affected report certification removedrepair: restore governed filter and rebuild semantic cacheverification required: 820 USD + Chapter 24 blockers PASSnext update trigger: repair verification complete or scope changes
A production communication channel may be ticketing/chat/status-page software. The content contract is vendor-neutral: what is affected, what is known/unknown, what consumers should do, and what event triggers the next update.
7. Safe backfill rules
| Rule | Reason |
|---|---|
| Declare affected partition/window | prevents accidental full-history rewrite |
| Use preserved raw/source evidence | keeps repair reproducible |
| Use same governed transformation version or explicit corrected version | avoids hidden semantic drift |
| Isolate resources when current loads must continue | limits operational blast radius |
| Re-run reconciliation and history/idempotency tests | proves convergence, not merely completion |
| Record batch/run/correction IDs | supports audit and rollback analysis |
8. Postmortem: facts, mechanisms, and actions—not blame
SOURCE LATENESSowner: ERP Operationsaction: publish manifest readiness metric and escalate at source-ready SLO breachdue: 2026-10-05SCHEMA DRIFTowner: Data Platformaction: require contract compatibility check before transform; retain raw payload for replaydue: 2026-10-03SEMANTIC BUGowner: Analytics Governanceaction: compile dashboard revenue from governed metric contract; add cross-tool golden testdue: 2026-10-07
The dates are synthetic lab commitments. What matters is that each prevention item is concrete, owned, and verifiable.
9. Controlled failure: postmortem with no owner
Wrong: “We should improve monitoring and be more careful.” No owner, due date, verification, or failure mode means no control changes. Repair: link each action to a detected gap (missing source-readiness escalation, missing compatibility gate, duplicated metric logic), assign an accountable owner, and define completion evidence.
10. End-to-end repaired state
canonical paid lines 10canonical paid orders 8canonical units 12canonical revenue 820.00 USDcanonical cost 495.00 USDcanonical gross profit 325.00 USDcommitted source seq 208source lateness repair PASSschema-drift repair PASSsemantic-bug repair PASSincident evidence hash 16659742da1bd9e56bf933f9748b7dc21c0b8ce823dae96d1866416ad1f4ba1d
The same controls after repair demonstrate convergence for this fixture. They do not prove that all unknown business facts are correct; they prove the encoded contracts and incident scenarios return to the accepted state.
11. What the runbook must preserve for rollback
Keep immutable/raw evidence, contract/schema snapshots, transformation/metric version, lineage graph version, failed test results, incident timeline, correction/backfill run IDs, and last-known-good certified artifacts where policy permits. Rollback may mean restoring the last certified product while a corrected partition is rebuilt; it should not mean erasing the evidence of what happened.
12. Security and privacy during incidents
Debug logs, quarantined rows, exports, screenshots, backups, and copied SQL can broaden exposure during response. Apply Chapter 22 least privilege to incident data, avoid posting raw PII/secrets in chat/tickets, and follow organization/jurisdiction policy for retention and erasure. Operational urgency does not nullify access controls.
13. Final production judgment and bridge to Chapter 26
Detectability: monitor signals that map to consumer risk. SLO impact: state the missed objective precisely. Communication: notify affected consumers before and after repair. Diagnose/repair: preserve evidence and match repair to mechanism. Backfill safety: scope partitions/assets and prove replay convergence. Learning: prevention actions need owner/due/evidence. Performance: incident response may reveal queue/scan/join bottlenecks, but optimization must preserve logical semantics. Chapter 26 continues with measured performance engineering and workload management.
14. Verification checklist
- Confirm the runbook distinguishes detection, containment, diagnosis, repair, backfill, verification, and re-certification.
- Confirm schema drift names 15 impacted assets and semantic drift names 9.
- Confirm the late source is communicated as stale rather than replaced with fabricated current data.
- Confirm every isolated repair reconciles to 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit.
- Confirm the postmortem contains three concrete owners/actions/due dates.
-
Confirm the incident evidence fingerprint is
16659742da1bd9e56bf933f9748b7dc21c0b8ce823dae96d1866416ad1f4ba1d. - Confirm cleanup/reset leaves no external service dependency.
Knowledge check
Acceptance questions
- Why is communication part of correctness operations?
- What evidence justifies a partition-scoped backfill?
- What makes a postmortem action useful?
- Why preserve failed evidence after resolution?
Review the answers
1. Consumers may already have used stale/wrong data and need explicit scope, mitigation, and corrected-state information.
2. Lineage/history evidence showing the defect is limited to that partition/window plus unaffected controls elsewhere.
3. It targets a specific failure mechanism, has an owner/due date, and has observable completion evidence.
4. It supports audit, root-cause validation, regression tests, and safe future replay/rollback analysis.
Authoritative references
- OpenTelemetry — Metrics specificationAuthoritative metric concepts and separation of instrumentation from collection/export behavior.
- OpenTelemetry — Metrics data modelUseful for reasoning about time-series measurements, aggregation, temporality, and preserving metric semantics.
- Google SRE Book — Service Level ObjectivesDefines SLIs/SLOs and emphasizes selecting indicators from user-relevant service behavior rather than everything that is easy to measure.
- Google SRE Book — Monitoring Distributed SystemsBackground on actionable monitoring signals and avoiding monitoring that produces noise without operator decisions.
- SQLite — TransactionsSupports the local demonstration of atomic repair/backfill state and safe commit boundaries.
- Python — statisticsUsed for small deterministic baseline examples; production anomaly detection needs workload-specific statistical validation.
- Python — hashlibUsed to fingerprint deterministic incident evidence and repaired outputs.
15. Lab cleanup/reset
The mandatory lab is entirely local. Close the process and delete any optional scratch outputs; rerunning the internal fixture reproduces the same incident evidence hash. No paid monitoring, incident-management, cloud warehouse, or managed orchestration product is required.