Operate an AtlasMart data-incident runbook that detects, communicates, repairs, backfills, reconciles, and learns from failures with owned postmortem actions.

Build a Data Incident Runbook with Detection, Blast Radius, Consumer Communication, Repair, Backfill, and Postmortem

Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.

Intermediate → Advanced150–190 minutesIncident runbook labDetection + communication + postmortemLast reviewed: September 2026

Learning outcomes

01

Build a runbook with detection, containment, blast-radius analysis, communication, repair, backfill, and verification.

02

Record incident timelines and last-known-good state.

03

Produce concise postmortems with concrete owners and due dates.

04

Demonstrate safe partition-scoped repair and re-certification.

05

Turn incident learning into test/contract/observability improvements.

Continuity: observability watches the governed system; it does not redefine it

Chapter 25 begins from the accepted AtlasMart state established through Chapters 01–24: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current accepted paid order line, source progress remains committed through sequence 208, Chapter 20 metric contracts remain authoritative, Chapter 21 certified marts remain dependent on conformed assets, Chapter 22 security/privacy controls remain in force, Chapter 23 lineage/ownership metadata supplies blast-radius context, and Chapter 24 tests remain the correctness gates. Observability adds continuous evidence and incident handling; it must not silently reinterpret business rules merely to make a dashboard look healthy.

Executed local reliability fixture

Runtime: Python 3.13.5 + SQLite 3.46.1. Storage: local in-memory structures plus SQLite semantics where transactions matter. Clock: UTC with explicit timestamps. Security: synthetic identifiers only. Cost: free/local. Healthy run: 10 records in/out, 4,820 bytes, 7.0 minutes, zero rejects/retries, 860 ms fixture CPU, 34 MB fixture peak memory, and certification 12 minutes after source readiness. Important limitation: these CPU/memory numbers are fixture measurements used to teach signal relationships; they are not performance recommendations for any warehouse engine or cloud service.

1. The realistic problem: reliability depends on what happens after the alert

AtlasMart can detect a freshness miss, schema drift, or revenue discrepancy. Reliability still fails if nobody knows who owns the response, consumers are not warned, raw evidence is overwritten, a backfill corrupts current partitions, or the postmortem produces no corrective action. A runbook is an operational decision path, not a list of vendor buttons.

2. AtlasMart runbook: the invariant sequence

  1. Detect and timestamp: capture the SLO/test/metric that fired plus run/source/partition identifiers.
  2. Contain: mark affected products stale/unsafe; stop further certification if correctness is uncertain.
  3. Assess blast radius: traverse lineage to models, metrics, marts, reports, and consumers.
  4. Communicate: state observed impact, affected period/products, last-known-good state, and next update trigger.
  5. Diagnose: classify source, pipeline/schema, warehouse performance, or semantic mechanism from evidence.
  6. Repair: change the smallest justified component; preserve raw evidence and audit trail.
  7. Backfill/replay: process only affected partitions/assets when evidence supports scoping.
  8. Verify/reconcile: rerun critical Chapter 24 tests and canonical controls.
  9. Re-certify and communicate resolution.
  10. Postmortem: record cause, contributing conditions, detection gaps, and owned prevention actions.

3. Incident 1 runbook: source lateness

Step Evidence/action
Detect source manifest absent at 08:10; stale-product SLO opens
Contain keep previous certified partition labeled with its as-of time; do not fabricate current rows
Communicate finance/operations told 2026-09-21 partition is delayed
Repair source manifest arrives 09:05; normal validated load begins
Backfill process only 2026-09-21 partition
Verify 10/8/12/820/495/325 controls pass
Resolve certified 09:18; record 48-minute certified SLO miss

4. Incident 2 runbook: schema drift

Step Evidence/action
Detect line_amount_usd missing; gross_line_amount_usd added
Contain contract gate blocks transform/publication; retain raw payload
Blast radius 15 dependent catalog assets from source field to consumer groups
Diagnose producer schema/contract change; semantics not assumed
Repair owner-approved compatibility mapping and contract v2
Backfill replay 2026-09-21 from preserved raw
Verify schema + lineage + metric controls pass before certification

5. Incident 3 runbook: semantic metric bug

Step Evidence/action
Detect cross-tool reconciliation: dashboard 665 USD vs governed 820 USD
Contain invalidate affected semantic cache/report certification
Blast radius 9 metric/mart/report/consumer assets
Diagnose dashboard-local channel filter excluded 155 USD sales population
Repair compile/consume governed metric definition
Backfill recompute semantic cache/product for affected partition/window
Verify 820 USD plus lineage/version checks; notify consumers of corrected result

6. Consumer communication is evidence-bearing, not vague

Minimal incident status record
incident: INC-SEM-20260921status: identified -> repairingaffected: gross_revenue_usd.v1 consumers for 2026-09-21observed: dashboard 665 USD; governed control 820 USDlast known good: prior certified metric contract/resultcontainment: affected report certification removedrepair: restore governed filter and rebuild semantic cacheverification required: 820 USD + Chapter 24 blockers PASSnext update trigger: repair verification complete or scope changes

A production communication channel may be ticketing/chat/status-page software. The content contract is vendor-neutral: what is affected, what is known/unknown, what consumers should do, and what event triggers the next update.

7. Safe backfill rules

Rule Reason
Declare affected partition/window prevents accidental full-history rewrite
Use preserved raw/source evidence keeps repair reproducible
Use same governed transformation version or explicit corrected version avoids hidden semantic drift
Isolate resources when current loads must continue limits operational blast radius
Re-run reconciliation and history/idempotency tests proves convergence, not merely completion
Record batch/run/correction IDs supports audit and rollback analysis

8. Postmortem: facts, mechanisms, and actions—not blame

Owned prevention actions from the fixture
SOURCE LATENESSowner: ERP Operationsaction: publish manifest readiness metric and escalate at source-ready SLO breachdue: 2026-10-05SCHEMA DRIFTowner: Data Platformaction: require contract compatibility check before transform; retain raw payload for replaydue: 2026-10-03SEMANTIC BUGowner: Analytics Governanceaction: compile dashboard revenue from governed metric contract; add cross-tool golden testdue: 2026-10-07

The dates are synthetic lab commitments. What matters is that each prevention item is concrete, owned, and verifiable.

9. Controlled failure: postmortem with no owner

Wrong: “We should improve monitoring and be more careful.” No owner, due date, verification, or failure mode means no control changes. Repair: link each action to a detected gap (missing source-readiness escalation, missing compatibility gate, duplicated metric logic), assign an accountable owner, and define completion evidence.

10. End-to-end repaired state

Reconciliation after every isolated repair
canonical paid lines     10canonical paid orders     8canonical units           12canonical revenue         820.00 USDcanonical cost            495.00 USDcanonical gross profit    325.00 USDcommitted source seq      208source lateness repair    PASSschema-drift repair       PASSsemantic-bug repair       PASSincident evidence hash    16659742da1bd9e56bf933f9748b7dc21c0b8ce823dae96d1866416ad1f4ba1d

The same controls after repair demonstrate convergence for this fixture. They do not prove that all unknown business facts are correct; they prove the encoded contracts and incident scenarios return to the accepted state.

11. What the runbook must preserve for rollback

Keep immutable/raw evidence, contract/schema snapshots, transformation/metric version, lineage graph version, failed test results, incident timeline, correction/backfill run IDs, and last-known-good certified artifacts where policy permits. Rollback may mean restoring the last certified product while a corrected partition is rebuilt; it should not mean erasing the evidence of what happened.

12. Security and privacy during incidents

Debug logs, quarantined rows, exports, screenshots, backups, and copied SQL can broaden exposure during response. Apply Chapter 22 least privilege to incident data, avoid posting raw PII/secrets in chat/tickets, and follow organization/jurisdiction policy for retention and erasure. Operational urgency does not nullify access controls.

13. Final production judgment and bridge to Chapter 26

Detectability: monitor signals that map to consumer risk. SLO impact: state the missed objective precisely. Communication: notify affected consumers before and after repair. Diagnose/repair: preserve evidence and match repair to mechanism. Backfill safety: scope partitions/assets and prove replay convergence. Learning: prevention actions need owner/due/evidence. Performance: incident response may reveal queue/scan/join bottlenecks, but optimization must preserve logical semantics. Chapter 26 continues with measured performance engineering and workload management.

14. Verification checklist

  1. Confirm the runbook distinguishes detection, containment, diagnosis, repair, backfill, verification, and re-certification.
  2. Confirm schema drift names 15 impacted assets and semantic drift names 9.
  3. Confirm the late source is communicated as stale rather than replaced with fabricated current data.
  4. Confirm every isolated repair reconciles to 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit.
  5. Confirm the postmortem contains three concrete owners/actions/due dates.
  6. Confirm the incident evidence fingerprint is 16659742da1bd9e56bf933f9748b7dc21c0b8ce823dae96d1866416ad1f4ba1d.
  7. Confirm cleanup/reset leaves no external service dependency.

Knowledge check

Acceptance questions

  1. Why is communication part of correctness operations?
  2. What evidence justifies a partition-scoped backfill?
  3. What makes a postmortem action useful?
  4. Why preserve failed evidence after resolution?
Review the answers

1. Consumers may already have used stale/wrong data and need explicit scope, mitigation, and corrected-state information.

2. Lineage/history evidence showing the defect is limited to that partition/window plus unaffected controls elsewhere.

3. It targets a specific failure mechanism, has an owner/due date, and has observable completion evidence.

4. It supports audit, root-cause validation, regression tests, and safe future replay/rollback analysis.

Authoritative references

15. Lab cleanup/reset

The mandatory lab is entirely local. Close the process and delete any optional scratch outputs; rerunning the internal fixture reproduces the same incident evidence hash. No paid monitoring, incident-management, cloud warehouse, or managed orchestration product is required.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.