Combine fast deterministic CI checks with production data tests, severity-driven quarantine and alerting, and governed time-bounded waivers that cannot suppress semantic blockers silently.
Create CI + Production Data Tests with Severity, Quarantine, Alerting, and Waiver Governance
Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.
Learning outcomes
Separate fast deterministic CI tests from time-relative production tests.
Assign severity to failure modes and map severity to block/quarantine/alert actions.
Govern waivers with owner, reason, approval, scope, and expiry.
Prove quarantine/reprocess does not corrupt accepted controls.
Build a test evidence record suitable for incident response and Chapter 25 observability.
Chapter 24 begins from the governed AtlasMart state produced by Chapters 01–23: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The current fact grain remains one accepted current paid order line; Chapter 20 metric definitions remain authoritative; Chapter 21 certified marts remain dependent on conformed assets; Chapter 22 security controls remain in force; and Chapter 23 ownership/lineage metadata supplies the change context. A test may block, quarantine, alert, or document a failure, but it must not silently change business semantics to make the suite green.
Runtime:
Python 3.13.5 + SQLite 3.46.1.
Storage: in-memory SQLite.
Time zone: UTC. Currency: USD.
Source/warehouse grain: one paid order line
identified by (order_id, line_id), with a separate
immutable delivery/event ID.
Security: synthetic customer IDs only.
Cost: free/local, no managed service.
Non-guarantee: passing these fixtures proves
the encoded contracts for this dataset/runtime; it does not
prove unknown business rules, external source truth, or
vendor-specific behavior not represented by the fixture.
1. The realistic problem: not every failure should do the same thing
An invalid negative sales row must not enter certified facts. A source freshness lag during a planned maintenance window may justify a bounded warning. A missing customer relationship may require quarantine. A changed revenue total against the golden fixture must block a release. If every test merely prints “FAILED,” operators either ignore alerts or overreact to harmless conditions.
2. CI tests and production tests answer different questions
| Surface | Good candidates | Avoid |
|---|---|---|
| CI / pull request | Unit transforms; schema contracts; negative fixtures; SCD boundaries; idempotency; golden metrics | Wall-clock freshness against live source; huge full-history scans |
| Pre-deploy/integration | Representative batch/restart/replay; migration compatibility; reconciliation on fixed fixture | Moving baselines without version/snapshot |
| Production | Freshness, volume, reject rate, relationships, reconciliation, SLOs, drift | Tests that mutate production evidence to make themselves pass |
Some checks can run in both places with different fixtures/severity. The design goal is high signal at the point where action can still prevent harm.
3. Severity must connect to action
| Severity | Example | Action in this chapter |
|---|---|---|
| Blocker | Grain duplicate, invalid range, orphan customer, golden metric mismatch | Stop certification/release; quarantine rows where row-scoped |
| Warning | 70-minute freshness lag during documented maintenance | Alert owner; allow only with valid bounded waiver |
| Info | Non-breaking metadata observation | Record for trend/review; no block |
These labels are AtlasMart fixture policy, not universal industry standards. The organization must define its own severity/action contract and review it against consumer risk.
4. Executed production gate before repair
blocker event_id_unique observed=1 expected=0blocker declared_grain_unique observed=2 expected=0blocker accepted_channel_values observed=1 expected=0blocker business_measure_ranges observed=1 expected=0blocker customer_relationship observed=1 expected=0warning freshness_under_60m observed=70 expected<=60
The gate is BLOCKED while any unwaived blocker exists. A warning does not turn blockers into warnings; severity is test-specific and cannot be downgraded merely because a release is inconvenient.
5. Quarantine is an action, not a synonym for ignore
Delivered staging rows: 13Quarantined: 3 row 11 -> DUPLICATE_EVENT+GRAIN_CONFLICT row 12 -> GRAIN_CONFLICT row 13 -> INVALID_CHANNEL+INVALID_RANGE+ORPHAN_CUSTOMERCertified accepted rows: 10Certified controls: (10, 8, 12, 820.0, 495.0, 325.0)
Each quarantined row needs a reason, raw evidence, owner/workflow, and reprocess path. The certified set reconciles exactly after quarantine; that demonstrates rejected records were not silently counted or partially applied.
6. Governed waivers require owner and expiry
required = [waiver.owner, waiver.reason, waiver.expires_at, waiver.approved_by]assert all(required)assert waiver.expires_at > run_timeassert test.severity != "blocker"# expiry means the next run fails/alerts normally unless renewed by review.
The negative fixture intentionally constructs a waiver with missing owner and expiry; governance rejects it. The valid fixture records:
W-20260921-01 test: freshness_under_60m owner: data-platform reason: planned ERP maintenance window expires_at: 2026-09-22T09:00:00Z approved_by: analytics-owner
A waiver is not a permanent configuration flag. It is a time-bounded risk decision with accountable ownership.
7. Alerting should carry diagnostic context
A useful alert includes test name/version, dataset/partition, severity, observed/expected values, source/batch/run IDs, owner, lineage/consumer blast radius, quarantine count, and runbook link. “Data test failed” is insufficient for operator ergonomics. Chapter 23 metadata supplies owners/consumers; Chapter 25 will turn this evidence into broader observability and incident response.
8. Deterministic evidence fingerprint
SHA-256 4dd1e1e2d460842058876b4d74a7fab4b23291acd272f9102f88fa3ae06085faFingerprint inputs: canonical controls quarantine rows + reason codes ordered test result records governed waiver record final CDC watermark
The fingerprint is a compact reproducibility check for this fixture. It does not replace the readable test log; a hash cannot explain a failure.
9. What not to waive
Do not use waivers to bypass unknown grain corruption, unresolved metric mismatches, security exposure, or irreversible history damage merely to meet a schedule. If business leadership accepts a known semantic exception, record it as a governed contract/version/risk decision with scope and migration plan—not as an invisible test suppression.
10. CI log versus production evidence retention
CI evidence can be short-lived if the source code/change record remains traceable. Production test results, quarantine events, waivers, and reconciliation decisions may require longer operational/audit retention. Exact retention depends on organization/jurisdiction/security policy; this lesson does not prescribe a universal period. Logs must avoid secrets and unnecessary PII.
11. Full acceptance matrix
| Failure mode | Test | Expected repair evidence |
|---|---|---|
| Schema drift | Contract/schema test | Versioned compatible schema or blocked migration |
| Bad row | Range/code/relationship tests | Quarantine + corrected source/mapping + successful reprocess |
| SCD overlap | History interval test | Non-overlapping half-open versions + boundary tests |
| Late-data misclassification | Event-time lookup test | Correct historical SK/segment |
| Crash/retry | CDC transaction/replay test | Watermark never leads target; duplicate ignored |
| Metric drift | Golden/reconciliation test | Governed metric output matches 820 USD fixture |
| Freshness lag | Production SLO test | Source recovery or bounded approved waiver |
12. Production judgment and bridge to Chapter 25
Signal quality: every test needs a failure mode and action. False positives: thresholds/expectations need ownership and review. Blast radius: use lineage/consumer inventory to prioritize incidents. CI speed: keep deterministic fixtures small; run heavier checks where they add signal. Security: test/quarantine evidence is governed data. Rollback: a failed release should preserve the last certified version and raw evidence for replay. Next: Chapter 25 observes freshness, volume, schema, lineage, quality, retries, resource use, and incidents continuously rather than only at test boundaries.
13. Verification checklist
- Run the lab and confirm the canonical controls are 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit.
- Confirm exactly three staging rows are quarantined and accepted controls remain unchanged.
- Confirm crash/retry ends with one E208 target effect and watermark 208.
- Confirm governed revenue is 820 USD and drifted dashboard result is 665 USD.
- Confirm the invalid waiver is rejected and the valid freshness waiver has owner/reason/approval/expiry.
-
Confirm the evidence fingerprint is
4dd1e1e2d460842058876b4d74a7fab4b23291acd272f9102f88fa3ae06085fa. - Rerun from a new process and confirm deterministic output.
Knowledge check
Acceptance questions
- Why separate CI tests from production tests?
- What makes a waiver governed rather than a suppression?
- Why should blockers not be downgraded to meet a release date?
- What information should a production alert carry?
Review the answers
1. CI protects deterministic code/contracts quickly; production tests observe time-relative real data/system conditions.
2. Explicit scope, owner, reason, approval, expiry, and auditability.
3. Severity maps to consumer/data risk, not schedule pressure.
4. Test/version, dataset/partition, observed/expected, severity, run/source IDs, owner, lineage/blast radius, and runbook context.
Authoritative references
- Python — unittestStandard-library support for deterministic fixtures, assertions, setup/teardown, and automated test suites.
- SQLite — CREATE TABLE / constraintsAuthoritative syntax and semantics for NOT NULL, CHECK, UNIQUE, PRIMARY KEY, and table constraints used in the local lab.
- SQLite — Foreign Key SupportExplains foreign-key behavior and the need to enable enforcement explicitly in SQLite connections.
- SQLite — TransactionsUsed to demonstrate that target writes, delivery ledger, and watermark advancement must commit atomically for safe restart.
- SQLite — UPSERTReference for conflict-aware idempotent target application in the CDC fixture.
- Python — hashlibUsed to fingerprint deterministic test evidence rather than compare unstable timestamps or environment-specific logs.
14. Lab cleanup/reset
All mandatory lab state is in memory. Close the process to reset accepted/quarantine/test/waiver/CDC state; rerun to reproduce the same evidence hash. No CI provider, production alerting system, or cloud data platform is required.