Chapter 12 · Source System Profiling, Data Contracts, Lineage, and Ingestion Readiness

Inventory Sources, Owners, Extraction Methods, Change Mechanisms, SLAs, and Failure Modes

Build a source inventory that records ownership, extraction and change semantics, freshness commitments, security boundaries, and failure modes before a warehouse pipeline is allowed to depend on a source.

Intermediate → Advanced105–125 minutesSource-inventory labPython 3.13.5 stdlib · local/syntheticLast reviewed: September 2026

Learning outcomes

AtlasMart is ready to add automated ingestion, but the team has only connection strings and table names. No one has written down who owns each source, whether an extract is a full snapshot or an increment, how updates/deletes become observable, which clock defines freshness, or what a partial export looks like. Coding first would convert unknown source behavior into hidden pipeline behavior.

01

Define a source inventory as an engineering contract rather than a list of database names.

02

Separate source owner, dataset owner/steward, extraction method, and change mechanism so accountability and mechanics are not conflated.

03

Document freshness/SLA assumptions, time zones, security boundaries, expected completeness, and failure modes before transformation code.

04

Explain why a high-water mark, snapshot timestamp, CDC log position, and file arrival are different evidence of change.

05

Use a deterministic inventory validator to block ingestion when ownership or change semantics are missing.

Chapter 12 continuity contract

Chapter 12 preserves the accepted AtlasMart warehouse contracts from Chapters 01–11. The current published sales state after Chapter 11 contains eight current paid order-line facts, five paid orders, ten units, 690 paid GMV, 425 cost-at-sale, and 265 gross profit. Customer history still uses half-open business-effective intervals, source-qualified durable identity, governed unknown members, and auditable corrections. This chapter moves upstream: it defines what evidence a source must provide before new ingestion or transformation code is trusted. The synthetic source extracts therefore reflect the accepted Chapter 11 corrections rather than silently reverting to the original 625-GMV fixture.

Execution and evidence boundary

The mandatory exercises are free, local, synthetic, and deterministic. They were validated with Python 3.13.5 using only standard-library data structures/JSON/hash functions. The examples prove the stated fixture, contract, and gate behavior; they do not prove production-source correctness, network extraction behavior, organizational ownership, legal compliance, or the behavior of a managed ingestion service. Treat every source statement as a contract that must be verified against the actual producer.

Important boundary

A source inventory describes how evidence can be obtained; it does not prove the producer actually satisfies the contract. Production readiness requires producer confirmation, monitored observations, access review, and failure/recovery tests.

1. Start with a source, owner, and business dependency—not a connector

A source system is an operational or upstream analytical system from which the warehouse receives evidence. A source dataset is the specific table, file, API resource, topic, or snapshot being depended on. An owner is the accountable team that can explain the source's semantics and authorize change. These definitions matter because a connector can successfully copy bytes from a dataset whose meaning is unknown.

AtlasMart needs sales, customer, and inventory evidence. The ingestion team should be able to answer: who confirms an order status definition, who explains a reused customer identifier, and who can tell whether an inventory file is complete? If the answer is “the pipeline code,” ownership is missing.

Source Owner Extract/change mechanism Freshness contract Primary risks
ERP orders/order lines Order Platform 15-minute incremental; updated_at high-water mark + overlap/reconciliation 95% certified ≤30 min after commit missed/duplicate/partial parent-child changes
CRM customers Customer Platform daily snapshot + updated_at evidence certified by 03:00 UTC late snapshot, code drift, deletion state loss
Inventory snapshot Supply Chain 30-minute full product snapshot; snapshot_ts certified ≤45 min from snapshot partial snapshot, duplicate product, unit change

2. Extraction method and change mechanism are separate contracts

An extraction method is how data is retrieved: full snapshot, incremental query, API page, file export, or log/CDC stream. A change mechanism is the evidence used to decide what changed: an updated_at column, monotonically increasing sequence, source log position, snapshot replacement, or explicit tombstone/delete event. Saying “we query ERP every 15 minutes” does not explain how the second query avoids missing or duplicating updates.

For the Chapter 12 fixture the ERP uses an updated_at high-water mark with overlap and reconciliation. That is a declared simulation, not a claim of exactly-once delivery. Chapter 15 will treat incremental/CDC mechanics in depth.

source_inventory.json
[  {    "source_id": "src.erp.orders",    "owner": "Order Platform",    "dataset": "ERP orders + order lines",    "extract": "15-minute incremental export",    "change_mechanism": "updated_at high-water mark with overlap/reconciliation",    "freshness": "95% of committed order changes certified within 30 minutes",    "timezone": "UTC/RFC3339",    "security": "synthetic order/customer identifiers; service identity read-only",    "failure_modes": [      "missed update on bad watermark",      "duplicate delivery after retry",      "partial header/line extract"    ]  },  {    "source_id": "src.crm.customers",    "owner": "Customer Platform",    "dataset": "CRM customers/segments",    "extract": "daily snapshot at 02:00 UTC",    "change_mechanism": "snapshot comparison + source updated_at evidence",    "freshness": "daily extract certified by 03:00 UTC",    "timezone": "UTC/RFC3339",    "security": "customer attributes classified restricted in production; synthetic fixture here",    "failure_modes": [      "late snapshot",      "silent code-set change",      "privacy/deletion state omitted"    ]  },  {    "source_id": "src.inventory.snapshot",    "owner": "Supply Chain",    "dataset": "inventory product snapshot",    "extract": "30-minute full product snapshot",    "change_mechanism": "snapshot_ts identifies state observation",    "freshness": "certified within 45 minutes of snapshot time",    "timezone": "UTC/RFC3339",    "security": "non-PII product/inventory state",    "failure_modes": [      "partial snapshot mistaken for complete",      "duplicate product row",      "unit change from units to cases"    ]  }]

3. SLA/freshness and failure modes turn expectations into observable statements

A freshness commitment needs a clock and endpoints. “ERP is real time” is ambiguous. AtlasMart instead states that 95% of committed order changes should be certified within 30 minutes. That still requires production telemetry to measure source commit time, extraction time, load time, and certification time. The local fixture records the contract but does not invent an observed p95.

Failure modes belong in the inventory because recovery design depends on them. A retry can duplicate an incremental export; a partial parent/child export can orphan order lines temporarily; a daily CRM snapshot can arrive late; an inventory snapshot can contain all valid rows but still be incomplete if the producer silently truncated the file.

4. Deliberately wrong approach: “SELECT * every 15 minutes”

Wrong

The team schedules SELECT * FROM orders every 15 minutes and treats each result as “changes.” Rows are copied repeatedly, deletes are invisible unless diffed against prior state, and there is no committed boundary between headers and lines. A successful query proves connectivity—not incremental correctness.

The repair is to document the source extraction/change contract before implementing it. For ERP that means a stable business key, updated_at semantics, overlap/reconciliation policy, delete representation, parent/child completeness rule, and recovery path. Later chapters implement those mechanics; Chapter 12 decides whether the source is sufficiently specified to begin.

5. Local lab — validate the source inventory gate

Save the script below as inventory_gate.py. It requires no external package. First run uses the valid inventory. The second run removes the CRM owner to demonstrate the gate.

inventory_gate.py
inventory = [ {"source_id":"src.erp.orders","owner":"Order Platform","extract":"15-minute incremental export","change_mechanism":"updated_at + overlap/reconciliation","freshness":"95% certified within 30 minutes","timezone":"UTC","security":"synthetic fixture"}, {"source_id":"src.crm.customers","owner":"Customer Platform","extract":"daily snapshot","change_mechanism":"snapshot comparison + updated_at evidence","freshness":"certified by 03:00 UTC","timezone":"UTC","security":"restricted in production; synthetic here"}, {"source_id":"src.inventory.snapshot","owner":"Supply Chain","extract":"30-minute full snapshot","change_mechanism":"snapshot_ts state observation","freshness":"certified within 45 minutes","timezone":"UTC","security":"non-PII product state"},]required=("source_id","owner","extract","change_mechanism","freshness","timezone","security")def gate(rows):    missing=[]    for row in rows:        for field in required:            if not row.get(field): missing.append((row.get("source_id","?"),field))    return "READY" if not missing else f"BLOCKED missing={missing}"print(gate(inventory))broken=[dict(x) for x in inventory]broken[1]["owner"]=""print(gate(broken))

Expected output:

expected output
READYBLOCKED missing=[('src.crm.customers', 'owner')]

The first line proves only that the local inventory artifact contains the required fields. The second proves that ownership is treated as a blocking dependency rather than documentation debt.

6. Security, replay, and observability consequences

Source access should be least privilege: extraction identities read only the required datasets and must not inherit broad administrative rights. CRM fields would be restricted in a real system; this course uses synthetic values. The inventory should also record whether a source can replay a historical interval, whether snapshots can be re-requested, and what evidence identifies a batch/version. Those properties determine future backfill and incident-response options.

Do not confuse scheduling success with source availability. A job can run on time against yesterday's stale snapshot. Production observability should measure source availability, extract completeness, lag, rejects, and contract/version identity independently.

7. Verification and cleanup

  • All three source IDs have an accountable owner and extraction/change description.
  • Every freshness statement identifies a time expectation; no fabricated runtime measurements are claimed.
  • CRM production sensitivity is distinguished from the synthetic fixture.
  • The broken-owner case returns BLOCKED.
  • Cleanup: delete inventory_gate.py; there is no database or cloud resource.

Knowledge check

Check your understanding

  1. Why is “query every 15 minutes” not a change mechanism?
  2. What does an owner add that a schema cannot?
  3. Does a documented 30-minute freshness contract prove the source meets it?
  4. Why record failure modes before implementation?
  5. What is the readiness decision if a critical source has no owner?
Review the answers

1. It states a schedule, not how inserts, updates, deletes, duplicates, or missed intervals are identified.

2. An accountable authority for business meaning, change communication, and ambiguity resolution.

3. No. It is an expectation that must be measured from production timestamps/telemetry.

4. Recovery, idempotency, completeness, and reconciliation designs depend on how the source can fail.

5. BLOCKED; unresolved ownership is a semantic/change-management risk, not merely missing prose.

Summary and next step

A usable source inventory records who owns meaning, how evidence is extracted, how change is recognized, when it is expected, and how it can fail. Lesson 2 now tests the source records themselves: keys, nulls, ranges, distributions, relationships, duplicates, and drift.

Authoritative references

  • W3C — PROV Overview — Official W3C overview of provenance concepts used to reason about entities, activities, agents, and lineage relationships.
  • OpenLineage — Specification — Open specification for dataset/job/run lineage events; referenced as a later operationalization option, not a prerequisite for this local lab.
  • JSON Schema — Specification — Official JSON Schema specification; useful when a source contract is represented as JSON, while the business semantics still require explicit agreement beyond structure.
  • RFC 3339 — Date and Time on the Internet — Authoritative timestamp format reference used for the UTC timestamp contract in the synthetic fixture.
  • Python documentation — csv — Standard-library CSV support suitable for free/local deterministic source fixtures.
  • Python documentation — json — Standard-library JSON support used for contract and readiness artifacts.
  • Python documentation — hashlib — Standard-library hashing API used only for reproducibility/evidence fingerprints, not as a semantic validation substitute.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.