Present evidence, ownership, limits, and evolution—not a feature checklist.

Present the Warehouse as an Evidence-Based Data Product with Lineage, Costs, Known Limitations, Ownership, and Evolution Roadmap

Package AtlasMart as an evidence-based governed data product with lineage, cost assumptions, known limits, accountable ownership, and evolution criteria.

Intermediate → Advanced160–210 minutesdata-product handofflineage + cost + roadmapLast reviewed: September 2026

Learning outcomes

01

Present AtlasMart as an owned analytical data product with explicit consumers, contracts, SLOs, lineage, access, runbooks, and evidence.

02

Distinguish measured/verified guarantees from simulations, assumptions, and known limitations.

03

Publish a cost model without confusing fictional teaching rates with provider pricing.

04

Define evolution triggers, deprecation/versioning, and rollback criteria so future change is governed.

05

Complete the course with an evidence pack that another engineer can review, reproduce, and operate.

Capstone continuity: production meaning is frozen while operations are exercised

Chapter 30 begins from the accepted AtlasMart state produced by the earlier chapters: 10 paid order lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The accepted source waterline remains 208. The atomic sales grain remains one paid order line identified by (order_id, line_no); gross_revenue_usd.v1 remains the governed paid-line revenue metric in USD with UTC time semantics. CDC, defect, backfill, performance, security, and disaster-recovery exercises run against disposable copies and must reconcile back to this baseline.

Executed local harness and exact scope

The capstone evidence was executed with Python 3.13.5 and SQLite 3.46.1 on synthetic data in UTC. SQLite supplies relational constraints, transactions, indexes, query-plan evidence, and file backup/restore. It does not reproduce distributed shuffle, cloud IAM, managed row policies, autoscaling, object-store catalogs, multi-region DR, or provider billing. Those concepts are represented only as explicit policy fixtures, scheduler/cost calculations, or architecture decisions. No production credentials, personal data, paid services, or network dependencies are required.

1. Problem frame: a warehouse is not finished when the demo ends

The capstone now works locally, but a durable data product needs more than a successful demonstration. Six months later, a new engineer must know what “revenue” means, which source/waterline is trusted, who owns failures, which dashboards depend on the metric, what the current freshness objective is, what tuning evidence exists, which security paths are allowed, how to backfill/restore, and which claims were never proven. That information must live with the product.

2. The AtlasMart data-product card

Field Capstone declaration
Product AtlasMart governed analytical sales product
Primary consumers Finance, merchandising, marketing, operations
Atomic grain one paid order line per (order_id, line_no)
Current controls 10 lines · 8 orders · 12 units · 820 revenue · 495 cost · 325 profit
Source progress accepted waterline 208
Metric gross_revenue_usd.v1; USD; UTC; governed filters
Service objective certified complete data within 30 minutes of source readiness (teaching fixture policy)
Semantic owner finance semantic owner
Operational owner data-platform on-call
Security model least-privilege governed surfaces; raw bypass denied for BI reader
Recovery bounded replay/backfill + tested local backup/restore runbook

The card is concise because detailed evidence lives behind each field. It is an entry point, not a substitute for contracts/tests/runbooks.

3. Lineage: trace a source field to decisions and consumers

The final lineage graph reaches beyond the warehouse table. A source amount flows through raw/integration fact state into the governed metric, dependent marts, dashboards, and named consumers. That end-to-end graph supports both impact analysis and incident blast-radius communication.

lineage excerpt
source.erp.sales  -> raw.sales_events  -> integration.fact_sales       -> semantic.gross_revenue_usd.v1            -> mart.finance_daily_sales                 -> dashboard.finance_daily            -> dashboard.executive_revenue            -> consumer.finance_controller       -> mart.marketing_sales

Lineage freshness matters. A graph that was correct three months ago but misses a new dashboard can understate blast radius. Automated extraction should be supplemented by human business context, ownership, certification, and review.

4. Evidence pack: every claim points to an artifact

Claim Evidence artifact Current result
Atomic correctness evidence_summary.json + reconciliation query (10,8,12,820,495,325)
History SCD test evidence C001 at Sep18 noon = Growth; no overlap; one current
Quality quarantine log 13 staged -> 10 accepted + 3 quarantined
Restart/idempotency CDC state trace rollback -> apply -> duplicate ignore -> delete -> duplicate ignore
Semantic consistency metric golden test 820 governed vs 665 deliberate drift
Performance benchmark distributions + plans 120k fact / 180 aggregate; equivalent result
Workload isolation scheduler simulation shared p95 wait 1.0s; isolated 0.0s; not a DB benchmark
Security access_matrix.json BI raw bypass denied
Backfill bounded Sep20 repair 185/130/55 restored
Recovery backup/restore hash + controls damaged 635 revenue -> restored 820; hash equal
Operations runbook.md contain -> classify -> repair -> replay -> reconcile -> certify

5. Cost evidence: show assumptions beside the number

The capstone’s teaching model computes 72.275 USD/month from fictional rates: 0.10 TB-month storage × 0.25 USD/TB-month, two compute units × 3 hours/day × 0.40 USD/unit-hour × 30 days, plus 5 GB egress × 0.05 USD/GB. It is useful because the units and formula are explicit; it is not a quote, forecast, or provider comparison.

Cost dimension Teaching assumption Production replacement
Storage 0.10 TB-month × 0.25 USD/TB-month current provider/region/table-format/compression/storage-tier pricing
Compute 2 units × 3 h/day × 0.40 × 30 actual warehouse/slot/credit/serverless billing semantics
Egress 5 GB × 0.05 USD/GB current source/destination region and network path
Refresh/materialization implicitly included only in teaching compute hours measure actual refresh/query schedules and concurrency

Cost must be reconciled with performance and freshness. A materialization that reduces query latency can add refresh compute/storage. Federation can reduce copies but increase remote scans/latency. No single architecture is “cheapest” independent of workload.

6. Known-limit register: uncertainty is part of the product

Known limitation Operational consequence / next validation
SQLite is the local harness; no distributed shuffle/autoscaling/cloud IAM/managed row policies/provider billing retest plans, concurrency, access, and cost on the selected production engine
Warm-cache timings are machine-specific; OS cache is uncontrolled establish production benchmark environment and tail-latency regression thresholds
Cost rates are fictional teaching inputs replace with current region/contract pricing and measured workload consumption
Security identities are synthetic policy fixtures test effective inherited privileges, secret rotation, audit retention, exports/backups under organization policy
DR restores one local database file run separate object store/catalog/identity/key/multi-region/BI dependency recovery exercises

Hiding these limits would make the data product less trustworthy, not more polished. A limitation becomes actionable when it names the unproven behavior, the risk, the owner, and the evidence required to close it.

7. Ownership and operating cadence

The product needs regular evidence refresh. Source-contract owners review breaking changes. The semantic owner approves metric versions/deprecation. Platform on-call owns pipeline incidents and runbooks. Security owners review access/bypass evidence. Domain consumers own acceptance of dashboard/report migrations. Cost/performance owners review workload changes against benchmark and budget data.

Cadence/event Required review
every certified load freshness/completeness, quality, reconciliation, watermark, publish version
schema/contract change impact analysis, compatibility, owner approval, replay/backfill plan
metric definition change new version, golden tests, effective date, consumer migration/deprecation
material workload change representative benchmark, p95/p99, queue/scan/cost evidence
security/role change effective permission and bypass-path regression tests
incident/DR exercise runbook evidence, blast radius, restoration, corrective-owner follow-through

8. Evolution roadmap: change when a measurable trigger appears

An evolution roadmap should not be a shopping list of engines. AtlasMart defines triggers that justify a decision review.

Trigger Evidence threshold/question Candidate response
Freshness need tightens 30-minute fixture objective no longer meets consumer decision timing evaluate lower-latency ingestion/CDC/orchestration; preserve idempotency/history
Interactive p95/p99 regresses under ELT load queue/scan/resource evidence shows noisy-neighbor impact isolate workloads or scale resources; remeasure cost
Repeated BI scans dominate work query corpus shows stable higher-grain reuse materialize/accelerate with freshness/reconciliation contract
Open-table interoperability becomes material multiple engines need governed shared storage and snapshot semantics evaluate Iceberg/Delta/Hudi separately against required guarantees
Security/privacy obligation changes new policy/jurisdiction/data classification alters permitted use/retention update purpose/access/erasure/backup controls with legal/organizational review
Current engine cost/operations no longer fit measured workload + region/contract costs exceed accepted envelope rehost/replatform/refactor with Chapter 29 acceptance and rollback discipline

9. Release, deprecation, and rollback are product lifecycle controls

New model/metric versions should coexist long enough for consumer migration when semantics change. Deprecation needs owner, announcement, usage inventory, deadline, and rollback/compatibility path. Physical changes that preserve semantics can be rolled back by routing to the prior fact/materialization. A semantic change cannot be “rolled back” safely if history was destructively overwritten; versioning and reproducible source evidence are therefore prerequisites.

10. Controlled failure: hide the limitations and declare victory

Wrong approach

The final presentation shows a polished architecture diagram, the 0.009 ms materialized p50, and “DR ready.” It omits that the timing is local warm-cache SQLite, the cost rates are fictional, the workload-isolation result is simulated, and the DR drill covers only one database file.

Diagnosis: evidence has been stripped of scope, turning true local observations into false general guarantees.

Repair: present every result with dataset/runtime/cache/concurrency/storage/security scope, pair claims with known limits, name owners and next validation, and distinguish verified guarantees from simulations or assumptions.

11. Final capstone acceptance record

capstone acceptance
SEMANTICS     PASS  grain + metric v1 + history + controls preservedQUALITY       PASS  13 staged -> 10 accepted / 3 quarantinedCDC/RESTART   PASS  crash rollback, retry, duplicate and delete convergeRECONCILE     PASS  10 / 8 / 12 / 820 / 495 / 325SECURITY      PASS  BI raw/PII bypass denied in policy fixturePERFORMANCE   PASS  equivalent results; distributions + cache scope recordedBACKFILL      PASS  Sep20 aggregate repaired from atomic factRECOVERY      PASS  damaged primary restored; fact hash equalLINEAGE       PASS  source -> fact -> metric -> marts/reports/consumerCOST          PASS  formula + fictional assumptions explicitly labeledLIMITS        PASS  unproven cloud/distributed/DR/security behavior documentedOWNERSHIP     PASS  source/model/semantic/operations/security/consumer owners named

“PASS” here means the deterministic local acceptance criteria were met. It is not a certification for an unspecified production cloud. The evidence pack tells the next team exactly what must be revalidated when the engine, workload, region, security context, policy, or source changes.

12. Reproducible local evidence and cleanup

The harness writes its artifacts under atlasmart_ch30_lab. After reviewing evidence_summary.json, bus_matrix.json, source_contract.json, architecture_decisions.json, known_limits.json, lineage.json, access_matrix.json, and runbook.md, reset only the disposable fixture:

bash
python -c "import shutil; shutil.rmtree('atlasmart_ch30_lab', ignore_errors=True)"
Safety boundary

Do not reuse the destructive DR/delete/backfill commands against a real warehouse. Production exercises require approved scope, backups, access controls, change windows, consumer communication, and organization-specific recovery policy.

13. Production judgment: what this course has actually proven

The AtlasMart capstone demonstrates the full reasoning loop of production analytical warehousing: start from decisions and business processes; declare grain; design facts/dimensions/history; load incrementally with restart semantics; quarantine and reconcile; orchestrate based on readiness rather than clocks; govern metric definitions; choose physical/materialized/domain/cloud/lakehouse patterns from evidence; secure every access path; catalog/lineage/own the product; test and observe it; rehearse backfill, incident response, migration, and recovery; then communicate costs and limitations honestly.

The most important outcome is not a particular schema or engine. It is the discipline to make analytical meaning and operational guarantees observable, reproducible, owned, and reversible.

Knowledge check

Checkpoint

What is the strongest reason to keep a known-limit register?

Show answer

It prevents scoped evidence from becoming an unqualified guarantee and turns uncertainty into owned validation work.

Checkpoint

Why should an evolution roadmap use triggers rather than a list of technologies?

Show answer

Technology changes are justified when measured freshness, performance, cost, interoperability, security, or operational requirements no longer fit the current design.

Checkpoint

What remains invariant when the physical platform changes?

Show answer

Declared business processes/grain, governed metric/history semantics, source contracts, security intent, reconciliation expectations, ownership, and acceptance evidence unless explicitly versioned/migrated.

Checkpoint

When is the warehouse a data product rather than a collection of tables?

Show answer

When consumers can discover its meaning, owners, contracts, freshness, lineage, access, quality/reconciliation, incident/recovery procedures, costs, limitations, and change/deprecation process.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.