Separate warehouse semantics from lakehouse storage and transaction mechanisms.
Warehouse vs Lakehouse: Governance, Transactions, Performance, Open Formats, and Ecosystem Tradeoffs
Build a hybrid analytical architecture whose storage, transaction, semantic, and serving responsibilities remain independently testable.
Learning outcomes
Distinguish a warehouse serving model from a lakehouse storage/transaction architecture without reducing either to a product label.
Separate governance, transaction semantics, performance, open formats, and ecosystem interoperability into independently testable concerns.
Explain why Parquet files alone are not a transactional table and why an open table format does not replace dimensional or semantic modeling.
Trace AtlasMart from operational system of record through bronze/silver state to warehouse serving while preserving grain and metric controls.
Evaluate a hybrid design by explicit copies, freshness, rollback, security, and consumer behavior rather than architecture fashion.
Chapter 28 begins from the accepted state through Chapter 27: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit, with source progress through sequence 208. The declared sales grain remains one paid order line. The ERP remains the operational system of record for order state; bronze stores received replay evidence; silver holds validated analytical atomic state; the dimensional warehouse serves governed BI; and the semantic layer still owns metric meaning. A lakehouse or open table format changes storage/transaction/interop mechanisms, not those business contracts.
Mandatory work is local, free, synthetic, and executed with Python. It writes JSONL object-store fixtures plus synthetic metadata examples explicitly labeled not real Iceberg/Delta/Hudi tables. All copies reconcile to 820 / 495 / 325 USD. The federation/materialization comparison is an explicit calculation: 120 dashboard queries/day × 320 MiB remote scan = 37.5 GiB/day federated remote reads; hourly materialization moves 28.8 GiB/day and scans 4.688 GiB/day locally. Modeled latency is 2,500 ms vs 180 ms and freshness is 2 minutes vs at most 60 minutes. These are teaching assumptions, not cloud measurements.
1. Problem frame
AtlasMart wants data science teams to read open object-store tables while finance requires stable dimensional queries and governed revenue. A proposal says “move everything to Parquet and the warehouse is obsolete.” That confuses a file representation, a transactional table abstraction, an analytical serving model, and a semantic contract. The business decision is whether one storage/transaction substrate can coexist with a dimensional serving layer without losing correctness, governance, latency, or portability.
Warehouse and lakehouse are overlapping architectural choices, not mutually exclusive religions.
2. Define the layers before comparing them
A data warehouse is an analytical system organized to deliver governed historical data and repeatable business queries. A lakehouse is an architectural approach that brings table-like transaction/metadata capabilities to data stored in lake-style object storage. A file format such as Parquet defines how bytes/columns are encoded in files. A table format defines table metadata, versions/snapshots, schema/partition evolution, and commit semantics across many files. A dimensional model declares business grain, facts, dimensions, history, and aggregation behavior. These are different layers.
| Concern | Warehouse serving | Lakehouse/open-table substrate | Still needs explicit design? |
|---|---|---|---|
| Business grain / metric meaning | Star/semantic contract | Not defined by table format | Yes |
| Atomic commit / versioned table state | Engine dependent | Core table-format concern | Yes, per format/engine |
| Open object files | May be internal/managed | Common | Yes |
| BI latency / workload isolation | Often optimized for serving | Depends on engine/cache/layout | Yes |
| Governance / access / lineage | Required | Required | Yes |
3. AtlasMart state surfaces
ERP order state -> bronze receipt log (replay evidence) -> silver validated atomic sales (one paid order line) -> dimensional warehouse fact_sales + conformed dimensions -> semantic metric gross_revenue_usd.v1 -> finance / operations / marketing consumersOpen-table metadata may wrap bronze/silver/gold storage,but it does not change the declared fact grain or metric formula.
The word gold is not a synonym for star schema. A gold table can be an aggregate, feature table, wide report, or dimensional mart depending on the product and organization. Conversely, a star schema can be physically stored in a warehouse engine or in open-table files. Model the consumer contract independently from the storage label.
4. Transactions and openness are format-specific
Iceberg, Delta Lake, and Hudi all manage table state above data files, but they do so with different metadata/protocol semantics. Iceberg defines snapshots, metadata files, manifests, field IDs, and partition/spec evolution. Delta Lake records table versions and actions in its transaction log and documents schema enforcement/evolution, history, restore, and retention-sensitive cleanup. Hudi centers table actions around its timeline and exposes Copy-on-Write and Merge-on-Read storage/query tradeoffs. Therefore “supports ACID/time travel/schema evolution” is only a category-level statement; exact isolation, retention, rename/evolution behavior, reader/writer compatibility, and maintenance must be verified per project and engine.
5. Controlled failure: call a folder of Parquet files a lakehouse
Wrong approach: write daily Parquet files to object storage and let every consumer list the directory. A writer partially replaces files while one BI query is scanning them. There is no shared table commit boundary, no durable schema/version contract, and no guarantee that two readers enumerate the same logical table state.
Diagnosis: columnar files improve scan economics, but file format metadata does not by itself supply table-level commit/version semantics. Repair: use a table format/catalog/engine combination whose documented transaction semantics fit the workload, or keep the managed warehouse as the transactional/serving owner. Then pin snapshot/version semantics where reproducibility matters and reconcile results to the dimensional controls.
6. Local lab: make boundaries executable
# Run locally with Python 3.x; standard library only.controls = {"paid_lines": 10, "orders": 8, "units": 12, "revenue_usd": 820, "cost_usd": 495, "gross_profit_usd": 325, "source_sequence": 208}boundaries = { "operational_system_of_record": "ERP", "bronze": "received source events / replay evidence", "silver": "validated atomic analytical state", "warehouse": "dimensional BI serving copy", "semantic": "governed metric definitions and filters"}assert controls["revenue_usd"] - controls["cost_usd"] == controls["gross_profit_usd"]print(controls)print(boundaries)
Expected control output is 10 / 8 / 12 / 820 / 495 / 325. The evidence proves the architecture did not redefine the business state. It does not prove any cloud engine’s transaction isolation, optimizer behavior, or throughput.
7. Production judgment
Choose a hybrid only when each state surface has an owner, freshness contract, security boundary, retention policy, lineage edge, reconciliation test, and rollback path. Open formats can reduce storage/engine coupling and improve interoperability, but every extra reader increases compatibility/testing obligations. A managed warehouse may provide stronger integrated serving/operations at the cost of provider coupling. The decision surface is governance + transaction needs + interoperability + BI latency + data-science access + operational ownership—not “lakehouse newer than warehouse.”
Knowledge check
Why is Parquet not equivalent to Iceberg/Delta/Hudi?
Show answer
Parquet is a file format. Iceberg, Delta Lake, and Hudi add table-level metadata and transaction/version semantics above data files; their exact semantics differ.
Can a star schema live in a lakehouse?
Show answer
Yes. Dimensional modeling is a logical/semantic design. Its tables can be implemented on an open-table substrate if the engine and table format meet correctness and serving requirements.
What does the 820 USD reconciliation prove?
Show answer
Only that the derived surfaces in the local fixture preserved the governed revenue control. It does not validate a real cloud table-format implementation.
Authoritative references
- Apache Iceberg — Table specificationOfficial specification for snapshot metadata, schema/partition evolution, manifests, and table-state commits.
- Apache Iceberg — Queries and metadata tablesOfficial documentation for snapshots, history, metadata logs, and time-travel query behavior.
- Delta Lake — Official documentationOfficial project documentation for Delta transaction-log tables, ACID semantics, schema enforcement/evolution, history, and time travel.
- Delta Lake — Table utility commandsOfficial details for table history, table metadata, restore, and retention-sensitive cleanup behavior.
- Apache Hudi — TimelineOfficial description of Hudi table actions/instants and the timeline as table-state metadata.
- Apache Hudi — Schema evolutionOfficial project documentation for supported schema-evolution behavior and version-specific caveats.
- Apache Hudi — Table and query typesOfficial description of Copy-on-Write and Merge-on-Read tradeoffs.
- Databricks — Medallion architectureA current product documentation example of bronze/silver/gold layering; used here as one implementation pattern, not as a prerequisite or universal definition.
- Kimball Group — Dimensional modeling techniquesReference for dimensional grain, fact/dimension modeling, conformance, and BI-serving semantics that remain separate from table-format mechanics.
8. Lab cleanup/reset
Delete atlasmart_ch28_lab/ and rerun the
standard-library fixture. No cloud bucket, metastore, warehouse,
or table-format service is created.