Verify each open table format’s actual metadata and compatibility semantics.
Apache Iceberg/Delta/Hudi Table-Format Concepts, Snapshot Metadata, Schema Evolution, and Warehouse Interoperability
Build a hybrid analytical architecture whose storage, transaction, semantic, and serving responsibilities remain independently testable.
Learning outcomes
Explain snapshots/versions, metadata commits, schema evolution, and interoperability as table-format mechanisms rather than generic “lakehouse magic.”
Contrast Iceberg, Delta Lake, and Hudi at the level of their documented metadata/protocol concepts without pretending the guarantees are identical.
Read synthetic metadata examples and identify what would need to be verified against a real engine/catalog implementation.
Explain why schema evolution compatibility includes reader/writer versions and downstream semantic contracts, not just whether a write succeeds.
Preserve AtlasMart grain, metric controls, and history while adding a nullable channel attribute.
Chapter 28 begins from the accepted state through Chapter 27: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit, with source progress through sequence 208. The declared sales grain remains one paid order line. The ERP remains the operational system of record for order state; bronze stores received replay evidence; silver holds validated analytical atomic state; the dimensional warehouse serves governed BI; and the semantic layer still owns metric meaning. A lakehouse or open table format changes storage/transaction/interop mechanisms, not those business contracts.
Mandatory work is local, free, synthetic, and executed with Python. It writes JSONL object-store fixtures plus synthetic metadata examples explicitly labeled not real Iceberg/Delta/Hudi tables. All copies reconcile to 820 / 495 / 325 USD. The federation/materialization comparison is an explicit calculation: 120 dashboard queries/day × 320 MiB remote scan = 37.5 GiB/day federated remote reads; hourly materialization moves 28.8 GiB/day and scans 4.688 GiB/day locally. Modeled latency is 2,500 ms vs 180 ms and freshness is 2 minutes vs at most 60 minutes. These are teaching assumptions, not cloud measurements.
1. Problem frame
AtlasMart wants to add a nullable channel field to
its validated atomic sales table. One engineer says, “Iceberg,
Delta, and Hudi all support schema evolution, so the migration
is identical.” That is precisely the kind of category-level
shortcut this lesson rejects.
Table formats solve overlapping problems with different metadata models and compatibility rules.
2. Snapshot/version metadata is the table
In an open-table architecture, the logical table is not “all files under a prefix.” Readers use table metadata to identify the committed files/state. Iceberg documents metadata files, snapshots, manifest lists/manifests, schema IDs and field IDs. Delta Lake documents table versions/history in its transaction log. Hudi records table actions as instants on a timeline and exposes storage/query choices such as Copy-on-Write and Merge-on-Read. The identifiers and commit protocols differ, so application code must use the actual project/engine semantics.
3. Synthetic metadata fixture — intentionally not executable product metadata
{ "iceberg_style": { "snapshot_id": 102, "parent_snapshot_id": 101, "schema_id": 2, "fields": [ {"id": 1, "name": "order_id", "type": "string"}, {"id": 3, "name": "line_amount_usd", "type": "decimal"}, {"id": 9, "name": "channel", "type": "string", "optional": true} ] }, "delta_style": { "table_version": 8, "previous_version": 7, "actions": ["metaData(schema adds nullable channel)", "add(part-000.parquet)"] }, "hudi_style": { "instant": "20260922080000", "state": "COMPLETED", "action": "commit", "table_type": "COW" }}
The fixture teaches what categories of evidence to look for. It does not reproduce any project’s serialization, catalog, concurrency, or engine integration.
4. Schema evolution is more than adding a column
Iceberg’s specification tracks fields by IDs and documents metadata-only schema operations such as add/drop/rename/reorder with defined type-promotion rules. Delta Lake documents schema enforcement/evolution and table history/versioning, with behavior depending on Delta version/features such as column mapping and the specific DML operation. Hudi’s current schema-evolution documentation distinguishes supported backward-compatible changes and version/table-type caveats. Therefore a migration record should state: writer version, reader versions, catalog/engine versions, operation, whether old snapshots remain readable, and what downstream SQL/metrics expect.
| Question | Iceberg | Delta Lake | Hudi |
|---|---|---|---|
| Versioned table state | Snapshots / metadata files | Transaction-log versions/history | Timeline instants/actions |
| Schema identity/evolution | Field IDs + schema IDs/spec rules | Schema enforcement/evolution + protocol/features | Schema-on-write/read behavior with documented compatibility limits |
| Update-oriented tradeoff | Engine/format implementation dependent | MERGE/update/delete supported by Delta semantics | CoW vs MoR is an explicit table-type tradeoff |
5. Controlled failure: copy an Iceberg rename guarantee into Hudi/Delta
Wrong approach: because one format tracks
stable field identity for renames, declare that renaming
line_amount_usd is metadata-only and safe
everywhere. A reader on another engine interprets the schema
differently or requires a feature/configuration that was never
enabled.
Repair: test the exact operation on the target project/engine versions, inspect old and new snapshots/versions, run representative readers, verify lineage/metric SQL, and keep the old field/version available until consumers migrate. Treat compatibility as an end-to-end contract, not merely a table-writer success.
6. Reconcile after evolution
before = {"revenue_usd":820, "cost_usd":495, "gross_profit_usd":325}after = {"revenue_usd":820, "cost_usd":495, "gross_profit_usd":325, "new_nullable_attribute":"channel"}assert before["revenue_usd"] == after["revenue_usd"]assert before["gross_profit_usd"] == after["gross_profit_usd"]print("schema evolution changed representation, not business meaning")
7. Interoperability boundaries
An “open” format improves the possibility that multiple engines can read the same table, but interoperability is versioned. A reader may lag a table-format feature, data type, deletion mechanism, catalog API, encryption feature, or protocol version. Certification must therefore include a reader/writer compatibility matrix and negative tests. Also remember that multiple engines reading the same table still need one governed metric contract; openness at storage does not guarantee semantic consistency.
Knowledge check
Why are the three JSON blocks labeled synthetic?
Show answer
Because they illustrate evidence categories only. They are not valid serialized table metadata and do not prove real transaction semantics.
Does a successful schema-evolution write prove all consumers are compatible?
Show answer
No. Readers, catalogs, streaming jobs, semantic models, lineage, and historical snapshots/versions can have separate compatibility constraints.
What must remain unchanged when adding channel?
Show answer
The declared fact grain, existing history/key rules, and governed metric results unless the business contract explicitly versions a change.
Summary and next step
This lesson established the mechanism and production boundaries for Apache Iceberg/Delta/Hudi Table-Format Concepts, Snapshot Metadata, Schema Evolution, and Warehouse Interoperability while preserving AtlasMart’s declared grain, governed metrics, history, and reconciliation evidence. Continue to Federated Query vs Data Movement, One-Copy Dreams, Performance, Consistency, and Governance with those contracts unchanged unless an explicit, tested migration says otherwise.
Authoritative references
- Apache Iceberg — Table specificationOfficial specification for snapshot metadata, schema/partition evolution, manifests, and table-state commits.
- Apache Iceberg — Queries and metadata tablesOfficial documentation for snapshots, history, metadata logs, and time-travel query behavior.
- Delta Lake — Official documentationOfficial project documentation for Delta transaction-log tables, ACID semantics, schema enforcement/evolution, history, and time travel.
- Delta Lake — Table utility commandsOfficial details for table history, table metadata, restore, and retention-sensitive cleanup behavior.
- Apache Hudi — TimelineOfficial description of Hudi table actions/instants and the timeline as table-state metadata.
- Apache Hudi — Schema evolutionOfficial project documentation for supported schema-evolution behavior and version-specific caveats.
- Apache Hudi — Table and query typesOfficial description of Copy-on-Write and Merge-on-Read tradeoffs.
- Databricks — Medallion architectureA current product documentation example of bronze/silver/gold layering; used here as one implementation pattern, not as a prerequisite or universal definition.
- Kimball Group — Dimensional modeling techniquesReference for dimensional grain, fact/dimension modeling, conformance, and BI-serving semantics that remain separate from table-format mechanics.
8. Lab cleanup/reset
Delete metadata/table_format_examples.json and
recreate it from the local script. To test a real format, use
that project’s official quickstart in an isolated disposable
environment and record the exact engine/format versions.