Make federation versus movement a workload decision with explicit consistency and freshness.
Federated Query vs Data Movement, One-Copy Dreams, Performance, Consistency, and Governance
Build a hybrid analytical architecture whose storage, transaction, semantic, and serving responsibilities remain independently testable.
Learning outcomes
Contrast federated reads with deliberate data movement/materialization using freshness, latency, scan/movement, consistency, and governance evidence.
Explain why “one copy” is not automatically cheaper, faster, simpler, or more consistent.
Identify remote dependency, snapshot/version, cache, egress, and access-control boundaries in a federated query.
Design a materialized serving copy with refresh, reconciliation, lineage, invalidation, and rollback contracts.
Choose per workload rather than enforcing federation or copying as a universal architecture rule.
Chapter 28 begins from the accepted state through Chapter 27: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit, with source progress through sequence 208. The declared sales grain remains one paid order line. The ERP remains the operational system of record for order state; bronze stores received replay evidence; silver holds validated analytical atomic state; the dimensional warehouse serves governed BI; and the semantic layer still owns metric meaning. A lakehouse or open table format changes storage/transaction/interop mechanisms, not those business contracts.
Mandatory work is local, free, synthetic, and executed with Python. It writes JSONL object-store fixtures plus synthetic metadata examples explicitly labeled not real Iceberg/Delta/Hudi tables. All copies reconcile to 820 / 495 / 325 USD. The federation/materialization comparison is an explicit calculation: 120 dashboard queries/day × 320 MiB remote scan = 37.5 GiB/day federated remote reads; hourly materialization moves 28.8 GiB/day and scans 4.688 GiB/day locally. Modeled latency is 2,500 ms vs 180 ms and freshness is 2 minutes vs at most 60 minutes. These are teaching assumptions, not cloud measurements.
1. Problem frame
AtlasMart has a certified dashboard queried 120 times per day. The team can query the silver open table directly through a warehouse federation connector, or refresh a warehouse-serving copy hourly. “One copy” sounds elegant, but the actual decision depends on remote scan work, dependency latency, freshness, consistent snapshot reads, egress/transfer, consumer concurrency, and governance.
Federation saves a copy; materialization buys isolation/locality at the cost of another state surface.
2. Define federation and movement
Federated query executes a query across data that remains under another system/storage owner, typically relying on a connector/catalog and remote reads. Materialization deliberately creates a derived local copy or aggregate for serving. Neither choice is inherently more correct. Federation reduces copy count but couples queries to remote availability, formats, metadata, network paths, and pushdown capabilities. Materialization adds refresh/staleness and storage/invalidation responsibilities but can isolate BI from remote variance.
3. Executed local calculation
| Property | Federation model | Materialized serving model |
|---|---|---|
| Dashboard queries/day | 120 | 120 |
| Read work | 320 MiB remote/query = 37.5 GiB/day | 40 MiB local/query = 4.688 GiB/day |
| Refresh movement | 0 serving-copy refresh | 1.2 GiB × 24 = 28.8 GiB/day |
| Modeled latency | 2,500 ms | 180 ms |
| Freshness | 2 min assumption | ≤60 min assumption |
| Extra serving copy | No | Yes |
queries = 120federated_gib = queries * 320 / 1024materialized_refresh_gib = 24 * 1.2materialized_local_scan_gib = queries * 40 / 1024print(federated_gib) # 37.5print(materialized_refresh_gib) # 28.8print(materialized_local_scan_gib) # 4.6875
These calculations compare work under stated assumptions. They do not claim a specific cloud bill or measured latency. A real design must add provider/network pricing, compression, caching, partition pruning, concurrency, and actual query profiles.
4. Consistency and snapshot pinning
A federated query that reads multiple tables must know what “same point in time” means. If one table advances between subqueries, the result can combine inconsistent states unless the engine/catalog/table format offers a documented snapshot/version mechanism and the query pins it correctly. Materialization can instead publish an atomic serving batch after all source dependencies reconcile—but that only works if the refresh commit and consumer switch are atomic for the serving engine.
5. Governance and security
Federation often requires the query engine/service identity to reach remote object storage/catalogs. The semantic/BI user should not automatically inherit direct bronze/silver object access. Materialization reduces some remote read paths but creates an additional governed copy whose retention, masking, row/column policies, backups, lineage, and deletion propagation must be managed. “Fewer copies” is not equivalent to “fewer security controls.”
6. Controlled failure: one-copy federation is free and consistent
Wrong approach: route every dashboard to remote silver because it avoids duplication. During peak BI load the connector loses predicate pushdown and scans broad files; one remote table advances while another remains at an older snapshot; latency and results become unstable.
Repair: define a query corpus, inspect remote pushdown/scan evidence, pin consistent snapshot/version semantics where supported, set freshness/availability SLOs, and materialize high-value repeatable workloads when the isolation/locality benefit exceeds refresh/copy cost. Keep ad-hoc/audit workloads federated when freshness/open access dominates latency.
7. Decision record
dashboard_revenue: mode: materialized reason: high repetition + strict latency + governed hourly freshness reconciliation: must equal gross_revenue_usd.v1 rollback: previous certified serving versionaudit_trace: mode: federated reason: low frequency + need newest accepted silver snapshot snapshot_rule: pin one documented table version per query fallback: export bounded snapshot for incident analysis
Knowledge check
What hidden cost can federation introduce even with no serving copy?
Show answer
Remote scans, network/egress, connector limits, weaker pushdown, dependency latency, and governance/credential complexity.
What new failure mode does materialization introduce?
Show answer
Staleness and divergence: the serving copy can lag or fail to reconcile to the source/atomic state, so it needs refresh/invalidation/lineage tests.
Why is one-copy not automatically more consistent?
Show answer
If a query spans independently advancing remote tables or snapshots, it can still observe inconsistent logical states unless a documented consistency mechanism pins them.
Summary and next step
This lesson established the mechanism and production boundaries for Federated Query vs Data Movement, One-Copy Dreams, Performance, Consistency, and Governance while preserving AtlasMart’s declared grain, governed metrics, history, and reconciliation evidence. Continue to Design a Hybrid Lakehouse+Warehouse Architecture with Clear System of Record, Serving, and Data-Movement Boundaries with those contracts unchanged unless an explicit, tested migration says otherwise.
Authoritative references
- Apache Iceberg — Table specificationOfficial specification for snapshot metadata, schema/partition evolution, manifests, and table-state commits.
- Apache Iceberg — Queries and metadata tablesOfficial documentation for snapshots, history, metadata logs, and time-travel query behavior.
- Delta Lake — Official documentationOfficial project documentation for Delta transaction-log tables, ACID semantics, schema enforcement/evolution, history, and time travel.
- Delta Lake — Table utility commandsOfficial details for table history, table metadata, restore, and retention-sensitive cleanup behavior.
- Apache Hudi — TimelineOfficial description of Hudi table actions/instants and the timeline as table-state metadata.
- Apache Hudi — Schema evolutionOfficial project documentation for supported schema-evolution behavior and version-specific caveats.
- Apache Hudi — Table and query typesOfficial description of Copy-on-Write and Merge-on-Read tradeoffs.
- Databricks — Medallion architectureA current product documentation example of bronze/silver/gold layering; used here as one implementation pattern, not as a prerequisite or universal definition.
- Kimball Group — Dimensional modeling techniquesReference for dimensional grain, fact/dimension modeling, conformance, and BI-serving semantics that remain separate from table-format mechanics.
8. Lab cleanup/reset
Delete
reports/federation_vs_materialization.json and
rerun the deterministic calculation. No network benchmark or
cloud bill is implied.