Separate warehouse semantics from cloud execution architecture and metering.
Shared-Nothing vs Disaggregated Storage/Compute vs Serverless Warehouses: Architectural Consequences
Build a representative workload profile before changing indexes, SQL, materializations, or concurrency policy.
Learning outcomes
Distinguish classic shared-nothing systems, disaggregated storage/compute systems, and serverless warehouse operating models without treating the labels as exact synonyms.
Trace one AtlasMart query through storage, metadata/control, compute, network/shuffle, cache, and result delivery boundaries.
Explain which responsibilities move from the data team to the platform and which remain with the data team.
Compare two documented cloud-warehouse styles using the same workload and explicit local assumptions.
Reject architecture comparisons that mix billing units or silently change the logical dimensional model.
Chapter 27 begins from the accepted AtlasMart state through Chapter 26: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit, with source progress committed through sequence 208. The governed metric contracts, security policies, lineage, tests, SLOs, incident practices, and representative performance corpus remain authoritative. The cloud calculations in this chapter alter only execution and cost assumptions; they never redefine grain, history, keys, or business metrics.
Execution: Python standard library only; no cloud account, paid service, or vendor SDK. Scale reference: Chapter 26’s deterministic 360,000-row performance fixture is used only to motivate workload shape. Cost workload: 30-day synthetic trace with BI, ad-hoc, ELT, and reconciliation classes. Storage: 500 GiB teaching assumption. Egress: 50 GiB/day teaching assumption. Price inputs: 5 USD/TiB scanned, 0.020 USD/GiB-month storage, 0.090 USD/GiB egress, 3 USD/credit, and 0.040 USD/slot-hour are hypothetical teaching inputs, not vendor price quotes. Time zone: UTC. Security: synthetic data only. Cloud limitation: no provisioning, cold-start, cache, network, or remote-shuffle behavior is claimed as measured; those are modeled or described from current official documentation.
1. Realistic problem: AtlasMart needs burst capacity without changing revenue
AtlasMart has one governed sales model but three execution pressures: a predictable nightly ELT window, bursty morning BI traffic, and exploratory analyst scans. The business question is not “which cloud warehouse is best?” It is whether a particular execution architecture can satisfy freshness, concurrency, latency, security, and cost constraints while preserving the same fact grain and metric semantics.
A shared-nothing warehouse distributes storage and compute across nodes whose local resources participate in execution. A disaggregated storage/compute warehouse persists data in a storage service that can be accessed by separately scalable compute. A serverless warehouse goes further operationally: the service provider hides most cluster provisioning and capacity-management mechanics behind a query/service interface. These are architectural/operating concepts, not mutually exclusive product labels.
2. Architecture comparison: what actually moves?
| Concern | Shared-nothing style | Disaggregated storage/compute | Serverless operating model |
|---|---|---|---|
| Persistent data | Often distributed with/among warehouse nodes | Central or managed remote storage independent of a compute pool | Managed remote storage; user often does not provision query workers |
| Compute lifecycle | Team sizes/rebalances clusters | Compute pools/warehouses can scale independently | Provider allocates compute; user controls quotas/reservations/policies to varying degrees |
| Failure domain | Node/storage balance matters | Compute can be replaced without moving the durable copy | Provider abstracts most worker lifecycle |
| Concurrency | Bound by cluster + WLM | Can isolate or multiply compute pools | May autoscale or allocate shared/dedicated capacity; quotas still exist |
| Cost unit | Node/cluster uptime common | Compute time + storage + transfer common | Bytes processed, slot/capacity time, request/compute units, storage, transfer, or hybrids |
| Portability | Distribution keys/layout tightly coupled | Logical model portable; physical layout still engine-specific | Logical model portable; billing/tuning/quotas highly provider-specific |
3. Two current documented styles
BigQuery-style serverless/disaggregated execution. Current Google documentation describes separate storage and compute layers, managed columnar storage, and workload management that can be billed on-demand by bytes processed or capacity-based by slot-hours. That means “serverless” does not mean “no capacity model”; it means the service abstracts infrastructure provisioning while still exposing workload/cost controls.
Snowflake-style disaggregated virtual warehouses. Current Snowflake documentation describes persistent cloud storage, independent virtual warehouses for compute, and a cloud-services layer for coordination. Separate warehouses can isolate compute, while auto-suspend/auto-resume can change idle cost and resume latency. This is not the same metering unit as bytes processed.
4. Observe one query end to end
consumer -> authentication / semantic policy -> metadata + optimizer -> choose/allocate compute -> read required columns/partitions from storage or cache -> local operators + network exchange/shuffle as needed -> aggregate / sort / finalize -> result cache or result delivery -> usage, lineage, audit, and cost telemetry
Separation of storage and compute changes where these steps occur and who provisions them; it does not remove them. A query that joins skewed dimensions can still create uneven work. A stale semantic aggregate can still return the wrong answer. A broad export can still create egress and privacy risk.
5. Controlled failure: treat bytes, slots, and credits as one “compute unit”
A migration spreadsheet says “1 TiB scanned ≈ 1 slot-hour ≈ 1 credit” and compares the numeric prices. This is dimensionally invalid. Bytes processed measure data volume under one billing model; slot-hours measure capacity over time; credits are a provider-defined consumption unit tied to compute/service usage. The same query can consume very different combinations depending on layout, cache, concurrency, and provider architecture.
Repair: preserve the workload trace, calculate each architecture in its native metering units, then convert to currency only after applying an explicit rate and region/contract assumption. Compare service-level outcomes alongside cost.
6. Local architecture/cost fixture
| Workload | Runs/day | Unoptimized scan GiB/run | Optimized scan GiB/run | Service-time assumption |
|---|---|---|---|---|
| certified_dashboard | 400 | 2.0 | 0.25 | 3.0 s |
| analyst_slice | 50 | 8.0 | 2.0 | 8.0 s |
| daily_product_refresh | 10 | 50.0 | 20.0 | 45.0 s |
| month_end_reconcile | 2 | 200.0 | 100.0 | 180.0 s |
The 30-day trace scans 63,000 GiB before the physical-design assumptions and 18,000 GiB after them—a calculated 71.43% reduction. This is a local model, not a cloud profile. It exists so every architecture receives the same workload and semantics.
7. What architecture does not decide for you
Storage/compute separation does not choose the right fact grain, SCD policy, revenue definition, row-level security rule, retention policy, data SLO, or incident owner. Serverless execution does not certify source completeness. Independent compute pools do not automatically prevent semantic drift. Keep these concerns above the engine so a migration can replace physical execution without rewriting business meaning.
Why is “serverless” not equivalent to “unlimited and maintenance-free”?
Show answer
Because serverless mainly changes infrastructure ownership and provisioning. Quotas, concurrency, queueing, cold starts, data layout, security, semantic correctness, cost controls, and incident responsibilities still exist, though their interfaces differ by service.
8. Production judgment and bridge
Choose a cloud execution model only after the logical warehouse, metric contracts, security boundaries, and reliability SLOs are fixed. Then compare workload-specific latency, concurrency, cold-start exposure, scan/compute/storage/egress units, region placement, ownership burden, and rollback options. A cloud service can automate provisioning without automating semantic correctness, cost governance, incident response, or workload prioritization. Keep vendor-specific physical settings in a decision record so a migration can preserve the logical model while replacing the execution strategy.
Next: Elastic Compute, Autoscaling, Concurrency, Pausing, Warehouses/Clusters, and Cold-Start Tradeoffs.
Knowledge check
Check your understanding
- What is the central mechanism in “Shared-Nothing vs Disaggregated Storage/Compute vs Serverless Warehouses: Architectural Consequences”, and which AtlasMart grain or metric contract must remain unchanged?
- Which observable evidence in this lesson distinguishes the correct design from the controlled failure?
- Which assumptions are local or engine-specific, and what must be re-checked before production use?
Review the answers
1. Preserve the lesson’s declared business grain, history semantics, governed metric definitions, and reconciliation controls while changing only the mechanism under study.
2. Use the lesson’s counts, sums, checksums, plans, traces, timing/cost calculations, or failure-state evidence—not a green task status or naming convention alone.
3. Re-check runtime/version, data scale and distribution, cache/concurrency, storage layout, security context, pricing/region where relevant, and the exact product guarantees before production adoption.
Authoritative references
- Google Cloud — BigQuery overviewCurrent documentation for BigQuery's serverless model and separation of storage and compute.
- Google Cloud — BigQuery storage overviewCurrent documentation for managed columnar storage, independent storage/compute scaling, and analytical scan behavior.
- Google Cloud — BigQuery workload managementCurrent documentation distinguishing on-demand bytes processed from capacity-based slot-hour models.
- Google Cloud — BigQuery pricingCurrent pricing-model documentation; this chapter deliberately uses hypothetical local rates instead of freezing region-dependent price numbers.
- Snowflake — Key concepts and architectureCurrent documentation for Snowflake storage, compute, and cloud-services layers and independent virtual warehouses.
- Snowflake — Warehouses overviewCurrent documentation for virtual warehouses, auto-suspend, and auto-resume.
- Snowflake — Warehouse considerationsCurrent documentation for warehouse credit metering, per-second billing after a 60-second minimum, and suspension tradeoffs.
- Snowflake — Understanding overall costCurrent documentation separating compute, storage, and data-transfer costs.
9. Lab cleanup/reset
The mandatory lab is local and synthetic. Delete
ch27_lab/ and rerun the embedded Python calculation
to reproduce the workload/cost evidence. No cloud warehouse,
object-store bucket, reservation, virtual warehouse, billing
account, or credential is created. If you optionally reproduce
examples in a real cloud, use a separate bounded sandbox and
follow that provider’s current cleanup and billing guidance.