Balance queueing, cold starts, idle cost, and workload isolation with measured traces.
Elastic Compute, Autoscaling, Concurrency, Pausing, Warehouses/Clusters, and Cold-Start Tradeoffs
Build a representative workload profile before changing indexes, SQL, materializations, or concurrency policy.
Learning outcomes
Define elasticity, autoscaling, concurrency, pausing/suspension, compute pools/warehouses/clusters, and cold start in operational terms.
Separate queue wait from service time and reason about scaling as a latency/cost tradeoff.
Calculate the idle-cost effect of auto-suspend policy and the latency effect of repeated resumes.
Explain why autoscaling does not eliminate workload management or noisy-neighbor controls.
Design different policies for interactive BI and batch ELT while preserving one governed warehouse model.
Chapter 27 begins from the accepted AtlasMart state through Chapter 26: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit, with source progress committed through sequence 208. The governed metric contracts, security policies, lineage, tests, SLOs, incident practices, and representative performance corpus remain authoritative. The cloud calculations in this chapter alter only execution and cost assumptions; they never redefine grain, history, keys, or business metrics.
Execution: Python standard library only; no cloud account, paid service, or vendor SDK. Scale reference: Chapter 26’s deterministic 360,000-row performance fixture is used only to motivate workload shape. Cost workload: 30-day synthetic trace with BI, ad-hoc, ELT, and reconciliation classes. Storage: 500 GiB teaching assumption. Egress: 50 GiB/day teaching assumption. Price inputs: 5 USD/TiB scanned, 0.020 USD/GiB-month storage, 0.090 USD/GiB egress, 3 USD/credit, and 0.040 USD/slot-hour are hypothetical teaching inputs, not vendor price quotes. Time zone: UTC. Security: synthetic data only. Cloud limitation: no provisioning, cold-start, cache, network, or remote-shuffle behavior is claimed as measured; those are modeled or described from current official documentation.
1. Realistic problem: scale for the 09:00 burst or pay all day?
AtlasMart sees six short BI bursts per day, a nightly ELT refresh, and a late reconciliation run. Keeping peak compute permanently warm may waste money. Suspending aggressively may add resume latency and minimum-billing events. Elasticity is the ability to add or remove capacity as demand changes. Autoscaling is an automated policy for doing that. Concurrency is simultaneous work admitted to the system. Cold start is added latency when compute, cache, or execution state must be re-established.
2. Capacity is not the same as admission policy
A system can have enough theoretical compute yet still queue work because of admission limits, workload isolation, quotas, or a cold pool. Conversely, admitting more work can increase service time through contention. That is why Chapter 26’s decomposition still applies:
consumer_elapsed = queue_wait + service_time + result_delivery# cold/resume delay can appear before admission or inside service,# depending on the product's reporting model.
3. Local concurrency simulation
| Policy | Workers | Burst makespan | p95 queue |
|---|---|---|---|
| Fixed capacity | 2 | 144 s | 54 s |
| Scaled capacity | 4 | 108 s | 18 s |
Doubling the simulated workers reduces the burst p95 queue from 54 s to 18 s and makespan from 144 s to 108 s. The model does not claim linear speedup in a real warehouse: shared storage, network, cache, shuffle, and per-query parallelism can become bottlenecks.
4. Auto-suspend/resume calculation
| Teaching policy | Resume windows / month | Billed compute seconds | Credits at 1 credit/hour | Compute cost at hypothetical 3 USD/credit |
|---|---|---|---|---|
| Suspend after 60 s | 240 | 36000 | 10.00 | $30.00 |
| Suspend after 600 s | 240 | 165600 | 46.00 | $138.00 |
The modeled 10-minute idle policy consumes 4.6× the compute credits of the 60-second policy for this burst pattern. Both have the same 240 resume windows in this deliberately separated schedule, so our assumed 3-second resume penalty totals 720 seconds either way. Real provisioning latency and billing behavior are provider- and workload-specific; Snowflake documentation, for example, states per-second warehouse billing with a 60-second minimum each time compute is provisioned and notes that resumption can take time.
5. Controlled failure: “autoscaling means workload management is obsolete”
Suppose ELT and BI share one autoscaling pool. ELT can consume newly added capacity just as fast as BI, so the platform scales cost without restoring interactive latency. The repair is to combine elasticity with admission/isolation: separate workload pools or reservations where justified, cap runaway exploratory work, give certified BI an explicit service objective, and isolate backfills from current loads.
6. Pausing and cache state
Suspension can save compute spend, but it may discard or reduce warm execution/cache state depending on the engine. A “first query after resume” should therefore be measured separately from steady-state warm performance. Never compare a warm benchmark from one architecture to a cold benchmark from another and call the difference architectural superiority.
7. Production policy record
| Workload | Latency sensitivity | Suggested control surface | What to monitor |
|---|---|---|---|
| Certified BI | High | interactive pool/reservation; bounded autoscale | queue p95/p99, cold-start rate, freshness |
| Ad-hoc | Medium/variable | budget/concurrency cap; separate queue if noisy | scan/compute units, queue, cancellation |
| Daily ELT | Freshness deadline | batch pool/warehouse; scheduled elasticity | duration, spill/shuffle, retries, freshness |
| Backfill | Low immediate latency; high resource demand | isolated pool and explicit cap | resource share, current-load impact, cost |
The exact names—warehouse, reservation, resource group, cluster—are product vocabulary. The portable decision is which workload gets which capacity, isolation, queueing, and budget contract.
What evidence would justify a longer auto-suspend interval?
Show answer
A measured workload with short recurrent idle gaps where repeated resume/minimum-billing or cold-start penalties cost more than remaining warm, while service objectives and budget remain satisfied. The threshold is workload- and provider-specific.
8. Production judgment and bridge
Choose a cloud execution model only after the logical warehouse, metric contracts, security boundaries, and reliability SLOs are fixed. Then compare workload-specific latency, concurrency, cold-start exposure, scan/compute/storage/egress units, region placement, ownership burden, and rollback options. A cloud service can automate provisioning without automating semantic correctness, cost governance, incident response, or workload prioritization. Keep vendor-specific physical settings in a decision record so a migration can preserve the logical model while replacing the execution strategy.
Next: Object Storage, Columnar Formats, Metadata Services, Caching, and Remote Shuffle Concepts.
Knowledge check
Check your understanding
- What is the central mechanism in “Elastic Compute, Autoscaling, Concurrency, Pausing, Warehouses/Clusters, and Cold-Start Tradeoffs”, and which AtlasMart grain or metric contract must remain unchanged?
- Which observable evidence in this lesson distinguishes the correct design from the controlled failure?
- Which assumptions are local or engine-specific, and what must be re-checked before production use?
Review the answers
1. Preserve the lesson’s declared business grain, history semantics, governed metric definitions, and reconciliation controls while changing only the mechanism under study.
2. Use the lesson’s counts, sums, checksums, plans, traces, timing/cost calculations, or failure-state evidence—not a green task status or naming convention alone.
3. Re-check runtime/version, data scale and distribution, cache/concurrency, storage layout, security context, pricing/region where relevant, and the exact product guarantees before production adoption.
Authoritative references
- Google Cloud — BigQuery overviewCurrent documentation for BigQuery's serverless model and separation of storage and compute.
- Google Cloud — BigQuery storage overviewCurrent documentation for managed columnar storage, independent storage/compute scaling, and analytical scan behavior.
- Google Cloud — BigQuery workload managementCurrent documentation distinguishing on-demand bytes processed from capacity-based slot-hour models.
- Google Cloud — BigQuery pricingCurrent pricing-model documentation; this chapter deliberately uses hypothetical local rates instead of freezing region-dependent price numbers.
- Snowflake — Key concepts and architectureCurrent documentation for Snowflake storage, compute, and cloud-services layers and independent virtual warehouses.
- Snowflake — Warehouses overviewCurrent documentation for virtual warehouses, auto-suspend, and auto-resume.
- Snowflake — Warehouse considerationsCurrent documentation for warehouse credit metering, per-second billing after a 60-second minimum, and suspension tradeoffs.
- Snowflake — Understanding overall costCurrent documentation separating compute, storage, and data-transfer costs.
9. Lab cleanup/reset
The mandatory lab is local and synthetic. Delete
ch27_lab/ and rerun the embedded Python calculation
to reproduce the workload/cost evidence. No cloud warehouse,
object-store bucket, reservation, virtual warehouse, billing
account, or credential is created. If you optionally reproduce
examples in a real cloud, use a separate bounded sandbox and
follow that provider’s current cleanup and billing guidance.