Automatic capacity changes who operates the scaler, not the workload physics or billing inputs.
Provisioned Throughput, Autoscaling, Serverless Billing, Storage, Egress, and Request Pricing
Model provisioned, autoscaling, and serverless billing together with storage, indexes, backups, egress, scale lag, throttling, and hot-partition limits instead of assuming automatic scaling removes capacity engineering.
Distinguish provisioned throughput, autoscaling, and pay-per-request/serverless billing models without assuming any one model is intrinsically cheaper.
Build a complete cost model including requests/compute, storage, indexes, backups, replication, egress, minimums, and idle periods.
Observe autoscaling lag and hot-partition throttling even when aggregate service capacity eventually becomes sufficient.
Use workload measurements and sensitivity analysis rather than copying example provider prices into architecture decisions.
1. Capacity mode is both a queueing policy and a billing policy
AtlasMart has three demand shapes: a catalog with a stable daytime baseline, flash-sale bursts, and low-volume internal tools. Provisioned throughput reserves a configured service rate and usually creates an idle-capacity trade. Autoscaling changes provisioned capacity according to observed usage or policy. Serverless/pay-per-request removes explicit capacity provisioning and meters work performed. These terms are service-specific; they do not standardize the unit, scale speed, burst budget, partition ceiling, or minimum charge.
Cost must be joined to performance evidence. A plan that appears inexpensive because requests are throttled is not a valid comparison. Record useful completed operations, p95/p99 latency, errors/throttles, consistency/durability mode, payload size, index fan-out, replication regions, and retries. Retries can increase billed operations while useful throughput falls.
2. Build the bill from workload dimensions, not one monthly number
A request-unit service may charge a normalized unit based on CPU, memory, I/O, payload and query work. Another service may charge read/write operations by size. A provisioned product may charge capacity-hours, node/vCPU/memory, or a mixture. Therefore AtlasMart’s worksheet keeps the formula abstract: compute/request work + stored data/indexes + backup/PITR + replicated copies + network egress + minimum/idle commitments + support/adjacent services.
| Dimension | Questions | Failure/cost coupling |
|---|---|---|
| Request/compute | Read/write mix? point reads or scans? item size? consistency? | poor query/index design can multiply metered work |
| Autoscaling | Reaction time? min/max? cooldown? per-partition behavior? | burst arrives before scale; quota or hot key throttles |
| Storage/indexes | Raw bytes plus index/metadata/versions? | alternate views improve reads but amplify storage/write cost |
| Replication | How many regions/copies? writes replicated/billed? | global durability/latency choice changes cost and conflict model |
| Egress | Cross-region, internet, analytics/backup movement? | architecture can shift cost outside database line item |
A service can increase table/account throughput while one partition or logical key remains the bottleneck. Current DynamoDB documentation explicitly discusses adaptive capacity and partition maximums; current Cosmos DB autoscale guidance also calls out scaling behavior around active partitions. Exact numbers and algorithms are implementation details and must be rechecked.
3. Deliberately wrong approach: extrapolate average traffic and assume the scaler is instantaneous
If AtlasMart sizes from an hourly mean, a four-minute flash sale can exceed the currently active capacity before the scaler reacts. Even after table capacity rises, a celebrity SKU can remain hot. The resulting throttles trigger retries, which can make the cost and load worse. A second error is comparing only request price while ignoring indexes, backup retention and cross-region egress.
4. AtlasMart lab: scale lag, hot-key throttling, and normalized cost
Python 3.13+ standard library only. Every dollar coefficient in this lab is explicitly hypothetical and exists only to exercise the formula. Do not use these values for a cloud quote.
from math import ceil
# Hypothetical teaching coefficients only -- NOT a cloud-provider price quote.
price = {
"provisioned_unit_hour": 0.00012,
"autoscale_peak_unit_hour": 0.00016,
"million_request_units": 0.28,
"gb_month_storage": 0.22,
"gb_month_backup": 0.06,
"gb_egress": 0.08,
}
# One hour, per-minute demand. A 4-minute flash sale creates a sharp spike.
demand = [180] * 20 + [900, 1300, 1500, 1100] + [260] * 36
# One hot key can consume no more than this synthetic per-partition ceiling.
hot_partition_demand = [70] * 20 + [300, 520, 680, 450] + [90] * 36
partition_ceiling = 400
# Synthetic autoscaler: starts at 300 units and reacts to previous minute with a 2-minute lag.
capacity = []
current = 300
for minute, d in enumerate(demand):
observed = demand[max(0, minute - 2)]
target = max(300, ceil(observed / 100) * 100)
current = min(max(current, target), 1600)
capacity.append(current)
throttled_total = 0
hot_throttled_total = 0
for d, h, c in zip(demand, hot_partition_demand, capacity):
throttled_total += max(0, d - c)
hot_throttled_total += max(0, h - partition_ceiling)
storage_gb = 240
index_gb = 95
backup_gb = 335
egress_gb = 480
hours_month = 730
provisioned_units = 900
autoscale_hourly_peak = max(capacity)
request_units_month = 860_000_000
provisioned_cost = provisioned_units * hours_month * price["provisioned_unit_hour"]
autoscale_cost = autoscale_hourly_peak * hours_month * price["autoscale_peak_unit_hour"]
serverless_cost = request_units_month / 1_000_000 * price["million_request_units"]
common = ((storage_gb + index_gb) * price["gb_month_storage"]
+ backup_gb * price["gb_month_backup"]
+ egress_gb * price["gb_egress"])
print("AUTOSCALING TIMELINE AROUND FLASH SALE")
for i in range(18, 27):
print(f"minute={i:02d} demand={demand[i]:4} capacity={capacity[i]:4} hot_key={hot_partition_demand[i]:3}")
print("table-level units throttled by scale lag:", throttled_total)
print("hot-partition units throttled despite table capacity:", hot_throttled_total)
print("\nHYPOTHETICAL MONTHLY COST MODEL (teaching inputs, not provider prices)")
print(f"common storage/index/backup/egress = ${common:,.2f}")
print(f"provisioned = ${provisioned_cost + common:,.2f}")
print(f"autoscale = ${autoscale_cost + common:,.2f}")
print(f"serverless = ${serverless_cost + common:,.2f}")
print("lesson: capacity mode changes billing mechanics, not hot-key physics or cost accounting")
The synthetic scaler lags the flash-sale demand, so aggregate work is throttled before capacity catches up. Separately, the hot key exceeds its synthetic partition ceiling even when total table capacity is sufficient. The printed monthly totals demonstrate how a capacity mode changes the formula, not which mode is universally cheapest.
5. Production judgment: price the SLO, not the raw operation
Use current provider calculators/APIs only after measuring the workload. Price several scenarios: normal month, launch burst, region failover, index rebuild/backfill, replay after outage, backup restore and unexpected egress. Include quotas and throttling in acceptance tests. Budget alarms are useful but do not replace hard safety controls such as maximum autoscale settings, request admission, tenant quotas, and retry backoff.
For serverless, low idle cost can be compelling for intermittent workloads. For sustained predictable traffic, committed/provisioned models can be more efficient. Neither statement is universal because price schedules, regions, editions and discounts change. Keep usage units and formulas in source control; resolve current prices at decision time. Next, Chapter 24 moves from one-region capacity to global placement, where latency, conflict semantics and residency can outweigh pure request cost.
Check your understanding
- Why can an autoscaling table still throttle?
- What costs belong in a serverless request model besides request units?
- Why is idle cost not the whole comparison?
- Why are example provider prices dangerous in a durable architecture document?
- What metric should accompany a cost-per-request calculation?
Review the answers
1. Scale can lag demand, quotas can cap growth, and a single hot partition/key can hit a local ceiling even when aggregate capacity is available.
2. Storage, indexes, backups/PITR, replication, egress, optional support/observability and adjacent services.
3. Provisioned modes may pay for reserved throughput while pay-per-request modes may cost more under sustained high utilization; workload shape determines the trade.
4. Rates vary by region, edition and time; record formulas and usage assumptions, then pull current price data during procurement.
5. At minimum latency/error/throttling under the same workload, because a cheap request that violates the SLO is not equivalent capacity.
References
Provider-specific claims below are current implementation anchors reviewed in August 2026. They are not universal NoSQL definitions, and mandatory labs do not require the services.
- Amazon DynamoDB Pricing — Official current pricing dimensions for reads/writes, storage and optional features.
- DynamoDB — On-demand capacity mode — Official scaling and table-level quota behavior for on-demand mode.
- Azure Cosmos DB — Request Units — Official RU abstraction and provisioned/serverless/autoscale throughput models.
- Azure Cosmos DB — Plan and manage costs — Official cost-planning guidance covering RU, storage and additional meters.
- Google Cloud Firestore Pricing — Official billing dimensions including document operations, index reads, storage and network bandwidth.