Automatic capacity changes who operates the scaler, not the workload physics or billing inputs.

Provisioned Throughput, Autoscaling, Serverless Billing, Storage, Egress, and Request Pricing

Model provisioned, autoscaling, and serverless billing together with storage, indexes, backups, egress, scale lag, throttling, and hot-partition limits instead of assuming automatic scaling removes capacity engineering.

Advanced135–175 minutesCost & autoscaling labPython 3.13+ · standard libraryHypothetical prices · no cloud chargesLast reviewed: August 2026
01

Distinguish provisioned throughput, autoscaling, and pay-per-request/serverless billing models without assuming any one model is intrinsically cheaper.

02

Build a complete cost model including requests/compute, storage, indexes, backups, replication, egress, minimums, and idle periods.

03

Observe autoscaling lag and hot-partition throttling even when aggregate service capacity eventually becomes sufficient.

04

Use workload measurements and sensitivity analysis rather than copying example provider prices into architecture decisions.

1. Capacity mode is both a queueing policy and a billing policy

AtlasMart has three demand shapes: a catalog with a stable daytime baseline, flash-sale bursts, and low-volume internal tools. Provisioned throughput reserves a configured service rate and usually creates an idle-capacity trade. Autoscaling changes provisioned capacity according to observed usage or policy. Serverless/pay-per-request removes explicit capacity provisioning and meters work performed. These terms are service-specific; they do not standardize the unit, scale speed, burst budget, partition ceiling, or minimum charge.

Cost must be joined to performance evidence. A plan that appears inexpensive because requests are throttled is not a valid comparison. Record useful completed operations, p95/p99 latency, errors/throttles, consistency/durability mode, payload size, index fan-out, replication regions, and retries. Retries can increase billed operations while useful throughput falls.

2. Build the bill from workload dimensions, not one monthly number

A request-unit service may charge a normalized unit based on CPU, memory, I/O, payload and query work. Another service may charge read/write operations by size. A provisioned product may charge capacity-hours, node/vCPU/memory, or a mixture. Therefore AtlasMart’s worksheet keeps the formula abstract: compute/request work + stored data/indexes + backup/PITR + replicated copies + network egress + minimum/idle commitments + support/adjacent services.

Dimension Questions Failure/cost coupling
Request/compute Read/write mix? point reads or scans? item size? consistency? poor query/index design can multiply metered work
Autoscaling Reaction time? min/max? cooldown? per-partition behavior? burst arrives before scale; quota or hot key throttles
Storage/indexes Raw bytes plus index/metadata/versions? alternate views improve reads but amplify storage/write cost
Replication How many regions/copies? writes replicated/billed? global durability/latency choice changes cost and conflict model
Egress Cross-region, internet, analytics/backup movement? architecture can shift cost outside database line item
Autoscaling does not repair a pathological key

A service can increase table/account throughput while one partition or logical key remains the bottleneck. Current DynamoDB documentation explicitly discusses adaptive capacity and partition maximums; current Cosmos DB autoscale guidance also calls out scaling behavior around active partitions. Exact numbers and algorithms are implementation details and must be rechecked.

3. Deliberately wrong approach: extrapolate average traffic and assume the scaler is instantaneous

If AtlasMart sizes from an hourly mean, a four-minute flash sale can exceed the currently active capacity before the scaler reacts. Even after table capacity rises, a celebrity SKU can remain hot. The resulting throttles trigger retries, which can make the cost and load worse. A second error is comparing only request price while ignoring indexes, backup retention and cross-region egress.

4. AtlasMart lab: scale lag, hot-key throttling, and normalized cost

Mandatory lab environment

Python 3.13+ standard library only. Every dollar coefficient in this lab is explicitly hypothetical and exists only to exercise the formula. Do not use these values for a cloud quote.

python · AtlasMart deterministic simulation
from math import ceil

# Hypothetical teaching coefficients only -- NOT a cloud-provider price quote.
price = {
    "provisioned_unit_hour": 0.00012,
    "autoscale_peak_unit_hour": 0.00016,
    "million_request_units": 0.28,
    "gb_month_storage": 0.22,
    "gb_month_backup": 0.06,
    "gb_egress": 0.08,
}

# One hour, per-minute demand. A 4-minute flash sale creates a sharp spike.
demand = [180] * 20 + [900, 1300, 1500, 1100] + [260] * 36
# One hot key can consume no more than this synthetic per-partition ceiling.
hot_partition_demand = [70] * 20 + [300, 520, 680, 450] + [90] * 36
partition_ceiling = 400

# Synthetic autoscaler: starts at 300 units and reacts to previous minute with a 2-minute lag.
capacity = []
current = 300
for minute, d in enumerate(demand):
    observed = demand[max(0, minute - 2)]
    target = max(300, ceil(observed / 100) * 100)
    current = min(max(current, target), 1600)
    capacity.append(current)

throttled_total = 0
hot_throttled_total = 0
for d, h, c in zip(demand, hot_partition_demand, capacity):
    throttled_total += max(0, d - c)
    hot_throttled_total += max(0, h - partition_ceiling)

storage_gb = 240
index_gb = 95
backup_gb = 335
egress_gb = 480
hours_month = 730
provisioned_units = 900
autoscale_hourly_peak = max(capacity)
request_units_month = 860_000_000

provisioned_cost = provisioned_units * hours_month * price["provisioned_unit_hour"]
autoscale_cost = autoscale_hourly_peak * hours_month * price["autoscale_peak_unit_hour"]
serverless_cost = request_units_month / 1_000_000 * price["million_request_units"]
common = ((storage_gb + index_gb) * price["gb_month_storage"]
          + backup_gb * price["gb_month_backup"]
          + egress_gb * price["gb_egress"])

print("AUTOSCALING TIMELINE AROUND FLASH SALE")
for i in range(18, 27):
    print(f"minute={i:02d} demand={demand[i]:4} capacity={capacity[i]:4} hot_key={hot_partition_demand[i]:3}")
print("table-level units throttled by scale lag:", throttled_total)
print("hot-partition units throttled despite table capacity:", hot_throttled_total)

print("\nHYPOTHETICAL MONTHLY COST MODEL (teaching inputs, not provider prices)")
print(f"common storage/index/backup/egress = ${common:,.2f}")
print(f"provisioned = ${provisioned_cost + common:,.2f}")
print(f"autoscale   = ${autoscale_cost + common:,.2f}")
print(f"serverless  = ${serverless_cost + common:,.2f}")
print("lesson: capacity mode changes billing mechanics, not hot-key physics or cost accounting")
Expected evidence

The synthetic scaler lags the flash-sale demand, so aggregate work is throttled before capacity catches up. Separately, the hot key exceeds its synthetic partition ceiling even when total table capacity is sufficient. The printed monthly totals demonstrate how a capacity mode changes the formula, not which mode is universally cheapest.

5. Production judgment: price the SLO, not the raw operation

Use current provider calculators/APIs only after measuring the workload. Price several scenarios: normal month, launch burst, region failover, index rebuild/backfill, replay after outage, backup restore and unexpected egress. Include quotas and throttling in acceptance tests. Budget alarms are useful but do not replace hard safety controls such as maximum autoscale settings, request admission, tenant quotas, and retry backoff.

For serverless, low idle cost can be compelling for intermittent workloads. For sustained predictable traffic, committed/provisioned models can be more efficient. Neither statement is universal because price schedules, regions, editions and discounts change. Keep usage units and formulas in source control; resolve current prices at decision time. Next, Chapter 24 moves from one-region capacity to global placement, where latency, conflict semantics and residency can outweigh pure request cost.

Check your understanding

  1. Why can an autoscaling table still throttle?
  2. What costs belong in a serverless request model besides request units?
  3. Why is idle cost not the whole comparison?
  4. Why are example provider prices dangerous in a durable architecture document?
  5. What metric should accompany a cost-per-request calculation?
Review the answers

1. Scale can lag demand, quotas can cap growth, and a single hot partition/key can hit a local ceiling even when aggregate capacity is available.

2. Storage, indexes, backups/PITR, replication, egress, optional support/observability and adjacent services.

3. Provisioned modes may pay for reserved throughput while pay-per-request modes may cost more under sustained high utilization; workload shape determines the trade.

4. Rates vary by region, edition and time; record formulas and usage assumptions, then pull current price data during procurement.

5. At minimum latency/error/throttling under the same workload, because a cheap request that violates the SLO is not equivalent capacity.

References

Provider-specific claims below are current implementation anchors reviewed in August 2026. They are not universal NoSQL definitions, and mandatory labs do not require the services.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.