Chapter 25 · Cost Engineering, Quotas, Limits, Billing, and Capacity/Usage Forecasting
Load-Test and Cost-Test a Feature Before Launch with Explicit Volume, Growth, and Worst-Case Assumptions
Combine load testing and cost testing into a pre-launch gate with latency percentiles, errors, billable-work traces and worst-case growth/abuse scenarios.
1. AtlasMart launch gate: performance test without cost test is incomplete
A load test reports acceptable p99 latency for 15 minutes, but it does not record documents returned, index entries scanned, listener updates, retries, network bytes, storage growth or recovery overhead. The team cannot tell whether the feature is economically safe at 10× traffic. Chapter 25 ends by combining load evidence and billable-work evidence into one launch gate.
- Design a bounded local load/cost harness without mistaking emulator throughput for production capacity.
- Capture p50/p95/p99, throughput, errors/retries and billable-unit proxies from the same run.
- Run small/growth/worst-case scenarios including listener and abuse tails.
- Define explicit launch thresholds and remediation owners.
- Plan an optional billed production test safely when managed-service latency/query-plan evidence is required.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project ID
demo-atlasmart-firestore, Standard-edition Native
mode, database (default), Node.js 22+, Firebase
CLI course baseline 15.30.0, Firebase JavaScript
SDK 12.19.0, Firebase Admin Node SDK
14.4.0 (with
@google-cloud/firestore 9.1.0), Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. Mandatory exercises are
local/no-cost and measure application operation counts, bytes,
retry assumptions, listener events and forecast arithmetic—not
production billing or capacity. Dated concrete USD examples
use on-demand us-central1 prices observed on 17
September 2026; always re-check the official pricing page,
your billing currency/SKU and discounts before a real
decision. Enterprise examples are analytical unless you
explicitly provision a billed Enterprise database.
2. Two test environments answer different questions
| Environment | Safe questions | Questions it cannot answer |
|---|---|---|
| Emulator/local deterministic harness | Does instrumentation work? Are operation counts, feature flows, retries and cleanup correct? Can the cost model consume the trace? | Production p95/p99, autoscaling/hotspots, regional networking, Enterprise RU, invoice SKUs, Query Explain scan bytes. |
| Disposable billed cloud project | What are real query plans/scans, managed-service latency/error behavior, network path and billed usage for a bounded test? | Future traffic distribution, long-term incident frequency, all regions/devices unless explicitly tested. |
3. Define the test matrix before generating load
| Scenario | Traffic multiplier | Retry rate | Listener reconnects | Vector scan assumption | Purpose |
|---|---|---|---|---|---|
| Best / warm | 0.5× | 0.2% | 1% | measured p50 | Lower envelope; caches/warm paths. |
| Base | 1× | 1% | 5% | measured median | Expected product forecast. |
| Growth | 10× | 2% | 10% | measured p95 | Near-term capacity/economics. |
| Worst / abuse | 25× | 8% | 35% | p99 or bounded maximum | Guardrails, incident budget and graceful degradation. |
Replace these illustrative multipliers with AtlasMart telemetry when available. They are scenario inputs, not universal defaults.
4. Local harness: one trace feeds latency and cost models
import { performance } from "node:perf_hooks";const samples=[];export async function measured(name, accounting, fn) { const t0=performance.now(); try { const out=await fn(); samples.push({name, ok:true, ms:performance.now()-t0, ...accounting(out)}); return out; } catch (e) { samples.push({name, ok:false, code:e.code ?? "UNKNOWN", ms:performance.now()-t0}); throw e; }}export function percentile(values,p){ const a=[...values].sort((x,y)=>x-y); return a[Math.max(0, Math.ceil(a.length*p/100)-1)] ?? 0;}export function report(){ const ok=samples.filter(x=>x.ok); return { requests:samples.length, errors:samples.length-ok.length, p50_ms:percentile(ok.map(x=>x.ms),50), p95_ms:percentile(ok.map(x=>x.ms),95), p99_ms:percentile(ok.map(x=>x.ms),99), documentReads:ok.reduce((s,x)=>s+(x.documentReads||0),0), writes:ok.reduce((s,x)=>s+(x.writes||0),0), vectorEntries:ok.reduce((s,x)=>s+(x.vectorEntries||0),0), bytesReturned:ok.reduce((s,x)=>s+(x.bytesReturned||0),0) };}
For emulator runs, the accounting function uses fixture knowledge (documents returned, local candidate count, serialized bytes). In managed cloud, replace approximations with Query Explain/Monitoring/Billing-export evidence where available.
5. Mandatory no-cost rehearsal
- Start Firestore/Auth emulators and seed deterministic AtlasMart catalog, carts, orders and listener fixtures.
- Run 500 browse requests, 200 cart mutations, 50 order workflows, 100 listener lifecycles and 50 vector-simulation queries.
- Inject bounded failures: 2% retriable backend simulation, 10 reconnect cycles, one listener leak that the test must catch, and one broad-query variant.
-
Write
trace.jsonlplussummary.jsoncontaining p50/p95/p99, errors/retries, reads/writes/vector entries/bytes and cleanup state. - Feed the summary into the Standard/Enterprise cost models for best/base/growth/worst cases.
- Assert emulator listeners are cleaned up and all test data is reset.
A 2 ms emulator p99 is not evidence that production p99 will be 2 ms. Use it only to catch regressions in your local harness and to validate accounting. Production performance requires a bounded managed test.
6. Optional billed managed test: bounded by design
If launch requires production evidence, create a disposable project/database in the intended edition/location, define a maximum request count and test duration, enable required observability, seed synthetic non-sensitive data, and run only the agreed matrix. Capture Query Explain for representative queries, Monitoring latency/errors, billing export/usage evidence and network path. Tear down test data/resources afterward.
| Before test | During test | After test |
|---|---|---|
| Record edition/mode/location, price-source timestamp, maximum spend/request budget, rollback/cleanup owner. | Watch errors, retries, p95/p99, scans/read units, listener counts and budget/usage signals. | Export evidence, delete synthetic resources, verify no listeners/workers remain, compare invoice when available. |
| Use synthetic data; no customer PII. | Do not “see how far it goes” against production. | Update forecast assumptions and acceptance decisions. |
7. Launch gate: combine SLO and cost envelopes
| Gate | Example evidence | Decision |
|---|---|---|
| Correctness/security | Rules/IAM/App Check tests, tenant isolation, idempotency | Must pass; never trade for lower bill. |
| Latency | p95/p99 by feature and region/device cohort | Must meet product SLO with margin. |
| Error/retry | Error codes, retry amplification, contention | Must remain inside incident budget. |
| Cost | Best/base/growth/worst billable dimensions + dated price inputs | Must fit unit-economics envelope with identified dominant drivers. |
| Recovery/governance | Backup/PITR/retention/network controls included | Must meet RPO/RTO/compliance, not be omitted from model. |
| Observability | Query Explain/Insights/Monitoring/billing alerts/runbook | Must make future drift diagnosable. |
8. Controlled failure: benchmark average latency and average user only
Run two forecasts with identical mean requests: one evenly distributed, one with a 2% heavy cohort and synchronized listener reconnect storm. The second can have much worse p99 and cost amplification despite the same average. Repair by keeping percentile/cohort and worst-case scenarios in the launch gate.
Production judgment: the forecast becomes a living operational contract
After launch, replace assumptions with observed distributions: per-feature DAU, documents/query, index entries scanned, bytes, listener churn, retry rate, storage/index growth, egress and recovery footprint. Compare actuals to the forecast monthly and after every data-model/query/index/retention change. A cost model that is never reconciled is documentation, not control.
Chapter 26 will take this same discipline into CI/CD: Rules/index deployment, emulator tests, schema migrations and release gates.
Knowledge check
- Why can’t emulator throughput certify production capacity?
- Why should latency and accounting use the same request trace?
- What makes a managed load test safe?
- Why does the worst-case scenario include abuse/reconnect tails?
- What should happen to the forecast after launch?
Review the answers
1. It does not reproduce production autoscaling, regional network, hotspots, query planner/billing or managed-service contention.
2. It lets you correlate expensive work with user-visible performance instead of optimizing disconnected measurements.
3. An isolated synthetic target, explicit request/time/spend bounds, observability, owner, rollback/cleanup and no customer PII.
4. Rare cohorts/incidents can dominate both p99 and cost despite a benign average.
5. Replace assumptions with measured distributions, reconcile actual usage/cost, and update after architecture changes.
Summary and next step
This lesson established the working contract for Load-Test and Cost-Test a Feature Before Launch with Explicit Volume, Growth, and Worst-Case Assumptions. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Emulator Suite Test Strategy for Data, Rules, Auth Context, Indexes, and Failure Cases.
Authoritative references
- Google Cloud: Firestore Standard edition pricing — location-specific document operations, storage, PITR, backups/restores and network pricing.
- Google Cloud: Firestore Enterprise edition pricing — Read Units, Write Units, realtime updates, storage, networking, Query Explain and recovery pricing.
- Firebase: Understand Cloud Firestore billing — index-entry reads, aggregations, listeners, offsets, Rules-dependent reads and free quota.
- Firebase: Firestore usage and limits — Standard free quota and current hard/configuration limits.
- Firebase: Enterprise Native mode quotas and limits — Enterprise free-tier units and limits.
- Firebase: Compare Standard and Enterprise Native mode — billing/index/realtime differences.
- Google Cloud Billing budgets — alert budgets notify; ordinary alert budgets do not automatically cap spend.