Chapter 09 · Transactions, Batched Writes, Atomic Field Operations, Retries, and Contention
Load-Test a Contended Workflow and Redesign It to Meet Throughput Without Violating Correctness
Load-test AtlasMart contention safely, record retries and tail latency, and redesign throughput bottlenecks without weakening correctness.
Learning outcomes
Build a bounded, repeatable contention harness that records success, errors, retries and latency without fabricating production claims.
Compare a hot transactional document, atomic transform and sharded counter under the same local fixture.
Preserve a strict inventory invariant while moving non-invariant metrics away from the hot transaction path.
Write a production validation plan that separates emulator correctness evidence from real-region performance/cost evidence.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart continues the same environment used in Chapters
01–08: project ID demo-atlasmart-firestore,
Standard edition / Native mode /
(default) database for the mandatory lab,
Firestore emulator 127.0.0.1:8080, Authentication
emulator 127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0 (bundling
@google-cloud/firestore 9.1.0), and Node.js 22+.
The mandatory exercises use only local/demo resources.
Production location, IAM credentials, billing, real lock
scheduling, multi-region latency and tail-throughput behavior
are not inferred from emulator results.
The Emulator Suite is appropriate for deterministic
correctness tests, Security Rules tests, retry-safe
application logic and controlled conflicting writes, but it is
not a production contention benchmark. Record your own attempt
counts and latencies. Do not publish emulator p95/p99 as a
Firestore service SLO. For Enterprise comparisons, the
mandatory path is a deterministic analysis: Standard server
libraries default to pessimistic concurrency, Enterprise
server libraries default to optimistic concurrency, while
mobile/web transactions emulate optimistic concurrency
regardless of the database setting. Pipeline DML is not a
replacement for Core transactions: current Pipeline
update/delete stages execute outside
transactions and can partially succeed across documents.
1. The AtlasMart performance question must start with correctness
“Make it faster” is underspecified. AtlasMart has two different workloads: inventory reservation, where stock must never go below zero, and order-count telemetry, where increments commute and can be sharded. A redesign that makes the benchmark green by weakening the invariant is a failure.
| Metric | Why record it | Correct interpretation |
|---|---|---|
| Committed reservations | Business correctness | Must never exceed initial stock |
| Transaction callback attempts | Contention/retry pressure | More attempts can mean conflicts; emulator count is not production forecast |
| ABORTED/other failures | Operational reliability | Classify by code and workload stage |
| p50/p95/p99 duration | Tail behavior for this exact run | Only meaningful with sample size/environment disclosed |
| Final shard histogram | Write distribution | Shows whether writes actually spread |
| Reads/writes per logical operation | Amplification/cost proxy | Translate to billing only with the correct edition pricing model |
2. Harness design: one fixture, several concurrency levels
Run short levels such as 1, 4, 8 and 16 concurrent workers with a fixed number of logical operations. Reset data between scenarios. Warm up once before recording. Record raw samples to JSON/CSV. The emulator is the mandatory safe target; an optional production test requires an isolated billed project, explicit region, cost cap and cleanup.
function percentile(sorted, p) { if (!sorted.length) return null; const i=Math.min(sorted.length-1, Math.ceil(p*sorted.length)-1); return sorted[i];}function summarize(samples) { const ok=samples.filter(x=>x.ok); const ms=ok.map(x=>x.ms).sort((a,b)=>a-b); return { count:samples.length, ok:ok.length, failed:samples.length-ok.length, p50:percentile(ms,0.50), p95:percentile(ms,0.95), p99:percentile(ms,0.99) };}
3. Scenario A: strict inventory transaction
async function runInventoryScenario({workers, initialStock}) { const ref=db.doc("catalogItems/p-load"); await ref.set({stock:initialStock,sku:"P-LOAD",schemaVersion:2}); let attempts=0; const samples=await Promise.all(Array.from({length:workers},(_,i)=>(async()=>{ const t0=performance.now(); try { const reserved=await db.runTransaction(async tx=>{ attempts++; const s=await tx.get(ref); const stock=s.get("stock"); if (stock < 1) return false; tx.update(ref,{stock:stock-1}); return true; }); return {ok:true,reserved,ms:performance.now()-t0}; } catch(e) { return {ok:false,code:e.code??"unknown",ms:performance.now()-t0}; } })())); const final=(await ref.get()).get("stock"); const committed=samples.filter(x=>x.ok&&x.reserved).length; if (committed > initialStock || final !== initialStock-committed) throw new Error("inventory invariant violated"); return {attempts,final,committed,samples};}
This is deliberately a hot document because all reservations for one SKU coordinate there. Keep it hot if that is the true invariant; do not “fix” the benchmark by allowing oversell.
4. Scenario B: hot telemetry transform, then sharded telemetry
async function runHotCounter(workers) { const ref=db.doc("metrics/hot-orders"); await ref.set({count:0}); const samples=await Promise.all(Array.from({length:workers},async()=>{ const t0=performance.now(); try { await ref.update({count:FieldValue.increment(1)}); return {ok:true,ms:performance.now()-t0}; } catch(e){ return {ok:false,code:e.code??"unknown",ms:performance.now()-t0}; } })); return {samples,total:(await ref.get()).get("count")};}async function runShardedCounter(workers, shards=10) { for(let i=0;i<shards;i++) await db.doc(`counters/load/shards/${i}`).set({count:0}); const samples=await Promise.all(Array.from({length:workers},async(_,i)=>{ const shard=i % shards; // deterministic distribution for the benchmark const ref=db.doc(`counters/load/shards/${shard}`); const t0=performance.now(); try { await ref.update({count:FieldValue.increment(1)}); return {ok:true,shard,ms:performance.now()-t0}; } catch(e){ return {ok:false,shard,code:e.code??"unknown",ms:performance.now()-t0}; } })); return {samples,total:await readExactCounter("load",shards)};}
The benchmark asks a narrower question: does spreading commutative telemetry writes reduce concentration under this fixture? It does not imply that inventory should use the same model.
5. Generate a reproducible report instead of narrating impressions
import { writeFile } from "node:fs/promises";const levels=[1,4,8,16];const report=[];for (const workers of levels) { const inv=await runInventoryScenario({workers,initialStock:Math.ceil(workers/2)}); const hot=await runHotCounter(workers*10); const sharded=await runShardedCounter(workers*10,10); report.push({ workers, inventory:{attempts:inv.attempts,committed:inv.committed,final:inv.final,...summarize(inv.samples)}, hotCounter:{total:hot.total,...summarize(hot.samples)}, shardedCounter:{total:sharded.total,...summarize(sharded.samples)} });}await writeFile("ch09-contention-report.json", JSON.stringify({ generatedAt:new Date().toISOString(), target:"Firestore Emulator Suite", projectId:"demo-atlasmart-firestore", firestorePort:8080, node:process.version, firebaseAdmin:"14.4.0", firebaseTools:"15.30.0", report},null,2));console.table(report.map(r=>({workers:r.workers,inventoryP95:r.inventory.p95,hotP95:r.hotCounter.p95,shardedP95:r.shardedCounter.p95})));
Keep the raw report. If you later run an optional real-project benchmark, produce a separate file that includes edition, mode, database ID, location, concurrency mode, client library version, billing dimensions, warm-up, document/index shape and cleanup timestamp. Never merge emulator and production samples into one percentile.
6. Interpret tail latency without cargo-cult thresholds
A p95 increase under contention means that the slowest 5% of
observed successful operations in that run crossed the
p95 boundary. It does not establish a universal Firestore
threshold. Look for correlated retry counts,
ABORTED errors and hot-key concentration. If a
transaction remains correct but too contended, reduce the
read/write set, partition independent invariants, or change the
workflow so fewer requests require the same synchronous
lock/version check.
7. Wrong approach: retry every failure forever
Unbounded application retries can amplify a hotspot and repeat surrounding side effects. Firestore SDKs already perform bounded transaction retries. The application should classify errors, cap its own retries, use backoff/jitter when appropriate, preserve idempotency, and surface a recoverable business outcome such as “reservation could not be confirmed; please retry.”
A retry policy is part of the workload, not a magical reliability switch. Record retry count and elapsed time. Stop when the business deadline or client request budget is exhausted, and never repeat a non-idempotent external action simply because the database returned a transient error.
8. Reproducible AtlasMart lab
{ "name": "atlasmart-firestore-ch09", "private": true, "type": "module", "engines": { "node": ">=22" }, "dependencies": { "firebase-admin": "14.4.0" }, "devDependencies": { "firebase-tools": "15.30.0" }}
{ "firestore": { "rules": "firestore.rules", "indexes": "firestore.indexes.json" }, "emulators": { "firestore": { "port": 8080 }, "auth": { "port": 9099 }, "ui": { "enabled": true, "port": 4000 } }}
rules_version = '2';service cloud.firestore { match /databases/{database}/documents { match /catalogItems/{productId} { allow read: if true; allow write: if false; } match /profiles/{uid} { allow read, write: if request.auth != null && request.auth.uid == uid; } match /orders/{orderId} { allow read: if request.auth != null && resource.data.customerId == request.auth.uid; allow write: if false; } match /{document=**} { allow read, write: if false; } }}
mkdir atlasmart-firestore-ch09 && cd atlasmart-firestore-ch09npm init -ynpm install firebase-admin@14.4.0npm install --save-dev firebase-tools@15.30.0# Save firebase.json, firestore.rules and firestore.indexes.json from this lesson.printf '{"indexes":[],"fieldOverrides":[]}' > firestore.indexes.jsonnpx firebase-tools@15.30.0 emulators:start --project demo-atlasmart-firestore --only firestore,auth
Run the three scenarios at the bounded concurrency levels and
save ch09-contention-report.json. Assertions:
inventory never oversells; hot and sharded counter totals equal
the number of successful logical increments; every failure is
classified; no external side effect occurs inside a retryable
callback. Then deliberately increase contention locally and
observe how attempt counts/tail latency change on your machine.
Optional production exercise: only in an isolated project with billing intentionally enabled, a fixed database/region, a small operation cap, and immediate cleanup. Inspect the server concurrency mode first. Do not assume Standard emulator behavior predicts Enterprise optimistic-server behavior.
Production judgment and Chapter 10 bridge
The best design keeps strict invariants strict and moves commutative or eventually consistent work off the synchronous hot path. Inventory may remain transactional; analytics counters can be sharded; external effects need idempotency/outbox patterns. Chapter 10 broadens this from one transaction to complete business workflows: reservations, expiration, compensation, event retries and recovery.
Knowledge check
- Why can a sharded counter benchmark improve while inventory remains transactional?
- What must accompany any reported p95/p99?
- Should the application retry ABORTED forever?
- What does the emulator benchmark prove?
- What is a successful redesign?
Review the answers
1. They have different correctness requirements: counter increments commute, while inventory must preserve a strict stock invariant.
2. Raw sample size and environment details such as emulator/production target, versions, concurrency, warm-up, data/index shape, edition/mode/region as applicable.
3. No. Use bounded retries/backoff within a business deadline and keep surrounding actions idempotent.
4. Correctness and comparative behavior for the pinned local fixture, not production service throughput, lock scheduling or SLOs.
5. One that improves the required operational target without weakening the business invariant or hiding cost/amplification elsewhere.
Summary
Contention testing is useful only when correctness assertions travel with the performance metrics. AtlasMart now has a measured path from transaction retries to transforms, distributed counters and idempotent workflows—the foundation for the broader atomicity patterns in Chapter 10.
Authoritative references
- Transactions and batched writes — atomicity, retry rules, offline boundary, batch behavior and failure conditions.
- Transaction serializability and isolation — Standard/Enterprise concurrency defaults, mobile/web optimistic emulation, server locking and contention errors.
- Usage and limits — transaction time/request/field-transform and Security Rules access-call limits.
- Distributed counters — shard-based write distribution and read/cost tradeoffs.
- Firestore best practices — hotspot, transaction size and scale guidance.
- Enterprise Native Core/Pipeline overview — operation-family boundaries.
- Pipeline DML — Preview update/delete semantics and non-transactional partial-success boundary.
- Firebase JavaScript SDK release notes — 12.19.0 baseline.
-
Firebase Admin Node.js release notes
— 14.4.0,
@google-cloud/firestore9.1.0 and Node.js 22+ baseline. - Firebase CLI release notes — 15.30.0 baseline.