Chapter 09 · Transactions, Batched Writes, Atomic Field Operations, Retries, and Contention
Hot Documents, High-Contention Counters, Sharded / Distributed Counter Patterns, and Idempotency
Redesign hot AtlasMart counters using distributed shards and idempotency while making read amplification and consistency tradeoffs explicit.
Learning outcomes
Identify a hot-document contention problem separately from a sequential-index hotspot.
Implement a distributed counter with deterministic shard selection for tests and randomized selection for normal traffic.
Quantify the write-distribution versus read-amplification tradeoff and explain optional roll-up documents.
Add idempotency keys so retries do not silently double-count business events.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart continues the same environment used in Chapters
01–08: project ID demo-atlasmart-firestore,
Standard edition / Native mode /
(default) database for the mandatory lab,
Firestore emulator 127.0.0.1:8080, Authentication
emulator 127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0 (bundling
@google-cloud/firestore 9.1.0), and Node.js 22+.
The mandatory exercises use only local/demo resources.
Production location, IAM credentials, billing, real lock
scheduling, multi-region latency and tail-throughput behavior
are not inferred from emulator results.
The Emulator Suite is appropriate for deterministic
correctness tests, Security Rules tests, retry-safe
application logic and controlled conflicting writes, but it is
not a production contention benchmark. Record your own attempt
counts and latencies. Do not publish emulator p95/p99 as a
Firestore service SLO. For Enterprise comparisons, the
mandatory path is a deterministic analysis: Standard server
libraries default to pessimistic concurrency, Enterprise
server libraries default to optimistic concurrency, while
mobile/web transactions emulate optimistic concurrency
regardless of the database setting. Pipeline DML is not a
replacement for Core transactions: current Pipeline
update/delete stages execute outside
transactions and can partially succeed across documents.
1. The AtlasMart problem: one “orders today” document becomes a global lock point
AtlasMart initially stores
metrics/global.ordersToday and increments it for
every accepted order. Even with
FieldValue.increment(1), every write targets the
same document. Atomic transforms protect arithmetic correctness,
but they do not make a single document capable of unlimited
concurrent update throughput. A hot document is
a key receiving enough competing operations that contention and
tail latency become material.
Do not confuse this with the sequential indexed field hotspot discussed in Chapter 06: that problem comes from monotonically changing indexed values concentrating index traffic. Here, the same document itself is the shared write target.
2. Distributed counter model
A distributed counter stores N shard documents. Each increment chooses one shard and atomically increments only that shard. The logical total is the sum of the shard counts. More shards spread writes across more keys; more shards also make exact reads more expensive because calculating the total reads more documents.
counters/orders-today numShards: 10 schemaVersion: 1counters/orders-today/shards/0 { count: 0 }counters/orders-today/shards/1 { count: 0 }...counters/orders-today/shards/9 { count: 0 }
| Design dimension | Fewer shards | More shards |
|---|---|---|
| Write distribution | Less | More |
| Exact counter read documents | Fewer | More |
| Read cost/latency for exact sum | Lower | Higher |
| Operational complexity | Lower | Higher |
| Need for roll-up cache | Less likely | More attractive at high read volume |
3. Implement the shard increment
async function incrementDistributedCounter(counterId, numShards) { const shard = Math.floor(Math.random() * numShards); const ref = db.doc(`counters/${counterId}/shards/${shard}`); await ref.set({ count: FieldValue.increment(1) }, { merge: true }); return shard;}async function readExactCounter(counterId, numShards) { const refs = Array.from({length:numShards},(_,i)=>db.doc(`counters/${counterId}/shards/${i}`)); const snaps = await db.getAll(...refs); return snaps.reduce((sum,s)=>sum+(s.exists ? (s.get("count") ?? 0) : 0),0);}
The official pattern does not give a universal shard count. Start from measured contention and read requirements. Increasing shards without evidence creates permanent read amplification.
4. Idempotency: a retryable event must not increment twice
A distributed counter solves hot-key distribution, not duplicate
delivery. If an order-created event can be delivered twice, two
increments are both valid Firestore writes. AtlasMart therefore
gives each business event an idempotency key, such as
ORDER_ACCEPTED:ord-9001, and commits the “seen
event” marker together with the chosen shard increment in one
transaction.
async function countOrderOnce(orderId, numShards=10) { const eventId=`ORDER_ACCEPTED:${orderId}`; const eventRef=db.doc(`counterEvents/${eventId}`); // Deterministic shard from event ID makes retries target the same shard. let hash=0; for (const ch of eventId) hash=(hash*31+ch.charCodeAt(0))>>>0; const shard=hash % numShards; const shardRef=db.doc(`counters/orders-today/shards/${shard}`); return db.runTransaction(async tx=>{ const seen=await tx.get(eventRef); if (seen.exists) return {counted:false, shard}; tx.create(eventRef,{orderId,kind:"ORDER_ACCEPTED",createdAt:FieldValue.serverTimestamp()}); tx.set(shardRef,{count:FieldValue.increment(1)},{merge:true}); return {counted:true, shard}; });}
The unique event marker ensures a repeated delivery becomes a no-op after the first successful commit. The marker also provides an audit trail and a manual repair point. For very high-volume event streams, the idempotency-key collection itself must use well-distributed document IDs; do not place every event under one shared summary document.
5. Roll-up totals trade freshness for cheaper reads
Reading all shards on every dashboard refresh can dominate cost
and latency. A slower background process can periodically sum
shards into counters/orders-today.rollupTotal. The
roll-up is intentionally stale between refreshes. This is a
materialized view: it reduces read amplification by giving up
immediate exactness.
Display a roll-up as “recent total” unless your workflow proves a freshness bound. A stale roll-up is not appropriate for an invariant such as remaining inventory. Distributed counters are for aggregates, not reservation authority.
6. Wrong approach: shard the inventory itself just to increase throughput
If one SKU has five physical units, splitting its inventory arbitrarily across ten shard documents can make “is there at least one unit anywhere?” a multi-document coordination problem. Sharding is excellent for commutative aggregate counts; it can make strict inventory invariants harder. Preserve the real atomic boundary, and redesign the business workflow if one product legitimately receives extreme concurrent demand.
7. Reproducible AtlasMart lab
{ "name": "atlasmart-firestore-ch09", "private": true, "type": "module", "engines": { "node": ">=22" }, "dependencies": { "firebase-admin": "14.4.0" }, "devDependencies": { "firebase-tools": "15.30.0" }}
{ "firestore": { "rules": "firestore.rules", "indexes": "firestore.indexes.json" }, "emulators": { "firestore": { "port": 8080 }, "auth": { "port": 9099 }, "ui": { "enabled": true, "port": 4000 } }}
rules_version = '2';service cloud.firestore { match /databases/{database}/documents { match /catalogItems/{productId} { allow read: if true; allow write: if false; } match /profiles/{uid} { allow read, write: if request.auth != null && request.auth.uid == uid; } match /orders/{orderId} { allow read: if request.auth != null && resource.data.customerId == request.auth.uid; allow write: if false; } match /{document=**} { allow read, write: if false; } }}
mkdir atlasmart-firestore-ch09 && cd atlasmart-firestore-ch09npm init -ynpm install firebase-admin@14.4.0npm install --save-dev firebase-tools@15.30.0# Save firebase.json, firestore.rules and firestore.indexes.json from this lesson.printf '{"indexes":[],"fieldOverrides":[]}' > firestore.indexes.jsonnpx firebase-tools@15.30.0 emulators:start --project demo-atlasmart-firestore --only firestore,auth
const counter=db.doc("counters/orders-today");await counter.set({numShards:10,schemaVersion:1});for (let i=0;i<10;i++) await db.doc(`counters/orders-today/shards/${i}`).set({count:0});const ids=Array.from({length:40},(_,i)=>`ord-${String(i).padStart(4,"0")}`);// Deliver every event twice; idempotency should keep logical total at 40.await Promise.all(ids.flatMap(id=>[countOrderOnce(id,10),countOrderOnce(id,10)]));console.log("exact total", await readExactCounter("orders-today",10));const shardSnaps=await Promise.all(Array.from({length:10},(_,i)=>db.doc(`counters/orders-today/shards/${i}`).get()));console.table(shardSnaps.map((s,i)=>({shard:i,count:s.get("count")??0})));
Verify total=40, exactly 40 event-marker documents, and a non-uniform but distributed shard histogram. The deterministic hash makes repeated deliveries choose the same shard, which improves reproducibility without requiring random state to be persisted.
Production judgment
Shard only after you can name the hot key and show contention evidence. Pick shard count from measured write pressure and acceptable exact-read amplification. Add idempotency independently of sharding. If consumers can tolerate stale totals, use a roll-up document and monitor its lag. Lesson 5 turns these ideas into a bounded load experiment comparing a contended single-document workflow with a correctness-preserving redesign.
Knowledge check
- Does FieldValue.increment remove hot-document contention?
- Why do more shards help writes?
- What is the main cost of more shards?
- Why is an idempotency key still needed?
- Should strict SKU inventory automatically be sharded like a popularity counter?
Review the answers
1. No. It removes the read-modify-write race, but all increments can still compete on the same document.
2. They distribute increments across more document keys, reducing concentration on one document.
3. Exact reads need to read/sum more shard documents, increasing read amplification and latency.
4. Retries or duplicate event delivery can otherwise perform multiple valid increments.
5. No. Inventory is a business invariant; arbitrary sharding can make the invariant harder to enforce.
Summary and next step
Distributed counters trade exact-read simplicity for write distribution, while idempotency protects against duplicate business events. The final lesson measures contention and validates that a redesign improves throughput without weakening correctness.
Authoritative references
- Transactions and batched writes — atomicity, retry rules, offline boundary, batch behavior and failure conditions.
- Transaction serializability and isolation — Standard/Enterprise concurrency defaults, mobile/web optimistic emulation, server locking and contention errors.
- Usage and limits — transaction time/request/field-transform and Security Rules access-call limits.
- Distributed counters — shard-based write distribution and read/cost tradeoffs.
- Firestore best practices — hotspot, transaction size and scale guidance.
- Enterprise Native Core/Pipeline overview — operation-family boundaries.
- Pipeline DML — Preview update/delete semantics and non-transactional partial-success boundary.
- Firebase JavaScript SDK release notes — 12.19.0 baseline.
-
Firebase Admin Node.js release notes
— 14.4.0,
@google-cloud/firestore9.1.0 and Node.js 22+ baseline. - Firebase CLI release notes — 15.30.0 baseline.