Chapter 10 · Atomicity Across Business Workflows: Counters, Reservations, Idempotency, and Event-Driven Consistency
Distributed / Sharded Counters, Read Aggregation, Write Distribution, and Accuracy / Latency Tradeoffs
Distribute high-rate AtlasMart counters across shards while making read aggregation, freshness and cost tradeoffs measurable and explicit.
Learning outcomes
Explain why one frequently updated counter document becomes a contention concentration point and how counter shards distribute writes.
Quantify the corresponding read amplification: an exact total requires reading/aggregating shard state unless a separate cached total is maintained.
Distinguish exact business invariants from counters where eventual or slightly stale totals are acceptable.
Build a deterministic AtlasMart counter lab that records shard distribution and validates totals without inventing throughput claims.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart continues the same environment used in Chapters
01–09: project ID demo-atlasmart-firestore,
Standard edition / Native mode /
(default) database for mandatory labs, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0 with
@google-cloud/firestore 9.1.0, and Node.js 22+.
Mandatory work remains local/no-cost. Cloud Functions,
Eventarc, managed TTL deletion, production IAM, billing,
regional delivery latency and external payment systems are
discussed accurately but are not falsely claimed to have run
in the local emulator.
The local lab simulates duplicate and reordered events deterministically with ordinary Node code so the learner can prove idempotency, compensation and repair behavior without deploying cloud infrastructure. Firestore-triggered Cloud Functions and Eventarc Standard can deliver events at least once; Firestore event ordering is not guaranteed. Firestore TTL deletion is asynchronous and documents are typically removed within about 24 hours after expiration, so TTL is a retention mechanism—not an exact reservation scheduler. Any production p95/p99, event-delivery delay, TTL cleanup delay, trigger retry count or cost must be measured in the actual edition/region/billing configuration rather than inferred from emulator timing.
1. The AtlasMart problem: one “orders started today” counter is useful but not worth making checkout hot
AtlasMart wants a dashboard counter for how many checkout
workflows started today. A single
metrics/daily document updated by every checkout
centralizes all increments on one key. Firestore cannot update
one document at an unlimited rate; sufficiently frequent
concurrent updates eventually contend. If the counter is
observational telemetry rather than a strict stock invariant, it
can be distributed across shard documents.
A distributed counter stores the logical total across multiple shard documents. Each writer chooses a shard and atomically increments only that shard. The total is the sum of the shard counts. More shards spread writes over more documents, but exact reads become more expensive because more shard documents must be read or aggregated.
| Design | Write concentration | Exact read work | Consistency/complexity |
|---|---|---|---|
| One document | Highest concentration | One document read | Simple; can become hot under high write rate |
| N shard documents | Spread across N keys | Read/sum N shards | More write headroom; read amplification |
| N shards + cached total | Writes spread; async cache update | One cached total read for fast display | Cached total can lag; needs repair/reconciliation |
| Strict inventory quantity | Do not shard merely for speed | Transaction/read model based on invariant | Sharding can weaken oversell protection if applied blindly |
2. Create and update deterministic shards
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, FieldValue, Timestamp } from "firebase-admin/firestore";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();
const COUNTER_ID="orders-started-2026-09-16";const SHARDS=10;async function initCounter() { const batch=db.batch(); batch.set(db.doc(`counters/${COUNTER_ID}`), {numShards:SHARDS,schemaVersion:3}); for (let i=0;i<SHARDS;i++) batch.set(db.doc(`counters/${COUNTER_ID}/shards/${i}`), {count:0}); await batch.commit();}async function incrementCounter(logicalOperationId) { // Deterministic hash for the lab; production may choose an appropriate distribution method. let hash=0; for (const c of logicalOperationId) hash=(hash*31+c.charCodeAt(0))>>>0; const shard=hash % SHARDS; await db.doc(`counters/${COUNTER_ID}/shards/${shard}`).set( {count:FieldValue.increment(1)}, {merge:true}); return shard;}async function readExact() { const snap=await db.collection(`counters/${COUNTER_ID}/shards`).get(); return snap.docs.reduce((sum,d)=>sum+(d.get("count")??0),0);}
3. Measure distribution, not folklore
await initCounter();const operations=Array.from({length:200},(_,i)=>`checkout-${String(i).padStart(4,"0")}`);const chosen=await Promise.all(operations.map(incrementCounter));const histogram=Object.fromEntries(Array.from({length:SHARDS},(_,i)=>[i,chosen.filter(x=>x===i).length]));const exact=await readExact();console.log({histogram, exact, expected:operations.length});if (exact!==operations.length) process.exitCode=1;
The histogram shows whether your fixture actually spreads writes. It does not establish production throughput. Emulator scheduling, local CPU and process concurrency differ from a regional Firestore deployment. If you later benchmark production, record edition, mode, region, index shape, concurrency, sample count and billing units.
4. Accuracy and latency are application decisions
If the dashboard can tolerate a delayed total, an asynchronous
worker can periodically write a materialized total such as
counters/. That reduces display reads but creates a consistency window.
The exact shard sum remains the reconciliation source. If the
value is money, inventory or an authorization limit, first ask
whether the counter abstraction is even the right correctness
model.
A sharded “remaining inventory” counter is dangerous unless the reservation algorithm itself preserves the stock invariant. Write distribution is not a substitute for a correctness proof.
5. Read amplification and cost
With N shards, a direct exact read touches N shard documents. Increasing shard count can increase write capacity but also increases exact-read work and storage. The official distributed-counter guidance explicitly presents this tradeoff. Do not pick “10” or “100” as a universal value: choose a shard count from measured write demand and acceptable read cost, and adjust with evidence.
| Metric to record | Why it matters | What it does not prove |
|---|---|---|
| Shard histogram | Whether writes distribute in the fixture | Production service throughput |
| Exact shard reads per dashboard refresh | Read amplification | Billing total without edition pricing context |
| Cached-total age | Freshness window | Correctness of strict invariants |
| Reconciliation mismatch count | Repair need | Cause without logs/correlation IDs |
| Writer latency samples | Local comparative evidence | Universal p95/p99 SLO |
6. Enterprise and MongoDB-compatibility boundary
The counter idea—distribute commutative increments over multiple keys—is architectural, but the exact SDK methods, billing units, indexing defaults and transaction/concurrency behavior vary by edition and mode. The mandatory code is Standard Native Core. For Enterprise Native or MongoDB compatibility, re-run the workload with that product’s documented operations and pricing rather than copying Standard cost assumptions.
7. Reproducible AtlasMart lab
{ "name": "atlasmart-firestore-ch10", "private": true, "type": "module", "engines": { "node": ">=22" }, "dependencies": { "firebase-admin": "14.4.0" }, "devDependencies": { "firebase-tools": "15.30.0" }}
{ "firestore": { "rules": "firestore.rules", "indexes": "firestore.indexes.json" }, "emulators": { "firestore": { "port": 8080 }, "auth": { "port": 9099 }, "ui": { "enabled": true, "port": 4000 } }}
rules_version = '2';service cloud.firestore { match /databases/{database}/documents { match /catalogItems/{productId} { allow read: if true; allow write: if false; } match /profiles/{uid} { allow read, write: if request.auth != null && request.auth.uid == uid; } match /orders/{orderId} { allow read: if request.auth != null && resource.data.customerId == request.auth.uid; allow write: if false; } // Workflow state, reservations, outbox/inbox, dedupe and repair evidence are server-owned. match /workflowCommands/{id} { allow read, write: if false; } match /reservations/{id} { allow read, write: if false; } match /workflowEvents/{id} { allow read, write: if false; } match /workflowOutbox/{id} { allow read, write: if false; } match /workflowDeadLetters/{id} { allow read, write: if false; } match /counters/{counterId}/{document=**} { allow read, write: if false; } match /{document=**} { allow read, write: if false; } }}
mkdir atlasmart-firestore-ch10 && cd atlasmart-firestore-ch10npm init -ynpm install firebase-admin@14.4.0npm install --save-dev firebase-tools@15.30.0# Save firebase.json, firestore.rules and firestore.indexes.json from this lesson.printf '{"indexes":[],"fieldOverrides":[]}' > firestore.indexes.jsonnpx firebase-tools@15.30.0 emulators:start --project demo-atlasmart-firestore --only firestore,auth
Run the 200-operation deterministic fixture twice: first against
one hot document with FieldValue.increment(1), then
against 10 shards. Save raw duration samples and the shard
histogram. Verify both logical totals equal the number of
successful operations. Do not report a “10× throughput
improvement” merely because the documentation explains that more
shards increase write capacity; your own bounded test must
remain descriptive of its actual environment.
Production judgment
Use sharded counters for commutative, high-frequency counts when additional read/aggregation complexity is acceptable. Keep strict resource reservations in transactional state, and use the counter as telemetry or a derived view. Monitor shard skew and cached-total lag. Lesson 3 returns to strict correctness: inventory reservations, expiration and compensation.
Knowledge check
- What is the main benefit of counter shards?
- What is the main exact-read cost?
- Does a sharded counter automatically preserve inventory correctness?
- Why is a cached total useful?
- What should determine the shard count?
Review the answers
1. They spread writes across multiple documents instead of concentrating every increment on one document.
2. The exact total requires reading/aggregating the shard documents unless a separate cached total is maintained.
3. No. Inventory requires its own invariant-preserving reservation design.
4. It reduces read work for frequent displays, at the cost of freshness and reconciliation complexity.
5. Measured write demand, acceptable exact-read amplification and operational evidence—not a universal magic number.
Summary and next step
Distributed counters exchange concentrated writes for read and reconciliation work. Next, AtlasMart uses strict transactions for reservation creation while treating expiration and compensation as explicit workflow steps.
Authoritative references
- Transactions and batched writes — atomic transaction/batch boundaries and retry behavior.
- Distributed counters — shard-based write distribution and read aggregation tradeoffs.
- Manage data retention with TTL policies — asynchronous deletion behavior, limits, pricing and monitoring.
- Cloud Firestore triggers — at-least-once delivery, non-guaranteed ordering, trigger scope and idempotency requirement.
- Eventarc Standard retry events — at-least-once delivery, duplicate handling, idempotency and dead-letter guidance.
- Firestore best practices — hotspot and scaling guidance.
- Firestore Enterprise overview — edition/mode boundaries to re-check before porting workflow assumptions.
- Firebase release notes — current SDK/tool versions.
- Admin Node.js release notes — 14.4.0 and Firestore client dependency baseline.
- Firebase CLI release notes — 15.30.0 baseline.