Chapter 16 · Aggregation Queries, Server-Side Count/Sum/Avg, Materialized Aggregates, and Analytics Boundaries
Aggregations with Filters / Collection Groups and Consistency Expectations
Apply aggregations to filtered and collection-group queries, then verify query/rules/index compatibility and distinguish fresh server answers from realtime or offline state.
1. AtlasMart problem: a correct aggregate can still be an invalid or insecure query
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
Operations wants “total order value for seller-a” across every
user’s nested orders subcollection. That is not a
single collection read; it is a collection-group query. The
aggregation inherits that query’s scope, filters, index
requirements and authorization constraints. Firestore Security
Rules do not post-filter the aggregate. If the query could
return a document the caller is not allowed to read, the request
is denied rather than silently subtracting that document from
the total.
AtlasMart continues the same mandatory environment used in
Chapters 01–15: project ID
demo-atlasmart-firestore, Standard edition /
Native mode / (default) database, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase JavaScript SDK
12.19.0, Firebase Admin Node.js SDK
14.4.0 carrying
@google-cloud/firestore 9.1.0,
@firebase/rules-unit-testing 5.0.2,
and Node.js 22+. For continuity with Chapters 01–15 the lab
remains pinned to Firebase CLI 15.30.0; CLI
15.30.1 is now available, but its patch notes do
not change the aggregation semantics taught here. Mandatory
work remains local/no-cost. The emulator is useful for
deterministic behavior and rules tests, but it is not billing
evidence, production latency evidence, Enterprise scan-cost
evidence, or a warehouse.
Core aggregation queries support count(),
sum() and average()/avg()
naming depending on SDK. They execute on the backend, skip
local cache and pending local writes, do not support realtime
listeners or offline queries, and can return
DEADLINE_EXCEEDED when an aggregation cannot
complete within 60 seconds. In Standard Native, aggregation
pricing is based on index entries read, billed as one read for
each batch of up to 1,000 index entries with a minimum of one
document read. Enterprise billing uses read units based on
data processed rather than this Standard formula. Enterprise
Pipeline operations have a distinct
aggregate(...) stage with grouping and
accumulator capabilities. Firestore with MongoDB compatibility
has its own MongoDB aggregation surface and must not be
treated as the Native SDK API.
Learning outcomes
Aggregate filtered collection and collection-group query shapes without changing their authorization semantics.
Explain how Security Rules, IAM and indexes remain prerequisites for aggregation queries.
Distinguish server freshness from realtime/offline UX and reason about a refresh window explicitly.
Verify collection-group totals against a deterministic fixture and diagnose denied/broad query variants.
Separate Standard Core index requirements from Enterprise optional-index scan behavior.
2. Fixture truth across nested parents
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, Timestamp } from "firebase-admin/firestore";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();const t = (s) => Timestamp.fromDate(new Date(s));await db.doc("products/p-1001").set({ name: "Trail Camera", public: true, sellerId: "seller-a", schemaVersion: 4});const reviews = { "r-001": { userId: "u-alice", published: true, rating: 5, body: "Excellent", createdAt: t("2026-09-01T10:00:00Z") }, "r-002": { userId: "u-bob", published: true, rating: 4, body: "Good", createdAt: t("2026-09-02T10:00:00Z") }, "r-003": { userId: "u-cara", published: false, rating: 2, body: "Draft", createdAt: t("2026-09-03T10:00:00Z") }, "r-004": { userId: "u-dan", published: true, rating: 3, body: "Okay", createdAt: t("2026-09-04T10:00:00Z") }, "r-005": { userId: "u-erin", published: true, rating: "5", body: "Legacy", createdAt: t("2026-09-05T10:00:00Z") }, "r-006": { userId: "u-faye", published: true, body: "No score", createdAt: t("2026-09-06T10:00:00Z") }};for (const [id, data] of Object.entries(reviews)) { await db.doc(`products/p-1001/reviews/${id}`).set(data);}await db.doc("users/u-alice/orders/o-1001").set({ ownerUid: "u-alice", tenantId: "seller-a", status: "paid", totalCents: 19800, createdAt: t("2026-09-10T10:00:00Z")});await db.doc("users/u-alice/orders/o-1002").set({ ownerUid: "u-alice", tenantId: "seller-b", status: "shipped", totalCents: 4900, createdAt: t("2026-09-11T10:00:00Z")});await db.doc("users/u-bob/orders/o-1003").set({ ownerUid: "u-bob", tenantId: "seller-a", status: "paid", totalCents: 9900, createdAt: t("2026-09-12T10:00:00Z")});console.log(JSON.stringify({ expected: { publishedReviewDocuments: 5, publishedNumericRatings: 3, publishedRatingSum: 12, publishedRatingAverage: 4, sellerAOrders: 2, sellerAOrderTotalCents: 29700, sellerAOrderAverageCents: 14850 }}, null, 2));
The seller-a collection-group query spans
users/u-alice/orders/o-1001 and
users/u-bob/orders/o-1003. Its truth is
2 orders, 29,700 cents total
and 14,850 cents average. The seller-b order is
not part of that filter.
{ "indexes": [ { "collectionGroup": "orders", "queryScope": "COLLECTION_GROUP", "fields": [ { "fieldPath": "tenantId", "order": "ASCENDING" }, { "fieldPath": "createdAt", "order": "DESCENDING" } ] } ], "fieldOverrides": [ { "collectionGroup": "reviews", "fieldPath": "body", "indexes": [] } ]}
3. Filtered and collection-group aggregations reuse query semantics
import { collection, collectionGroup, query, where, getAggregateFromServer, count, sum, average} from "firebase/firestore";export async function publishedRatingTruth(db) { const q = query( collection(db, "products", "p-1001", "reviews"), where("published", "==", true) ); const snap = await getAggregateFromServer(q, { documents: count(), ratingSum: sum("rating"), ratingAverage: average("rating") }); return snap.data();}export async function sellerAOrderTruth(db) { const q = query(collectionGroup(db, "orders"), where("tenantId", "==", "seller-a")); const snap = await getAggregateFromServer(q, { orders: count(), totalCents: sum("totalCents"), averageCents: average("totalCents") }); return snap.data();}
Nothing about count() or sum() weakens
the query contract. A filter that needs a
composite/collection-group index in Standard still needs that
index in production. The emulator does not enforce compound
indexes, so a locally successful aggregation does not prove the
managed index exists or is READY.
4. Rules are not aggregate filters
Suppose Alice is allowed to read only documents whose
ownerUid == request.auth.uid. A client
collection-group query constrained only by
tenantId == "seller-a" could match Bob’s order as
well as Alice’s. Firestore cannot “sum Alice’s permitted row and
hide Bob’s.” The query must be provably compatible with the
rules, for example by including an ownership filter whose
semantics line up with the rule, or the aggregation must run
through a trusted backend that performs its own application
authorization under IAM.
An aggregate result can leak sensitive information even when no individual document body is returned. Treat counts and sums as data access. Design separate public/customer/seller/admin aggregate paths and test them adversarially just as you test document queries.
5. Allowed/denied query matrix
| Caller/query | Expected | Why |
|---|---|---|
| Unauthenticated collectionGroup(orders) | DENY | Rules require an authenticated owner |
| u-alice + ownerUid == u-alice | ALLOW for Alice-owned order set | Query can be proven compatible with ownership condition |
| u-alice + tenantId == seller-a only | DENY | Potential result includes Bob-owned order |
| Trusted Admin/server process after seller authorization | IAM/application decision | Server SDK bypasses Rules; backend must authorize seller scope itself |
| Emulator admin context | Can bypass Rules for fixture setup | Useful for tests, not evidence of production IAM |
6. Consistency expectation: fresh server answer, not a subscription
A direct aggregation returns a backend result for that request. It does not create a listener and does not update automatically when a new order arrives one second later. If the UI refreshes every 30 seconds, the product requirement is “server aggregate refreshed at most every 30 seconds,” not “realtime.” If users need immediate local feedback after placing an order, show pending local workflow state separately until the authoritative aggregate refreshes.
7. Boundary cases that change the population
| Change | Effect |
|---|---|
| Add where(status == "paid") | Only paid documents contribute; index/rules must support the narrower shape |
| Add orderBy(totalCents) | Documents missing totalCents no longer participate in that query population |
| Use sum(totalCents) where some value is non-numeric | Non-numeric values do not contribute to sum; schema validation should catch this earlier |
| Switch collection query to collectionGroup("orders") | Scope expands across all parents; rules/index assumptions must be re-reviewed |
| Run same concept with Enterprise unindexed query | Query may scan rather than fail for missing index; success is not proof of efficiency |
8. Deliberately wrong approach: aggregate broadly, then hide the result in UI
UI hiding is not authorization. An attacker can call the SDK directly. A broad aggregate can expose counts, averages or revenue even if the individual documents never render. Repair the design by making query scope/rules provably compatible or routing privileged aggregate work through a backend that verifies the caller and seller/tenant relationship before executing with its workload identity.
9. Emulator vs production evidence
The local emulator is ideal for deterministic fixtures and Security Rules assertions. It is not authoritative for compound-index enforcement, managed query planning, production billing, regional latency or Enterprise scan cost. A bounded optional managed verification should record the exact project, database, edition, location, index state, query shape and Query Explain evidence; do not paste production credentials into course code.
Production judgment and bridge to Lesson 3
Filtered aggregation is attractive when the query is bounded and the application can tolerate request/refresh semantics. Repeatedly computing the same product rating or dashboard total on every view eventually trades write simplicity for read cost and latency. Lesson 3 asks when to materialize the answer and accept write amplification, consistency/repair responsibility and potential hotspot risk.
Knowledge check
- Do Security Rules filter an aggregation result?
- Why is collectionGroup scope security-sensitive?
- Does a locally successful composite aggregation prove the production index exists?
- Is a server aggregation a realtime subscription?
- What changes in Enterprise if an index is missing?
Review the answers
1. No. The underlying query must itself be allowed; otherwise the request is denied.
2. It crosses many parent documents, so ownership/tenant assumptions tied to one parent path may no longer be sufficient.
3. No. The emulator does not enforce compound indexes like production Standard does.
4. No. It is a direct backend response for that invocation.
5. Enterprise can execute unindexed queries by scanning, so the query may succeed but become slower/more costly as data grows.
Summary
Aggregation inherits query scope. AtlasMart now treats filters, collection groups, rules, indexes and refresh semantics as part of the aggregate contract rather than as implementation details.
Authoritative references
- Summarize data with aggregation queries
- Understand Cloud Firestore billing
- Write-time aggregations
- Distributed counters
- Securely query data
- Connect to the Firestore Emulator and emulator differences
- Firestore Standard Core operations overview
- Firestore Enterprise Native Core/Pipeline overview
- Enterprise Pipeline aggregate stage
- Enterprise pricing examples
- MongoDB compatibility behavior differences
- Managed Firestore export/import
- Firebase Extensions migration guidance, including Stream Firestore to BigQuery
- Firebase JavaScript SDK release notes
- Firebase Admin Node.js SDK release notes