Chapter 18 · Firestore Enterprise Native Mode: Core vs Pipeline Operations and Advanced Querying
Relational-Style Join Capabilities Through Sub-Pipelines: Power, Cost, and Modeling Implications
Use correlated sub-pipelines to express relational-style joins over AtlasMart collections, then measure the modeling, security, read-unit, latency, and denormalization tradeoffs rather than assuming Firestore has become relational.
1. AtlasMart problem: seller dashboards need related seller and product data
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
A dashboard row needs seller metadata alongside matching products. AtlasMart previously duplicated seller display fields or made multiple reads because Core operations have no traditional server-side join. Enterprise Pipeline can use correlated subqueries/sub-pipelines to combine related data. That is powerful, but it does not remove data-model ownership, security, fan-out, or cost questions.
AtlasMart retains the course-wide project identity
demo-atlasmart-firestore, Node.js 22+, Firebase
CLI 15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node.js SDK
14.4.0, and the Admin-bundled
@google-cloud/firestore 9.1.0.
Chapters 01–17 used Standard edition / Native mode /
(default) as the canonical managed model. Chapter
18 adds an isolated Enterprise-edition emulator profile on
Firestore 127.0.0.1:8180 with Emulator UI
127.0.0.1:4100, so it cannot accidentally share
state with the Standard lab on port 8080. Mandatory exercises
are local/no-cost.
Current Local Emulator Suite documentation allows the
Firestore emulator to be configured with
edition: "enterprise". That proves local
Enterprise-edition configuration and lets us exercise ordinary
document/Core behavior. It does not establish
production latency, byte-based billing, index-build state,
Query Explain statistics, or complete Pipeline/search/DML
parity. Where the current production service is required, the
lesson uses a deterministic query-plan/byte-scan simulator and
marks the real managed command as optional. DML pipeline
stages and Pipeline text/geospatial search are explicitly
labeled Preview.
Learning outcomes
Explain a correlated subquery as a nested pipeline evaluated in outer-row context.
Distinguish join expressiveness from relational constraints and relational transaction semantics.
Identify N×inner-work risks and the index shapes that can reduce them.
Compare duplication, multiple reads, and server-side sub-pipeline joins for AtlasMart.
Preserve tenant authorization across both outer and inner data access.
2. Sub-pipeline mental model
A subquery is an expression that can appear inside stages such as selection or field addition. It executes a nested Pipeline in the context of the current outer document. That makes it possible to use an outer seller ID to retrieve related products. “Join” here means query-time combination, not foreign-key enforcement, cascade behavior, or cross-collection schema constraints.
// Shape follows the documented sub-pipeline model; run only with the pinned SDK/API.const productSubquery = db.pipeline() .collection("catalogItems") .where(field("sellerId").equal(variable("outerSellerId"))) .where(field("published").equal(true)) .select("name", "price", "stock");const sellerDashboard = db.pipeline() .collection("sellers") .define(field("__name__").as("outerSellerId")) .addFields(productSubquery.as("products"));
3. Join does not mean Firestore became relational
| Capability | Pipeline subquery gives you | It does not automatically give you |
|---|---|---|
| Combine related collections | Yes, query-time nested retrieval | Foreign-key constraint |
| Project related fields | Yes | Normalized schema is now always best |
| Filter related data | Yes | Free/constant cost |
| Business invariants | Can read multiple shapes | Cross-entity invariant enforcement by join itself |
| Deletion lifecycle | Can discover data | Cascade delete |
4. Cost model: outer rows multiply inner work
If the outer pipeline emits many sellers and the inner subquery scans products for each seller, work can grow rapidly. A good index on the correlated join field can transform that work; an unindexed scan can be costly even though the result is correct. Query Explain exposes operators such as nested-loop joins and reports data read, rows and memory on managed Enterprise.
const sellers = ["seller-a","seller-b","seller-c"];const products = PRODUCTS;for (const sellerId of sellers) { const unindexedInnerScans = products.length; const indexedCandidates = products.filter(p => p.sellerId === sellerId).length; console.log({sellerId, unindexedInnerScans, indexedCandidates});}// Teaching counts only. Managed Query Explain is the evidence for actual execution.
5. Modeling options are now a three-way comparison
For a seller name on every product card, denormalization may still be simplest because the field is small, read-heavy, and can tolerate controlled write fan-out. For a back-office report that combines seller policy and many product aggregates, a Pipeline join may reduce application round-trips. For a low-volume workflow, two explicit Core reads can remain clearer. Pipeline expands the option set; it does not make duplication obsolete.
6. Security across outer and inner data
A join can accidentally widen the blast radius of one query. On mobile/web, Security Rules must authorize the overall request and the query must contain constraints the Rules engine can prove. On a server path, IAM can permit broad database access; the backend must still enforce the user's tenant/seller membership before constructing either the outer source or correlated inner filter. Returning a correctly joined cross-tenant row is still a security incident.
function authorizeSellerDashboard(caller, requestedSellerId) { if (!caller?.uid) throw new Error("UNAUTHENTICATED"); if (!caller.sellerIds?.includes(requestedSellerId)) throw new Error("FORBIDDEN"); return { sellerId: requestedSellerId, tenantId: caller.tenantId };}// Bind these trusted values into both outer and sub-pipeline constraints.
7. Failure injection: remove the join-field index
In a managed isolated project, capture Query Explain with the join-field index, then remove it and repeat only on bounded fixture data. The expected result rows may remain the same while scanned bytes/read units and latency increase. Restore the index before broadening the dataset. The mandatory no-cost exercise simulates candidate counts rather than claiming emulator scans reproduce production.
“The join returned six rows in development, therefore it scales.” Development fixture size hides the multiplicative cost. Repair the acceptance test by recording outer cardinality, inner candidate cardinality, index state, Explain data bytes read, result count and p95/p99 under declared load.
8. Query Explain operators are evidence, not tuning folklore
Enterprise Query Explain can show an execution tree including join and aggregation operators and summary metrics such as results returned, peak memory and data bytes read. Treat that output as versioned evidence for a specific query/data/index state; do not turn one plan into a permanent optimizer guarantee.
Verification checklist
- Join key and tenant key are explicit fields.
- Fixture truth is asserted before managed execution.
- Outer/inner cardinalities are recorded.
- Index-present/index-absent tests compare the same output rows.
- Server authorization constrains both halves of the join.
- No section claims foreign keys, cascades or relational transactions.
Production judgment and bridge to Lesson 4
Sub-pipeline joins are useful when they reduce application orchestration and the correlated work stays selective. They are a poor excuse to abandon access-pattern modeling. Lesson 4 turns the index question into an explicit measurement loop using Query Explain and Enterprise's sparse/non-sparse/unique indexing choices.
Knowledge check
- Does a Pipeline join create a foreign-key constraint?
- What is the main performance danger of a correlated subquery?
- Why might denormalization still be preferable?
- What should an index-removal experiment preserve?
- Does IAM alone enforce AtlasMart tenant membership on a privileged backend?
Review the answers
1. No. It combines data at query time; it does not enforce referential integrity.
2. Outer rows can multiply inner work, especially if the correlated lookup is unindexed or unselective.
3. It can keep hot read paths simple and predictable when duplicated data is small and write fan-out is manageable.
4. The same bounded fixture and expected result so only execution/index conditions change.
5. No. Application authorization must constrain the query to the caller's allowed tenant/seller scope.
Summary and next step
This lesson established the working contract for Relational-Style Join Capabilities Through Sub-Pipelines: Power, Cost, and Modeling Implications. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Index-Optional Querying in Enterprise: Query Explain, Performance, and When to Add Indexes.
Authoritative references
- Firebase · Overview of Firestore in Native mode (Core and Pipeline operations)
- Firebase · Standard vs Enterprise Native mode support
- Firebase · Get data with Pipeline operations
- Firebase · Perform joins with sub-pipelines
- Firebase · Query Explain for Enterprise
- Firebase · Enterprise Native index overview
- Firebase · Pipeline DML stages (Preview)
- Firebase · Pipeline search stage (Preview)
- Google Cloud · Firestore Enterprise pricing
- Firebase · Connect to the Firestore emulator / Enterprise edition configuration