Chapter 24 · Observability: Query Explain, Query Insights, Key Visualizer, Metrics, Logs, and Troubleshooting
Query Explain to Inspect Execution, Index Use, Scanned / Returned Work, and Optimization Evidence
Use Query Explain as a controlled experiment: separate planning from execution, inspect selected indexes and scan amplification, measure billable work, and verify that a proposed index or query rewrite actually improves the backend path.
1. AtlasMart problem: a query “works” but its work is opaque
AtlasMart’s operations console runs a query for in-stock cameras in a price range. The response is correct, but latency and read cost have grown with catalog size. The team proposes a new composite index because “indexes make queries faster.” Query Explain turns that guess into an experiment: what index was selected, how many entries/documents were scanned, how many results were returned, how long execution took, and what was billable?
- Use planning-only versus analyze/execution modes safely.
- Interpret indexes used, results returned, execution duration, index entries/documents scanned, and billing details.
- Recognize scan amplification and verify a query/index rewrite before/after.
- Distinguish Standard/Core Explain from Enterprise Pipeline and MongoDB-compatible explain.
- Avoid turning Explain into a high-volume observability API.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project ID
demo-atlasmart-firestore, Standard-edition Native
mode, database (default), Node.js 22+, Firebase
CLI course baseline 15.30.0, Firebase JavaScript
SDK 12.19.0, Firebase Admin Node SDK
14.4.0 (bundling
@google-cloud/firestore 9.1.0), Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. Firebase CLI
15.30.1 is the current patch release at review
time; no Chapter 24 lab depends on that patch, so the course
stays pinned to 15.30.0 for continuity. Mandatory exercises
are local/no-cost. Cloud Monitoring, Query Insights, Key
Visualizer, Cloud Audit Logs, production Query Explain billing
evidence, and real hotspot/capacity measurements require a
real Google Cloud/Firebase project and are optional bounded
verification steps. Emulator latency is never presented as
production capacity evidence. Standard Query Explain requests
use server/IAM authentication; Firebase Authentication is not
the authorization mechanism for these server Explain calls.
| Term | Operational meaning in Chapter 24 |
|---|---|
| Client timing | End-to-end elapsed time observed by the browser/mobile/backend caller. It includes client scheduling and network time that Firestore backend metrics do not. |
| Backend latency |
Firestore service processing time. Cloud Monitoring
api/request_latencies excludes
client-to-service round-trip time.
|
| Standard Native | Standard edition Core operations. Queries require indexes; Key Visualizer is currently documented for this edition/mode. |
| Enterprise Native Core | Familiar Core API on Enterprise. Realtime/offline remain Core features, but indexes are optional and billing is byte-unit based. |
| Enterprise Pipeline | Advanced stage/expression query interface. Query Explain exposes an execution tree, scanned records/bytes, memory and read units. |
| MongoDB compatibility |
Enterprise MongoDB-protocol interface. It has separate
explain/Query Insights behavior and must
not be diagnosed with Native assumptions.
|
| Query Explain | Per-query planner/execution evidence. Planning-only and execute/analyze modes have different cost and side-effect implications. |
| Query Insights | Aggregated normalized-query statistics over time, useful for frequency/latency/read-load prioritization rather than one-off diagnosis. |
| Key Visualizer | Standard Native key-range/index-range heatmaps for hotspot diagnosis. It is not a generic trace viewer or a capacity benchmark. |
| Correlation ID | Application-generated opaque identifier propagated through logs/timers so one user action can be linked across client, backend and Firestore evidence without logging tokens or PII. |
2. Planning and execution are different experiments
| Mode/interface | Executes query? | Returns results? | Cost/side effect |
|---|---|---|---|
| Standard/Core default Explain | No | No | One document read is charged; no index/read operations are executed. |
| Standard/Core analyze | Yes | Yes through Explain result | Query executes and is billed normally; analysis itself adds no separate surcharge. |
Enterprise Pipeline explain |
No | No | Planner information only. |
Enterprise Pipeline stats |
Yes | No | Executes for runtime statistics without returning query results. |
Enterprise Pipeline analyze |
Yes | Yes | Executes and returns result plus runtime/billing evidence. |
MongoDB compatibility queryPlanner |
No execution stats | No normal result set | Mongo-compatible planner view. |
MongoDB executionStats/allPlansExecution
|
Yes | Explain document | Execution/billing/memory evidence for supported commands. |
Planning-only Standard Explain still has a one-read charge. Analyze executes the query and incurs its normal read/index work. Use Explain for controlled analysis, not on every production request.
3. Standard/Core Query Explain fields that change decisions
| Evidence | Interpretation | Red flag |
|---|---|---|
indexes_used |
Planner-selected index structures | Unexpected index or index intersection relative to intended access path. |
results_returned |
Useful output rows/documents | Tiny result set with huge scanned work. |
execution_duration |
Backend execution time | Regression after data/index shape change. |
index_entries_scanned |
Index entries inspected | High scan-to-result ratio. |
documents_scanned |
Documents inspected | Filtering after broad index scan or poor selectivity. |
read_operations/billing_details
|
Billable work | Cost unexpectedly high for a frequent query. |
The useful ratio is not a universal threshold; it is a comparison across your own stable workload. A scan ratio of 20:1 may be acceptable for a rare admin report and unacceptable for a query executed thousands of times per minute.
4. Server-side code: plan first, analyze second
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";initializeApp({projectId: process.env.GCLOUD_PROJECT});const db=getFirestore();const q=db.collection("catalogItems") .where("category","==","camera") .where("price",">=",80) .where("price","<=",250) .orderBy("price") .limit(20);// Production server library: planning-only first.const plan=await q.explain({analyze:false});console.log(plan.metrics.planSummary);// Optional bounded production analysis; this executes the query.// const run=await q.explain({analyze:true});// console.log(run.metrics.executionStats);
The emulator is suitable for code-path and index/rules development, but do not claim that local Explain output—if a specific capability is unavailable or differs—matches managed execution. The mandatory exercise therefore uses a deterministic saved Explain fixture for runtime statistics and an optional real-project step for actual Explain.
5. Deterministic saved trace: diagnose scan amplification
{ "query":"catalog camera price range", "results_returned":20, "index_entries_scanned":2400, "documents_scanned":400, "read_operations":401, "execution_duration_ms":92, "indexes_used":["(category ASC, price ASC, __name__ ASC)"], "source":"DETERMINISTIC_TRAINING_FIXTURE"}
{ "query":"catalog camera price range + published equality", "results_returned":20, "index_entries_scanned":80, "documents_scanned":20, "read_operations":21, "execution_duration_ms":18, "indexes_used":["(category ASC, published ASC, price ASC, __name__ ASC)"], "source":"DETERMINISTIC_TRAINING_FIXTURE"}
The numbers are explicitly synthetic training evidence, not claimed Firestore output. The lab assertion is that the mechanism detects lower scan/read amplification after the query contract and index are changed.
6. Mandatory lab: compare evidence, not anecdotes
import fs from "node:fs";const before=JSON.parse(fs.readFileSync("evidence/explain-before.json"));const after=JSON.parse(fs.readFileSync("evidence/explain-after.json"));const ratio=x => x.index_entries_scanned/Math.max(1,x.results_returned);console.log({beforeScanPerResult:ratio(before), afterScanPerResult:ratio(after)});if (!(after.index_entries_scanned < before.index_entries_scanned)) throw new Error("scan-not-improved");if (!(after.read_operations < before.read_operations)) throw new Error("read-work-not-improved");console.log("QUERY_EVIDENCE_IMPROVED");
- Seed the same deterministic catalog fixture used in Lesson 1.
- Run the query contract and verify result correctness first.
- Evaluate saved before/after Explain fixtures with the comparator.
- Optional cloud step: collect real planning-only Explain first; only run analyze on a bounded dataset after estimating read cost.
- Store timestamp, database edition/mode, index definition, SDK version, dataset size and query text with the evidence.
7. Enterprise Pipeline and MongoDB compatibility: same goal, different report shape
Enterprise Pipeline Explain returns a summary and execution tree
with records scanned, bytes read, memory, latency and read
units. Because Enterprise indexes are optional, a successful
query can contain a TableScan; “it ran” is not
evidence that it is economical at scale. MongoDB compatibility
supports explain on documented commands such as
find, aggregate, count,
distinct, update,
delete and findAndModify, with
Mongo-style verbosity modes. Do not compare field names across
these interfaces as if they were identical.
8. Boundary: Query Explain is not streaming/listener profiling
Standard Query Explain currently supports polled queries, not streaming queries. If a realtime screen is expensive, use listener counts, document changes, Rules behavior, application timing and billing/read evidence; do not expect Query Explain to profile an active snapshot listener directly.
Production judgment
Use Explain when a query has enough frequency, latency, cost or uncertainty to justify a targeted experiment. Preserve correctness tests before changing index/order/filter semantics. An index can reduce read work while increasing write/storage overhead, so Chapter 24 records both sides of the change and Chapter 25 will cost-model them explicitly.
Knowledge check
- Does Standard planning-only Explain execute the query?
-
What is the difference between
documents_scannedandresults_returned? - Why is a successful Enterprise unindexed query not proof of good performance?
- Can Standard Query Explain profile snapshot listeners directly?
- What evidence should be stored with an Explain result?
Review the answers
1. No; it plans only, with the documented one-read charge.
2. Scanned is work inspected; returned is useful output.
3. Enterprise can fall back to scans because indexes are optional.
4. No; current Standard Explain supports polled queries, not streaming queries.
5. Query/index definition, mode/edition, SDK, dataset size/time, and whether execution occurred.
Summary and next step
This lesson established the working contract for Query Explain to Inspect Execution, Index Use, Scanned/Returned Work, and Optimization Evidence. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Query Insights to Identify Expensive/Slow Query Patterns and Prioritize Tuning.
Authoritative references
- Firebase: Monitor Cloud Firestore activity — usage dashboards, Monitoring metrics, Security Rules metrics, listener/connections, and dashboard-versus-billing caveats.
-
Google Cloud Monitoring metrics reference
— current
firestore.googleapis.com/*metric types, labels, sampling and launch stages. - Firebase: Query Explain for Standard/Core queries — planner vs analyze behavior, scan statistics, IAM authentication, and billing semantics.
- Firebase: Enterprise Native Pipeline Query Explain — explain/analyze/stats modes and execution-tree evidence.
- Firebase: Query Insights — normalized queries, latency/read-load statistics, retention, delay and IAM.
- Firebase: Key Visualizer overview — Standard Native hotspot heatmaps, eligibility, metrics, limits, and data retention.
-
Firebase: Cloud Firestore audit logging
— Data Access/Admin Activity audit evidence and
processing_duration. - Firebase: Firestore editions overview — current observability availability and Standard/Enterprise distinctions.