Chapter 24 · Observability: Query Explain, Query Insights, Key Visualizer, Metrics, Logs, and Troubleshooting
Query Insights to Identify Expensive / Slow Query Patterns and Prioritize Tuning
Use normalized Query Insights statistics to prioritize work by frequency, latency, scanned work, and read load instead of tuning whichever query produced the last anecdotal complaint.
1. AtlasMart problem: optimize the fleet, not the loudest query
A support engineer captures one 600 ms query and asks for an emergency index. Meanwhile another normalized query executes 300,000 times per hour at 35 ms and scans far more index entries in aggregate. Query Insights changes the prioritization question from “which query looked bad once?” to “which query shapes contribute the most latency, errors, scanned work, and billable read load over time?”
- Understand normalized query grouping and why literal values should not fragment query families.
- Use execution count, error count, average duration/results/scanned documents/index entries and read load to prioritize tuning.
- Account for Query Insights delay and retention instead of expecting a live trace.
- Separate Standard/Enterprise Native and MongoDB-compatible query-stat surfaces.
- Combine Insights with Explain and application correlation evidence.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project ID
demo-atlasmart-firestore, Standard-edition Native
mode, database (default), Node.js 22+, Firebase
CLI course baseline 15.30.0, Firebase JavaScript
SDK 12.19.0, Firebase Admin Node SDK
14.4.0 (bundling
@google-cloud/firestore 9.1.0), Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. Firebase CLI
15.30.1 is the current patch release at review
time; no Chapter 24 lab depends on that patch, so the course
stays pinned to 15.30.0 for continuity. Mandatory exercises
are local/no-cost. Cloud Monitoring, Query Insights, Key
Visualizer, Cloud Audit Logs, production Query Explain billing
evidence, and real hotspot/capacity measurements require a
real Google Cloud/Firebase project and are optional bounded
verification steps. Emulator latency is never presented as
production capacity evidence. Current Query Insights data can
be delayed by roughly one to two hours. It keeps 10-minute
granularity for recent intervals and hourly granularity out to
30 days; it has no separate Query Insights surcharge.
| Term | Operational meaning in Chapter 24 |
|---|---|
| Client timing | End-to-end elapsed time observed by the browser/mobile/backend caller. It includes client scheduling and network time that Firestore backend metrics do not. |
| Backend latency |
Firestore service processing time. Cloud Monitoring
api/request_latencies excludes
client-to-service round-trip time.
|
| Standard Native | Standard edition Core operations. Queries require indexes; Key Visualizer is currently documented for this edition/mode. |
| Enterprise Native Core | Familiar Core API on Enterprise. Realtime/offline remain Core features, but indexes are optional and billing is byte-unit based. |
| Enterprise Pipeline | Advanced stage/expression query interface. Query Explain exposes an execution tree, scanned records/bytes, memory and read units. |
| MongoDB compatibility |
Enterprise MongoDB-protocol interface. It has separate
explain/Query Insights behavior and must
not be diagnosed with Native assumptions.
|
| Query Explain | Per-query planner/execution evidence. Planning-only and execute/analyze modes have different cost and side-effect implications. |
| Query Insights | Aggregated normalized-query statistics over time, useful for frequency/latency/read-load prioritization rather than one-off diagnosis. |
| Key Visualizer | Standard Native key-range/index-range heatmaps for hotspot diagnosis. It is not a generic trace viewer or a capacity benchmark. |
| Correlation ID | Application-generated opaque identifier propagated through logs/timers so one user action can be linked across client, backend and Firestore evidence without logging tokens or PII. |
2. Normalized query statistics: structure, not one literal
Queries with equivalent structure are grouped under a normalized query. That lets an operations team see the shape “orders by tenant and createdAt range” instead of hundreds of separate rows for each tenant/date. A normalized query is an observability grouping, not a security boundary: tenant authorization still belongs in Rules/IAM/application checks.
| Statistic | Operational use | Caveat |
|---|---|---|
| Execution count | Find high-frequency query shapes | High frequency is not bad by itself. |
| Error count | Spot broken deploy/query contracts | Separate auth/index/invalid-argument/transient errors. |
| Average execution duration | Find persistently slow shapes | Average can still hide tail latency. |
| Average results returned | Estimate useful output size | Not a billing statement. |
| Average documents scanned | Detect document scan amplification | Interpret with mode/index state. |
| Average index entries scanned | Detect index scan amplification | Real-time queries are not represented in per-query Monitoring metrics. |
| Load by read operations | Prioritize financially material queries | Confirm final cost in billing reports. |
3. Delay and retention determine what questions Insights can answer
Query Insights is retrospective, not request tracing. Current documentation states an approximately one-to-two-hour delay, 10-minute granularity for data up to four days old, and hourly granularity out to 30 days. For an incident happening now, use application timing, request/error metrics and logs first; use Insights later to quantify the query family across the incident window.
You will be waiting on a delayed aggregate while the system is changing. Use it to establish frequency/load and to validate post-incident tuning at fleet scale, not as the sole live incident console.
4. Mandatory local simulation: rank normalized query families
[ {"id":"Q-catalog-search","exec":120000,"errors":30,"avg_ms":42,"avg_results":18,"avg_docs_scanned":22,"avg_index_scanned":160,"read_ops":260000}, {"id":"Q-seller-report","exec":600,"errors":1,"avg_ms":410,"avg_results":80,"avg_docs_scanned":5400,"avg_index_scanned":8800,"read_ops":3250000}, {"id":"Q-order-detail","exec":300000,"errors":12,"avg_ms":18,"avg_results":1,"avg_docs_scanned":1,"avg_index_scanned":1,"read_ops":300000}, {"id":"Q-admin-audit","exec":40,"errors":0,"avg_ms":900,"avg_results":500,"avg_docs_scanned":15000,"avg_index_scanned":0,"read_ops":20000}]
import fs from "node:fs";const rows=JSON.parse(fs.readFileSync("evidence/query-insights.json"));for (const r of rows) { r.scanAmp=r.avg_docs_scanned/Math.max(1,r.avg_results); r.errorRate=r.errors/Math.max(1,r.exec); // Training score only: never ship universal weights as an SRE policy. r.trainingPriority=(Math.log10(r.exec+1)*r.avg_ms)+(Math.log10(r.read_ops+1)*r.scanAmp);}rows.sort((a,b)=>b.trainingPriority-a.trainingPriority);console.table(rows.map(({id,exec,avg_ms,read_ops,scanAmp,errorRate,trainingPriority})=>({id,exec,avg_ms,read_ops,scanAmp,errorRate,trainingPriority})));
The score is deliberately labeled a training heuristic, not a recommended universal formula. In production, prioritize against your SLOs, cost budgets, product criticality, security risk and change safety.
5. From Insights to Explain: fleet signal → query experiment
- Use Query Insights to select a high-impact normalized query family.
- Reproduce the exact access pattern with representative parameter ranges.
- Run planning-only Explain; inspect selected indexes/plan.
- Bound and run analyze where cost permits; capture scanned/returned/read evidence.
- Change one variable: index, filter/order contract, materialized view, data model or call frequency.
- Verify result correctness, then compare Explain evidence.
- After deployment, wait for Insights delay and compare the normalized family over equivalent windows.
6. Standard vs Enterprise vs MongoDB compatibility
| Surface | What to expect |
|---|---|
| Standard Native | Current editions documentation lists Query Insights alongside Query Explain and Key Visualizer. Use document/read-operation semantics and required indexes. |
| Enterprise Native Core/Pipeline |
Query Insights can include methods such as
listDocuments, runQuery,
runAggregationQuery,
partitionQuery, and
executePipeline; indexes are optional, so
scan evidence is particularly important.
|
| MongoDB compatibility | Separate Query Insights/Explain documentation and MongoDB-compatible query methods. Diagnose using MongoDB-compatible normalized operations and byte-unit billing—not Standard document-read assumptions. |
| Realtime listeners | Insights is not a per-listener lifecycle profiler. Use listener/connection metrics and application instrumentation. |
7. Security and privacy of query observability
Normalized query text and operational metadata can still reveal
collection/field names or application structure. Grant
datastore.insights.get via a least-privilege role
such as Datastore Viewer or a custom role. Do not copy Query
Insights screenshots into public tickets if field names are
sensitive. Observability data itself belongs in the governance
matrix from Chapter 23.
8. Failure injection: optimize the wrong query
Choose the query with the largest average latency only.
In the fixture, Q-admin-audit is slowest but rare;
Q-seller-report has much greater aggregate read
load and scan amplification. Repair the process by recording
product criticality, frequency, read load and scan amplification
before choosing the first tuning candidate.
Production judgment
Query Insights is a prioritization surface. Query Explain is a query experiment. Monitoring is a service-level signal. Application logs/traces connect symptoms to product actions. A mature workflow uses them together; none is a replacement for correctness tests or billing reports.
Knowledge check
- Why normalize queries?
- How delayed is Query Insights?
- Why can the highest-latency query be a poor first optimization target?
- What should follow a Query Insights finding?
- Does Query Insights replace billing reports?
Review the answers
1. To aggregate equivalent query structures across literal parameter values.
2. Roughly one to two hours per current documentation.
3. It may be rare and low-cost; frequency/load/product criticality matter.
4. Reproduce, Explain, correctness-test, change one variable, then validate over comparable windows.
5. No; it supports operational/cost prioritization, not invoice truth.
Summary and next step
This lesson established the working contract for Query Insights to Identify Expensive/Slow Query Patterns and Prioritize Tuning. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Key Visualizer/Hotspot Diagnostics Where Available, Application Tracing, Logs, and Correlation IDs.
Authoritative references
- Firebase: Monitor Cloud Firestore activity — usage dashboards, Monitoring metrics, Security Rules metrics, listener/connections, and dashboard-versus-billing caveats.
-
Google Cloud Monitoring metrics reference
— current
firestore.googleapis.com/*metric types, labels, sampling and launch stages. - Firebase: Query Explain for Standard/Core queries — planner vs analyze behavior, scan statistics, IAM authentication, and billing semantics.
- Firebase: Enterprise Native Pipeline Query Explain — explain/analyze/stats modes and execution-tree evidence.
- Firebase: Query Insights — normalized queries, latency/read-load statistics, retention, delay and IAM.
- Firebase: Key Visualizer overview — Standard Native hotspot heatmaps, eligibility, metrics, limits, and data retention.
-
Firebase: Cloud Firestore audit logging
— Data Access/Admin Activity audit evidence and
processing_duration. - Firebase: Firestore editions overview — current observability availability and Standard/Enterprise distinctions.