Chapter 24 · Observability: Query Explain, Query Insights, Key Visualizer, Metrics, Logs, and Troubleshooting
Build a Troubleshooting Runbook for Permission Errors, Missing Indexes, Hotspots, Latency, Quota, and Cost Spikes
Turn recurring Firestore incidents into a deterministic evidence-first runbook that classifies permission, index, hotspot, latency, quota, listener, and cost failures before changing production configuration.
1. AtlasMart problem: stop debugging by reflex
AtlasMart has accumulated “fixes” from past incidents: retry everything, add indexes whenever a query fails, raise quotas on spikes, and restart listeners when users complain. Those actions can amplify outages. The final lesson converts Chapters 05–24 into an evidence-first runbook: classify the failure, collect the minimum safe evidence, test a hypothesis, change one control, verify, and preserve rollback.
- Classify permission, missing-index, hotspot, latency, quota, listener and cost incidents.
- Map each class to evidence and a safe first action.
- Use correlation IDs, metrics, Explain/Insights/Key Visualizer and billing evidence without mixing their meanings.
- Define stop conditions, rollback gates and post-change validation.
- Create a machine-checkable runbook contract.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project ID
demo-atlasmart-firestore, Standard-edition Native
mode, database (default), Node.js 22+, Firebase
CLI course baseline 15.30.0, Firebase JavaScript
SDK 12.19.0, Firebase Admin Node SDK
14.4.0 (bundling
@google-cloud/firestore 9.1.0), Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. Firebase CLI
15.30.1 is the current patch release at review
time; no Chapter 24 lab depends on that patch, so the course
stays pinned to 15.30.0 for continuity. Mandatory exercises
are local/no-cost. Cloud Monitoring, Query Insights, Key
Visualizer, Cloud Audit Logs, production Query Explain billing
evidence, and real hotspot/capacity measurements require a
real Google Cloud/Firebase project and are optional bounded
verification steps. Emulator latency is never presented as
production capacity evidence. The local drill uses saved
traces and emulator failures so every branch is repeatable
without billable cloud incidents.
| Term | Operational meaning in Chapter 24 |
|---|---|
| Client timing | End-to-end elapsed time observed by the browser/mobile/backend caller. It includes client scheduling and network time that Firestore backend metrics do not. |
| Backend latency |
Firestore service processing time. Cloud Monitoring
api/request_latencies excludes
client-to-service round-trip time.
|
| Standard Native | Standard edition Core operations. Queries require indexes; Key Visualizer is currently documented for this edition/mode. |
| Enterprise Native Core | Familiar Core API on Enterprise. Realtime/offline remain Core features, but indexes are optional and billing is byte-unit based. |
| Enterprise Pipeline | Advanced stage/expression query interface. Query Explain exposes an execution tree, scanned records/bytes, memory and read units. |
| MongoDB compatibility |
Enterprise MongoDB-protocol interface. It has separate
explain/Query Insights behavior and must
not be diagnosed with Native assumptions.
|
| Query Explain | Per-query planner/execution evidence. Planning-only and execute/analyze modes have different cost and side-effect implications. |
| Query Insights | Aggregated normalized-query statistics over time, useful for frequency/latency/read-load prioritization rather than one-off diagnosis. |
| Key Visualizer | Standard Native key-range/index-range heatmaps for hotspot diagnosis. It is not a generic trace viewer or a capacity benchmark. |
| Correlation ID | Application-generated opaque identifier propagated through logs/timers so one user action can be linked across client, backend and Firestore evidence without logging tokens or PII. |
2. Triage matrix: symptom → evidence → action
| Symptom | First evidence | Likely branches | Unsafe reflex |
|---|---|---|---|
PERMISSION_DENIED |
Client auth state, Rules test/evaluation metrics; server IAM/Audit logs | Rules/query incompatibility, stale claims, IAM/service perimeter |
Open Rules to allow read, write: if true.
|
| Missing-index error (Standard) | Exact query + index error + deployed index config | Required composite/index still building or undeployed | Create every suggested index without query inventory. |
| High latency | Client p95/p99 + backend p95/p99 + errors | Network, query work, hotspot, contention, dependency, retry storm | Raise quota or add retries before finding layer. |
| ABORTED/DEADLINE_EXCEEDED | Request status + transaction attempt counts + hotspot evidence | Contention/hot document/key range, overload, slow dependency | Infinite retry. |
| Quota/resource exhaustion | Quota metric + request rate + caller/operation mix | Legitimate growth, bug, abuse, burst | Increase quota before backpressure/abuse checks. |
| Cost spike | Billing export/report + read/write/listener/query evidence | New listeners, scan amplification, index reads, batch job, TTL/export | Trust Usage dashboard alone. |
| Realtime stale/expensive | Listener/connection metrics + lifecycle registry + snapshot metadata | Leaked listener, broad query, reconnect churn, offline/cache semantics | Poll faster and keep listener too. |
3. Branch 1: permission errors—identify the authorization plane
- Is this mobile/web client or trusted server?
- Client: verify Firebase Auth identity/claims, exact query shape, Rules test, Rules evaluation ALLOW/DENY/ERROR trend.
- Server/Admin: Rules are bypassed; inspect ADC/workload identity, IAM, application authorization, VPC-SC/perimeter policy and audit log principal.
- MongoDB compatibility: inspect SCRAM/OIDC/IAM auth and Mongo-compatible command support rather than Firebase client Rules assumptions.
A server PERMISSION_DENIED is not repaired by
changing Firestore Security Rules. A client query denied
because Rules cannot prove its constraints is not repaired by
granting a server service account.
4. Branch 2: missing index or high scan work
In Standard, missing required indexes can fail a query. In Enterprise, the same conceptual access pattern may run by scanning because indexes are optional. Therefore the runbook has two different branches:
| Edition/interface | Failure/performance signal | Evidence |
|---|---|---|
| Standard Core | Missing-index error or unexpectedly high index scan | Index config + Query Explain + query correctness test. |
| Enterprise Core/Pipeline | Query succeeds but scans many rows/bytes | Query Explain tree/read units + Query Insights. |
| MongoDB compatibility | Mongo-compatible command/plan behavior |
Mongo explain + compatibility matrix +
Query Insights.
|
5. Branch 3: hotspot/contention
Look for hot-document updates, sequential document IDs, indexed sequential fields, sudden traffic ramps and transaction retry pressure. For Standard Native, eligible Key Visualizer heatmaps can corroborate document/index key concentration. For all modes, combine request status/latency with data-model knowledge. Do not shard a strict invariant such as inventory merely to reduce contention without redesigning correctness.
6. Branch 4: quota and cost spikes
Quota exhaustion and cost spikes can share a root cause—accidental fan-out, listener leaks, abuse, batch jobs—but they are not the same. A quota is a service limit/control; billing is usage multiplied by SKU/pricing. First isolate operation/caller/query family and determine whether work is legitimate. Add backpressure, bounded concurrency, cache/materialization, query/index redesign or abuse protection before requesting a higher ceiling.
7. Mandatory incident drill: saved evidence bundle
{ "incident":"atlasmart-obs-001", "window":"2026-09-17T00:00:00Z/2026-09-17T00:30:00Z", "edition":"Standard", "mode":"Native/Core", "symptom":"p99 checkout latency + read spike", "evidence":[ "app-latency.json", "firestore-latency.json", "query-insights.json", "explain-before.json", "listener-gauge.json", "billing-note.txt" ], "sensitivePayloadsIncluded":false}
import fs from "node:fs";const app=JSON.parse(fs.readFileSync("incident/app-latency.json"));const db=JSON.parse(fs.readFileSync("incident/firestore-latency.json"));const listeners=JSON.parse(fs.readFileSync("incident/listener-gauge.json"));const insights=JSON.parse(fs.readFileSync("incident/query-insights.json"));const findings=[];if (app.p99_ms > db.p99_ms*2) findings.push("client/network/application-overhead-suspect");if (listeners.after > listeners.before*1.5) findings.push("listener-growth-suspect");const top=[...insights].sort((a,b)=>b.read_ops-a.read_ops)[0];if (top && top.avg_docs_scanned > top.avg_results*20) findings.push(`scan-amplification:${top.id}`);console.log({findings});if (!findings.length) process.exitCode=2;
Run the drill twice. First with a listener-growth fixture; second with a scan-amplification fixture. The diagnostic must choose different hypotheses. That proves the runbook is evidence-driven rather than hard-coded to the last incident.
8. Change gate: one hypothesis, one reversible change
| Before change | After change |
|---|---|
| State query/data/rules/index configuration and expected effect. | Re-run correctness tests before performance conclusions. |
| Record baseline p50/p95/p99, error rate, scan/read/listener evidence. | Compare equivalent window/workload; do not cherry-pick. |
| Define rollback trigger. | Preserve rollback until metrics stabilize. |
| Estimate billing/operational impact. | Confirm billing separately from usage-dashboard estimates. |
| Identify mode/edition dependencies. | Check feature status/version compatibility before generalizing. |
9. Machine-check the runbook
{ "branches":[ {"id":"permission","requires":["principal","auth_plane","error_code","policy_evidence"],"rollback":true}, {"id":"index-scan","requires":["query_shape","edition_mode","index_state","explain_or_fixture"],"rollback":true}, {"id":"hotspot","requires":["key_shape","contention_errors","latency_distribution"],"rollback":true}, {"id":"quota-cost","requires":["operation_mix","caller","quota_or_billing_evidence"],"rollback":true} ]}
import fs from "node:fs";const r=JSON.parse(fs.readFileSync("runbook.json"));for(const b of r.branches){ if(!b.requires?.length) throw new Error(`${b.id}:no-evidence-contract`); if(b.rollback !== true) throw new Error(`${b.id}:no-rollback-gate`);}console.log(`RUNBOOK_OK branches=${r.branches.length}`);
10. Operational non-guarantees
- A fast Explain sample does not guarantee fleet p99.
- A clean Key Visualizer scan does not exclude network/client bottlenecks.
- Query Insights is delayed and aggregated.
- Monitoring usage metrics are sampled/delayed and not final billing records.
- Emulator timings are not production performance.
- Retries can increase load and cost; retry only documented transient/idempotent operations with bounds.
- Backups/PITR protect recoverability, not live latency or authorization correctness.
11. Bridge to Chapter 25
Chapter 24 established measured work: requests, scans, reads, writes, listeners, storage, errors and latency. Chapter 25 converts those measured dimensions into cost engineering, quotas, forecasts and capacity/usage budgets without mistaking “cheap per operation” for “cheap at fleet scale.”
Knowledge check
- Why must the runbook identify the authorization plane for a permission error?
- How does Standard missing-index behavior differ from Enterprise index-optional behavior?
- What is the first response to a quota spike?
- Why should each tuning change have a rollback trigger?
- Which source is financial truth when the usage dashboard disagrees with billing?
Review the answers
1. Client Rules and server IAM/application authorization are different enforcement layers.
2. Standard required-index queries can fail; Enterprise can scan and become expensive/slow.
3. Identify operation/caller/root cause and apply backpressure/abuse/query fixes before simply raising limits.
4. Performance changes can break correctness, cost or other workloads; reversibility contains blast radius.
5. Billing reports/billing export.
Summary and next step
This lesson established the working contract for Build a Troubleshooting Runbook for Permission Errors, Missing Indexes, Hotspots, Latency, Quota, and Cost Spikes. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Read/Write/Delete/Index/Storage/Network/Backup/PITR Billing Dimensions and Edition Differences.
Authoritative references
- Firebase: Monitor Cloud Firestore activity — usage dashboards, Monitoring metrics, Security Rules metrics, listener/connections, and dashboard-versus-billing caveats.
-
Google Cloud Monitoring metrics reference
— current
firestore.googleapis.com/*metric types, labels, sampling and launch stages. - Firebase: Query Explain for Standard/Core queries — planner vs analyze behavior, scan statistics, IAM authentication, and billing semantics.
- Firebase: Enterprise Native Pipeline Query Explain — explain/analyze/stats modes and execution-tree evidence.
- Firebase: Query Insights — normalized queries, latency/read-load statistics, retention, delay and IAM.
- Firebase: Key Visualizer overview — Standard Native hotspot heatmaps, eligibility, metrics, limits, and data retention.
-
Firebase: Cloud Firestore audit logging
— Data Access/Admin Activity audit evidence and
processing_duration. - Firebase: Firestore editions overview — current observability availability and Standard/Enterprise distinctions.