Chapter 24 · Observability: Query Explain, Query Insights, Key Visualizer, Metrics, Logs, and Troubleshooting

Build a Troubleshooting Runbook for Permission Errors, Missing Indexes, Hotspots, Latency, Quota, and Cost Spikes

Turn recurring Firestore incidents into a deterministic evidence-first runbook that classifies permission, index, hotspot, latency, quota, listener, and cost failures before changing production configuration.

Advanced · 170–230 minutesrunbook · permission · indexes · hotspots · quota · costNode 22+ · Firebase CLI course baseline 15.30.0 · JS SDK 12.19.0 · Admin SDK 14.4.0Mandatory evidence lab local/no-cost · Monitoring/Insights/Key Visualizer/Audit optional managed verificationLast reviewed: 17 September 2026

1. AtlasMart problem: stop debugging by reflex

AtlasMart has accumulated “fixes” from past incidents: retry everything, add indexes whenever a query fails, raise quotas on spikes, and restart listeners when users complain. Those actions can amplify outages. The final lesson converts Chapters 05–24 into an evidence-first runbook: classify the failure, collect the minimum safe evidence, test a hypothesis, change one control, verify, and preserve rollback.

Learning outcomes
  • Classify permission, missing-index, hotspot, latency, quota, listener and cost incidents.
  • Map each class to evidence and a safe first action.
  • Use correlation IDs, metrics, Explain/Insights/Key Visualizer and billing evidence without mixing their meanings.
  • Define stop conditions, rollback gates and post-change validation.
  • Create a machine-checkable runbook contract.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 24 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps project ID demo-atlasmart-firestore, Standard-edition Native mode, database (default), Node.js 22+, Firebase CLI course baseline 15.30.0, Firebase JavaScript SDK 12.19.0, Firebase Admin Node SDK 14.4.0 (bundling @google-cloud/firestore 9.1.0), Firestore emulator 127.0.0.1:8080, Auth emulator 127.0.0.1:9099, and Emulator UI 127.0.0.1:4000. Firebase CLI 15.30.1 is the current patch release at review time; no Chapter 24 lab depends on that patch, so the course stays pinned to 15.30.0 for continuity. Mandatory exercises are local/no-cost. Cloud Monitoring, Query Insights, Key Visualizer, Cloud Audit Logs, production Query Explain billing evidence, and real hotspot/capacity measurements require a real Google Cloud/Firebase project and are optional bounded verification steps. Emulator latency is never presented as production capacity evidence. The local drill uses saved traces and emulator failures so every branch is repeatable without billable cloud incidents.

Term Operational meaning in Chapter 24
Client timing End-to-end elapsed time observed by the browser/mobile/backend caller. It includes client scheduling and network time that Firestore backend metrics do not.
Backend latency Firestore service processing time. Cloud Monitoring api/request_latencies excludes client-to-service round-trip time.
Standard Native Standard edition Core operations. Queries require indexes; Key Visualizer is currently documented for this edition/mode.
Enterprise Native Core Familiar Core API on Enterprise. Realtime/offline remain Core features, but indexes are optional and billing is byte-unit based.
Enterprise Pipeline Advanced stage/expression query interface. Query Explain exposes an execution tree, scanned records/bytes, memory and read units.
MongoDB compatibility Enterprise MongoDB-protocol interface. It has separate explain/Query Insights behavior and must not be diagnosed with Native assumptions.
Query Explain Per-query planner/execution evidence. Planning-only and execute/analyze modes have different cost and side-effect implications.
Query Insights Aggregated normalized-query statistics over time, useful for frequency/latency/read-load prioritization rather than one-off diagnosis.
Key Visualizer Standard Native key-range/index-range heatmaps for hotspot diagnosis. It is not a generic trace viewer or a capacity benchmark.
Correlation ID Application-generated opaque identifier propagated through logs/timers so one user action can be linked across client, backend and Firestore evidence without logging tokens or PII.

2. Triage matrix: symptom → evidence → action

Symptom First evidence Likely branches Unsafe reflex
PERMISSION_DENIED Client auth state, Rules test/evaluation metrics; server IAM/Audit logs Rules/query incompatibility, stale claims, IAM/service perimeter Open Rules to allow read, write: if true.
Missing-index error (Standard) Exact query + index error + deployed index config Required composite/index still building or undeployed Create every suggested index without query inventory.
High latency Client p95/p99 + backend p95/p99 + errors Network, query work, hotspot, contention, dependency, retry storm Raise quota or add retries before finding layer.
ABORTED/DEADLINE_EXCEEDED Request status + transaction attempt counts + hotspot evidence Contention/hot document/key range, overload, slow dependency Infinite retry.
Quota/resource exhaustion Quota metric + request rate + caller/operation mix Legitimate growth, bug, abuse, burst Increase quota before backpressure/abuse checks.
Cost spike Billing export/report + read/write/listener/query evidence New listeners, scan amplification, index reads, batch job, TTL/export Trust Usage dashboard alone.
Realtime stale/expensive Listener/connection metrics + lifecycle registry + snapshot metadata Leaked listener, broad query, reconnect churn, offline/cache semantics Poll faster and keep listener too.

3. Branch 1: permission errors—identify the authorization plane

  1. Is this mobile/web client or trusted server?
  2. Client: verify Firebase Auth identity/claims, exact query shape, Rules test, Rules evaluation ALLOW/DENY/ERROR trend.
  3. Server/Admin: Rules are bypassed; inspect ADC/workload identity, IAM, application authorization, VPC-SC/perimeter policy and audit log principal.
  4. MongoDB compatibility: inspect SCRAM/OIDC/IAM auth and Mongo-compatible command support rather than Firebase client Rules assumptions.
Do not weaken the wrong plane.

A server PERMISSION_DENIED is not repaired by changing Firestore Security Rules. A client query denied because Rules cannot prove its constraints is not repaired by granting a server service account.

4. Branch 2: missing index or high scan work

In Standard, missing required indexes can fail a query. In Enterprise, the same conceptual access pattern may run by scanning because indexes are optional. Therefore the runbook has two different branches:

Edition/interface Failure/performance signal Evidence
Standard Core Missing-index error or unexpectedly high index scan Index config + Query Explain + query correctness test.
Enterprise Core/Pipeline Query succeeds but scans many rows/bytes Query Explain tree/read units + Query Insights.
MongoDB compatibility Mongo-compatible command/plan behavior Mongo explain + compatibility matrix + Query Insights.

5. Branch 3: hotspot/contention

Look for hot-document updates, sequential document IDs, indexed sequential fields, sudden traffic ramps and transaction retry pressure. For Standard Native, eligible Key Visualizer heatmaps can corroborate document/index key concentration. For all modes, combine request status/latency with data-model knowledge. Do not shard a strict invariant such as inventory merely to reduce contention without redesigning correctness.

6. Branch 4: quota and cost spikes

Quota exhaustion and cost spikes can share a root cause—accidental fan-out, listener leaks, abuse, batch jobs—but they are not the same. A quota is a service limit/control; billing is usage multiplied by SKU/pricing. First isolate operation/caller/query family and determine whether work is legitimate. Add backpressure, bounded concurrency, cache/materialization, query/index redesign or abuse protection before requesting a higher ceiling.

7. Mandatory incident drill: saved evidence bundle

incident/manifest.json
{  "incident":"atlasmart-obs-001",  "window":"2026-09-17T00:00:00Z/2026-09-17T00:30:00Z",  "edition":"Standard",  "mode":"Native/Core",  "symptom":"p99 checkout latency + read spike",  "evidence":[    "app-latency.json",    "firestore-latency.json",    "query-insights.json",    "explain-before.json",    "listener-gauge.json",    "billing-note.txt"  ],  "sensitivePayloadsIncluded":false}
diagnose.mjs
import fs from "node:fs";const app=JSON.parse(fs.readFileSync("incident/app-latency.json"));const db=JSON.parse(fs.readFileSync("incident/firestore-latency.json"));const listeners=JSON.parse(fs.readFileSync("incident/listener-gauge.json"));const insights=JSON.parse(fs.readFileSync("incident/query-insights.json"));const findings=[];if (app.p99_ms > db.p99_ms*2) findings.push("client/network/application-overhead-suspect");if (listeners.after > listeners.before*1.5) findings.push("listener-growth-suspect");const top=[...insights].sort((a,b)=>b.read_ops-a.read_ops)[0];if (top && top.avg_docs_scanned > top.avg_results*20) findings.push(`scan-amplification:${top.id}`);console.log({findings});if (!findings.length) process.exitCode=2;

Run the drill twice. First with a listener-growth fixture; second with a scan-amplification fixture. The diagnostic must choose different hypotheses. That proves the runbook is evidence-driven rather than hard-coded to the last incident.

8. Change gate: one hypothesis, one reversible change

Before change After change
State query/data/rules/index configuration and expected effect. Re-run correctness tests before performance conclusions.
Record baseline p50/p95/p99, error rate, scan/read/listener evidence. Compare equivalent window/workload; do not cherry-pick.
Define rollback trigger. Preserve rollback until metrics stabilize.
Estimate billing/operational impact. Confirm billing separately from usage-dashboard estimates.
Identify mode/edition dependencies. Check feature status/version compatibility before generalizing.

9. Machine-check the runbook

runbook.json
{  "branches":[    {"id":"permission","requires":["principal","auth_plane","error_code","policy_evidence"],"rollback":true},    {"id":"index-scan","requires":["query_shape","edition_mode","index_state","explain_or_fixture"],"rollback":true},    {"id":"hotspot","requires":["key_shape","contention_errors","latency_distribution"],"rollback":true},    {"id":"quota-cost","requires":["operation_mix","caller","quota_or_billing_evidence"],"rollback":true}  ]}
validate-runbook.mjs
import fs from "node:fs";const r=JSON.parse(fs.readFileSync("runbook.json"));for(const b of r.branches){  if(!b.requires?.length) throw new Error(`${b.id}:no-evidence-contract`);  if(b.rollback !== true) throw new Error(`${b.id}:no-rollback-gate`);}console.log(`RUNBOOK_OK branches=${r.branches.length}`);

10. Operational non-guarantees

  • A fast Explain sample does not guarantee fleet p99.
  • A clean Key Visualizer scan does not exclude network/client bottlenecks.
  • Query Insights is delayed and aggregated.
  • Monitoring usage metrics are sampled/delayed and not final billing records.
  • Emulator timings are not production performance.
  • Retries can increase load and cost; retry only documented transient/idempotent operations with bounds.
  • Backups/PITR protect recoverability, not live latency or authorization correctness.

11. Bridge to Chapter 25

Chapter 24 established measured work: requests, scans, reads, writes, listeners, storage, errors and latency. Chapter 25 converts those measured dimensions into cost engineering, quotas, forecasts and capacity/usage budgets without mistaking “cheap per operation” for “cheap at fleet scale.”

Knowledge check

  1. Why must the runbook identify the authorization plane for a permission error?
  2. How does Standard missing-index behavior differ from Enterprise index-optional behavior?
  3. What is the first response to a quota spike?
  4. Why should each tuning change have a rollback trigger?
  5. Which source is financial truth when the usage dashboard disagrees with billing?
Review the answers

1. Client Rules and server IAM/application authorization are different enforcement layers.

2. Standard required-index queries can fail; Enterprise can scan and become expensive/slow.

3. Identify operation/caller/root cause and apply backpressure/abuse/query fixes before simply raising limits.

4. Performance changes can break correctness, cost or other workloads; reversibility contains blast radius.

5. Billing reports/billing export.

Summary and next step

This lesson established the working contract for Build a Troubleshooting Runbook for Permission Errors, Missing Indexes, Hotspots, Latency, Quota, and Cost Spikes. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Read/Write/Delete/Index/Storage/Network/Backup/PITR Billing Dimensions and Edition Differences.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.