Chapter 24 · Observability: Query Explain, Query Insights, Key Visualizer, Metrics, Logs, and Troubleshooting

Cloud Monitoring Metrics for Reads / Writes / Deletes, Latency, Errors, Quotas, Storage, and Listener Behavior

Instrument AtlasMart with server and application evidence so throughput, tail latency, errors, storage, and realtime listener behavior can be diagnosed without confusing dashboard estimates with billing truth.

Advanced · 170–230 minutesmetrics · latency · errors · quotas · listeners · storageNode 22+ · Firebase CLI course baseline 15.30.0 · JS SDK 12.19.0 · Admin SDK 14.4.0Mandatory evidence lab local/no-cost · Monitoring/Insights/Key Visualizer/Audit optional managed verificationLast reviewed: 17 September 2026

1. AtlasMart incident: “Firestore is slow” is not a diagnosis

At 10:15 UTC the AtlasMart operations channel reports three symptoms at once: checkout p99 latency increased, document reads doubled, and seller dashboards show delayed low-stock updates. One engineer blames Firestore capacity, another blames the network, and a third proposes raising quotas. None of those actions is justified yet. Chapter 24 starts from an operational rule: measure the layer that owns the symptom before changing the system.

Learning outcomes
  • Distinguish application end-to-end latency from Firestore backend latency.
  • Read operations, errors, listener/connections, storage, Rules and quota signals without treating any one dashboard as billing truth.
  • Compute p50/p95/p99 from local measurements and interpret tail latency rather than averages only.
  • Separate emulator evidence from managed-cloud evidence.
  • Build privacy-safe structured logs and a reusable incident evidence bundle.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 24 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps project ID demo-atlasmart-firestore, Standard-edition Native mode, database (default), Node.js 22+, Firebase CLI course baseline 15.30.0, Firebase JavaScript SDK 12.19.0, Firebase Admin Node SDK 14.4.0 (bundling @google-cloud/firestore 9.1.0), Firestore emulator 127.0.0.1:8080, Auth emulator 127.0.0.1:9099, and Emulator UI 127.0.0.1:4000. Firebase CLI 15.30.1 is the current patch release at review time; no Chapter 24 lab depends on that patch, so the course stays pinned to 15.30.0 for continuity. Mandatory exercises are local/no-cost. Cloud Monitoring, Query Insights, Key Visualizer, Cloud Audit Logs, production Query Explain billing evidence, and real hotspot/capacity measurements require a real Google Cloud/Firebase project and are optional bounded verification steps. Emulator latency is never presented as production capacity evidence. Cloud Monitoring metrics are sampled about every minute and can take several minutes to become visible, so dashboards are not an instantaneous request tracer.

Term Operational meaning in Chapter 24
Client timing End-to-end elapsed time observed by the browser/mobile/backend caller. It includes client scheduling and network time that Firestore backend metrics do not.
Backend latency Firestore service processing time. Cloud Monitoring api/request_latencies excludes client-to-service round-trip time.
Standard Native Standard edition Core operations. Queries require indexes; Key Visualizer is currently documented for this edition/mode.
Enterprise Native Core Familiar Core API on Enterprise. Realtime/offline remain Core features, but indexes are optional and billing is byte-unit based.
Enterprise Pipeline Advanced stage/expression query interface. Query Explain exposes an execution tree, scanned records/bytes, memory and read units.
MongoDB compatibility Enterprise MongoDB-protocol interface. It has separate explain/Query Insights behavior and must not be diagnosed with Native assumptions.
Query Explain Per-query planner/execution evidence. Planning-only and execute/analyze modes have different cost and side-effect implications.
Query Insights Aggregated normalized-query statistics over time, useful for frequency/latency/read-load prioritization rather than one-off diagnosis.
Key Visualizer Standard Native key-range/index-range heatmaps for hotspot diagnosis. It is not a generic trace viewer or a capacity benchmark.
Correlation ID Application-generated opaque identifier propagated through logs/timers so one user action can be linked across client, backend and Firestore evidence without logging tokens or PII.

2. Four clocks answer four different questions

Clock/evidence What it measures Typical mistake
Client/backend app timer Full caller-observed latency around the SDK call Blaming Firestore for DNS, mobile radio, proxy, serialization or application queueing.
Cloud Monitoring api/request_latencies Completed non-streaming request latency from Firestore frontend through response production Treating it as end-user RTT.
Audit Log processing_duration Database-side processing time for audited data reads/writes Assuming Data Access logs are always enabled or free.
Query Explain execution duration/tree One explained query’s backend plan/work Generalizing one run into fleet-wide frequency or p99 behavior.

If client p99 rises while backend p99 stays stable, investigate the path before Firestore: network, connection churn, client CPU, proxy, retries, serialization, or application queueing. If both rise, continue into query/index/hotspot evidence.

3. Cloud Monitoring metrics: build a database-level signal map

Signal family Current metric examples Question answered
Requests/latency api/request_count, api/request_latencies Which API methods/status codes are slow or failing?
Documents document/read_ops_count, write_ops_count, delete_ops_count Is workload shape changing?
Realtime network/active_connections, network/snapshot_listeners Did listener/connection count change?
Query work query_stat/per_query/scanned_documents_counts, scanned_index_entries_counts, result_counts Is scan amplification growing? Realtime queries are excluded from these per-query metrics.
Rules rules/evaluation_count with ALLOW/DENY/ERROR Did client Security Rules behavior change?
Storage storage/data_and_index_storage_bytes, backup/PITR bytes Is data/index/recovery footprint changing?
Enterprise billing units api/billable_read_units, billable_write_units, realtime units What byte-unit work is generated on Enterprise?
Dashboard is not invoice

The Firestore usage dashboard can omit or collapse work that billing counts: index-entry reads, zero-result minimum reads, imports/exports, some no-op/collapsed writes, and other documented cases. Use the billing report for financial truth; use Monitoring for operational shape.

4. Realtime listeners: connection count is not listener count

A mobile/web SDK normally maintains one active connection that can multiplex several snapshot listeners. Server client libraries create a connection per snapshot listener. Therefore a “connections doubled” alert and a “listeners doubled” alert imply different things. A screen accidentally mounting ten listeners per user can inflate network/snapshot_listeners without multiplying mobile connections tenfold.

listener-registry.mjs
const active = new Map();export function trackedListen(key, subscribe) {  if (active.has(key)) throw new Error(`duplicate-listener:${key}`);  const unsubscribe = subscribe();  active.set(key, { startedAt: Date.now(), unsubscribe });  return () => {    active.get(key)?.unsubscribe();    active.delete(key);  };}export function listenerGauge() { return active.size; }

Local listener registries provide immediate application evidence; Cloud Monitoring gives managed aggregate evidence later. Use both rather than treating either as the whole truth.

5. Mandatory local lab: timing, percentiles, error codes, and correlation IDs

observe-query.mjs
import crypto from "node:crypto";import { performance } from "node:perf_hooks";const samples = [];export async function observed(name, fn) {  const correlationId = crypto.randomUUID();  const start = performance.now();  try {    const value = await fn(correlationId);    samples.push({name, correlationId, ok:true, ms:performance.now()-start});    return value;  } catch (error) {    samples.push({name, correlationId, ok:false, code:error.code ?? "UNKNOWN", ms:performance.now()-start});    throw error;  }}function pct(values, p) {  const a=[...values].sort((x,y)=>x-y);  const i=Math.min(a.length-1, Math.ceil((p/100)*a.length)-1);  return a[Math.max(0,i)] ?? 0;}export function report() {  const ok=samples.filter(x=>x.ok).map(x=>x.ms);  return {count:samples.length, errors:samples.length-ok.length,    p50_ms:pct(ok,50), p95_ms:pct(ok,95), p99_ms:pct(ok,99)};}
  1. Seed a bounded AtlasMart fixture in the emulator: 200 catalogItems, 40 sellers, deterministic prices/categories.
  2. Run 100 bounded queries through observed(); save raw per-request JSON and a percentile summary.
  3. Inject ten deterministic denied client reads through Rules and record the error code separately from latency.
  4. Mount and unmount the same listener 100 times; assert the local registry returns to zero.
  5. Do not infer production p99 or capacity from these numbers. They prove instrumentation and regression logic only.

6. Controlled failure: averages hide the tail

Create 95 synthetic samples at 10 ms and five at 500 ms. The average looks modest compared with the user-visible tail. A production SLO expressed only as mean latency can therefore remain “healthy” while a small but important cohort is badly affected.

tail-latency-demo.mjs
const x=[...Array(95).fill(10), ...Array(5).fill(500)];const mean=x.reduce((a,b)=>a+b,0)/x.length;const sorted=[...x].sort((a,b)=>a-b);const p95=sorted[Math.ceil(.95*x.length)-1];const p99=sorted[Math.ceil(.99*x.length)-1];console.log({mean,p95,p99});// Expected shape: mean is far below p99; investigate tail cohorts.

7. Optional managed verification: Monitoring and audit evidence

In a disposable project, filter api/request_latencies by database, API method and response code; compare p50/p95/p99 with application timing over the same interval. Data Access audit logs can provide processing_duration for audited reads/writes, but enabling and retaining those logs has cost/privacy implications. Never enable verbose payload logging as an incident shortcut.

Sampling reality

Firestore metrics are typically sampled every 60 seconds and can take up to roughly four minutes to appear. Query Insights is delayed much longer. For a live incident, pair aggregate managed signals with application-local correlation/timing evidence.

Production judgment: alert on decisions, not curiosity

Useful alerts have an owner and a decision: sustained backend p99 degradation, rising error-code rate, unexpected listener growth, Rules denials after a deploy, or storage/index growth beyond a forecast envelope. Avoid alerting on every metric independently. Quota increases are not a first response to abuse, retry storms, accidental listeners, or an inefficient query shape.

Knowledge check

  1. Why can application p99 be higher than Firestore api/request_latencies?
  2. Why is the usage dashboard not financial truth?
  3. What is the relationship between mobile connections and snapshot listeners?
  4. Why is p99 often more actionable than a mean?
  5. What does emulator latency prove?
Review the answers

1. App timing includes network/client/application overhead omitted by backend metrics.

2. Documented billing work can be omitted/collapsed in usage dashboards; billing reports win.

3. Mobile/web can multiplex listeners over one connection; server listeners use separate connections.

4. It exposes tail cohorts hidden by averages.

5. Instrumentation/correctness under local conditions—not production capacity, regional latency, autoscaling, or billing.

Summary and next step

This lesson established the working contract for Cloud Monitoring Metrics for Reads/Writes/Deletes, Latency, Errors, Quotas, Storage, and Listener Behavior. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Query Explain to Inspect Execution, Index Use, Scanned/Returned Work, and Optimization Evidence.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.