Chapter 24 · Observability: Query Explain, Query Insights, Key Visualizer, Metrics, Logs, and Troubleshooting
Cloud Monitoring Metrics for Reads / Writes / Deletes, Latency, Errors, Quotas, Storage, and Listener Behavior
Instrument AtlasMart with server and application evidence so throughput, tail latency, errors, storage, and realtime listener behavior can be diagnosed without confusing dashboard estimates with billing truth.
1. AtlasMart incident: “Firestore is slow” is not a diagnosis
At 10:15 UTC the AtlasMart operations channel reports three symptoms at once: checkout p99 latency increased, document reads doubled, and seller dashboards show delayed low-stock updates. One engineer blames Firestore capacity, another blames the network, and a third proposes raising quotas. None of those actions is justified yet. Chapter 24 starts from an operational rule: measure the layer that owns the symptom before changing the system.
- Distinguish application end-to-end latency from Firestore backend latency.
- Read operations, errors, listener/connections, storage, Rules and quota signals without treating any one dashboard as billing truth.
- Compute p50/p95/p99 from local measurements and interpret tail latency rather than averages only.
- Separate emulator evidence from managed-cloud evidence.
- Build privacy-safe structured logs and a reusable incident evidence bundle.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project ID
demo-atlasmart-firestore, Standard-edition Native
mode, database (default), Node.js 22+, Firebase
CLI course baseline 15.30.0, Firebase JavaScript
SDK 12.19.0, Firebase Admin Node SDK
14.4.0 (bundling
@google-cloud/firestore 9.1.0), Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. Firebase CLI
15.30.1 is the current patch release at review
time; no Chapter 24 lab depends on that patch, so the course
stays pinned to 15.30.0 for continuity. Mandatory exercises
are local/no-cost. Cloud Monitoring, Query Insights, Key
Visualizer, Cloud Audit Logs, production Query Explain billing
evidence, and real hotspot/capacity measurements require a
real Google Cloud/Firebase project and are optional bounded
verification steps. Emulator latency is never presented as
production capacity evidence. Cloud Monitoring metrics are
sampled about every minute and can take several minutes to
become visible, so dashboards are not an instantaneous request
tracer.
| Term | Operational meaning in Chapter 24 |
|---|---|
| Client timing | End-to-end elapsed time observed by the browser/mobile/backend caller. It includes client scheduling and network time that Firestore backend metrics do not. |
| Backend latency |
Firestore service processing time. Cloud Monitoring
api/request_latencies excludes
client-to-service round-trip time.
|
| Standard Native | Standard edition Core operations. Queries require indexes; Key Visualizer is currently documented for this edition/mode. |
| Enterprise Native Core | Familiar Core API on Enterprise. Realtime/offline remain Core features, but indexes are optional and billing is byte-unit based. |
| Enterprise Pipeline | Advanced stage/expression query interface. Query Explain exposes an execution tree, scanned records/bytes, memory and read units. |
| MongoDB compatibility |
Enterprise MongoDB-protocol interface. It has separate
explain/Query Insights behavior and must
not be diagnosed with Native assumptions.
|
| Query Explain | Per-query planner/execution evidence. Planning-only and execute/analyze modes have different cost and side-effect implications. |
| Query Insights | Aggregated normalized-query statistics over time, useful for frequency/latency/read-load prioritization rather than one-off diagnosis. |
| Key Visualizer | Standard Native key-range/index-range heatmaps for hotspot diagnosis. It is not a generic trace viewer or a capacity benchmark. |
| Correlation ID | Application-generated opaque identifier propagated through logs/timers so one user action can be linked across client, backend and Firestore evidence without logging tokens or PII. |
2. Four clocks answer four different questions
| Clock/evidence | What it measures | Typical mistake |
|---|---|---|
| Client/backend app timer | Full caller-observed latency around the SDK call | Blaming Firestore for DNS, mobile radio, proxy, serialization or application queueing. |
Cloud Monitoring api/request_latencies
|
Completed non-streaming request latency from Firestore frontend through response production | Treating it as end-user RTT. |
Audit Log processing_duration |
Database-side processing time for audited data reads/writes | Assuming Data Access logs are always enabled or free. |
| Query Explain execution duration/tree | One explained query’s backend plan/work | Generalizing one run into fleet-wide frequency or p99 behavior. |
If client p99 rises while backend p99 stays stable, investigate the path before Firestore: network, connection churn, client CPU, proxy, retries, serialization, or application queueing. If both rise, continue into query/index/hotspot evidence.
3. Cloud Monitoring metrics: build a database-level signal map
| Signal family | Current metric examples | Question answered |
|---|---|---|
| Requests/latency |
api/request_count,
api/request_latencies
|
Which API methods/status codes are slow or failing? |
| Documents |
document/read_ops_count,
write_ops_count,
delete_ops_count
|
Is workload shape changing? |
| Realtime |
network/active_connections,
network/snapshot_listeners
|
Did listener/connection count change? |
| Query work |
query_stat/per_query/scanned_documents_counts, scanned_index_entries_counts,
result_counts
|
Is scan amplification growing? Realtime queries are excluded from these per-query metrics. |
| Rules |
rules/evaluation_count with
ALLOW/DENY/ERROR
|
Did client Security Rules behavior change? |
| Storage |
storage/data_and_index_storage_bytes,
backup/PITR bytes
|
Is data/index/recovery footprint changing? |
| Enterprise billing units |
api/billable_read_units,
billable_write_units, realtime units
|
What byte-unit work is generated on Enterprise? |
The Firestore usage dashboard can omit or collapse work that billing counts: index-entry reads, zero-result minimum reads, imports/exports, some no-op/collapsed writes, and other documented cases. Use the billing report for financial truth; use Monitoring for operational shape.
4. Realtime listeners: connection count is not listener count
A mobile/web SDK normally maintains one active connection that
can multiplex several snapshot listeners. Server client
libraries create a connection per snapshot listener. Therefore a
“connections doubled” alert and a “listeners doubled” alert
imply different things. A screen accidentally mounting ten
listeners per user can inflate
network/snapshot_listeners without multiplying
mobile connections tenfold.
const active = new Map();export function trackedListen(key, subscribe) { if (active.has(key)) throw new Error(`duplicate-listener:${key}`); const unsubscribe = subscribe(); active.set(key, { startedAt: Date.now(), unsubscribe }); return () => { active.get(key)?.unsubscribe(); active.delete(key); };}export function listenerGauge() { return active.size; }
Local listener registries provide immediate application evidence; Cloud Monitoring gives managed aggregate evidence later. Use both rather than treating either as the whole truth.
5. Mandatory local lab: timing, percentiles, error codes, and correlation IDs
import crypto from "node:crypto";import { performance } from "node:perf_hooks";const samples = [];export async function observed(name, fn) { const correlationId = crypto.randomUUID(); const start = performance.now(); try { const value = await fn(correlationId); samples.push({name, correlationId, ok:true, ms:performance.now()-start}); return value; } catch (error) { samples.push({name, correlationId, ok:false, code:error.code ?? "UNKNOWN", ms:performance.now()-start}); throw error; }}function pct(values, p) { const a=[...values].sort((x,y)=>x-y); const i=Math.min(a.length-1, Math.ceil((p/100)*a.length)-1); return a[Math.max(0,i)] ?? 0;}export function report() { const ok=samples.filter(x=>x.ok).map(x=>x.ms); return {count:samples.length, errors:samples.length-ok.length, p50_ms:pct(ok,50), p95_ms:pct(ok,95), p99_ms:pct(ok,99)};}
-
Seed a bounded AtlasMart fixture in the emulator: 200
catalogItems, 40 sellers, deterministic prices/categories. -
Run 100 bounded queries through
observed(); save raw per-request JSON and a percentile summary. - Inject ten deterministic denied client reads through Rules and record the error code separately from latency.
- Mount and unmount the same listener 100 times; assert the local registry returns to zero.
- Do not infer production p99 or capacity from these numbers. They prove instrumentation and regression logic only.
6. Controlled failure: averages hide the tail
Create 95 synthetic samples at 10 ms and five at 500 ms. The average looks modest compared with the user-visible tail. A production SLO expressed only as mean latency can therefore remain “healthy” while a small but important cohort is badly affected.
const x=[...Array(95).fill(10), ...Array(5).fill(500)];const mean=x.reduce((a,b)=>a+b,0)/x.length;const sorted=[...x].sort((a,b)=>a-b);const p95=sorted[Math.ceil(.95*x.length)-1];const p99=sorted[Math.ceil(.99*x.length)-1];console.log({mean,p95,p99});// Expected shape: mean is far below p99; investigate tail cohorts.
7. Optional managed verification: Monitoring and audit evidence
In a disposable project, filter
api/request_latencies by database, API method and
response code; compare p50/p95/p99 with application timing over
the same interval. Data Access audit logs can provide
processing_duration for audited reads/writes, but
enabling and retaining those logs has cost/privacy implications.
Never enable verbose payload logging as an incident shortcut.
Firestore metrics are typically sampled every 60 seconds and can take up to roughly four minutes to appear. Query Insights is delayed much longer. For a live incident, pair aggregate managed signals with application-local correlation/timing evidence.
Production judgment: alert on decisions, not curiosity
Useful alerts have an owner and a decision: sustained backend p99 degradation, rising error-code rate, unexpected listener growth, Rules denials after a deploy, or storage/index growth beyond a forecast envelope. Avoid alerting on every metric independently. Quota increases are not a first response to abuse, retry storms, accidental listeners, or an inefficient query shape.
Knowledge check
-
Why can application p99 be higher than Firestore
api/request_latencies? - Why is the usage dashboard not financial truth?
- What is the relationship between mobile connections and snapshot listeners?
- Why is p99 often more actionable than a mean?
- What does emulator latency prove?
Review the answers
1. App timing includes network/client/application overhead omitted by backend metrics.
2. Documented billing work can be omitted/collapsed in usage dashboards; billing reports win.
3. Mobile/web can multiplex listeners over one connection; server listeners use separate connections.
4. It exposes tail cohorts hidden by averages.
5. Instrumentation/correctness under local conditions—not production capacity, regional latency, autoscaling, or billing.
Summary and next step
This lesson established the working contract for Cloud Monitoring Metrics for Reads/Writes/Deletes, Latency, Errors, Quotas, Storage, and Listener Behavior. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Query Explain to Inspect Execution, Index Use, Scanned/Returned Work, and Optimization Evidence.
Authoritative references
- Firebase: Monitor Cloud Firestore activity — usage dashboards, Monitoring metrics, Security Rules metrics, listener/connections, and dashboard-versus-billing caveats.
-
Google Cloud Monitoring metrics reference
— current
firestore.googleapis.com/*metric types, labels, sampling and launch stages. - Firebase: Query Explain for Standard/Core queries — planner vs analyze behavior, scan statistics, IAM authentication, and billing semantics.
- Firebase: Enterprise Native Pipeline Query Explain — explain/analyze/stats modes and execution-tree evidence.
- Firebase: Query Insights — normalized queries, latency/read-load statistics, retention, delay and IAM.
- Firebase: Key Visualizer overview — Standard Native hotspot heatmaps, eligibility, metrics, limits, and data retention.
-
Firebase: Cloud Firestore audit logging
— Data Access/Admin Activity audit evidence and
processing_duration. - Firebase: Firestore editions overview — current observability availability and Standard/Enterprise distinctions.