Chapter 24 · Observability: Query Explain, Query Insights, Key Visualizer, Metrics, Logs, and Troubleshooting

Key Visualizer / Hotspot Diagnostics Where Available, Application Tracing, Logs, and Correlation IDs

Correlate Standard Native Key Visualizer heatmaps with client/backend timing and privacy-safe structured logs so sequential-key, index-key, sudden-ramp, network, and application bottlenecks can be distinguished.

Advanced · 170–230 minutesKey Visualizer · tracing · logs · correlation IDs · hotspotsNode 22+ · Firebase CLI course baseline 15.30.0 · JS SDK 12.19.0 · Admin SDK 14.4.0Mandatory evidence lab local/no-cost · Monitoring/Insights/Key Visualizer/Audit optional managed verificationLast reviewed: 17 September 2026

1. AtlasMart problem: the hotspot is visible only when evidence is correlated

AtlasMart receives a flash-sale burst. Client p99 rises, Firestore backend write latency rises, and a few request IDs dominate error logs. The data team suspects sequential event IDs, while the platform team suspects an external API. Key Visualizer, application correlation IDs and structured logs answer different parts: where key-range/index heat concentrates, which user action generated the requests, and whether time was spent before or inside Firestore.

Learning outcomes
  • Interpret Standard Native Key Visualizer heatmap patterns and eligibility boundaries.
  • Distinguish document-key hotspots from index-key hotspots.
  • Propagate correlation IDs without logging credentials or sensitive payloads.
  • Build structured application logs that support latency/error investigations.
  • Separate storage-layer hotspot evidence from network/application latency.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 24 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps project ID demo-atlasmart-firestore, Standard-edition Native mode, database (default), Node.js 22+, Firebase CLI course baseline 15.30.0, Firebase JavaScript SDK 12.19.0, Firebase Admin Node SDK 14.4.0 (bundling @google-cloud/firestore 9.1.0), Firestore emulator 127.0.0.1:8080, Auth emulator 127.0.0.1:9099, and Emulator UI 127.0.0.1:4000. Firebase CLI 15.30.1 is the current patch release at review time; no Chapter 24 lab depends on that patch, so the course stays pinned to 15.30.0 for continuity. Mandatory exercises are local/no-cost. Cloud Monitoring, Query Insights, Key Visualizer, Cloud Audit Logs, production Query Explain billing evidence, and real hotspot/capacity measurements require a real Google Cloud/Firebase project and are optional bounded verification steps. Emulator latency is never presented as production capacity evidence. Key Visualizer is currently documented for Firestore Standard edition in Native mode. It is not listed as the Enterprise observability tool; use Enterprise Query Explain/Insights/Monitoring instead.

Term Operational meaning in Chapter 24
Client timing End-to-end elapsed time observed by the browser/mobile/backend caller. It includes client scheduling and network time that Firestore backend metrics do not.
Backend latency Firestore service processing time. Cloud Monitoring api/request_latencies excludes client-to-service round-trip time.
Standard Native Standard edition Core operations. Queries require indexes; Key Visualizer is currently documented for this edition/mode.
Enterprise Native Core Familiar Core API on Enterprise. Realtime/offline remain Core features, but indexes are optional and billing is byte-unit based.
Enterprise Pipeline Advanced stage/expression query interface. Query Explain exposes an execution tree, scanned records/bytes, memory and read units.
MongoDB compatibility Enterprise MongoDB-protocol interface. It has separate explain/Query Insights behavior and must not be diagnosed with Native assumptions.
Query Explain Per-query planner/execution evidence. Planning-only and execute/analyze modes have different cost and side-effect implications.
Query Insights Aggregated normalized-query statistics over time, useful for frequency/latency/read-load prioritization rather than one-off diagnosis.
Key Visualizer Standard Native key-range/index-range heatmaps for hotspot diagnosis. It is not a generic trace viewer or a capacity benchmark.
Correlation ID Application-generated opaque identifier propagated through logs/timers so one user action can be linked across client, backend and Firestore evidence without logging tokens or PII.

2. Key Visualizer: what it can and cannot show

A Key Visualizer scan covers a two-hour period divided into 10-second segments. Current scan eligibility requires traffic exceeding 3,000 document operations in any minute of that period. Scan data is retained for 14 days. Heatmaps bucket contiguous document or index keys; they do not show every individual request and are not available simply because a database contains lots of data.

Heatmap pattern Likely mechanism Corroborating evidence
Bright diagonal document-key band Sequential increasing/decreasing document IDs Write/lookup/query latency rises on same range; inspect ID generation.
Bright diagonal index-key band Indexed monotonically increasing field such as timestamp Index Write Ops/s + index config; consider exemption if field is not queried.
Sudden vertical/bright change Traffic ramp faster than service can adapt Request rate + latency metrics around the same time.
Single persistent hot range Hot document/small key range Contention/ABORTED/deadline evidence; data model concentration.
Broad latency band across ranges May be network/service/application-wide rather than one key hotspot Client RTT, backend latency, status dashboard, dependency traces.

3. Key Visualizer metrics are storage-layer evidence

Document-key scans include Ops/s, write/lookup/query rates, average latency and tail latency metrics. Index-key scans expose Index Write Ops/s. Because these are storage-layer measurements, they can be lower than total API-call latency. A clean heatmap does not prove the client network is healthy; an ugly heatmap does not identify the original product action by itself.

4. Correlation IDs: connect product action → backend → Firestore

correlation.mjs
import crypto from "node:crypto";export function makeContext(incoming={}) {  return {    correlationId: incoming.correlationId || crypto.randomUUID(),    operation: incoming.operation || "unknown",    tenantHash: incoming.tenantId ? crypto.createHash("sha256").update(incoming.tenantId).digest("hex").slice(0,12) : undefined  };}export function safeLog(level, ctx, fields={}) {  const forbidden=["authorization","idToken","refreshToken","email","phone","address"];  for (const k of forbidden) if (k in fields) throw new Error(`forbidden-log-field:${k}`);  console.log(JSON.stringify({severity:level, ...ctx, ...fields, ts:new Date().toISOString()}));}

Hashing a tenant identifier is not automatically sufficient anonymization; the point is to avoid raw business identifiers unless policy explicitly allows them. Never log Firebase ID tokens, refresh tokens, service-account credentials, App Check tokens or full user payloads for convenience.

5. Instrument one AtlasMart request path

seller-low-stock-handler.mjs
import { performance } from "node:perf_hooks";import { makeContext, safeLog } from "./correlation.mjs";export async function lowStock(req, db) {  const ctx=makeContext({correlationId:req.headers["x-correlation-id"], operation:"seller-low-stock", tenantId:req.tenantId});  const t0=performance.now();  safeLog("INFO",ctx,{phase:"start"});  try {    const snap=await db.collection("catalogItems")      .where("sellerId","==",req.sellerId)      .where("stock","<=",5)      .limit(50).get();    safeLog("INFO",ctx,{phase:"firestore-done",ms:performance.now()-t0,results:snap.size});    return snap.docs.map(d=>({id:d.id,...d.data()}));  } catch (e) {    safeLog("ERROR",ctx,{phase:"firestore-error",ms:performance.now()-t0,code:e.code ?? "UNKNOWN"});    throw e;  }}

6. Mandatory local hotspot simulation: deterministic, not performance theater

The emulator cannot reproduce managed key-range splitting or Key Visualizer, so the mandatory exercise simulates key-distribution evidence rather than claiming a real hotspot benchmark.

key-distribution.mjs
const sequential=Array.from({length:1000},(_,i)=>`event-${String(i).padStart(8,"0")}`);const scattered=Array.from({length:1000},(_,i)=>`s${(i*7919)%9973}-${i}`);function prefixHistogram(ids,n=2){  const m=new Map(); for(const id of ids){const p=id.slice(0,n);m.set(p,(m.get(p)||0)+1);} return [...m.entries()].sort((a,b)=>b[1]-a[1]);}console.log("sequential top",prefixHistogram(sequential,3).slice(0,5));console.log("scattered top",prefixHistogram(scattered,3).slice(0,5));// This illustrates concentration only; it is NOT a Firestore throughput benchmark.

Then compare this static distribution with the Chapter 15 ramp/hot-document reasoning. Optional production verification can use Key Visualizer only on an isolated workload that naturally meets scan eligibility; do not generate billable traffic solely to unlock a heatmap.

7. Tracing boundaries

OpenTelemetry/Cloud Trace or another tracing system can connect backend spans, HTTP dependencies and Firestore call wrappers, but the Firestore service’s own backend execution evidence still comes from Monitoring/Explain/Insights. Keep a correlation ID in logs and attach trace/span IDs where available. Do not claim a trace span equals Firestore internal processing time unless the instrumentation specifically measures that layer.

8. Controlled failure: verbose logging leaks the incident

Inject an authorization field into safeLog(). The logger must reject it. A common incident anti-pattern is enabling raw request logging to diagnose latency and accidentally persisting credentials or PII. The repair is field allowlisting/redaction, not “remember to delete logs later.”

Production judgment

Use Key Visualizer when Standard Native key/index hotspotting is plausible and scan eligibility exists. Use Query Explain/Insights for query work, Monitoring for fleet-level latency/errors, and app tracing/logging for request causality. A hotspot fix—random IDs, sharding, index exemption, ramp control—must preserve query requirements and correctness.

Knowledge check

  1. What current traffic threshold makes a two-hour Key Visualizer scan eligible?
  2. What does a diagonal bright band commonly suggest?
  3. Why can Key Visualizer latency be lower than user-observed latency?
  4. Should you generate production traffic solely to unlock Key Visualizer?
  5. What fields should never be casually logged?
Review the answers

1. More than 3,000 document operations in any minute of that period.

2. Sequential document/index keys and a moving hotspot.

3. It reflects storage-layer behavior, not all network/application overhead.

4. No; use it when legitimate workload qualifies.

5. Credentials/tokens/secrets and policy-prohibited PII.

Summary and next step

This lesson established the working contract for Key Visualizer/Hotspot Diagnostics Where Available, Application Tracing, Logs, and Correlation IDs. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Build a Troubleshooting Runbook for Permission Errors, Missing Indexes, Hotspots, Latency, Quota, and Cost Spikes.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.