Chapter 17 · Vector Search: Embeddings, KNN Indexes, Distance Measures, Filters, and Retrieval Design

Build and Evaluate a Semantic Retrieval Feature with Recall / Quality, Latency, Security, and Cost Measurements

Evaluate AtlasMart semantic retrieval against a judged ground truth using recall@k, latency distributions, filter/security tests, Standard billing estimates, model-version drift checks, and an explicit decision boundary for dedicated search systems.

Advanced · 170–220 minutesrecall@k · latency · security · costFirebase JS 12.19.0 · Admin 14.4.0 · @google-cloud/firestore 9.1.0CLI 15.30.0 lab pin · Standard Native canonical lab · Enterprise differences explicitLast reviewed: 17 September 2026

1. AtlasMart problem: “the demos look good” is not a retrieval acceptance test

Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Before shipping semantic search, AtlasMart needs evidence across quality, latency, authorization, model freshness and cost. A search feature can return plausible demos while failing rare categories, crossing tenant boundaries, scanning too much of the vector index, or degrading after a model/backfill change.

Chapter 17 reproducibility baseline · reviewed 17 September 2026

AtlasMart continues the same canonical lab environment used in Chapters 01–16: project ID demo-atlasmart-firestore, Standard edition / Native mode / (default) database, Firestore emulator 127.0.0.1:8080, Authentication emulator 127.0.0.1:9099, Emulator UI 127.0.0.1:4000, Firebase JavaScript SDK 12.19.0, Firebase Admin Node.js SDK 14.4.0 carrying @google-cloud/firestore 9.1.0, @firebase/rules-unit-testing 5.0.2, Firebase CLI 15.30.0, and Node.js 22+. Mandatory work is local/no-cost. Managed vector indexes, managed KNN billing, production latency, Query Explain, Enterprise Pipeline execution, and external embedding services are optional bounded exercises only.

Emulator evidence boundary

The emulator is suitable for AtlasMart documents, Security Rules, deterministic fixture management, and local application tests, but it does not enforce production composite/vector-index requirements and cannot certify managed vector-query latency, index-build status, billing, availability, or Enterprise execution. Therefore the mandatory lab stores deterministic vectors as ordinary numeric arrays and runs an exact local KNN oracle; the lesson also shows the real managed gcloud/@google-cloud/firestore commands, clearly labeled as optional production/demo-project steps.

Learning outcomes

01

Build a judged query set and compute recall@k instead of relying on screenshots.

02

Measure p50/p95/p99 local execution without mislabeling emulator/local numbers as production latency.

03

Test tenant/category filters and backend authorization alongside relevance.

04

Estimate Standard vector-query read charges from measured returned documents and vector-index entries scanned.

05

Define release/rollback gates and when to move to a dedicated search/vector system.

2. Evaluation dataset: write the expected neighbors down

judged-queries.json
[  {"id":"q-camera","query":[0.90,0.06,0.02,0.02,0.24,0.05,0.01,0.13],"relevant":["p-1001","p-1005"],"tenantId":null},  {"id":"q-camera-tenant-a","query":[0.90,0.06,0.02,0.02,0.24,0.05,0.01,0.13],"relevant":["p-1001"],"tenantId":"tenant-a"},  {"id":"q-power","query":[0.02,0.20,0.09,0.16,0.01,0.94,0.05,0.01],"relevant":["p-1006"],"tenantId":"tenant-a"}]

3. Quality harness: recall@k plus “no unauthorized result”

evaluate-quality.mjs
import { performance } from "node:perf_hooks";import rows from "./atlasmart-vectors.json" with { type:"json" };import judged from "./judged-queries.json" with { type:"json" };import { rank } from "./vector-math.mjs";function recallAtK(ids, truth, k) {  const t = new Set(truth);  return ids.slice(0,k).filter(id => t.has(id)).length / Math.max(1,t.size);}const latencies=[]; const reports=[];for (let warm=0; warm<25; warm++) for (const q of judged) rank(rows,q.query);for (let round=0; round<200; round++) {  for (const q of judged) {    const pool = q.tenantId ? rows.filter(x => x.tenantId === q.tenantId) : rows;    const t0=performance.now(); const ranked=rank(pool,q.query); const ms=performance.now()-t0;    latencies.push(ms);    const ids=ranked.map(x=>x.id);    if (q.tenantId && ranked.some(x=>x.tenantId!==q.tenantId)) throw new Error("tenant leak");    reports.push({id:q.id, recall2:recallAtK(ids,q.relevant,2)});  }}latencies.sort((a,b)=>a-b);const p = x => latencies[Math.min(latencies.length-1, Math.floor((x/100)*latencies.length))];console.log({p50:p(50),p95:p(95),p99:p(99), sampleCount:latencies.length});console.log(reports.slice(0,judged.length));console.log("LOCAL CPU timing only: not Firestore service latency");

4. Latency: report what you actually measured

The harness measures local exact-distance CPU time over six vectors. It teaches percentile calculation and reproducibility, not Firestore performance. For managed production evidence, record region, edition/mode, SDK/runtime, vector dimension, K, filters, index configuration, dataset cardinality, warmup, concurrency, Query Explain/index entries scanned where available, and billing dimensions. Never substitute emulator/local numbers for p95/p99 service latency.

5. Standard vector-query billing model

Current Standard pricing charges document reads for returned documents plus one read for each batch of up to 100 KNN vector-index entries read. The service’s actual scanned-entry count depends on the managed query/index; the local emulator cannot produce it. The following calculator turns measured managed evidence into billable read units without inventing the scan count.

standard-vector-cost.mjs
export function standardVectorReadCharges({ returnedDocuments, vectorIndexEntriesRead }) {  if (returnedDocuments < 0 || vectorIndexEntriesRead < 0) throw new Error("non-negative inputs required");  return returnedDocuments + Math.ceil(vectorIndexEntriesRead / 100);}for (const sample of [  { returnedDocuments:5, vectorIndexEntriesRead:0 },  { returnedDocuments:5, vectorIndexEntriesRead:100 },  { returnedDocuments:5, vectorIndexEntriesRead:1550 }]) console.log({ ...sample, billedReads: standardVectorReadCharges(sample) });// Pricing model only. Supply real managed scan evidence before using for forecasting.

6. Filter selectivity and cost/security evidence

Track candidate population before/after metadata filters. A tenant filter can reduce candidates and is required for isolation, but it is not enough to say “filters make vector search cheaper.” Measure the actual managed vector-index entries scanned. Likewise, a server service account may have database-wide IAM; application authorization still must prevent a user from asking the backend to substitute another tenant ID.

authorization-contract.mjs
export function authorizeSearch({ caller, requestedTenantId }) {  if (!caller?.uid) throw Object.assign(new Error("unauthenticated"), { status:401 });  if (caller.tenantId !== requestedTenantId) throw Object.assign(new Error("forbidden tenant"), { status:403 });  return { tenantId: caller.tenantId };}// Build the Firestore filter from the returned trusted object, not raw request JSON.

7. Model/version drift gate

Quality reports must be segmented by embeddingModel/embeddingVersion. If a v2 backfill is incomplete, either restrict v2 queries to v2-complete documents with a deliberate coverage tradeoff or keep v1 primary. Mixing vector versions into one index without proof they share the same space invalidates evaluation.

Release gate Evidence Rollback trigger
Quality Recall@k / judged-query regression report Material regression on critical query sets
Security Cross-tenant/visibility negative tests Any unauthorized result
Freshness Backfill coverage + source-hash reconciliation Coverage below declared launch threshold
Latency Managed p50/p95/p99 under declared load SLO breach after controlled rollout
Cost Returned docs + vector-index entries scanned + request volume Budget envelope exceeded

8. Failure injection matrix

Injection Expected evidence Repair
Query vector has seven values Preflight dimension error Reject before database call
Caller requests another tenant 403 before vector query Derive tenant from verified identity
v2 vector missing on one product Reconciliation lists product ID Replay idempotent backfill
Model changed but index still v1 Version gate blocks cutover Build/evaluate v2 index separately
No relevant neighbor under threshold Empty/fallback UX Do not force low-quality K results

9. When Firestore remains a good fit—and when to separate search

Firestore is attractive when vector retrieval is close to the transactional documents, query shapes fit supported KNN/filter capabilities, managed index/billing are acceptable, and one operational datastore simplifies consistency. A dedicated search/vector system can be warranted when retrieval needs exceed current Firestore features—for example advanced hybrid ranking, very specialized ANN controls, search-specific observability/relevance tooling, or independent scaling/cost boundaries. That decision should come from measured requirements, not fashion.

10. Edition/mode/client matrix

Surface Vector capability Important boundary
Standard Native Core Managed nearest-neighbor search through supported server libraries; flat vector indexes; metadata prefilters Maximum 2,048 dimensions; maximum 1,000 returned documents; no realtime snapshot listeners for vector search; current documented client-library support is Python, Node.js, Go, Java
Enterprise Native Core Native Core semantics in Enterprise context Enterprise indexing/billing differ from Standard; do not copy Standard billing equations blindly
Enterprise Native Pipeline findNearest transformation stage with broader Pipeline client surface Pipeline API and supported distance options are distinct; verify current stage and SDK behavior rather than treating it as Standard Core
MongoDB compatibility Separate MongoDB-compatible query/index surface Do not assume Native findNearest, vector-index commands, Rules, or pricing semantics map one-for-one
Emulator Useful for fixture/security/application plumbing; recent emulator versions can store vector values Composite/vector-index enforcement, production latency, billing, index build state, and all service limits are not production evidence

11. Final Chapter 17 verification checklist

  • Five lesson labs use the same deterministic six-product/eight-dimensional fixture.
  • Every vector carries model/version/dimension lineage.
  • Local quality/latency are explicitly labeled local evidence.
  • Managed Standard cost uses returned documents plus measured vector-index entries scanned.
  • Tenant authorization is tested independently from relevance.
  • Vector generation is external to Firestore and external providers are optional/governed.
  • Standard Core, Enterprise Native Pipeline and MongoDB compatibility are not conflated.

Production judgment and bridge to Chapter 18

AtlasMart now has a complete vector-retrieval engineering loop: versioned embeddings, managed index contract, distance/threshold/filter semantics, replayable generation, judged quality, authorization tests, latency methodology and cost evidence. Chapter 18 moves to Firestore Enterprise Native mode and compares familiar Core operations with the separate Pipeline execution model—including where findNearest participates in that broader query engine.

Knowledge check

  1. What two inputs are required for a credible Standard vector cost estimate?
  2. Can local p99 from the six-product oracle be called Firestore p99?
  3. What is the strongest security acceptance criterion for tenant vector search?
  4. Why segment evaluation by embedding version?
  5. When should AtlasMart consider a dedicated search/vector system?
Review the answers

1. Returned document count and measured vector-index entries read, plus request volume for forecasting.

2. No. It measures local CPU only.

3. Negative tests proving no cross-tenant result can be produced, with tenant identity derived from trusted authentication context.

4. Different model spaces/versions can have different quality and are not safe to mix without proof.

5. When measured product requirements exceed Firestore’s retrieval/ranking/operational boundaries enough to justify separate complexity and cost.

Summary and next step

This lesson established the working contract for Build and Evaluate a Semantic Retrieval Feature with Recall/Quality, Latency, Security, and Cost Measurements. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Enterprise Edition Architecture/Query Engine and How Native Core Semantics Differ from Standard.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.