Chapter 17 · Vector Search: Embeddings, KNN Indexes, Distance Measures, Filters, and Retrieval Design
Build and Evaluate a Semantic Retrieval Feature with Recall / Quality, Latency, Security, and Cost Measurements
Evaluate AtlasMart semantic retrieval against a judged ground truth using recall@k, latency distributions, filter/security tests, Standard billing estimates, model-version drift checks, and an explicit decision boundary for dedicated search systems.
1. AtlasMart problem: “the demos look good” is not a retrieval acceptance test
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
Before shipping semantic search, AtlasMart needs evidence across quality, latency, authorization, model freshness and cost. A search feature can return plausible demos while failing rare categories, crossing tenant boundaries, scanning too much of the vector index, or degrading after a model/backfill change.
AtlasMart continues the same canonical lab environment used in
Chapters 01–16: project ID
demo-atlasmart-firestore, Standard edition /
Native mode / (default) database, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase JavaScript SDK
12.19.0, Firebase Admin Node.js SDK
14.4.0 carrying
@google-cloud/firestore 9.1.0,
@firebase/rules-unit-testing 5.0.2,
Firebase CLI 15.30.0, and Node.js 22+. Mandatory
work is local/no-cost. Managed vector indexes, managed KNN
billing, production latency, Query Explain, Enterprise
Pipeline execution, and external embedding services are
optional bounded exercises only.
The emulator is suitable for AtlasMart documents, Security
Rules, deterministic fixture management, and local application
tests, but it does not enforce production
composite/vector-index requirements and cannot certify managed
vector-query latency, index-build status, billing,
availability, or Enterprise execution. Therefore the mandatory
lab stores deterministic vectors as ordinary numeric arrays
and runs an exact local KNN oracle; the lesson also shows the
real managed gcloud/@google-cloud/firestore
commands, clearly labeled as optional production/demo-project
steps.
Learning outcomes
Build a judged query set and compute recall@k instead of relying on screenshots.
Measure p50/p95/p99 local execution without mislabeling emulator/local numbers as production latency.
Test tenant/category filters and backend authorization alongside relevance.
Estimate Standard vector-query read charges from measured returned documents and vector-index entries scanned.
Define release/rollback gates and when to move to a dedicated search/vector system.
2. Evaluation dataset: write the expected neighbors down
[ {"id":"q-camera","query":[0.90,0.06,0.02,0.02,0.24,0.05,0.01,0.13],"relevant":["p-1001","p-1005"],"tenantId":null}, {"id":"q-camera-tenant-a","query":[0.90,0.06,0.02,0.02,0.24,0.05,0.01,0.13],"relevant":["p-1001"],"tenantId":"tenant-a"}, {"id":"q-power","query":[0.02,0.20,0.09,0.16,0.01,0.94,0.05,0.01],"relevant":["p-1006"],"tenantId":"tenant-a"}]
3. Quality harness: recall@k plus “no unauthorized result”
import { performance } from "node:perf_hooks";import rows from "./atlasmart-vectors.json" with { type:"json" };import judged from "./judged-queries.json" with { type:"json" };import { rank } from "./vector-math.mjs";function recallAtK(ids, truth, k) { const t = new Set(truth); return ids.slice(0,k).filter(id => t.has(id)).length / Math.max(1,t.size);}const latencies=[]; const reports=[];for (let warm=0; warm<25; warm++) for (const q of judged) rank(rows,q.query);for (let round=0; round<200; round++) { for (const q of judged) { const pool = q.tenantId ? rows.filter(x => x.tenantId === q.tenantId) : rows; const t0=performance.now(); const ranked=rank(pool,q.query); const ms=performance.now()-t0; latencies.push(ms); const ids=ranked.map(x=>x.id); if (q.tenantId && ranked.some(x=>x.tenantId!==q.tenantId)) throw new Error("tenant leak"); reports.push({id:q.id, recall2:recallAtK(ids,q.relevant,2)}); }}latencies.sort((a,b)=>a-b);const p = x => latencies[Math.min(latencies.length-1, Math.floor((x/100)*latencies.length))];console.log({p50:p(50),p95:p(95),p99:p(99), sampleCount:latencies.length});console.log(reports.slice(0,judged.length));console.log("LOCAL CPU timing only: not Firestore service latency");
4. Latency: report what you actually measured
The harness measures local exact-distance CPU time over six vectors. It teaches percentile calculation and reproducibility, not Firestore performance. For managed production evidence, record region, edition/mode, SDK/runtime, vector dimension, K, filters, index configuration, dataset cardinality, warmup, concurrency, Query Explain/index entries scanned where available, and billing dimensions. Never substitute emulator/local numbers for p95/p99 service latency.
5. Standard vector-query billing model
Current Standard pricing charges document reads for returned documents plus one read for each batch of up to 100 KNN vector-index entries read. The service’s actual scanned-entry count depends on the managed query/index; the local emulator cannot produce it. The following calculator turns measured managed evidence into billable read units without inventing the scan count.
export function standardVectorReadCharges({ returnedDocuments, vectorIndexEntriesRead }) { if (returnedDocuments < 0 || vectorIndexEntriesRead < 0) throw new Error("non-negative inputs required"); return returnedDocuments + Math.ceil(vectorIndexEntriesRead / 100);}for (const sample of [ { returnedDocuments:5, vectorIndexEntriesRead:0 }, { returnedDocuments:5, vectorIndexEntriesRead:100 }, { returnedDocuments:5, vectorIndexEntriesRead:1550 }]) console.log({ ...sample, billedReads: standardVectorReadCharges(sample) });// Pricing model only. Supply real managed scan evidence before using for forecasting.
6. Filter selectivity and cost/security evidence
Track candidate population before/after metadata filters. A tenant filter can reduce candidates and is required for isolation, but it is not enough to say “filters make vector search cheaper.” Measure the actual managed vector-index entries scanned. Likewise, a server service account may have database-wide IAM; application authorization still must prevent a user from asking the backend to substitute another tenant ID.
export function authorizeSearch({ caller, requestedTenantId }) { if (!caller?.uid) throw Object.assign(new Error("unauthenticated"), { status:401 }); if (caller.tenantId !== requestedTenantId) throw Object.assign(new Error("forbidden tenant"), { status:403 }); return { tenantId: caller.tenantId };}// Build the Firestore filter from the returned trusted object, not raw request JSON.
7. Model/version drift gate
Quality reports must be segmented by
embeddingModel/embeddingVersion. If a
v2 backfill is incomplete, either restrict v2 queries to
v2-complete documents with a deliberate coverage tradeoff or
keep v1 primary. Mixing vector versions into one index without
proof they share the same space invalidates evaluation.
| Release gate | Evidence | Rollback trigger |
|---|---|---|
| Quality | Recall@k / judged-query regression report | Material regression on critical query sets |
| Security | Cross-tenant/visibility negative tests | Any unauthorized result |
| Freshness | Backfill coverage + source-hash reconciliation | Coverage below declared launch threshold |
| Latency | Managed p50/p95/p99 under declared load | SLO breach after controlled rollout |
| Cost | Returned docs + vector-index entries scanned + request volume | Budget envelope exceeded |
8. Failure injection matrix
| Injection | Expected evidence | Repair |
|---|---|---|
| Query vector has seven values | Preflight dimension error | Reject before database call |
| Caller requests another tenant | 403 before vector query | Derive tenant from verified identity |
| v2 vector missing on one product | Reconciliation lists product ID | Replay idempotent backfill |
| Model changed but index still v1 | Version gate blocks cutover | Build/evaluate v2 index separately |
| No relevant neighbor under threshold | Empty/fallback UX | Do not force low-quality K results |
9. When Firestore remains a good fit—and when to separate search
Firestore is attractive when vector retrieval is close to the transactional documents, query shapes fit supported KNN/filter capabilities, managed index/billing are acceptable, and one operational datastore simplifies consistency. A dedicated search/vector system can be warranted when retrieval needs exceed current Firestore features—for example advanced hybrid ranking, very specialized ANN controls, search-specific observability/relevance tooling, or independent scaling/cost boundaries. That decision should come from measured requirements, not fashion.
10. Edition/mode/client matrix
| Surface | Vector capability | Important boundary |
|---|---|---|
| Standard Native Core | Managed nearest-neighbor search through supported server libraries; flat vector indexes; metadata prefilters | Maximum 2,048 dimensions; maximum 1,000 returned documents; no realtime snapshot listeners for vector search; current documented client-library support is Python, Node.js, Go, Java |
| Enterprise Native Core | Native Core semantics in Enterprise context | Enterprise indexing/billing differ from Standard; do not copy Standard billing equations blindly |
| Enterprise Native Pipeline |
findNearest transformation stage with broader
Pipeline client surface
|
Pipeline API and supported distance options are distinct; verify current stage and SDK behavior rather than treating it as Standard Core |
| MongoDB compatibility | Separate MongoDB-compatible query/index surface |
Do not assume Native findNearest,
vector-index commands, Rules, or pricing semantics map
one-for-one
|
| Emulator | Useful for fixture/security/application plumbing; recent emulator versions can store vector values | Composite/vector-index enforcement, production latency, billing, index build state, and all service limits are not production evidence |
11. Final Chapter 17 verification checklist
- Five lesson labs use the same deterministic six-product/eight-dimensional fixture.
- Every vector carries model/version/dimension lineage.
- Local quality/latency are explicitly labeled local evidence.
- Managed Standard cost uses returned documents plus measured vector-index entries scanned.
- Tenant authorization is tested independently from relevance.
- Vector generation is external to Firestore and external providers are optional/governed.
- Standard Core, Enterprise Native Pipeline and MongoDB compatibility are not conflated.
Production judgment and bridge to Chapter 18
AtlasMart now has a complete vector-retrieval engineering loop:
versioned embeddings, managed index contract,
distance/threshold/filter semantics, replayable generation,
judged quality, authorization tests, latency methodology and
cost evidence. Chapter 18 moves to Firestore Enterprise Native
mode and compares familiar Core operations with the separate
Pipeline execution model—including where
findNearest participates in that broader query
engine.
Knowledge check
- What two inputs are required for a credible Standard vector cost estimate?
- Can local p99 from the six-product oracle be called Firestore p99?
- What is the strongest security acceptance criterion for tenant vector search?
- Why segment evaluation by embedding version?
- When should AtlasMart consider a dedicated search/vector system?
Review the answers
1. Returned document count and measured vector-index entries read, plus request volume for forecasting.
2. No. It measures local CPU only.
3. Negative tests proving no cross-tenant result can be produced, with tenant identity derived from trusted authentication context.
4. Different model spaces/versions can have different quality and are not safe to mix without proof.
5. When measured product requirements exceed Firestore’s retrieval/ranking/operational boundaries enough to justify separate complexity and cost.
Summary and next step
This lesson established the working contract for Build and Evaluate a Semantic Retrieval Feature with Recall/Quality, Latency, Security, and Cost Measurements. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Enterprise Edition Architecture/Query Engine and How Native Core Semantics Differ from Standard.
Authoritative references
- Firebase · Search with vector embeddings
- Google Cloud · Firestore Standard pricing
- Firebase · Connect to the Firestore emulator and understand differences from production
- Google Cloud · Enterprise Native Pipeline findNearest stage
- Firebase · Enterprise Native supported data types
- Google Cloud · Vertex AI embedding sample