Chapter 17 · Vector Search: Embeddings, KNN Indexes, Distance Measures, Filters, and Retrieval Design
Firestore Does Not Generate Embeddings: Vertex AI / External Generation Pipelines, Backfills, and Updates
Design an idempotent embedding-generation and backfill pipeline around Firestore, with versioned vectors, retries, governance, privacy boundaries, and optional Vertex AI integration kept separate from the mandatory local lab.
1. AtlasMart problem: Firestore can search a vector only after another system creates it
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
Product descriptions change, model versions change, and some content must never leave AtlasMart’s trust boundary. Firestore is the vector storage/retrieval layer here; embedding generation is a separate pipeline with its own authentication, quotas, privacy, retries, cost and model lifecycle.
AtlasMart continues the same canonical lab environment used in
Chapters 01–16: project ID
demo-atlasmart-firestore, Standard edition /
Native mode / (default) database, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase JavaScript SDK
12.19.0, Firebase Admin Node.js SDK
14.4.0 carrying
@google-cloud/firestore 9.1.0,
@firebase/rules-unit-testing 5.0.2,
Firebase CLI 15.30.0, and Node.js 22+. Mandatory
work is local/no-cost. Managed vector indexes, managed KNN
billing, production latency, Query Explain, Enterprise
Pipeline execution, and external embedding services are
optional bounded exercises only.
The emulator is suitable for AtlasMart documents, Security
Rules, deterministic fixture management, and local application
tests, but it does not enforce production
composite/vector-index requirements and cannot certify managed
vector-query latency, index-build status, billing,
availability, or Enterprise execution. Therefore the mandatory
lab stores deterministic vectors as ordinary numeric arrays
and runs an exact local KNN oracle; the lesson also shows the
real managed gcloud/@google-cloud/firestore
commands, clearly labeled as optional production/demo-project
steps.
Learning outcomes
Design a versioned embedding-generation state machine outside Firestore transactions.
Backfill safely with idempotency, progress markers and retryable batches.
Separate deterministic local/precomputed vectors from optional Vertex AI calls.
Govern sensitive text before it is sent to any external model service.
Cut over model versions with dual fields/indexes and measurable rollback.
2. Mandatory local generator: deterministic and free
The course does not require an external model. The mandatory
path uses the checked-in vectors from
atlasmart-vectors.json. That guarantees every
learner sees the same ranking and allows model-lifecycle
mechanics to be tested without network, billing, model drift or
rate limits.
{ "name": "atlasmart-firestore-ch17", "private": true, "type": "module", "engines": { "node": ">=22" }, "dependencies": { "firebase": "12.19.0", "firebase-admin": "14.4.0" }, "devDependencies": { "firebase-tools": "15.30.0" }, "scripts": { "emulators": "firebase emulators:start --only firestore,auth --project demo-atlasmart-firestore", "seed": "node seed-ch17.mjs", "evaluate": "node evaluate-ch17.mjs", "reset": "node reset-ch17.mjs" }}
[ {"id":"p-1001","tenantId":"tenant-a","category":"cameras","name":"Trail Camera","embeddingV1":[0.93,0.05,0.02,0.00,0.20,0.06,0.01,0.12]}, {"id":"p-1002","tenantId":"tenant-a","category":"accessories","name":"USB-C Hub","embeddingV1":[0.04,0.91,0.08,0.15,0.01,0.18,0.04,0.02]}, {"id":"p-1003","tenantId":"tenant-b","category":"sensors","name":"Temp Sensor","embeddingV1":[0.09,0.14,0.86,0.05,0.05,0.18,0.02,0.03]}, {"id":"p-1004","tenantId":"tenant-b","category":"gateways","name":"Edge Gateway","embeddingV1":[0.10,0.55,0.31,0.51,0.04,0.28,0.08,0.04]}, {"id":"p-1005","tenantId":"tenant-c","category":"cameras","name":"PoE Camera","embeddingV1":[0.86,0.09,0.02,0.06,0.30,0.07,0.01,0.14]}, {"id":"p-1006","tenantId":"tenant-a","category":"power","name":"Bench PSU","embeddingV1":[0.02,0.22,0.11,0.18,0.01,0.89,0.06,0.02]}]
3. Pipeline state is data, not log text
{ "productId": "p-1001", "targetModel": "atlasmart-local-v2", "targetVersion": 2, "targetDimension": 8, "sourceHash": "sha256:...", "status": "queued", "attempt": 0, "nextAttemptAt": null, "lastError": null, "generatedAt": null, "indexedAt": null, "correlationId": "embed-p-1001-v2"}
4. Idempotent backfill algorithm
for (const product of products) { const sourceHash = sha256(canonicalText(product)); const ref = db.doc(`embeddingJobs/${product.id}-v2`); const existing = await ref.get(); if (existing.exists && existing.get("sourceHash") === sourceHash && existing.get("status") === "complete") continue; await ref.set({ productId: product.id, targetVersion: 2, sourceHash, status: "running" }, { merge: true }); try { const vector = deterministicLocalEmbeddingV2(product); // mandatory no-cost path if (vector.length !== 8) throw new Error("dimension mismatch"); await db.doc(`vectorProducts/${product.id}`).set({ embeddingV2: vector, embeddingModelV2: "atlasmart-local-v2", embeddingVersionV2: 2, embeddingDimensionV2: 8, embeddingSourceHashV2: sourceHash }, { merge: true }); await ref.set({ status: "complete", completedAt: new Date().toISOString() }, { merge: true }); } catch (error) { await ref.set({ status: "retryable_error", lastError: String(error) }, { merge: true }); throw error; }}
5. Why generation must not run inside a retryable Firestore transaction
A transaction callback can rerun after contention. Calling an external embedding API inside that callback can create duplicate charges, repeated sensitive-data transfer and long transaction duration. Repair the design by making the transaction commit only durable intent/state; a worker then performs the external call with an idempotency key/source hash and writes the result separately.
Make the local generator throw after writing the vector but before marking the job complete. On retry, compare source hash and stored target version so the same vector write is safe. The exercise proves replayability without requiring an external API.
6. Optional Vertex AI path: explicitly separate and governed
If AtlasMart chooses Vertex AI, current Google Cloud samples
show gemini-embedding-001 as a stable embedding
model. That choice is not part of the mandatory lab and can
change over time. Record the exact model ID, task type, output
dimension, location, library/API version and source hash. If the
model emits more than Firestore’s 2,048 supported dimensions,
request a compatible output dimension where the model supports
it or apply a governed dimensionality-reduction step; never
truncate blindly.
{ "provider": "vertex-ai", "model": "gemini-embedding-001", "taskType": "RETRIEVAL_DOCUMENT", "requestedOutputDimension": 768, "firestoreVectorField": "embeddingVertex768V1", "firestoreIndexDimension": 768, "contentPolicy": "public product catalog text only", "piiAllowed": false, "secretsAllowed": false}
7. Privacy and data governance gate
| Source text | Default decision | Reason |
|---|---|---|
| Public product name/description | Potentially eligible | Still record provider/model/location and retention policy |
| Customer email/order notes | Do not send by default | May contain PII or confidential information |
| Support ticket with credentials | Never send raw | Secrets must be removed; consider local/private model or no embedding |
| Regulated tenant data | Policy/legal review | Residency, processor and retention obligations may dominate architecture |
8. Cutover state machine
A robust migration has observable phases:
v2_index_building → v2_backfilling → v2_shadow_read →
v2_quality_approved → v2_primary → v1_retiring. Shadow reads compare v1/v2 on judged traffic without changing
user-visible results. Rollback means switching the read contract
back to v1 while v1 index/data still exist.
9. Wrong approach: “backfill until the script exits 0”
A one-shot script with no per-document state cannot prove completeness after throttling, partial failure or source updates. Repair it with idempotent job records, source hashes, pagination/checkpoints, retry classes, completion counts, and a reconciliation query that identifies products missing the target model/version.
const missing = [];for (const doc of (await db.collection("vectorProducts").get()).docs) { const d = doc.data(); if (d.searchable && (d.embeddingVersionV2 !== 2 || !Array.isArray(d.embeddingV2) || d.embeddingV2.length !== 8)) { missing.push(doc.id); }}console.log({ missingCount: missing.length, missing });if (missing.length) process.exitCode = 1;
10. Edition/mode/client matrix
| Surface | Vector capability | Important boundary |
|---|---|---|
| Standard Native Core | Managed nearest-neighbor search through supported server libraries; flat vector indexes; metadata prefilters | Maximum 2,048 dimensions; maximum 1,000 returned documents; no realtime snapshot listeners for vector search; current documented client-library support is Python, Node.js, Go, Java |
| Enterprise Native Core | Native Core semantics in Enterprise context | Enterprise indexing/billing differ from Standard; do not copy Standard billing equations blindly |
| Enterprise Native Pipeline |
findNearest transformation stage with broader
Pipeline client surface
|
Pipeline API and supported distance options are distinct; verify current stage and SDK behavior rather than treating it as Standard Core |
| MongoDB compatibility | Separate MongoDB-compatible query/index surface |
Do not assume Native findNearest,
vector-index commands, Rules, or pricing semantics map
one-for-one
|
| Emulator | Useful for fixture/security/application plumbing; recent emulator versions can store vector values | Composite/vector-index enforcement, production latency, billing, index build state, and all service limits are not production evidence |
Verification checklist
- Generation is outside retryable transaction callbacks.
- Every target vector is tied to a source hash and model/version/dimension.
- Backfill retry after injected failure is idempotent.
- External generation is optional and has an explicit content-governance gate.
- Cutover and rollback states are documented before retiring v1.
Production judgment and bridge to Lesson 5
A reliable vector pipeline can still deliver a bad search product. Lesson 5 closes the loop with judged relevance, recall, latency distribution, authorization tests, cost accounting, drift detection and a decision on whether Firestore remains the right retrieval engine.
Knowledge check
- Does Firestore generate embeddings?
- Why is an external embedding call unsafe inside a retryable transaction callback?
- What makes a backfill replayable?
- What should happen before sending customer text to an embedding provider?
- Why keep v1 during v2 shadow evaluation?
Review the answers
1. No. Embeddings are generated externally/local to Firestore and then stored/indexed.
2. The callback may rerun, duplicating cost/data transfer/side effects and extending transaction duration.
3. Per-item durable state, source/model versioning, idempotent writes, checkpoints and reconciliation.
4. A data-classification/governance decision, including PII/secrets/residency/retention constraints.
5. It provides a rollback path and a stable baseline for measured quality comparison.
Summary and next step
This lesson established the working contract for Firestore Does Not Generate Embeddings: Vertex AI/External Generation Pipelines, Backfills, and Updates. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Build and Evaluate a Semantic Retrieval Feature with Recall/Quality, Latency, Security, and Cost Measurements.
Authoritative references
- Firebase · Search with vector embeddings
- Google Cloud · Firestore Standard pricing
- Firebase · Connect to the Firestore emulator and understand differences from production
- Google Cloud · Enterprise Native Pipeline findNearest stage
- Firebase · Enterprise Native supported data types
- Google Cloud · Vertex AI embedding sample