Chapter 10 · Vector Sets, Vector Search, Hybrid Retrieval, and AI Workloads
Design Semantic Search, Recommendations, or RAG Retrieval with Security Filters and Evaluation
Assemble semantic search, recommendations, and RAG retrieval with explicit security, evaluation, lifecycle, and rollback gates.
Learning outcomes
AtlasMart's final Chapter 10 task is architectural: choose and operate a retrieval design for semantic search, recommendations, or retrieval-augmented generation (RAG) without leaking tenant data or confusing nearest-neighbor fidelity with product usefulness.
Design semantic search, recommendation, and RAG retrieval as explicit pipelines with source, embedding, retrieval, filtering, and evaluation stages.
Apply mandatory security/tenant constraints before results leave retrieval.
Separate ANN recall metrics from business relevance and RAG answer-quality metrics.
Plan embedding/model version migrations, rollback, reindex/rebuild, and idempotent ingestion.
Define observability for latency, recall, under-fill, memory, stale embeddings, filter selectivity, and failures.
All Chapter 10 mandatory labs reuse the disposable Chapter 01
environment: Redis Open Source 8.10.1 from Docker
Official Image redis:8.10.1, container
atlasmart-redis-ch01, standalone topology, host
publication 127.0.0.1:6379, TLS disabled only
because traffic stays on loopback, default ACL user disabled,
named users atlasmart-app and
academy-admin, logical database 0, AOF with
appendfsync everysec plus RDB snapshots,
persistent /data, and no explicit
maxmemory limit or eviction policy. Redis 8
integrates Vector Sets and the Redis Query Engine into Redis
Open Source. Mandatory examples use synthetic numeric vectors
created locally—Redis stores/searches vectors but does not
generate embeddings. Fixtures stay under
atlasmart:ch10:*. Vector Set commands use the
restricted application user where allowed; Search index
administration uses the disposable
academy-admin user. No paid embedding API,
managed service, production endpoint, or real credential is
required.
Redis 8.0 introduced Vector Sets as a beta data type. The
current Redis Open Source 8.10 command reference documents
VADD, VSIM, VINFO,
filtering, quantization, and related commands as available
since 8.0, with standard Redis Software/Redis Cloud
compatibility. The official sources checked for this lesson do
not provide a separate explicit “Vector Sets became GA on
version X” declaration. Treat Vector Set API/product status,
client coverage, managed-service support, and Active-Active
compatibility as version-sensitive and verify the exact target
rather than inventing a GA date.
1. A production retrieval system is a pipeline, not one VSIM call
| Stage | Input/output | Failure to plan |
|---|---|---|
| source record | product/document with tenant + lifecycle | stale/deleted content can remain retrievable |
| embedding generation | text/image → versioned vector | model drift or API failure |
| vector storage/index | Vector Set or Search vector field | dimension/quantization/index mismatch |
| retrieval + security | query vector + mandatory filters → candidates | cross-tenant leakage or under-fill |
| reranking/task logic | candidates → task score | ANN-nearest is not necessarily useful |
| response/RAG generation | authorized evidence → output | hallucination or unsupported claims |
| evaluation/telemetry | queries + labels + metrics | no detection of quality regression |
2. Semantic search: measure retrieval and click/task quality separately
For semantic product search, exact/approximate recall tells you whether the index recovered its mathematical neighbors. Offline relevance judgments, click-through, conversion, add-to-cart, or human grading tell you whether those neighbors satisfy users. A high recall@10 system can still retrieve semantically adjacent but commercially useless products.
3. Recommendations: identity and business constraints matter
Recommendation vectors may encode product similarity, user preference, or session state. Never let a recommendation model bypass stock, age/region rules, tenant visibility, blocked sellers, or privacy boundaries. Store stable element IDs and fetch authoritative current state before final action when staleness matters.
4. RAG: retrieval quality does not guarantee answer correctness
RAG means retrieval-augmented generation: retrieve supporting documents/chunks, then give them to a language model as context. Redis can retrieve vectors; it does not guarantee the model cites correctly, follows instructions, or avoids hallucination. Evaluate retrieval recall/precision, context coverage, answer groundedness, citation correctness, latency, and security independently.
The mandatory lab stops at deterministic retrieval/evaluation. If you later connect an embedding or language-model API, keep credentials outside lesson code and provide local/offline fixtures for regression tests.
5. Security filter first, then rerank/return
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch10:capstone:vectorsdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:capstone:vectors VALUES 3 1 0 0 doc:a1 SETATTR '{"tenant":"tenant-a","visibility":"public","active":true,"embeddingVersion":"demo-v1"}'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:capstone:vectors VALUES 3 0.99 0.01 0 doc:b1 SETATTR '{"tenant":"tenant-b","visibility":"private","active":true,"embeddingVersion":"demo-v1"}'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:capstone:vectors VALUES 3 0.90 0.10 0.02 doc:a2 SETATTR '{"tenant":"tenant-a","visibility":"public","active":true,"embeddingVersion":"demo-v1"}'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VSIM atlasmart:ch10:capstone:vectors VALUES 3 0.995 0.005 0 WITHSCORES WITHATTRIBS COUNT 5 FILTER '.tenant == "tenant-a" && .visibility == "public" && .active == true && .embeddingVersion == "demo-v1"'
The tenant-b vector is intentionally very similar so an unfiltered query would expose it. The filtered query is the acceptance test.
6. Version embeddings as data, not tribal knowledge
Changing an embedding model can change dimension and geometry.
Record embeddingVersion, model identifier,
preprocessing, normalization policy, and dimension. During
migration, dual-write/dual-index or rebuild into a new
key/index, compare quality, switch reads, retain rollback, then
remove the old representation after evidence.
| Change | Safe migration shape |
|---|---|
| same dimension, new model | new Vector Set/index namespace; compare before cutover |
| new dimension | must use new Vector Set/index because dimensional contract changes |
| quantization change | build separate representation; compare memory/recall/latency |
| tenant policy change | re-evaluate key/index isolation and filter tests before rollout |
7. Idempotent ingestion and deletion lifecycle
Stable element labels let VADD update an existing
vector instead of creating a duplicate label. That helps
idempotent ingestion, but you still need source version checks,
retries, tombstones/deletion handling, and reconciliation. When
a product is deleted or access revoked, remove or disable its
retrievable vector promptly.
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:capstone:vectors VALUES 3 0.92 0.08 0.01 doc:a2docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VISMEMBER atlasmart:ch10:capstone:vectors doc:a2docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VREM atlasmart:ch10:capstone:vectors doc:a2docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VISMEMBER atlasmart:ch10:capstone:vectors doc:a2
8. Evaluation matrix: retrieval fidelity versus task quality
| Metric | Question | Do not confuse with |
|---|---|---|
| recall@k | Did ANN recover exact top-k neighbors? | human relevance |
| precision/nDCG/MRR | Are ranked results judged useful? | authorization correctness |
| security leakage rate | Did any forbidden item escape? | semantic quality |
| under-fill rate | Did filters leave fewer than requested k? | low recall alone |
| p50/p95/p99 | How fast is retrieval at median/tail? | quality |
| groundedness/citation accuracy | For RAG, does answer follow retrieved evidence? | nearest-neighbor fidelity |
9. Offline benchmark harness must freeze ground truth
Keep a versioned query set with expected relevant IDs/grades and
tenant constraints. For ANN recall, compute exact neighbors with
VSIM TRUTH or a Search FLAT baseline over the same
vectors. For task quality, use labeled relevance. Re-run after
model, quantization, EF, M, filtering, Redis patch, client, or
topology changes.
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VSIM atlasmart:ch10:capstone:vectors VALUES 3 0.995 0.005 0 WITHSCORES COUNT 2 FILTER '.tenant == "tenant-a" && .active == true' TRUTHdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VSIM atlasmart:ch10:capstone:vectors VALUES 3 0.995 0.005 0 WITHSCORES COUNT 2 FILTER '.tenant == "tenant-a" && .active == true' EF 100
Use the exact result as neighbor ground truth and compare the approximate result for the same authorized subset. On this tiny fixture they may match; that is not a production recall claim.
10. Online observability signals
- request count, errors, timeouts, retries, p50/p95/p99;
- vector cardinality, dimension, memory, quantization/index algorithm;
- filter selectivity, top-k under-fill, unauthorized-result test failures;
- embedding-version distribution and stale/missing-vector count;
- ANN recall on canary/offline probes and task relevance dashboards;
- ingestion lag, reconciliation drift, deletion/tombstone lag;
- Redis memory, CPU, latency events, persistence/replication/failover health.
11. Wrong approach: let the LLM retrieve “whatever is closest”
This merges retrieval, authorization, and task relevance into one opaque step. A close vector may be forbidden, stale, or irrelevant. Repair by pre-authorizing/structurally filtering, retrieving bounded candidates, optionally reranking, fetching authoritative source state, and only then constructing model context.
12. Wrong approach: store API secrets in vector attributes
Vector attributes can be returned with
WITHATTRIBS and may appear in dumps/logs/backups.
Never store embedding-provider keys, bearer tokens, passwords,
or private prompts there. Use a secret manager/environment
mechanism appropriate to deployment.
13. Failure injection: stale vector and cross-tenant adversary
Two safe local drills: (1) update a source fixture without
updating its vector and verify reconciliation detects stale
embeddingVersion; (2) insert a very-similar
tenant-b element and prove every tenant-a retrieval excludes it.
These are more valuable than happy-path demos because they test
the correctness boundary.
14. Capstone acceptance checklist
| Area | Acceptance evidence |
|---|---|
| version | server/redis-cli command metadata captured; target client support recorded |
| dimension/type | ingest rejects mismatches/zero-invalid vectors per model contract |
| security | adversarial cross-tenant candidate never returned |
| quality | recall@k and task metric exceed project-specific measured thresholds |
| latency | p50/p95/p99 measured under disclosed load |
| memory | whole-key/index memory and headroom budgeted |
| lifecycle | updates/deletes/rebuild/reconciliation tested |
| rollback | previous embedding/index namespace remains recoverable during cutover |
15. Cleanup
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch10:capstone:vectors
16. Production judgment
Choose the simplest retrieval surface that satisfies query and governance needs: Vector Sets for Redis-native similarity with light attributes; Redis Search vector fields when richer hybrid schemas/querying matter. For production AI workloads, model cost and availability are outside Redis but still part of end-to-end SLOs. Account for patching, persistence/recovery, replication/failover, Cluster routing, memory headroom, tenant isolation, observability, evaluation data governance, model licensing/terms, and rollback. Do not claim exactly reproducible semantic quality without freezing the model, preprocessing, dataset, and query labels.
17. Chapter summary and bridge
Chapter 10 established a complete vector discipline: embeddings are external representations; dimension/metric/normalization are contracts; Vector Sets and Search vector indexes are distinct tools; ANN must be judged against exact recall and task relevance; quantization/memory/latency are coupled; and security filters must constrain retrieval. Chapter 11 changes approximation domains—from nearest neighbors to probabilistic summaries such as HyperLogLog, Bloom filters, Count-Min Sketch, Top-K, and t-digest.
Check your understanding
- Why version embeddings explicitly?
- Can recall@k replace relevance evaluation?
- What does an adversarial tenant fixture test?
- Why keep source vectors/rebuild capability?
- What is RAG in one sentence?
Review the answers
Model/preprocessing/dimension changes alter vector meaning and require migration control.
No. Recall measures ANN fidelity to exact neighbors, not user/task usefulness.
That retrieval-time security filters/isolation prevent cross-tenant leakage even for highly similar items.
Quantized/index state may not be the right source artifact and indexes need migration/recovery.
Retrieve authorized relevant evidence, then provide it as context to a generative model while evaluating grounding separately.
Authoritative references
- Redis Vector Sets — native Vector Set data type, commands, filtering, and examples
- VADD — Vector Set insertion, quantization, REDUCE, EF, M, and attributes
- VSIM — similarity queries, scores, filters, EF, TRUTH, and NOTHREAD
- VINFO — Vector Set implementation and configuration evidence
- VEMB — stored/reconstructed vector evidence
- VSETATTR — JSON attributes attached to Vector Set elements
- Vector Set memory optimization — Q8/BIN/NOQUANT, dimensions, graph links, and memory tradeoffs
- Vector Set performance — quantization and vector-set performance considerations
- Redis vector search concepts — Search FLAT/HNSW/SVS-VAMANA vector indexes and runtime parameters
- Vector field options — Search vector field types, metrics, and index algorithms
- FT.CREATE — Search schema and vector field creation
- FT.SEARCH — KNN/hybrid query syntax and parameters
- Redis 8.10 commands — target-version command surface
- Redis 8.10 release notes — 8.10.1 security baseline including Vector Set fixes
- Redis 8 GA announcement — historical 8.0 Vector Set beta status and integrated Redis 8 capabilities