Chapter 10 · Vector Sets, Vector Search, Hybrid Retrieval, and AI Workloads

Vector Compression/Quantization Awareness, Memory Planning, and Accuracy Benchmarks

Benchmark quantization, dimension reduction, graph memory, recall, and tail latency as one coupled capacity problem.

Advanced175–210 minutesVector memory and accuracy labRedis Open Source 8.10.1Free/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart's prototype retrieves good neighbors, but production capacity depends on millions of dimensions, encoded vectors, graph links, attributes, persistence, and concurrency. Memory optimization is only useful if retrieval quality remains acceptable.

01

Compare Q8, BIN, and NOQUANT as measurable Vector Set representation choices.

02

Estimate raw coordinate memory and separate it from graph/label/attribute overhead.

03

Explain how M, EF, REDUCE, and dimension affect memory, ingestion, latency, and recall.

04

Build an exact ground-truth benchmark with VSIM TRUTH and compute recall@k.

05

Produce p50/p95/p99 latency plus memory-per-vector evidence without fabricated production numbers.

Exact lab baseline

All Chapter 10 mandatory labs reuse the disposable Chapter 01 environment: Redis Open Source 8.10.1 from Docker Official Image redis:8.10.1, container atlasmart-redis-ch01, standalone topology, host publication 127.0.0.1:6379, TLS disabled only because traffic stays on loopback, default ACL user disabled, named users atlasmart-app and academy-admin, logical database 0, AOF with appendfsync everysec plus RDB snapshots, persistent /data, and no explicit maxmemory limit or eviction policy. Redis 8 integrates Vector Sets and the Redis Query Engine into Redis Open Source. Mandatory examples use synthetic numeric vectors created locally—Redis stores/searches vectors but does not generate embeddings. Fixtures stay under atlasmart:ch10:*. Vector Set commands use the restricted application user where allowed; Search index administration uses the disposable academy-admin user. No paid embedding API, managed service, production endpoint, or real credential is required.

Feature-status discipline

Redis 8.0 introduced Vector Sets as a beta data type. The current Redis Open Source 8.10 command reference documents VADD, VSIM, VINFO, filtering, quantization, and related commands as available since 8.0, with standard Redis Software/Redis Cloud compatibility. The official sources checked for this lesson do not provide a separate explicit “Vector Sets became GA on version X” declaration. Treat Vector Set API/product status, client coverage, managed-service support, and Active-Active compatibility as version-sensitive and verify the exact target rather than inventing a GA date.

1. Start with raw coordinate arithmetic, then measure Redis

A 300-dimensional FP32 vector contains 300 × 4 = 1200 raw coordinate bytes before labels, object headers, allocator effects, HNSW links, attributes, and persistence buffers. Q8 uses approximately one byte/component (about 4× smaller than FP32 coordinates); BIN is approximately one bit/component (about 32× smaller). These are coordinate-level ratios, not complete key memory.

Mode Coordinate intuition Quality expectation
NOQUANT full FP32-style coordinate footprint highest representation fidelity, highest memory
Q8 ~1 byte/component; default high recall/efficiency balance
BIN ~1 bit/component lowest coordinate memory, lower recall

2. Measure whole-key memory with MEMORY USAGE

redis-cli · same tiny vectors, three quantization modes
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch10:bench:q8 atlasmart:ch10:bench:noq atlasmart:ch10:bench:bindocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:bench:q8 VALUES 4 1.262185 1.958231 0.4 0.9 item Q8docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:bench:noq VALUES 4 1.262185 1.958231 0.4 0.9 item NOQUANTdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:bench:bin VALUES 4 1.262185 1.958231 0.4 0.9 item BINdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app MEMORY USAGE atlasmart:ch10:bench:q8docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app MEMORY USAGE atlasmart:ch10:bench:noqdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app MEMORY USAGE atlasmart:ch10:bench:bin

A one-vector fixture is too small for production capacity inference because fixed overhead dominates. Repeat at representative N/dimension and divide measured whole-key memory by cardinality while also reporting fixed/key-level overhead.

3. Quantization changes vectors and may change neighbors

redis-cli · inspect reconstructed coordinates
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VEMB atlasmart:ch10:bench:q8 itemdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VEMB atlasmart:ch10:bench:noq itemdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VEMB atlasmart:ch10:bench:bin item

BIN can visibly collapse coordinate detail. The correct question is not “which representation looks closest?” but “does the representation preserve the required top-k neighbors and task quality under the target workload?”

4. Exact truth makes recall measurable

For each benchmark query, run an exact VSIM ... TRUTH result and an approximate result with identical COUNT k. Compute recall@k. Repeat across queries, not one cherry-picked vector.

python · aggregate recall over query results
def recall_at_k(exact, approx, k):    return len(set(exact[:k]) & set(approx[:k])) / kqueries=[ (["a","b","c","d"],["a","b","c","x"]), (["m","n","o","p"],["m","n","q","p"]),]vals=[recall_at_k(e,a,4) for e,a in queries]print(vals, sum(vals)/len(vals))

5. M buys graph connectivity with memory

M controls maximum neighbor links in the HNSW graph. Current Redis Vector Set docs note that layer 0 uses roughly 2*M links and higher layers roughly M, with pointer memory contributing significantly. Raising M without a recall problem wastes memory and ingestion work.

Capacity rule

Budget coordinates + graph links + labels + attributes + key/allocator overhead + persistence/replication/fork headroom. maxmemory is not the same as process RSS.

6. REDUCE changes geometry, not just bytes

Random projection through REDUCE lowers stored dimension and saves coordinate/graph-adjacent work, but it changes distances. Benchmark reduced and unreduced systems on the same query/relevance set.

redis-cli · compare stored dimension
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch10:bench:full atlasmart:ch10:bench:reduceddocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:bench:full VALUES 4 1 0.2 0.1 0.7 item-adocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VADD atlasmart:ch10:bench:reduced REDUCE 2 VALUES 4 1 0.2 0.1 0.7 item-adocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VDIM atlasmart:ch10:bench:fulldocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app VDIM atlasmart:ch10:bench:reduced

7. Search vector indexes have a separate compression design space

Redis Search vector fields support FLAT, HNSW, and SVS-VAMANA. Search also supports multiple numeric vector types and SVS-oriented compression options depending on release. Do not map Vector Set Q8/BIN knobs mechanically onto Search index settings; they are different APIs and representations.

System Exact option Approximate option Compression/representation
Vector Set VSIM TRUTH baseline built-in HNSW-style VSIM Q8 default, BIN, NOQUANT, REDUCE
Search vector field FLAT HNSW / SVS-VAMANA vector TYPE plus algorithm-specific compression options

8. Latency benchmark contract

For every configuration record server/patch, CPU/RAM, container/native, dimension, N, quantization, M/build-EF/query-EF, filter selectivity, query concurrency, warmup, persistence/fsync, pipeline/client connection model, and p50/p95/p99. Run enough samples for stable tails. Do not paste the demonstration numbers below into capacity documents.

python · percentiles from real samples
# Replace with timings captured from your benchmark harness.samples_ms=[1.02,1.07,1.08,1.11,1.15,1.18,1.24,1.31,1.52,2.10]def nearest_rank(xs,p):    xs=sorted(xs); return xs[max(0,min(len(xs)-1, math.ceil(p*len(xs))-1))]import mathfor p in (0.50,0.95,0.99): print(p, nearest_rank(samples_ms,p))

9. Memory-per-vector benchmark contract

Measure MEMORY USAGE after a warm, fully built fixture and VCARD. Report both total bytes and bytes/vector. For attributes, test representative lengths and selectivities. Account for temporary ingestion/build peaks and fork/AOF/replication headroom separately.

10. Wrong approach: publish “Q8 is 4× cheaper” as total Redis memory

The 4× figure refers primarily to coordinate representation versus FP32; graph links, labels, attributes, allocator metadata, and fixed key overhead do not shrink by the same factor. Repair by measuring whole-key memory at realistic cardinality and disclosing what is included.

11. Wrong approach: optimize latency with BIN and never re-check relevance

Binary quantization may be fast and memory-efficient while changing nearest-neighbor order. Any representation/tuning change that affects geometry requires regression on exact recall@k and downstream task metrics.

12. Reproducible cleanup

redis-cli · remove quantization benchmark fixtures
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch10:bench:q8 atlasmart:ch10:bench:noq atlasmart:ch10:bench:bin atlasmart:ch10:bench:full atlasmart:ch10:bench:reduced

13. Production judgment

Capacity planning couples memory, recall, and latency. Keep enough headroom for replication/AOF buffers and fork-based persistence; large Vector Sets can become hot keys on a standalone/Cluster slot; retries and background search threads affect tails; patch releases can contain Vector Set correctness/security fixes. Rebuild/migration plans must include source embeddings or a reproducible embedding pipeline because a highly quantized in-memory representation is not necessarily the right long-term source artifact.

14. Summary and next step

You can now benchmark representation choices instead of arguing from folklore. Lesson 5 assembles the pieces into production-shaped semantic search, recommendation, and retrieval-augmented generation (RAG) designs with evaluation and security gates.

Check your understanding

  1. Why is Q8 “4× smaller” not a whole-key memory guarantee?
  2. What is VSIM TRUTH for?
  3. What does M change?
  4. Why does REDUCE require a recall regression?
  5. Which latency percentiles should the prompt explicitly measure?
Review the answers

Graph, labels, attributes, allocator/key overhead do not shrink by the same coordinate ratio.

Exact linear-scan ground truth and recall benchmarking.

HNSW graph connectivity, affecting memory/build/search behavior.

It changes vector geometry through projection.

At least p50, p95, and p99 for this chapter.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.