Test filtered and multi-vector retrieval under tenant boundaries, then treat quantization and hardware choices as measured quality-capacity tradeoffs.

Filtering Vector Search, Multi-Vector/Nested Patterns, Quantization/Compression, and Hardware/Memory Planning

Teach vector search from embedding contracts through ANN indexing so quality, latency, memory, filtering, quantization, and model-version assumptions are measured rather than inferred.

Intermediate → Advanced130–175 minutesFiltered/nested/quantization planning lab · Chapter 25 · Lesson 04Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Explain how vector filters interact with ANN candidate generation and why filter selectivity belongs in recall tests.

02

Model passage/multi-vector content using explicit nested or child-document boundaries instead of pretending a vector field is inherently multi-valued.

03

Compare quantization/compression strategies as quality-memory-disk-build tradeoffs rather than “free” optimization.

04

Estimate first-order raw vector footprint and identify graph/native/off-heap/filesystem overhead that the estimate does not include.

05

Design tenant-safe filtered retrieval and verify that authorization and vector similarity remain separate controls.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned vector-search baseline

Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released September 3, 2026) and OpenSearch/OpenSearch Dashboards 3.8.0 (released August 4, 2026), using the course's established local TLS/auth conventions and bundled JVMs. The mandatory lab uses precomputed deterministic vectors, so no hosted model, GPU, paid inference endpoint, or proprietary API is required. Elasticsearch examples explicitly map dense_vector rather than relying on changing defaults; OpenSearch examples explicitly map knn_vector and use the built-in Lucene engine for the portable local path. No live cluster is available in this generation environment, so performance cells are labeled MEASURED and must be filled from the learner's run rather than invented.

1. AtlasMart problem: the unfiltered benchmark passes, production fails

Unfiltered ANN recall is excellent, but production queries always include tenant_id, category, stock region, and visibility filters. A selective filter changes the eligible set and may change how the engine traverses or falls back to exact work. AtlasMart therefore benchmarks the filter distributions users actually send.

Wrong approach: test recall only without filters.

An ANN configuration can look excellent on the global corpus and miss the service objective under selective filters. Filter semantics are part of the retrieval contract.

2. Efficient filter versus post-filter

A post-filter can retrieve k vector neighbors first and discard ineligible documents afterward, potentially returning fewer than k useful hits. A vector-aware efficient filter constrains search so the engine tries to return k neighbors from the eligible set when enough exist. Exact mechanics differ by engine and selectivity. OpenSearch documents dynamic exact/approximate decisions for supported Lucene/Faiss/JVector filtering; Elasticsearch kNN filters are part of the kNN request/query mechanism.

Elasticsearch: tenant + category inside kNN
POST /atlasmart-vector-es-v25/_search
{
  "size":3,
  "knn":{
    "field":"embedding",
    "query_vector":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20],
    "k":3,
    "num_candidates":20,
    "filter":{"bool":{"filter":[
      {"term":{"tenant_id":"tenant-a"}},
      {"term":{"category":"outdoor"}}
    ]}}
  }
}
# Eligible fixture docs are P-1001, P-1002, P-1005.
OpenSearch Lucene HNSW: inline efficient filter
POST /atlasmart-vector-os-v25/_search
{
  "size":3,
  "query":{"knn":{"embedding":{
    "vector":[0.70,0.42,0.04,0.05,0.05,0.55,0.08,0.20],
    "k":3,
    "filter":{"bool":{"filter":[
      {"term":{"tenant_id":"tenant-a"}},
      {"term":{"category":"outdoor"}}
    ]}}
  }}}
}
# Verify supported filter behavior for the exact engine/version; do not move this to a generic post_filter and assume equivalence.

3. Tenant metadata is not optional

Vector similarity is not authorization. Store and enforce tenant/visibility metadata through roles/DLS/application boundaries appropriate to the product, then use vector filters for retrieval semantics. Do not rely on “different tenants probably have dissimilar vectors.” That is data leakage by design.

Security test

Use a query vector intentionally close to a document from another tenant. The principal/filter must still prevent that document from being returned. Record a 403 or an absent hit as security evidence depending on the chosen enforcement layer.

4. Multi-vector and nested patterns

Long products, manuals, reviews, or passages often need more than one vector per logical product. Treat each passage vector as an addressable retrieval unit with its own text, model metadata, and position. Elasticsearch supports nested kNN patterns; OpenSearch supports nested vector structures with product-specific query syntax. Validate the returned document/passages semantics, not only vector distance.

Conceptual passage model
product_id: P-2001
passages:
  - passage_id: description
    text: "waterproof shell for alpine rain"
    embedding: [ ... 8 floats ... ]
  - passage_id: care
    text: "wash cold; reactivate DWR"
    embedding: [ ... 8 floats ... ]
embedding_model: atlasmart-product-v1
# Implement with the product's documented nested/multi-vector mapping; one dense_vector value is not an array-of-vectors shortcut.

5. Quantization and compression

Approach Benefit Cost/risk Verification
8-bit/scalar quantization lower vector memory some recall loss; raw/re-rank storage may remain recall + memory/disk
4-bit / more aggressive scalar more compression higher quality risk and dimension constraints judged recall + p99
binary/BBQ very large memory reduction larger approximation error; often needs oversampling/rescore candidate/rescore sweep
Faiss PQ strong compression training + engine-specific behavior training representativeness + recall/build cost

Elasticsearch and OpenSearch expose different quantization families and defaults. Do not infer equal recall from similar compression ratios.

6. First-order capacity math

A raw float32 vector uses approximately 4 × dims bytes before document, graph, segment, metadata, native-library, cache, and replication overhead. For 768 dimensions that is about 3072 bytes/vector before overhead. This is a lower-bound planning number, not a node-sizing formula.

Capacity worksheet
raw_vector_bytes = vector_count * dimensions * bytes_per_dimension
replicated_raw_bytes = raw_vector_bytes * (1 + replica_count)

# Then MEASURE, do not guess:
# - graph/index bytes
# - segment/store bytes
# - native/off-heap/vector memory
# - filesystem cache behavior
# - JVM heap and breakers
# - merge/recovery temporary headroom
# - quantized representation + raw vector/rescore overhead

7. Hardware and lifecycle stage

Vector search can stress memory bandwidth, CPU SIMD paths, filesystem cache, native memory, and storage differently from lexical search. Local SSD can help storage-bound modes, but benchmark the actual engine/mode and failure/recovery path. Quantization that makes steady-state search fit in memory can still leave index builds, merges, recovery, and rescoring as peak-resource events.

8. Production judgment

Choose filters, nested layout, and compression together. A tenant filter that is highly selective may make exact scoring practical; a low-selectivity filter at billion-vector scale may favor an ANN engine with efficient in-search filtering. Re-evaluate when tenant sizes, model dimensions, passage counts, or retention change.

Check your understanding

  1. Why can a post-filter return fewer than k usable hits?
  2. Is tenant_id in the document enough for security?
  3. Why are passage vectors modeled explicitly?
  4. What does 4*dims estimate?
  5. What must accompany a quantization claim?
Review the answers

1. It may discard ineligible documents only after vector neighbors were selected.

2. No. Authorization must be enforced by roles/DLS/application boundaries; retrieval filters alone are not a complete security model.

3. A vector field is not automatically multi-valued and passage identity/text must remain auditable.

4. Only the raw bytes for one float32 vector, excluding graph, segment, metadata, cache, replicas, and other overhead.

5. Measured recall/quality, latency, memory/disk, build cost, and the exact engine/index settings.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.