Chapter 19 · Vector Search in Cassandra 5.0
Combine Vector Retrieval with Metadata Predicates and Query-First Modeling
Combine vector ANN with tenant/bucket and metadata predicates while keeping authorization, fanout, and query-first partition modeling explicit.
Learning outcomes
AtlasMart's semantic search is useful until a test query from tenant A retrieves a highly similar private document owned by tenant B. “We'll remove unauthorized rows after ANN” is not an acceptable architecture: it can leak metadata/timing and may leave fewer than k usable results. This lesson makes structured scope part of retrieval design.
Combine ANN with primary-key/SAI metadata predicates and distinguish data filtering from authorization.
Design bounded tenant/corpus buckets so application fanout is explicit and measurable.
Explain why post-retrieval authorization is unsafe and can damage top-k quality.
Compare partition-constrained ANN with broader distributed SAI fanout.
Test filtered/unfiltered results and node failure without claiming row-level security from SAI.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 container, Java 17 inside the
image, cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. The chapter keyspace is
atlasmart_vector with
NetworkTopologyStrategy, RF=3, and normal
reads/writes at LOCAL_QUORUM. Tables explicitly
use UnifiedCompactionStrategy (UCS), no default TTL, and the
Cassandra default gc_grace_seconds unless a
lesson says otherwise. Authentication, client TLS, internode
TLS, and remote JMX are disabled only inside this isolated
local learning network; production systems must enforce
appropriate authentication, authorization, encryption, and
network boundaries. Apache Cassandra Java Driver 4.19.3 is
optional for application-level latency/vector-codec examples.
The mandatory embedding data is a deterministic
four-dimensional pedagogical fixture—free, local, and
intentionally not a production semantic model. No paid
embedding or LLM API is required.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and mental model
A vector is an ordered fixed-length array of numbers. An embedding is a vector produced by a model or deterministic transformation to represent an object in a geometric space. Its dimension is the number of elements. Embedding provenance means the model name/version, preprocessing, normalization policy, and source revision that explain how the vector was created. A similarity function or distance rule decides what “near” means. Cosine similarity compares direction, dot product combines direction and magnitude unless vectors are normalized, and Euclidean is based on geometric distance. Normalization usually means scaling a vector to unit length. Approximate Nearest Neighbor (ANN) retrieval searches an index for likely nearest vectors without guaranteeing the exact top-k ordering; exact retrieval evaluates every candidate under the chosen metric. Recall@k is the fraction of exact top-k items recovered by ANN. It measures retrieval fidelity, not whether the documents are useful to a user.
Storage-Attached Indexing (SAI) is Cassandra
5.0's storage-integrated secondary indexing framework; vector
ANN search uses an SAI index on a
vector<float,n> column. A
coordinator is the Cassandra node handling a
particular client request. A replica stores
data for a token range according to the keyspace replication
strategy. A partition is the Cassandra
storage/routing unit selected by a partition key; the
partitioner hashes that key to a token. An
immutable SSTable is an on-disk Sorted String
Table. A consistency level (CL) states the
replica acknowledgments/responses required by an operation.
RAG (Retrieval-Augmented Generation) is an
application architecture that retrieves evidence and supplies it
to a generative model; Cassandra can be the retrieval store but
does not generate embeddings or model answers.
1. Query-first vector modeling: scope before similarity
The course table uses
((tenant_id,corpus_bucket),document_id). That makes
tenant and bucket part of the physical routing key. The
application derives allowed tenant/buckets from an authenticated
principal before issuing ANN queries. Metadata such as
locale, doc_type, and
embedding_version can then narrow candidates
further. SAI accelerates those predicates; it does not
authenticate the caller.
# Verify the shared course lab first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the disposable course cluster does not exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three nodes are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_vectorWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_vector.documents ( tenant_id text, corpus_bucket tinyint, document_id uuid, title text, body text, locale text, doc_type text, embedding_model text, embedding_version text, embedding vector<float,4>, updated_at timestamp, PRIMARY KEY ((tenant_id,corpus_bucket),document_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;CREATE INDEX IF NOT EXISTS documents_embedding_annON atlasmart_vector.documents (embedding) USING 'sai'WITH OPTIONS = {'similarity_function':'COSINE'};CREATE INDEX IF NOT EXISTS documents_locale_saiON atlasmart_vector.documents (locale) USING 'sai';CREATE INDEX IF NOT EXISTS documents_type_saiON atlasmart_vector.documents (doc_type) USING 'sai';
INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001901,'Return policy','How to return an unopened product within the return window.','en','support','atlasmart-toy','v1',[1.0,0.0,0.0,0.0],'2026-09-08T07:00:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001902,'Refund delay','Why approved refunds can take several days to appear.','en','support','atlasmart-toy','v1',[0.9,0.2,0.1,0.0],'2026-09-08T07:01:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001903,'Shipping tracker','Track a parcel after warehouse dispatch.','en','support','atlasmart-toy','v1',[0.0,1.0,0.0,0.0],'2026-09-08T07:02:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001904,'Password reset','Recover access to an AtlasMart account.','en','security','atlasmart-toy','v1',[0.0,0.0,1.0,0.0],'2026-09-08T07:03:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001905,'Invoice copy','Download a VAT invoice for an order.','en','billing','atlasmart-toy','v1',[0.0,0.0,0.0,1.0],'2026-09-08T07:04:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001906,'Exchange item','Exchange an eligible item for another size.','en','support','atlasmart-toy','v1',[0.8,0.1,0.0,0.2],'2026-09-08T07:05:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001907,'Cancel order','Cancel before warehouse fulfillment begins.','en','support','atlasmart-toy','v1',[0.7,0.0,0.2,0.1],'2026-09-08T07:06:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001908,'Delivery delay','Investigate a shipment that missed its delivery date.','en','support','atlasmart-toy','v1',[0.1,0.9,0.0,0.0],'2026-09-08T07:07:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-b',0,00000000-0000-0000-0000-000000001909,'Tenant B private return note','Private tenant B support content.','en','support','atlasmart-toy','v1',[0.99,0.01,0.0,0.0],'2026-09-08T07:08:00Z');
SELECT document_id,title,doc_type,locale,embedding_version, similarity_cosine(embedding,[1.0,0.1,0.0,0.0]) AS scoreFROM atlasmart_vector.documentsWHERE tenant_id='tenant-a' AND corpus_bucket=0 AND locale='en' AND doc_type='support'ORDER BY embedding ANN OF [1.0,0.1,0.0,0.0]LIMIT 3;
The full partition key bounds storage access; SAI metadata indexes can intersect with ANN. The returned rows are still approximate nearest neighbors among the eligible candidate space. If the tenant spans several buckets, the application can query each authorized bucket and merge/trim by the same similarity function.
2. Why post-filtering is a failure, not a convenience
-- Deliberately unsafe application pattern: no tenant predicate.SELECT tenant_id,corpus_bucket,document_id,title, similarity_cosine(embedding,[1.0,0.1,0.0,0.0]) AS scoreFROM atlasmart_vector.documentsORDER BY embedding ANN OF [1.0,0.1,0.0,0.0]LIMIT 5;
The toy corpus includes a tenant-B private document intentionally close to the returns query. If it appears, dropping it in application code after retrieval means unauthorized data already crossed the database/application boundary. Even if the body is never displayed, identifiers, scores, timing, logs, caches, or traces can leak. It also means a requested top-3 may become top-2 after filtering.
Cassandra native roles grant permissions on database resources
such as keyspaces/tables, not arbitrary row-level predicates.
The application must derive and enforce allowed scope;
stronger isolation may require separate
tables/keyspaces/clusters or another security architecture.
Never promise tenant isolation merely because the CQL includes
tenant_id.
3. Bounded fanout and deterministic merge
# Each bucket query overfetches a small k under the SAME similarity metric.bucket0=[('1901',0.995),('1902',0.982),('1903',0.110)]bucket1=[('1906',0.979),('1907',0.955),('1908',0.220)]merged=sorted(bucket0+bucket1,key=lambda x:x[1],reverse=True)[:3]print(merged)# In production, preserve document ID + score + source bucket and deduplicate IDs.
Bucketing trades one unbounded distributed search for a known number of parallel requests. Measure how many buckets each business query fans out to, whether overfetch is sufficient to preserve recall, how duplicates are handled, and how failures/timeouts affect partial results. Random sharding without a routing plan merely moves the cost into application fanout.
4. Broader SAI model for comparison—not a recommendation
To understand distributed fanout, create a second disposable table keyed only by document ID, with SAI indexes on tenant, locale, type, and vector. This can answer flexible predicates without a partition key, but the coordinator may search token ranges/endpoints across the cluster. It is useful evidence for why primary-key modeling remains first.
CREATE TABLE IF NOT EXISTS atlasmart_vector.documents_global ( document_id uuid PRIMARY KEY, tenant_id text, locale text, doc_type text, embedding_version text, embedding vector<float,4>, title text) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE INDEX IF NOT EXISTS global_tenant_sai ON atlasmart_vector.documents_global (tenant_id) USING 'sai';CREATE INDEX IF NOT EXISTS global_locale_sai ON atlasmart_vector.documents_global (locale) USING 'sai';CREATE INDEX IF NOT EXISTS global_type_sai ON atlasmart_vector.documents_global (doc_type) USING 'sai';CREATE INDEX IF NOT EXISTS global_embedding_ann ON atlasmart_vector.documents_global (embedding) USING 'sai';-- After loading equivalent rows, trace a scoped query.TRACING ON;SELECT document_id,title FROM atlasmart_vector.documents_globalWHERE tenant_id='tenant-a' AND locale='en' AND doc_type='support'ORDER BY embedding ANN OF [1.0,0.1,0.0,0.0]LIMIT 3;TRACING OFF;
Capture participating endpoints/ranges, SAI indexes/segments touched, p95/p99 latency, and index bytes. The correct design depends on corpus size, filter selectivity, cluster size, and latency SLO; do not extrapolate from nine rows.
5. Verification and failure check
docker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status# Run only the tenant-a/bucket-0 query and record outcome/latency.docker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
- No mandatory query retrieves tenant B and filters it afterward.
- Allowed tenant/bucket scope is decided before ANN.
-
Metadata indexes are observable in
system_views.indexes. - Broad-search comparison is clearly labeled as a tradeoff experiment.
- All failure-injection state is restored.
Check your understanding
- Why is post-filtering tenant B after ANN unsafe?
- Does an SAI tenant_id predicate implement Cassandra row-level authorization?
- What is the purpose of corpus_bucket?
- Why might a global document_id-keyed SAI table be slower at scale?
- What should happen when one bucket times out?
Review the answers
1. Unauthorized rows have already influenced/crossed the retrieval boundary, and top-k can shrink after filtering.
2. No. It is a data filter; authorization must be enforced by the application/security architecture.
3. Bound partition size and make search fanout explicit/controllable.
4. Without partition routing, the coordinator may fan out across token ranges/endpoints depending on selectivity and LIMIT satisfaction.
5. Follow an explicit product contract—fail closed, return partial with disclosure, or retry if safe—rather than silently pretending a complete top-k.
Production judgment
Decide whether Cassandra vector search fits by measuring the whole retrieval system: corpus size and growth; vector dimension and bytes; embedding generation/update rate; embedding provenance; tenant and authorization model; query filters and bucket fanout; RF/CL and node/DC failures; ANN recall@k against a fixed exact baseline; application relevance metrics; p50/p95/p99 retrieval latency; result payload size; write amplification; SAI disk and memory footprint; SSTable count/compaction; vector overwrite/delete rate; index build/rebuild/streaming time; repair/backup/restore behavior; JVM/off-heap/chunk-cache pressure; driver timeouts/retries/idempotency; guardrails; observability; and operator skill. Do not report ANN latency without recall, and do not call recall “answer quality.”
Security is not a similarity metric. Cassandra role permissions are table/keyspace oriented, not row-level authorization; an application must derive allowed tenant/corpus scope before retrieval and encode that scope in the data model/query. Post-filtering unauthorized ANN results can leak information or reduce useful top-k results. Managed services and Cassandra-compatible APIs may differ in vector syntax, limits, index lifecycle, metrics, encryption, billing, and topology. Lesson 5 turns the chapter into an operating/evaluation lifecycle: capacity, builds/streaming, embedding-version migration, recall/relevance gates, and RAG boundaries.
Summary and next bridge
Vector retrieval must be query-first and authorization-aware. Similarity only ranks candidates inside an allowed scope; it must never decide who is allowed to see them. The final lesson adds lifecycle capacity, streaming/rebuild, embedding migrations, and RAG system boundaries.
Authoritative references
Vector search and SAI are version-sensitive. Re-check these sources when regenerating the lesson rather than freezing a 2026 assumption.