Chapter 19 · Vector Search in Cassandra 5.0
Create SAI Vector Indexes and Perform Approximate Nearest Neighbor Queries
Create and observe SAI vector indexes, ANN queries, similarity scores, build state, disk cost, documented limits, and replica-failure behavior.
Learning outcomes
The schema is now valid, but AtlasMart still has no search structure. This lesson attaches an SAI vector index to the embedding column, waits for it to become queryable, executes ANN requests, observes similarity scores and index disk/build state, and tests what a one-replica outage does—and does not—say about result correctness.
Create a Cassandra 5.0 SAI vector index with an explicit similarity function and explain why that metric is an index-level choice.
Use ORDER BY vector ANN OF ... LIMIT k and explain the documented LIMIT and least-similar boundaries.
Observe SAI build/queryable state and per-column/per-table disk usage through system_views.
Distinguish ANN candidates from exact neighbors and application relevance.
Exercise a reversible replica failure while preserving RF=3/LOCAL_QUORUM assumptions.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 container, Java 17 inside the
image, cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. The chapter keyspace is
atlasmart_vector with
NetworkTopologyStrategy, RF=3, and normal
reads/writes at LOCAL_QUORUM. Tables explicitly
use UnifiedCompactionStrategy (UCS), no default TTL, and the
Cassandra default gc_grace_seconds unless a
lesson says otherwise. Authentication, client TLS, internode
TLS, and remote JMX are disabled only inside this isolated
local learning network; production systems must enforce
appropriate authentication, authorization, encryption, and
network boundaries. Apache Cassandra Java Driver 4.19.3 is
optional for application-level latency/vector-codec examples.
The mandatory embedding data is a deterministic
four-dimensional pedagogical fixture—free, local, and
intentionally not a production semantic model. No paid
embedding or LLM API is required.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and mental model
A vector is an ordered fixed-length array of numbers. An embedding is a vector produced by a model or deterministic transformation to represent an object in a geometric space. Its dimension is the number of elements. Embedding provenance means the model name/version, preprocessing, normalization policy, and source revision that explain how the vector was created. A similarity function or distance rule decides what “near” means. Cosine similarity compares direction, dot product combines direction and magnitude unless vectors are normalized, and Euclidean is based on geometric distance. Normalization usually means scaling a vector to unit length. Approximate Nearest Neighbor (ANN) retrieval searches an index for likely nearest vectors without guaranteeing the exact top-k ordering; exact retrieval evaluates every candidate under the chosen metric. Recall@k is the fraction of exact top-k items recovered by ANN. It measures retrieval fidelity, not whether the documents are useful to a user.
Storage-Attached Indexing (SAI) is Cassandra
5.0's storage-integrated secondary indexing framework; vector
ANN search uses an SAI index on a
vector<float,n> column. A
coordinator is the Cassandra node handling a
particular client request. A replica stores
data for a token range according to the keyspace replication
strategy. A partition is the Cassandra
storage/routing unit selected by a partition key; the
partitioner hashes that key to a token. An
immutable SSTable is an on-disk Sorted String
Table. A consistency level (CL) states the
replica acknowledgments/responses required by an operation.
RAG (Retrieval-Augmented Generation) is an
application architecture that retrieves evidence and supplies it
to a generative model; Cassandra can be the retrieval store but
does not generate embeddings or model answers.
1. SAI turns the vector column into an ANN retrieval path
The vector index is maintained by Storage-Attached Indexing
alongside Cassandra storage. The index definition chooses one
similarity function: COSINE (default),
DOT_PRODUCT, or EUCLIDEAN. Cassandra
does not let each query silently reinterpret the same index
under another metric. To change the metric,
drop/recreate/rebuild and re-evaluate the search behavior.
# Verify the shared course lab first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -version# Recreate only if the disposable course cluster does not exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three nodes are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec -it atlasmart-cass-1 cqlsh
CREATE KEYSPACE IF NOT EXISTS atlasmart_vectorWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_vector.documents ( tenant_id text, corpus_bucket tinyint, document_id uuid, title text, body text, locale text, doc_type text, embedding_model text, embedding_version text, embedding vector<float,4>, updated_at timestamp, PRIMARY KEY ((tenant_id,corpus_bucket),document_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;CREATE INDEX IF NOT EXISTS documents_embedding_annON atlasmart_vector.documents (embedding) USING 'sai'WITH OPTIONS = {'similarity_function':'COSINE'};CREATE INDEX IF NOT EXISTS documents_locale_saiON atlasmart_vector.documents (locale) USING 'sai';CREATE INDEX IF NOT EXISTS documents_type_saiON atlasmart_vector.documents (doc_type) USING 'sai';
INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001901,'Return policy','How to return an unopened product within the return window.','en','support','atlasmart-toy','v1',[1.0,0.0,0.0,0.0],'2026-09-08T07:00:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001902,'Refund delay','Why approved refunds can take several days to appear.','en','support','atlasmart-toy','v1',[0.9,0.2,0.1,0.0],'2026-09-08T07:01:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001903,'Shipping tracker','Track a parcel after warehouse dispatch.','en','support','atlasmart-toy','v1',[0.0,1.0,0.0,0.0],'2026-09-08T07:02:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001904,'Password reset','Recover access to an AtlasMart account.','en','security','atlasmart-toy','v1',[0.0,0.0,1.0,0.0],'2026-09-08T07:03:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001905,'Invoice copy','Download a VAT invoice for an order.','en','billing','atlasmart-toy','v1',[0.0,0.0,0.0,1.0],'2026-09-08T07:04:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001906,'Exchange item','Exchange an eligible item for another size.','en','support','atlasmart-toy','v1',[0.8,0.1,0.0,0.2],'2026-09-08T07:05:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001907,'Cancel order','Cancel before warehouse fulfillment begins.','en','support','atlasmart-toy','v1',[0.7,0.0,0.2,0.1],'2026-09-08T07:06:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001908,'Delivery delay','Investigate a shipment that missed its delivery date.','en','support','atlasmart-toy','v1',[0.1,0.9,0.0,0.0],'2026-09-08T07:07:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-b',0,00000000-0000-0000-0000-000000001909,'Tenant B private return note','Private tenant B support content.','en','support','atlasmart-toy','v1',[0.99,0.01,0.0,0.0],'2026-09-08T07:08:00Z');
DESCRIBE TABLE atlasmart_vector.documents;SELECT keyspace_name,index_name,column_name,cell_count,indexed_sstable_count, is_building,is_queryable,per_column_disk_size,per_table_disk_sizeFROM system_views.indexesWHERE keyspace_name='atlasmart_vector';
Wait until documents_embedding_ann reports
is_queryable=true and
is_building=false before treating ANN queries as
ready. Creating an index after loading data may launch a
background build. Exact timing depends on data size, SSTables,
storage, compaction, and node resources.
2. First ANN query: candidate retrieval, not truth
TRACING ON;SELECT document_id,title,embedding_version, similarity_cosine(embedding,[1.0,0.1,0.0,0.0]) AS scoreFROM atlasmart_vector.documentsWHERE tenant_id='tenant-a' AND corpus_bucket=0ORDER BY embedding ANN OF [1.0,0.1,0.0,0.0]LIMIT 3;TRACING OFF;
On this transparent toy space, return/refund-oriented documents
should rank near the query. ANN is approximate: the result set
is a candidate top-k under the index metric, not a mathematical
guarantee of the exact best k. The CQL
similarity_cosine function lets you print a score
for returned rows, but it does not convert ANN into exhaustive
search.
ANN query LIMIT must be 1000 or less.
Least-similar retrieval is not supported. The vector index
works best when vector values are relatively stable; frequent
overwrites/deletes of the indexed vector can degrade search
performance.
3. Index state and physical cost
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_vector documents; donedocker exec atlasmart-cass-1 cqlsh -e "SELECT index_name,column_name,cell_count,indexed_sstable_count,is_building,is_queryable,per_column_disk_size,per_table_disk_size FROM system_views.indexes WHERE keyspace_name='atlasmart_vector';"docker exec atlasmart-cass-1 cqlsh -e "SELECT index_name,sstable_name,cell_count,per_column_disk_size,per_table_disk_size,start_token,end_token FROM system_views.sstable_indexes WHERE keyspace_name='atlasmart_vector';"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_vector.documents
The reported bytes are node-local SAI/SSTable state, not the full replicated cluster cost. Multiply nothing blindly: RF, data distribution, compaction, SSTable count, shared per-table SAI components, and transient compaction/build headroom all affect actual capacity.
4. Reversible replica failure
Pausing one replica demonstrates the distinction between
database availability and ANN quality. With RF=3 and
LOCAL_QUORUM, two live local replicas can usually
satisfy the read. That does not prove every ANN candidate set is
identical to a three-replica run, and it does not make the index
globally linearizable.
docker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status# Run the ANN query from cqlsh on node 1 while node 3 is paused.# Then restore the lab.docker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
Consistency level governs replica responses for Cassandra data operations; ANN remains approximate retrieval. Evaluate failure-mode recall/latency separately and use repair/convergence practices for replica data consistency.
5. Verification checklist
- The vector index is queryable on every node used for testing.
- The ANN query uses an explicit query vector with the schema's four dimensions.
LIMITis within the documented maximum.- Index disk/build state is captured before benchmark claims.
-
All nodes are restored to
UNafter failure injection.
Check your understanding
- Can the similarity metric be changed per ANN query without rebuilding the index?
- What does is_queryable=false mean after CREATE INDEX on existing data?
- Why is similarity_cosine on returned rows not an exact-search baseline?
- What does one paused replica test?
- What is Cassandra's documented ANN LIMIT ceiling in this version?
Review the answers
1. No. The index similarity_function is part of the index definition.
2. The index is not yet ready to serve queries; wait for the build rather than benchmarking incomplete state.
3. Because ANN already chose which rows were returned; exact evaluation must score the complete candidate corpus or a controlled ground-truth set.
4. Availability/failure behavior under the chosen RF/CL and the observed ANN result/latency—not mathematical exactness.
5. 1000.
Production judgment
Decide whether Cassandra vector search fits by measuring the whole retrieval system: corpus size and growth; vector dimension and bytes; embedding generation/update rate; embedding provenance; tenant and authorization model; query filters and bucket fanout; RF/CL and node/DC failures; ANN recall@k against a fixed exact baseline; application relevance metrics; p50/p95/p99 retrieval latency; result payload size; write amplification; SAI disk and memory footprint; SSTable count/compaction; vector overwrite/delete rate; index build/rebuild/streaming time; repair/backup/restore behavior; JVM/off-heap/chunk-cache pressure; driver timeouts/retries/idempotency; guardrails; observability; and operator skill. Do not report ANN latency without recall, and do not call recall “answer quality.”
Security is not a similarity metric. Cassandra role permissions are table/keyspace oriented, not row-level authorization; an application must derive allowed tenant/corpus scope before retrieval and encode that scope in the data model/query. Post-filtering unauthorized ANN results can leak information or reduce useful top-k results. Managed services and Cassandra-compatible APIs may differ in vector syntax, limits, index lifecycle, metrics, encryption, billing, and topology. Lesson 3 builds an exact local baseline and turns “fast search” into a recall-versus-latency measurement across cosine, normalized dot-product concepts, and Euclidean behavior.
Summary and next bridge
SAI supplies a storage-attached ANN path, observable build state, and measurable disk cost. ANN results are candidates under one index metric. Next, establish exact ground truth and evaluate recall and latency instead of celebrating one fast query.
Authoritative references
Vector search and SAI are version-sensitive. Re-check these sources when regenerating the lesson rather than freezing a 2026 assumption.