Chapter 19 · Vector Search in Cassandra 5.0

Capacity, Index Build / Streaming, Embedding Updates, Evaluation, and AI / RAG Architecture Tradeoffs

Plan vector capacity, build/rebuild/streaming, embedding-version migration, evaluation gates, rollback, and Cassandra's precise role in RAG.

Advanced130–180 minutesLifecycle + RAG architecture labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 optional · SAI Vector Search · RF=3 · LOCAL_QUORUMLast reviewed: September 2026

Learning outcomes

AtlasMart's prototype is ready for a production review. The final questions are operational: how much index capacity is required, what happens during bootstrap/rebuild, how embeddings are updated without mixing spaces, how recall/relevance regressions block rollout, and where Cassandra stops in a RAG architecture.

01

Estimate vector/SSTable/SAI storage and memory cost from measured node-local evidence rather than one magic percentage.

02

Observe index build/queryable state and understand SAI zero-copy streaming/rebuild behavior.

03

Design a dual-version embedding migration with evaluation gates and rollback.

04

Separate ANN recall from product relevance and RAG answer quality.

05

Define Cassandra's role and non-role in an AI/RAG architecture, including authorization and model-generation boundaries.

Chapter 19 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 container, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. The chapter keyspace is atlasmart_vector with NetworkTopologyStrategy, RF=3, and normal reads/writes at LOCAL_QUORUM. Tables explicitly use UnifiedCompactionStrategy (UCS), no default TTL, and the Cassandra default gc_grace_seconds unless a lesson says otherwise. Authentication, client TLS, internode TLS, and remote JMX are disabled only inside this isolated local learning network; production systems must enforce appropriate authentication, authorization, encryption, and network boundaries. Apache Cassandra Java Driver 4.19.3 is optional for application-level latency/vector-codec examples. The mandatory embedding data is a deterministic four-dimensional pedagogical fixture—free, local, and intentionally not a production semantic model. No paid embedding or LLM API is required.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and mental model

A vector is an ordered fixed-length array of numbers. An embedding is a vector produced by a model or deterministic transformation to represent an object in a geometric space. Its dimension is the number of elements. Embedding provenance means the model name/version, preprocessing, normalization policy, and source revision that explain how the vector was created. A similarity function or distance rule decides what “near” means. Cosine similarity compares direction, dot product combines direction and magnitude unless vectors are normalized, and Euclidean is based on geometric distance. Normalization usually means scaling a vector to unit length. Approximate Nearest Neighbor (ANN) retrieval searches an index for likely nearest vectors without guaranteeing the exact top-k ordering; exact retrieval evaluates every candidate under the chosen metric. Recall@k is the fraction of exact top-k items recovered by ANN. It measures retrieval fidelity, not whether the documents are useful to a user.

Storage-Attached Indexing (SAI) is Cassandra 5.0's storage-integrated secondary indexing framework; vector ANN search uses an SAI index on a vector<float,n> column. A coordinator is the Cassandra node handling a particular client request. A replica stores data for a token range according to the keyspace replication strategy. A partition is the Cassandra storage/routing unit selected by a partition key; the partitioner hashes that key to a token. An immutable SSTable is an on-disk Sorted String Table. A consistency level (CL) states the replica acknowledgments/responses required by an operation. RAG (Retrieval-Augmented Generation) is an application architecture that retrieves evidence and supplies it to a generative model; Cassandra can be the retrieval store but does not generate embeddings or model answers.

1. Capacity: vectors, replicas, indexes, and transient headroom

A raw vector<float,d> has roughly 4 × d bytes of float payload before row/storage/index overhead. A 1536-dimensional float vector therefore carries about 6 KiB of raw coordinates per row, while the lab's four-dimensional vector carries only 16 bytes. Production capacity must add primary-key/row metadata, text/metadata columns, SAI vector/index structures, RF replication, SSTable compression, compaction/rebuild headroom, snapshots/backups, and free-space reserve.

CQL · create the bounded AtlasMart vector schema
CREATE KEYSPACE IF NOT EXISTS atlasmart_vectorWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_vector.documents (    tenant_id text,    corpus_bucket tinyint,    document_id uuid,    title text,    body text,    locale text,    doc_type text,    embedding_model text,    embedding_version text,    embedding vector<float,4>,    updated_at timestamp,    PRIMARY KEY ((tenant_id,corpus_bucket),document_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;CREATE INDEX IF NOT EXISTS documents_embedding_annON atlasmart_vector.documents (embedding) USING 'sai'WITH OPTIONS = {'similarity_function':'COSINE'};CREATE INDEX IF NOT EXISTS documents_locale_saiON atlasmart_vector.documents (locale) USING 'sai';CREATE INDEX IF NOT EXISTS documents_type_saiON atlasmart_vector.documents (doc_type) USING 'sai';
CQL · load deterministic four-dimensional teaching vectors
INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001901,'Return policy','How to return an unopened product within the return window.','en','support','atlasmart-toy','v1',[1.0,0.0,0.0,0.0],'2026-09-08T07:00:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001902,'Refund delay','Why approved refunds can take several days to appear.','en','support','atlasmart-toy','v1',[0.9,0.2,0.1,0.0],'2026-09-08T07:01:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001903,'Shipping tracker','Track a parcel after warehouse dispatch.','en','support','atlasmart-toy','v1',[0.0,1.0,0.0,0.0],'2026-09-08T07:02:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001904,'Password reset','Recover access to an AtlasMart account.','en','security','atlasmart-toy','v1',[0.0,0.0,1.0,0.0],'2026-09-08T07:03:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001905,'Invoice copy','Download a VAT invoice for an order.','en','billing','atlasmart-toy','v1',[0.0,0.0,0.0,1.0],'2026-09-08T07:04:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001906,'Exchange item','Exchange an eligible item for another size.','en','support','atlasmart-toy','v1',[0.8,0.1,0.0,0.2],'2026-09-08T07:05:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001907,'Cancel order','Cancel before warehouse fulfillment begins.','en','support','atlasmart-toy','v1',[0.7,0.0,0.2,0.1],'2026-09-08T07:06:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001908,'Delivery delay','Investigate a shipment that missed its delivery date.','en','support','atlasmart-toy','v1',[0.1,0.9,0.0,0.0],'2026-09-08T07:07:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-b',0,00000000-0000-0000-0000-000000001909,'Tenant B private return note','Private tenant B support content.','en','support','atlasmart-toy','v1',[0.99,0.01,0.0,0.0],'2026-09-08T07:08:00Z');
CQL · inspect SAI vector build/queryable and disk state
DESCRIBE TABLE atlasmart_vector.documents;SELECT keyspace_name,index_name,column_name,cell_count,indexed_sstable_count,       is_building,is_queryable,per_column_disk_size,per_table_disk_sizeFROM system_views.indexesWHERE keyspace_name='atlasmart_vector';
bash · capture table/index/streaming/compaction evidence
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_vector.documentsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool netstatsdocker exec atlasmart-cass-1 sh -lc "du -sh /var/lib/cassandra/data/atlasmart_vector/* 2>/dev/null || true"

per_column_disk_size and per_table_disk_size are node-local SAI component evidence. Measure after representative flush/compaction state; do not multiply a tiny-table ratio into a petabyte forecast.

2. Build, rebuild, streaming, and readiness

Creating an SAI index on existing SSTables launches a build. During a drop/recreate, queries that do not depend on the index can continue, but the vector index itself is not usable until it reports queryable. Cassandra 5.0 SAI is compatible with zero-copy streaming: when that path is eligible during topology streaming, index components can move with SSTables rather than being rebuilt independently on the receiver. If an index must be rebuilt, nodetool rebuild_index is the supported operator command.

CQL · readiness gate before serving ANN traffic
SELECT index_name,is_building,is_queryable,indexed_sstable_count,       per_column_disk_size,per_table_disk_sizeFROM system_views.indexesWHERE keyspace_name='atlasmart_vector';
bash · rebuild syntax and streaming observation
# Run only in the disposable lab when intentionally testing rebuild cost.docker exec atlasmart-cass-1 nodetool rebuild_index atlasmart_vector documents documents_embedding_ann# Observe build and topology/network activity in separate terminals.docker exec atlasmart-cass-1 nodetool netstatsdocker exec atlasmart-cass-1 cqlsh -e "SELECT index_name,is_building,is_queryable,indexed_sstable_count FROM system_views.indexes WHERE keyspace_name='atlasmart_vector';"
Do not turn rebuild into a casual tuning step.

Rebuild consumes CPU, disk I/O, memory/cache, and time. Measure it under representative SSTable/data volume and preserve enough headroom. During topology change, monitor streaming plus compaction/index state together.

3. Embedding migration: dual space, evaluate, cut over, rollback

Updating every vector in place to a new model can mix old/new spaces while traffic is live and causes sustained vector-index update work. A safer pattern writes v2 into a separate versioned table/index (or otherwise isolates the new space), backfills, waits for queryability, evaluates the same fixed query/ground-truth set, shadows production queries, then cuts over. Keep v1 until rollback criteria expire.

CQL · explicit v2 table instead of semantic in-place mutation
CREATE TABLE IF NOT EXISTS atlasmart_vector.documents_v2 (  tenant_id text,  corpus_bucket tinyint,  document_id uuid,  title text,  locale text,  doc_type text,  embedding_model text,  embedding_version text,  embedding vector<float,4>,  updated_at timestamp,  PRIMARY KEY ((tenant_id,corpus_bucket),document_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CREATE INDEX IF NOT EXISTS documents_v2_embedding_annON atlasmart_vector.documents_v2 (embedding) USING 'sai'WITH OPTIONS = {'similarity_function':'COSINE'};CREATE INDEX IF NOT EXISTS documents_v2_locale_sai ON atlasmart_vector.documents_v2 (locale) USING 'sai';CREATE INDEX IF NOT EXISTS documents_v2_type_sai ON atlasmart_vector.documents_v2 (doc_type) USING 'sai';
Gate v1 baseline v2 candidate Decision
index readiness queryable must be queryable no traffic before ready
recall@k fixed ground truth same query set must meet retrieval target
business relevance human/labeled judgments same rubric ANN recall alone insufficient
p95/p99 latency representative load same load/topology respect SLO budget
disk/write cost measured measured capacity/headroom acceptable
failure mode node/DC drill same drill no hidden availability regression

4. RAG architecture: Cassandra retrieves; the application governs and generates

A robust Retrieval-Augmented Generation pipeline has separate stages: authenticate the caller; derive authorized corpus scope; embed the query with the same compatible model space; retrieve ANN candidates from Cassandra with structured filters; optionally rerank; fetch source text/metadata; construct a prompt/context with provenance; call a local or hosted generative model; validate/cite the answer; log evaluation/feedback. Cassandra handles durable distributed storage, structured metadata, SAI filtering, and vector ANN retrieval. It does not produce embeddings, rerank with a cross-encoder, enforce arbitrary row-level policy, create prompts, or generate the final answer.

text · RAG control-flow contract
request  -> authenticate user/service  -> derive allowed tenant/corpus/buckets  -> generate query embedding (same model/version/normalization)  -> Cassandra filtered ANN retrieval  -> optional rerank / exact rescoring  -> fetch source passages + provenance  -> prompt/context construction  -> local or hosted LLM generation  -> grounding/citation/policy checks  -> response + telemetry + evaluation
Wrong architecture: retrieve globally, then remove unauthorized chunks.

Authorization must constrain candidate retrieval before content leaves the authorized boundary. A paid embedding or LLM API can be an optional deployment choice, but the course lab remains free/local through deterministic vectors and retrieval-only evaluation.

5. Final Chapter 19 acceptance test

bash · final observable-state snapshot
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "SELECT index_name,column_name,is_building,is_queryable,indexed_sstable_count,per_column_disk_size,per_table_disk_size FROM system_views.indexes WHERE keyspace_name='atlasmart_vector';"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_vector.documentsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool netstats
  • Version/topology/RF/CL are recorded.
  • Embedding model/version/normalization contract is documented.
  • ANN index metric and readiness are recorded.
  • Recall@k uses fixed exact ground truth, while relevance is evaluated separately.
  • p50/p95/p99 benchmark context is documented.
  • Authorization constrains retrieval before ANN results leave the data layer.
  • v2 migration has a rollback plan; v1 is not destroyed before acceptance.

Check your understanding

  1. What is the first readiness gate after building/rebuilding a vector index?
  2. Why can in-place embedding overwrite be operationally/semantically risky?
  3. Does zero-copy streaming mean no monitoring is needed?
  4. What is the difference between recall@k and business relevance?
  5. What does Cassandra do in RAG?
Review the answers

1. Verify is_queryable=true and is_building=false on the relevant nodes before serving ANN traffic.

2. It can mix incompatible spaces during rollout and causes sustained index update work; dual-version isolation makes evaluation and rollback explicit.

3. No. Streaming still consumes network/disk/CPU resources and must be monitored with topology, compaction, index readiness, and latency signals.

4. Recall@k compares ANN to exact neighbors under a metric; relevance judges whether retrieved items satisfy the application's information need.

5. It can store metadata/content/vectors and perform structured/ANN retrieval; embedding generation, authorization policy, reranking, prompting, generation, and answer evaluation belong to other layers.

Production judgment

Decide whether Cassandra vector search fits by measuring the whole retrieval system: corpus size and growth; vector dimension and bytes; embedding generation/update rate; embedding provenance; tenant and authorization model; query filters and bucket fanout; RF/CL and node/DC failures; ANN recall@k against a fixed exact baseline; application relevance metrics; p50/p95/p99 retrieval latency; result payload size; write amplification; SAI disk and memory footprint; SSTable count/compaction; vector overwrite/delete rate; index build/rebuild/streaming time; repair/backup/restore behavior; JVM/off-heap/chunk-cache pressure; driver timeouts/retries/idempotency; guardrails; observability; and operator skill. Do not report ANN latency without recall, and do not call recall “answer quality.”

Security is not a similarity metric. Cassandra role permissions are table/keyspace oriented, not row-level authorization; an application must derive allowed tenant/corpus scope before retrieval and encode that scope in the data model/query. Post-filtering unauthorized ANN results can leak information or reduce useful top-k results. Managed services and Cassandra-compatible APIs may differ in vector syntax, limits, index lifecycle, metrics, encryption, billing, and topology. Chapter 20 leaves vector search and examines materialized views and explicit derived-table alternatives—another case where automatic derivation must be weighed against ownership, failure, and operational control.

Summary and next bridge

Cassandra 5.0 vector search is a storage-integrated ANN capability, not an AI stack by itself. Production readiness requires fixed embedding provenance, measurable recall and latency, authorized structured scope, capacity/build/streaming evidence, and a versioned migration strategy. Chapter 20 applies the same mechanism-first discipline to materialized views and derived-table choices.

Authoritative references

Vector search and SAI are version-sensitive. Re-check these sources when regenerating the lesson rather than freezing a 2026 assumption.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.