Chapter 19 · Vector Search in Cassandra 5.0
Cosine / Dot / Euclidean Similarity Concepts, Normalization, Recall, and Latency
Compare cosine, dot-product, and Euclidean geometry; build exact ground truth; compute recall@k; and measure p50/p95/p99 ANN latency.
Learning outcomes
AtlasMart's prototype returns plausible documents in a few milliseconds, and someone calls the project finished. That is not an evaluation. A vector system must answer two separate questions: how quickly does ANN return candidates, and how often do those candidates contain the exact nearest neighbors under the chosen metric? This lesson builds a fixed ground truth and quantifies both.
Explain cosine, dot product, Euclidean distance/similarity and how normalization changes their interpretation.
Build a deterministic exact top-k baseline outside Cassandra for the small teaching corpus.
Compute recall@k by comparing Cassandra ANN IDs with exact IDs.
Measure p50/p95/p99 application-level latency with a maintained driver while recording warmup/concurrency context.
Avoid metric mixing and latency-only vector benchmarks.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 container, Java 17 inside the
image, cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. The chapter keyspace is
atlasmart_vector with
NetworkTopologyStrategy, RF=3, and normal
reads/writes at LOCAL_QUORUM. Tables explicitly
use UnifiedCompactionStrategy (UCS), no default TTL, and the
Cassandra default gc_grace_seconds unless a
lesson says otherwise. Authentication, client TLS, internode
TLS, and remote JMX are disabled only inside this isolated
local learning network; production systems must enforce
appropriate authentication, authorization, encryption, and
network boundaries. Apache Cassandra Java Driver 4.19.3 is
optional for application-level latency/vector-codec examples.
The mandatory embedding data is a deterministic
four-dimensional pedagogical fixture—free, local, and
intentionally not a production semantic model. No paid
embedding or LLM API is required.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and mental model
A vector is an ordered fixed-length array of numbers. An embedding is a vector produced by a model or deterministic transformation to represent an object in a geometric space. Its dimension is the number of elements. Embedding provenance means the model name/version, preprocessing, normalization policy, and source revision that explain how the vector was created. A similarity function or distance rule decides what “near” means. Cosine similarity compares direction, dot product combines direction and magnitude unless vectors are normalized, and Euclidean is based on geometric distance. Normalization usually means scaling a vector to unit length. Approximate Nearest Neighbor (ANN) retrieval searches an index for likely nearest vectors without guaranteeing the exact top-k ordering; exact retrieval evaluates every candidate under the chosen metric. Recall@k is the fraction of exact top-k items recovered by ANN. It measures retrieval fidelity, not whether the documents are useful to a user.
Storage-Attached Indexing (SAI) is Cassandra
5.0's storage-integrated secondary indexing framework; vector
ANN search uses an SAI index on a
vector<float,n> column. A
coordinator is the Cassandra node handling a
particular client request. A replica stores
data for a token range according to the keyspace replication
strategy. A partition is the Cassandra
storage/routing unit selected by a partition key; the
partitioner hashes that key to a token. An
immutable SSTable is an on-disk Sorted String
Table. A consistency level (CL) states the
replica acknowledgments/responses required by an operation.
RAG (Retrieval-Augmented Generation) is an
application architecture that retrieves evidence and supplies it
to a generative model; Cassandra can be the retrieval store but
does not generate embeddings or model answers.
1. Same vectors, different geometry
| Metric | Interpretation | Normalization sensitivity | Cassandra vector-index option |
|---|---|---|---|
| Cosine | directional similarity | magnitude largely removed by formula | COSINE (default) |
| Dot product | alignment × magnitude | for unit vectors ranking aligns with cosine | DOT_PRODUCT |
| Euclidean | geometric distance / mapped similarity | scale directly affects distance | EUCLIDEAN |
Current Cassandra guidance recommends considering dot product for normalized embeddings because cosine and dot-product rankings align on unit vectors and dot product can be cheaper. But that is conditional: using dot product on embeddings whose magnitude carries uncontrolled scale can silently change the ranking. The embedding model's contract—not a generic tuning rule—must determine preprocessing.
from math import sqrtdef unit(v): n = sqrt(sum(x*x for x in v)) return [x/n for x in v]def dot(a,b): return sum(x*y for x,y in zip(a,b))def cosine(a,b): return dot(a,b)/(sqrt(dot(a,a))*sqrt(dot(b,b)))def euclid(a,b): return sqrt(sum((x-y)**2 for x,y in zip(a,b)))q = [1.0,0.1,0.0,0.0]a = [1.0,0.0,0.0,0.0]b = [0.9,0.2,0.1,0.0]print('cos(a)', cosine(q,a), 'cos(b)', cosine(q,b))print('dot normalized a/b', dot(unit(q),unit(a)), dot(unit(q),unit(b)))print('euclidean a/b', euclid(q,a), euclid(q,b))
2. Exact ground truth for the fixed eight-document corpus
CREATE KEYSPACE IF NOT EXISTS atlasmart_vectorWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_vector.documents ( tenant_id text, corpus_bucket tinyint, document_id uuid, title text, body text, locale text, doc_type text, embedding_model text, embedding_version text, embedding vector<float,4>, updated_at timestamp, PRIMARY KEY ((tenant_id,corpus_bucket),document_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;CREATE INDEX IF NOT EXISTS documents_embedding_annON atlasmart_vector.documents (embedding) USING 'sai'WITH OPTIONS = {'similarity_function':'COSINE'};CREATE INDEX IF NOT EXISTS documents_locale_saiON atlasmart_vector.documents (locale) USING 'sai';CREATE INDEX IF NOT EXISTS documents_type_saiON atlasmart_vector.documents (doc_type) USING 'sai';
INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001901,'Return policy','How to return an unopened product within the return window.','en','support','atlasmart-toy','v1',[1.0,0.0,0.0,0.0],'2026-09-08T07:00:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001902,'Refund delay','Why approved refunds can take several days to appear.','en','support','atlasmart-toy','v1',[0.9,0.2,0.1,0.0],'2026-09-08T07:01:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001903,'Shipping tracker','Track a parcel after warehouse dispatch.','en','support','atlasmart-toy','v1',[0.0,1.0,0.0,0.0],'2026-09-08T07:02:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',0,00000000-0000-0000-0000-000000001904,'Password reset','Recover access to an AtlasMart account.','en','security','atlasmart-toy','v1',[0.0,0.0,1.0,0.0],'2026-09-08T07:03:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001905,'Invoice copy','Download a VAT invoice for an order.','en','billing','atlasmart-toy','v1',[0.0,0.0,0.0,1.0],'2026-09-08T07:04:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001906,'Exchange item','Exchange an eligible item for another size.','en','support','atlasmart-toy','v1',[0.8,0.1,0.0,0.2],'2026-09-08T07:05:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001907,'Cancel order','Cancel before warehouse fulfillment begins.','en','support','atlasmart-toy','v1',[0.7,0.0,0.2,0.1],'2026-09-08T07:06:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-a',1,00000000-0000-0000-0000-000000001908,'Delivery delay','Investigate a shipment that missed its delivery date.','en','support','atlasmart-toy','v1',[0.1,0.9,0.0,0.0],'2026-09-08T07:07:00Z');INSERT INTO atlasmart_vector.documents (tenant_id,corpus_bucket,document_id,title,body,locale,doc_type,embedding_model,embedding_version,embedding,updated_at) VALUES ('tenant-b',0,00000000-0000-0000-0000-000000001909,'Tenant B private return note','Private tenant B support content.','en','support','atlasmart-toy','v1',[0.99,0.01,0.0,0.0],'2026-09-08T07:08:00Z');
from math import sqrtrows = { '1901':[1.0,0.0,0.0,0.0], '1902':[0.9,0.2,0.1,0.0], '1903':[0.0,1.0,0.0,0.0], '1904':[0.0,0.0,1.0,0.0], '1905':[0.0,0.0,0.0,1.0], '1906':[0.8,0.1,0.0,0.2], '1907':[0.7,0.0,0.2,0.1], '1908':[0.1,0.9,0.0,0.0],}q=[1.0,0.1,0.0,0.0]def cosine(a,b): d=sum(x*y for x,y in zip(a,b)) return d/(sqrt(sum(x*x for x in a))*sqrt(sum(x*x for x in b)))exact=[k for k,_ in sorted(rows.items(), key=lambda kv: cosine(q,kv[1]), reverse=True)[:3]]print('exact@3 =', exact)# Replace with the last four digits returned by your Cassandra ANN query.ann=['1901','1902','1906']recall=len(set(exact)&set(ann))/len(exact)print('recall@3 =', recall)
This exact script is intentionally tiny enough to score every row. Production corpora need a controlled evaluation sample or an exact reference system because exhaustive scoring over millions of vectors defeats the purpose of ANN. Keep the query set fixed when comparing model/index changes.
-- Query each authorized bucket separately; merge the small overfetch sets in the application.SELECT document_id,title, similarity_cosine(embedding,[1.0,0.1,0.0,0.0]) AS scoreFROM atlasmart_vector.documentsWHERE tenant_id='tenant-a' AND corpus_bucket=0ORDER BY embedding ANN OF [1.0,0.1,0.0,0.0]LIMIT 3;SELECT document_id,title, similarity_cosine(embedding,[1.0,0.1,0.0,0.0]) AS scoreFROM atlasmart_vector.documentsWHERE tenant_id='tenant-a' AND corpus_bucket=1ORDER BY embedding ANN OF [1.0,0.1,0.0,0.0]LIMIT 3;
The partition key includes corpus_bucket, so the
lab keeps routing/fanout explicit: search each authorized
bucket, overfetch a bounded k, then merge by the same metric.
This avoids hiding multi-partition fanout behind an uncertain
planner shortcut and mirrors a production design that can
budget parallelism and partial failures.
3. Measure latency with a driver, not cqlsh process startup
Launching docker exec ... cqlsh for every sample
mostly measures process/container overhead. For application
latency use a persistent driver session. Apache Java Driver
4.19.3 supports CQL vectors through CqlVector. Warm
up first, record concurrency and payload size, then report
percentiles rather than the minimum.
<dependency> <groupId>org.apache.cassandra</groupId> <artifactId>java-driver-core</artifactId> <version>4.19.3</version></dependency>
import com.datastax.oss.driver.api.core.CqlSession;import com.datastax.oss.driver.api.core.data.CqlVector;import com.datastax.oss.driver.api.core.cql.*;import java.util.*;try (CqlSession s = CqlSession.builder().withLocalDatacenter("dc1").build()) { PreparedStatement ps = s.prepare("SELECT document_id,title FROM atlasmart_vector.documents " + "WHERE tenant_id='tenant-a' AND corpus_bucket=0 " + "ORDER BY embedding ANN OF ? LIMIT 3"); CqlVector<Float> q = CqlVector.newInstance(1.0f,0.1f,0.0f,0.0f); BoundStatement bs = ps.bind(q).setConsistencyLevel(com.datastax.oss.driver.api.core.ConsistencyLevel.LOCAL_QUORUM); for (int i=0;i<50;i++) s.execute(bs); // warmup long[] us = new long[500]; for (int i=0;i<us.length;i++) { long t=System.nanoTime(); s.execute(bs); us[i]=(System.nanoTime()-t)/1000; } Arrays.sort(us); System.out.printf("p50=%dus p95=%dus p99=%dus%n", us[249],us[474],us[494]);}
Those percentiles are client-observed latency for this machine and fixture. They are not portable Cassandra constants. Repeat with representative dimensions, corpus size, filters, concurrency, RF/CL, compaction state, and cold/warm cache conditions.
4. Wrong benchmark: change metric and dataset at the same time
If v2 embeddings are normalized and v1 are not, then switching from cosine/v1 to dot-product/v2 changes two variables. You cannot attribute recall or latency change to the metric. Hold corpus, query set, embedding space, k, filters, and load constant when testing index settings. When evaluating a new model, call it a model migration experiment and compare end-to-end relevance as well as ANN recall.
Check your understanding
- Why can dot product and cosine produce similar rankings for unit vectors?
- What does ANN recall@k require?
- Why is cqlsh-per-query timing misleading?
- Does high recall@k prove RAG answer quality?
- What must stay fixed when comparing similarity-function performance?
Review the answers
1. For normalized vectors, dot product equals cosine similarity.
2. A fixed query plus an exact/ground-truth top-k set to compare against.
3. Process startup, Docker exec, parsing, and connection setup dominate the measurement.
4. No. Retrieval recall is one layer; document relevance, prompt construction, model behavior, grounding, and answer evaluation are separate.
5. At minimum corpus/query vectors, preprocessing, k, filters, workload, topology, and benchmark method.
Production judgment
Decide whether Cassandra vector search fits by measuring the whole retrieval system: corpus size and growth; vector dimension and bytes; embedding generation/update rate; embedding provenance; tenant and authorization model; query filters and bucket fanout; RF/CL and node/DC failures; ANN recall@k against a fixed exact baseline; application relevance metrics; p50/p95/p99 retrieval latency; result payload size; write amplification; SAI disk and memory footprint; SSTable count/compaction; vector overwrite/delete rate; index build/rebuild/streaming time; repair/backup/restore behavior; JVM/off-heap/chunk-cache pressure; driver timeouts/retries/idempotency; guardrails; observability; and operator skill. Do not report ANN latency without recall, and do not call recall “answer quality.”
Security is not a similarity metric. Cassandra role permissions are table/keyspace oriented, not row-level authorization; an application must derive allowed tenant/corpus scope before retrieval and encode that scope in the data model/query. Post-filtering unauthorized ANN results can leak information or reduce useful top-k results. Managed services and Cassandra-compatible APIs may differ in vector syntax, limits, index lifecycle, metrics, encryption, billing, and topology. Lesson 4 adds metadata/tenant predicates and shows why authorization scope and query-first bucket design must constrain vector retrieval before results reach the application.
Summary and next bridge
Vector evaluation needs geometry plus evidence: choose one metric contract, establish exact ground truth, compute recall@k, and report latency distributions under controlled conditions. Next, combine ANN with metadata and tenant scope without turning post-filtering into a security bug.
Authoritative references
Vector search and SAI are version-sensitive. Re-check these sources when regenerating the lesson rather than freezing a 2026 assumption.