Chapter 08 · CQL Data Types, Collections, Tuples, UDTs, Static Columns, and Frozen Values
Lists, Sets, Maps: Read / Write Semantics, Collection Growth, and Tombstone Risks
Use Cassandra lists, sets, and maps only for bounded values; observe element mutations, growth limits, and the tombstones deletions and expiry can create.
Learning outcomes
AtlasMart wants compact profile preferences, labels, and a few
notification channels beside one customer row. Collections are
attractive because they keep bounded related values together,
but the same convenience becomes dangerous when developers turn
them into an ever-growing event log. This lesson treats
list, set, and map as
storage structures with mutation and tombstone consequences, not
as generic application containers.
Distinguish list ordering/duplicates, set uniqueness, and map key-value semantics.
Choose collections only when cardinality and total value size are deliberately bounded.
Observe element-level mutations in non-frozen collections and the tombstones produced by removals/expiry.
Explain list operations that may require read-before-write or positional work and prefer simpler structures when possible.
Redesign a high-cardinality relationship as clustering rows instead of an unbounded collection.
The mandatory labs use the pinned
cassandra:5.0.9 image. Java 17,
cqlsh, and nodetool are the versions
bundled by that image. The course topology is three disposable
nodes (atlasmart-cass-1..3) in cluster
atlasmart-course, datacenter dc1,
racks rack1..rack3, 16 vnodes per node,
replication factor (RF) 3, and LOCAL_QUORUM for
consistency-sensitive examples. Authentication, client TLS,
internode TLS, and remote JMX are not enabled in this isolated
learning network; production must secure those boundaries
separately. Chapter 08 uses keyspace
atlasmart_types; new tables explicitly use
UnifiedCompactionStrategy (UCS), default table TTL is zero
unless stated, and gc_grace_seconds is not
changed. Storage-Attached Indexing (SAI) and vector search are
not required. Apache Cassandra Java Driver 4.19.3 is used only
in optional decoding snippets; the mandatory path remains
free/local with cqlsh.
The commands and CQL below are documentation- and syntax-reviewed but were not executed in this generation environment. Treat output as an expected shape, then capture exact UUIDs, time values, SSTable paths, tombstone counters, and driver-decoded values on your machine. If three nodes are too heavy, use one disposable node and RF=1 to learn type/mutation semantics, but do not treat that reduced topology as evidence about RF=3 availability or repair behavior.
1. Collections are bounded denormalized values
A non-frozen collection is stored as multiple cells so Cassandra can update individual elements. That is useful for a small set of labels, a bounded map of preferences, or a short list whose positional semantics are genuinely needed. It is not a replacement for a one-to-many table. Current Cassandra documentation explicitly warns that collections are not internally paged and should not be used for unbounded growth such as all messages or all sensor events.
| Type | Semantics | Good bounded example | Failure mode |
|---|---|---|---|
set<text> |
unique elements; query output is sorted | customer labels | high-cardinality relationship hidden in one row |
map<text,text> |
unique typed keys mapped to values | small preference dictionary | per-event map key growth |
list<text> |
duplicates + positional order | short ordered fallback channels | indexed/prepend operations and large positional mutations |
2. Lab: mutate elements, then delete them
# Verify the shared course lab if it already exists.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool status# Standalone local recreation path. Skip resources that already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait for node 1 to answer before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all nodes show UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_types WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE KEYSPACE atlasmart_types"
CREATE TABLE atlasmart_types.customer_preferences ( customer_id uuid PRIMARY KEY, labels set<text>, preferences map<text,text>, fallback_channels list<text>) WITH compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_types.customer_preferences(customer_id,labels,preferences,fallback_channels)VALUES(11111111-1111-1111-1111-111111111111, {'vip','newsletter'}, {'currency':'USD','locale':'en-US'}, ['email','sms']);SELECT * FROM atlasmart_types.customer_preferencesWHERE customer_id=11111111-1111-1111-1111-111111111111;
UPDATE atlasmart_types.customer_preferencesSET labels = labels + {'fraud-reviewed'}, preferences['currency'] = 'AZN', fallback_channels = fallback_channels + ['push']WHERE customer_id=11111111-1111-1111-1111-111111111111;UPDATE atlasmart_types.customer_preferencesSET labels = labels - {'newsletter'}WHERE customer_id=11111111-1111-1111-1111-111111111111;DELETE preferences['locale']FROM atlasmart_types.customer_preferencesWHERE customer_id=11111111-1111-1111-1111-111111111111;SELECT * FROM atlasmart_types.customer_preferencesWHERE customer_id=11111111-1111-1111-1111-111111111111;
The final row should contain the new label, no
newsletter, an updated currency, no
locale map entry, and the appended channel. The
removed elements are not simply erased from every
replica/SSTable immediately: Cassandra records deletion markers
so replicas can converge. TTL expiration has the same tombstone
consequence.
3. Observe tombstone evidence without manufacturing a huge workload
docker exec atlasmart-cass-1 nodetool flush atlasmart_types customer_preferences# Read the row so read-side tombstone statistics have a chance to update.docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_types.customer_preferences WHERE customer_id=11111111-1111-1111-1111-111111111111;"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_types customer_preferences# Optional low-level evidence; exact SSTable path/output varies by 5.0 patch/format.docker exec atlasmart-cass-1 bash -lc 'F=$(find /var/lib/cassandra/data/atlasmart_types/customer_preferences-* -name "*Data.db" | head -1); echo "$F"; test -n "$F" && sstabledump "$F" | grep -E "deletion|tombstone|markedForDeleteAt" | head -20 || true'
tablestats is operational evidence, not a proof
that every element tombstone is currently visible in one
SSTable. The optional sstabledump path is
deliberately version-sensitive and should be treated as forensic
learning output. The correct production response to collection
tombstones is usually better modeling and bounded mutation
patterns—not manual compaction as a routine cure.
That collection grows with business history, is not internally
paged, and turns one logical row into an ever-growing
read/mutation unit. Repair it with a table such as
orders_by_customer_bucket, using a bounded
partition key and clustering columns so reads can page and
retention can be modeled explicitly.
4. Production judgment
Collection choice changes write amplification, read payload, tombstone density, compaction work, and application concurrency. A set is usually preferable to a list when order/duplicates are not required. Maps are excellent for small keyed attributes but dangerous when keys become an event stream. Concurrent list position changes are particularly difficult to reason about because positional operations can involve read-before-write behavior. Set/map element updates are naturally idempotent when the new value is deterministic; list appends may not be safe to retry blindly. Model limits in business terms—maximum labels, channels, or settings per entity—then monitor actual cardinality rather than relying on theoretical guardrail maxima.
Verification checklist
- The collection cardinalities are bounded by an explicit product invariant.
- Element additions/removals produce the expected before/after row.
- You can explain why deletions/TTL create tombstones instead of immediate physical erasure.
- The optional SSTable/tombstone output is labeled version-dependent.
- The unbounded relationship redesign uses clustering rows instead of a collection.
Check your understanding
- Why are collections unsuitable for an unlimited event history?
- What is the key semantic difference between set and list?
- What does removing a collection element create?
- Why can list operations be operationally riskier?
- What is the modeling repair for thousands of child items?
Review the answers
1. They are read as a collection and are not internally paged; unbounded growth makes the row/mutation unit increasingly expensive.
2. Sets enforce uniqueness; lists preserve duplicates and positional order.
3. A deletion marker/tombstone that participates in normal replica convergence and later compaction.
4. Some positional/prepend operations can require read-before-write work, and retries/ordering are harder to reason about.
5. Use a separate query table with partition/clustering keys and explicit bucketing/retention.
# Destructive only to this disposable chapter keyspace.docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_types;"# Keep shared course containers for Chapter 09, or remove them for a full reset:# docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3# docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data# docker network rm atlasmart-cassandra
Summary and next bridge
Collections are useful only when their growth and mutation model remain small and predictable. Next we model richer structured values with tuples and user-defined types, where positional versus named semantics and schema evolution matter more than collection cardinality.
Authoritative references
- CQL data types — scalar, collection, tuple, user-defined type, frozen, and literal semantics.
- Creating collections — bounded collection guidance and current collection guardrails.
- CQL data definition — static-column behavior, table/type definitions, and schema restrictions.
- CREATE TABLE reference — frozen/non-frozen UDT and static-column examples.
- Tombstones — deletion markers, grace, reads, and compaction implications.
- Apache Cassandra downloads — current server and Java-driver release baselines.