Chapter 08 · CQL Data Types, Collections, Tuples, UDTs, Static Columns, and Frozen Values

Lists, Sets, Maps: Read / Write Semantics, Collection Growth, and Tombstone Risks

Use Cassandra lists, sets, and maps only for bounded values; observe element mutations, growth limits, and the tombstones deletions and expiry can create.

Intermediate90–120 minutesCollection mutation + tombstone labApache Cassandra 5.0.9 · cqlsh/nodetool · Java Driver 4.19.3 optional · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart wants compact profile preferences, labels, and a few notification channels beside one customer row. Collections are attractive because they keep bounded related values together, but the same convenience becomes dangerous when developers turn them into an ever-growing event log. This lesson treats list, set, and map as storage structures with mutation and tombstone consequences, not as generic application containers.

01

Distinguish list ordering/duplicates, set uniqueness, and map key-value semantics.

02

Choose collections only when cardinality and total value size are deliberately bounded.

03

Observe element-level mutations in non-frozen collections and the tombstones produced by removals/expiry.

04

Explain list operations that may require read-before-write or positional work and prefer simpler structures when possible.

05

Redesign a high-cardinality relationship as clustering rows instead of an unbounded collection.

Chapter 08 lab baseline

The mandatory labs use the pinned cassandra:5.0.9 image. Java 17, cqlsh, and nodetool are the versions bundled by that image. The course topology is three disposable nodes (atlasmart-cass-1..3) in cluster atlasmart-course, datacenter dc1, racks rack1..rack3, 16 vnodes per node, replication factor (RF) 3, and LOCAL_QUORUM for consistency-sensitive examples. Authentication, client TLS, internode TLS, and remote JMX are not enabled in this isolated learning network; production must secure those boundaries separately. Chapter 08 uses keyspace atlasmart_types; new tables explicitly use UnifiedCompactionStrategy (UCS), default table TTL is zero unless stated, and gc_grace_seconds is not changed. Storage-Attached Indexing (SAI) and vector search are not required. Apache Cassandra Java Driver 4.19.3 is used only in optional decoding snippets; the mandatory path remains free/local with cqlsh.

Execution disclosure and resource path

The commands and CQL below are documentation- and syntax-reviewed but were not executed in this generation environment. Treat output as an expected shape, then capture exact UUIDs, time values, SSTable paths, tombstone counters, and driver-decoded values on your machine. If three nodes are too heavy, use one disposable node and RF=1 to learn type/mutation semantics, but do not treat that reduced topology as evidence about RF=3 availability or repair behavior.

1. Collections are bounded denormalized values

A non-frozen collection is stored as multiple cells so Cassandra can update individual elements. That is useful for a small set of labels, a bounded map of preferences, or a short list whose positional semantics are genuinely needed. It is not a replacement for a one-to-many table. Current Cassandra documentation explicitly warns that collections are not internally paged and should not be used for unbounded growth such as all messages or all sensor events.

Type Semantics Good bounded example Failure mode
set<text> unique elements; query output is sorted customer labels high-cardinality relationship hidden in one row
map<text,text> unique typed keys mapped to values small preference dictionary per-event map key growth
list<text> duplicates + positional order short ordered fallback channels indexed/prepend operations and large positional mutations

2. Lab: mutate elements, then delete them

bash · verify or recreate the disposable course cluster
# Verify the shared course lab if it already exists.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool status# Standalone local recreation path. Skip resources that already exist.docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait for node 1 to answer before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all nodes show UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_types WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE KEYSPACE atlasmart_types"
sql · create a bounded preferences row
CREATE TABLE atlasmart_types.customer_preferences (    customer_id uuid PRIMARY KEY,    labels set<text>,    preferences map<text,text>,    fallback_channels list<text>) WITH compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_types.customer_preferences(customer_id,labels,preferences,fallback_channels)VALUES(11111111-1111-1111-1111-111111111111, {'vip','newsletter'}, {'currency':'USD','locale':'en-US'}, ['email','sms']);SELECT * FROM atlasmart_types.customer_preferencesWHERE customer_id=11111111-1111-1111-1111-111111111111;
sql · perform element-level mutations and removals
UPDATE atlasmart_types.customer_preferencesSET labels = labels + {'fraud-reviewed'},    preferences['currency'] = 'AZN',    fallback_channels = fallback_channels + ['push']WHERE customer_id=11111111-1111-1111-1111-111111111111;UPDATE atlasmart_types.customer_preferencesSET labels = labels - {'newsletter'}WHERE customer_id=11111111-1111-1111-1111-111111111111;DELETE preferences['locale']FROM atlasmart_types.customer_preferencesWHERE customer_id=11111111-1111-1111-1111-111111111111;SELECT * FROM atlasmart_types.customer_preferencesWHERE customer_id=11111111-1111-1111-1111-111111111111;

The final row should contain the new label, no newsletter, an updated currency, no locale map entry, and the appended channel. The removed elements are not simply erased from every replica/SSTable immediately: Cassandra records deletion markers so replicas can converge. TTL expiration has the same tombstone consequence.

3. Observe tombstone evidence without manufacturing a huge workload

bash · flush and inspect a small controlled fixture
docker exec atlasmart-cass-1 nodetool flush atlasmart_types customer_preferences# Read the row so read-side tombstone statistics have a chance to update.docker exec atlasmart-cass-1 cqlsh -e "SELECT * FROM atlasmart_types.customer_preferences WHERE customer_id=11111111-1111-1111-1111-111111111111;"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_types customer_preferences# Optional low-level evidence; exact SSTable path/output varies by 5.0 patch/format.docker exec atlasmart-cass-1 bash -lc 'F=$(find /var/lib/cassandra/data/atlasmart_types/customer_preferences-* -name "*Data.db" | head -1); echo "$F"; test -n "$F" && sstabledump "$F" | grep -E "deletion|tombstone|markedForDeleteAt" | head -20 || true'

tablestats is operational evidence, not a proof that every element tombstone is currently visible in one SSTable. The optional sstabledump path is deliberately version-sensitive and should be treated as forensic learning output. The correct production response to collection tombstones is usually better modeling and bounded mutation patterns—not manual compaction as a routine cure.

Wrong approach: one collection for every order a customer has ever placed.

That collection grows with business history, is not internally paged, and turns one logical row into an ever-growing read/mutation unit. Repair it with a table such as orders_by_customer_bucket, using a bounded partition key and clustering columns so reads can page and retention can be modeled explicitly.

4. Production judgment

Collection choice changes write amplification, read payload, tombstone density, compaction work, and application concurrency. A set is usually preferable to a list when order/duplicates are not required. Maps are excellent for small keyed attributes but dangerous when keys become an event stream. Concurrent list position changes are particularly difficult to reason about because positional operations can involve read-before-write behavior. Set/map element updates are naturally idempotent when the new value is deterministic; list appends may not be safe to retry blindly. Model limits in business terms—maximum labels, channels, or settings per entity—then monitor actual cardinality rather than relying on theoretical guardrail maxima.

Verification checklist

  • The collection cardinalities are bounded by an explicit product invariant.
  • Element additions/removals produce the expected before/after row.
  • You can explain why deletions/TTL create tombstones instead of immediate physical erasure.
  • The optional SSTable/tombstone output is labeled version-dependent.
  • The unbounded relationship redesign uses clustering rows instead of a collection.

Check your understanding

  1. Why are collections unsuitable for an unlimited event history?
  2. What is the key semantic difference between set and list?
  3. What does removing a collection element create?
  4. Why can list operations be operationally riskier?
  5. What is the modeling repair for thousands of child items?
Review the answers

1. They are read as a collection and are not internally paged; unbounded growth makes the row/mutation unit increasingly expensive.

2. Sets enforce uniqueness; lists preserve duplicates and positional order.

3. A deletion marker/tombstone that participates in normal replica convergence and later compaction.

4. Some positional/prepend operations can require read-before-write work, and retries/ordering are harder to reason about.

5. Use a separate query table with partition/clustering keys and explicit bucketing/retention.

bash · reset Chapter 08 data when desired
# Destructive only to this disposable chapter keyspace.docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_types;"# Keep shared course containers for Chapter 09, or remove them for a full reset:# docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3# docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data# docker network rm atlasmart-cassandra

Summary and next bridge

Collections are useful only when their growth and mutation model remain small and predictable. Next we model richer structured values with tuples and user-defined types, where positional versus named semantics and schema evolution matter more than collection cardinality.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.