Chapter 07 · Primary Keys, Partition Keys, Clustering Columns, and Ordering
Estimate Rows / Bytes per Partition Before Deploying the Schema
Build an estimate → load → measure → redesign loop instead of relying on folklore partition limits.
Learning outcomes
AtlasMart's schema review should reject dangerous partitions
before a production load test discovers them. This lesson builds
a sizing worksheet from peak row rate and row width, then
validates the assumptions with a synthetic Cassandra fixture and
nodetool tablehistograms/tablestats.
The goal is not a magic partition limit; it is a repeatable
estimate → measure → revise loop.
Estimate rows per partition from peak arrival rate, bucket duration, skew and retention.
Estimate logical partition bytes without pretending payload bytes equal SSTable bytes.
Relate RF, compression, tombstones, indexes and repair copies to cluster-wide storage cost.
Use tablehistograms and
tablestats to compare observed data with the
worksheet.
Convert an estimate that misses the latency/size objective into a bucket/shard redesign and migration plan.
The mandatory labs use the pinned
cassandra:5.0.9 image. Java 17,
cqlsh, and nodetool are the versions
bundled by that image. The course topology is three disposable
nodes (atlasmart-cass-1..3) in cluster
atlasmart-course, datacenter dc1,
racks rack1..rack3, 16 vnodes per node,
replication factor (RF) 3, and LOCAL_QUORUM for
the chapter's consistency-sensitive examples. Authentication,
client TLS, internode TLS, and remote JMX are not enabled in
this isolated learning network; do not copy that security
posture to production. New Chapter 07 tables explicitly use
UnifiedCompactionStrategy (UCS). Default table TTL is zero
unless a lesson says otherwise;
gc_grace_seconds is not changed. SAI/vector
features are not used. No application driver is required for
the mandatory lab; if you adapt the examples to an
application, re-check your chosen driver's current
compatibility and routing behavior.
These commands are documentation- and syntax-reviewed but were not executed in this generation environment. Treat shown output as an expected shape, then capture your own exact tokens, latencies, partition-size percentiles, and errors. A three-node local cluster commonly needs several GiB of RAM; if your machine cannot support it, use one disposable node and RF=1 to learn primary-key/clustering mechanics, but do not interpret that reduced topology as evidence about RF=3 availability or replica behavior.
1. Row-count estimate first
Assume the hottest AtlasMart tenant can sustain 20 events/second
and the design uses a time bucket. At one hour, the upper-order
estimate is 20 × 3600 = 72,000 rows per partition.
At one day it is 20 × 86,400 = 1,728,000 rows. If
peak skew can double that tenant's rate, include the skew factor
rather than hiding it inside an average.
| Bucket | 20 rows/s | With 2× skew | Operational question |
|---|---|---|---|
| 1 hour | 72,000 rows | 144,000 | can one partition meet slice-read/repair objectives? |
| 6 hours | 432,000 | 864,000 | is larger locality worth wider repair/compaction work? |
| 1 day | 1,728,000 | 3,456,000 | likely needs strong evidence before acceptance |
2. Byte estimates are ranges, not exact SSTable predictions
Suppose a representative logical row—including
partition/clustering keys, values and a conservative allowance
for cell/row metadata—averages about 600 bytes in the workload
model. Then 72,000 rows is roughly 43.2 MB of logical row
material before compression and storage-engine structure. RF=3
means three replica copies, so the cluster stores roughly three
times the live logical data before considering compression,
repair overlap, snapshots, compaction amplification, tombstones
and indexes. This arithmetic is planning evidence, not a promise
about Space used (live).
Workload shape matters: number of cells, row width, read slice width, tombstone density, compression, compaction strategy, disk, cache/JVM behavior, repair/streaming windows and tail-latency SLOs. Use guardrails/warnings as safety signals, not as a substitute for schema-specific load tests.
3. Lab: create a measurable sizing probe
# If you already have the Chapter 01-06 lab, verify it first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool status# Standalone recreation path (skip existing resources as needed).docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 answers before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three nodes report UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_keys WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE KEYSPACE atlasmart_keys"
CREATE TABLE atlasmart_keys.partition_size_probe ( tenant_id uuid, hour_bucket text, seq int, payload text, PRIMARY KEY ((tenant_id, hour_bucket), seq)) WITH CLUSTERING ORDER BY (seq ASC) AND compaction = {'class':'UnifiedCompactionStrategy'};
docker exec atlasmart-cass-1 bash -lc 'rm -f /tmp/ch07_probe.cql; for i in $(seq 1 2000); do printf "INSERT INTO atlasmart_keys.partition_size_probe (tenant_id,hour_bucket,seq,payload) VALUES (aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa,\0472026-09-07T10\047,%s,\047xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx\047);\n" "$i" >> /tmp/ch07_probe.cql; done'docker exec atlasmart-cass-1 cqlsh -f /tmp/ch07_probe.cqldocker exec atlasmart-cass-1 cqlsh -e "SELECT count(*) FROM atlasmart_keys.partition_size_probe WHERE tenant_id=aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa AND hour_bucket='2026-09-07T10';"docker exec atlasmart-cass-1 nodetool flush atlasmart_keys partition_size_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_keys partition_size_probedocker exec atlasmart-cass-1 nodetool tablestats -H atlasmart_keys.partition_size_probe
Expected evidence: the CQL count is 2,000. After flush,
tablehistograms should expose a partition-size
percentile and cell-count distribution that reflects the
fixture. Exact bytes vary with Cassandra version, serialization,
compression, SSTable state and the node's share of replicas. On
a tiny one-partition fixture, percentiles collapse around that
partition and are not production distributions.
4. Compare estimate to measurement correctly
The payload string above is 100 characters, but the measured
partition is not simply 2000 × 100 bytes. Keys,
clustering values, timestamps/metadata, compression blocks and
SSTable structures matter. Conversely,
tablestats “space used” is table/node scope, not a
direct per-partition serializer. Use the tools together: known
row count and payload composition explain the fixture; histogram
percentiles expose partition-size distribution; table stats
expose SSTable count, live space and related state.
For a production-like acceptance test, generate many partitions with the real key distribution, payload-size distribution, TTL/delete rate and skew. Warm the workload, flush/compact in a controlled way, then compare p50/p95/p99/max partition sizes and read/write p95/p99. Disclose RF/CL, hardware, disk, JVM, concurrency and retry policy. One average-size partition on a laptop cannot establish production capacity.
5. Broken estimates and redesign loop
Three common errors are: sizing from average event rate, multiplying only payload bytes, and forgetting that a “month bucket” for a hot tenant may still contain millions of rows. If the measured partition or tail latency is too large, shorten the time bucket, add a small deterministic shard dimension, reduce row width, split access patterns into separate tables, or revise retention. Every change affects read fan-out and migration.
A safe rollout treats the new primary key as a schema migration: create a new table, dual-write or stream changes with idempotent logic, backfill bounded ranges, reconcile counts/checksums/business invariants, cut reads gradually, and preserve a rollback window. Never mutate a production partition scheme in place by wishful thinking.
6. Production judgment
Partition-size review belongs in normal capacity planning. Track partition-size/cell-count percentiles, hot keys, local read/write tail latency, SSTable counts, compaction backlog, disk utilization, GC/JVM pressure and repair duration. Re-estimate when tenant size, retention, payloads or features change. SAI/vector indexes add separate storage and query cost; they do not remove primary-key partition economics. Security/tenancy controls must ensure a tenant cannot deliberately force pathological keys or oversized payloads. Managed services may expose different metrics/guardrails, but the physical need to bound partitions and test migration remains.
Check your understanding
- At 20 rows/s, how many rows enter one hourly partition before skew?
- Why is 72,000 × payload bytes not an exact SSTable-size prediction?
- What does RF=3 change in a sizing discussion?
- Why flush before using tablehistograms for this fixture?
- What is the correct response if the estimate fails the SLO?
Review the answers
1. About 72,000 rows.
2. Keys, metadata, compression and SSTable structures contribute, while compression can reduce on-disk bytes.
3. It creates three replica copies of the logical data and increases cluster-wide storage/repair/streaming work; it does not triple one logical partition’s row count.
4. Partition-size statistics are most meaningful once data is represented in SSTables rather than existing only in memtables.
5. Redesign bucket/shard/row width/access tables, then migrate and remeasure; do not rely on a universal tuning value.
Final Chapter 07 verification checklist
- You can parse partition versus clustering components from parentheses alone.
- You have estimated hottest-key rows per bucket before accepting a schema.
- You have captured real partition-size evidence from a representative load rather than using tutorial numbers as limits.
- You can state the maximum read fan-out introduced by each bucket/shard rule.
- You have a migration/reconciliation/rollback plan for any primary-key redesign.
# Destructive only to this disposable chapter keyspace.docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_keys;"# Keep the shared course containers for the next chapter, or remove them only if you want a full reset:# docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3# docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data# docker network rm atlasmart-cassandra
Summary and next bridge
Chapter 07 connected primary-key syntax to token placement, partition cardinality, clustering slices, bucketing/sharding, and measurable partition growth. Chapter 08 builds on that physical model by choosing CQL scalar/collection/tuple/UDT/static/frozen representations without letting convenience types create hidden growth or mutation costs.
Authoritative references
- CQL data definition — primary-key grammar, partition keys, clustering columns, and clustering order.
- CQL data manipulation — primary-key restrictions, contiguous clustering slices, token queries, and ordering behavior.
- Cassandra data-modeling introduction — partition-key and clustering-key physical meaning.
-
CREATE TABLE reference
— composite partition keys and
CLUSTERING ORDER BY. - nodetool tablehistograms — table-level percentile evidence including partition-size distribution.
- nodetool tablestats — table statistics and partition-size-related metrics.