Chapter 07 · Primary Keys, Partition Keys, Clustering Columns, and Ordering

Estimate Rows / Bytes per Partition Before Deploying the Schema

Build an estimate → load → measure → redesign loop instead of relying on folklore partition limits.

Intermediate90–120 minutesPartition sizing + tablehistograms labApache Cassandra 5.0.9 · cqlsh/nodetool · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's schema review should reject dangerous partitions before a production load test discovers them. This lesson builds a sizing worksheet from peak row rate and row width, then validates the assumptions with a synthetic Cassandra fixture and nodetool tablehistograms/tablestats. The goal is not a magic partition limit; it is a repeatable estimate → measure → revise loop.

01

Estimate rows per partition from peak arrival rate, bucket duration, skew and retention.

02

Estimate logical partition bytes without pretending payload bytes equal SSTable bytes.

03

Relate RF, compression, tombstones, indexes and repair copies to cluster-wide storage cost.

04

Use tablehistograms and tablestats to compare observed data with the worksheet.

05

Convert an estimate that misses the latency/size objective into a bucket/shard redesign and migration plan.

Chapter 07 lab baseline

The mandatory labs use the pinned cassandra:5.0.9 image. Java 17, cqlsh, and nodetool are the versions bundled by that image. The course topology is three disposable nodes (atlasmart-cass-1..3) in cluster atlasmart-course, datacenter dc1, racks rack1..rack3, 16 vnodes per node, replication factor (RF) 3, and LOCAL_QUORUM for the chapter's consistency-sensitive examples. Authentication, client TLS, internode TLS, and remote JMX are not enabled in this isolated learning network; do not copy that security posture to production. New Chapter 07 tables explicitly use UnifiedCompactionStrategy (UCS). Default table TTL is zero unless a lesson says otherwise; gc_grace_seconds is not changed. SAI/vector features are not used. No application driver is required for the mandatory lab; if you adapt the examples to an application, re-check your chosen driver's current compatibility and routing behavior.

Execution disclosure and resource path

These commands are documentation- and syntax-reviewed but were not executed in this generation environment. Treat shown output as an expected shape, then capture your own exact tokens, latencies, partition-size percentiles, and errors. A three-node local cluster commonly needs several GiB of RAM; if your machine cannot support it, use one disposable node and RF=1 to learn primary-key/clustering mechanics, but do not interpret that reduced topology as evidence about RF=3 availability or replica behavior.

1. Row-count estimate first

Assume the hottest AtlasMart tenant can sustain 20 events/second and the design uses a time bucket. At one hour, the upper-order estimate is 20 × 3600 = 72,000 rows per partition. At one day it is 20 × 86,400 = 1,728,000 rows. If peak skew can double that tenant's rate, include the skew factor rather than hiding it inside an average.

Bucket 20 rows/s With 2× skew Operational question
1 hour 72,000 rows 144,000 can one partition meet slice-read/repair objectives?
6 hours 432,000 864,000 is larger locality worth wider repair/compaction work?
1 day 1,728,000 3,456,000 likely needs strong evidence before acceptance

2. Byte estimates are ranges, not exact SSTable predictions

Suppose a representative logical row—including partition/clustering keys, values and a conservative allowance for cell/row metadata—averages about 600 bytes in the workload model. Then 72,000 rows is roughly 43.2 MB of logical row material before compression and storage-engine structure. RF=3 means three replica copies, so the cluster stores roughly three times the live logical data before considering compression, repair overlap, snapshots, compaction amplification, tombstones and indexes. This arithmetic is planning evidence, not a promise about Space used (live).

Do not publish “100 MB is safe” or any other folklore threshold.

Workload shape matters: number of cells, row width, read slice width, tombstone density, compression, compaction strategy, disk, cache/JVM behavior, repair/streaming windows and tail-latency SLOs. Use guardrails/warnings as safety signals, not as a substitute for schema-specific load tests.

3. Lab: create a measurable sizing probe

bash · verify or recreate the disposable course cluster
# If you already have the Chapter 01-06 lab, verify it first.docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool status# Standalone recreation path (skip existing resources as needed).docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 answers before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three nodes report UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_keys WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE KEYSPACE atlasmart_keys"
sql · create a simple partition-size probe
CREATE TABLE atlasmart_keys.partition_size_probe (    tenant_id uuid,    hour_bucket text,    seq int,    payload text,    PRIMARY KEY ((tenant_id, hour_bucket), seq)) WITH CLUSTERING ORDER BY (seq ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'};
bash · generate 2,000 deterministic rows inside the Cassandra container
docker exec atlasmart-cass-1 bash -lc 'rm -f /tmp/ch07_probe.cql; for i in $(seq 1 2000); do printf "INSERT INTO atlasmart_keys.partition_size_probe (tenant_id,hour_bucket,seq,payload) VALUES (aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa,\0472026-09-07T10\047,%s,\047xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx\047);\n" "$i" >> /tmp/ch07_probe.cql; done'docker exec atlasmart-cass-1 cqlsh -f /tmp/ch07_probe.cqldocker exec atlasmart-cass-1 cqlsh -e "SELECT count(*) FROM atlasmart_keys.partition_size_probe WHERE tenant_id=aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa AND hour_bucket='2026-09-07T10';"docker exec atlasmart-cass-1 nodetool flush atlasmart_keys partition_size_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_keys partition_size_probedocker exec atlasmart-cass-1 nodetool tablestats -H atlasmart_keys.partition_size_probe

Expected evidence: the CQL count is 2,000. After flush, tablehistograms should expose a partition-size percentile and cell-count distribution that reflects the fixture. Exact bytes vary with Cassandra version, serialization, compression, SSTable state and the node's share of replicas. On a tiny one-partition fixture, percentiles collapse around that partition and are not production distributions.

4. Compare estimate to measurement correctly

The payload string above is 100 characters, but the measured partition is not simply 2000 × 100 bytes. Keys, clustering values, timestamps/metadata, compression blocks and SSTable structures matter. Conversely, tablestats “space used” is table/node scope, not a direct per-partition serializer. Use the tools together: known row count and payload composition explain the fixture; histogram percentiles expose partition-size distribution; table stats expose SSTable count, live space and related state.

For a production-like acceptance test, generate many partitions with the real key distribution, payload-size distribution, TTL/delete rate and skew. Warm the workload, flush/compact in a controlled way, then compare p50/p95/p99/max partition sizes and read/write p95/p99. Disclose RF/CL, hardware, disk, JVM, concurrency and retry policy. One average-size partition on a laptop cannot establish production capacity.

5. Broken estimates and redesign loop

Three common errors are: sizing from average event rate, multiplying only payload bytes, and forgetting that a “month bucket” for a hot tenant may still contain millions of rows. If the measured partition or tail latency is too large, shorten the time bucket, add a small deterministic shard dimension, reduce row width, split access patterns into separate tables, or revise retention. Every change affects read fan-out and migration.

A safe rollout treats the new primary key as a schema migration: create a new table, dual-write or stream changes with idempotent logic, backfill bounded ranges, reconcile counts/checksums/business invariants, cut reads gradually, and preserve a rollback window. Never mutate a production partition scheme in place by wishful thinking.

6. Production judgment

Partition-size review belongs in normal capacity planning. Track partition-size/cell-count percentiles, hot keys, local read/write tail latency, SSTable counts, compaction backlog, disk utilization, GC/JVM pressure and repair duration. Re-estimate when tenant size, retention, payloads or features change. SAI/vector indexes add separate storage and query cost; they do not remove primary-key partition economics. Security/tenancy controls must ensure a tenant cannot deliberately force pathological keys or oversized payloads. Managed services may expose different metrics/guardrails, but the physical need to bound partitions and test migration remains.

Check your understanding

  1. At 20 rows/s, how many rows enter one hourly partition before skew?
  2. Why is 72,000 × payload bytes not an exact SSTable-size prediction?
  3. What does RF=3 change in a sizing discussion?
  4. Why flush before using tablehistograms for this fixture?
  5. What is the correct response if the estimate fails the SLO?
Review the answers

1. About 72,000 rows.

2. Keys, metadata, compression and SSTable structures contribute, while compression can reduce on-disk bytes.

3. It creates three replica copies of the logical data and increases cluster-wide storage/repair/streaming work; it does not triple one logical partition’s row count.

4. Partition-size statistics are most meaningful once data is represented in SSTables rather than existing only in memtables.

5. Redesign bucket/shard/row width/access tables, then migrate and remeasure; do not rely on a universal tuning value.

Final Chapter 07 verification checklist

  • You can parse partition versus clustering components from parentheses alone.
  • You have estimated hottest-key rows per bucket before accepting a schema.
  • You have captured real partition-size evidence from a representative load rather than using tutorial numbers as limits.
  • You can state the maximum read fan-out introduced by each bucket/shard rule.
  • You have a migration/reconciliation/rollback plan for any primary-key redesign.
bash · reset only Chapter 07 data
# Destructive only to this disposable chapter keyspace.docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_keys;"# Keep the shared course containers for the next chapter, or remove them only if you want a full reset:# docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3# docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data# docker network rm atlasmart-cassandra

Summary and next bridge

Chapter 07 connected primary-key syntax to token placement, partition cardinality, clustering slices, bucketing/sharding, and measurable partition growth. Chapter 08 builds on that physical model by choosing CQL scalar/collection/tuple/UDT/static/frozen representations without letting convenience types create hidden growth or mutation costs.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.