Chapter 11 · SSTables and On-Disk Storage Internals

Connect SSTable Count / Size to Compaction Strategy, Read Amplification, and Disk Headroom

Turn SSTable count and size into a measured read-amplification, compaction, snapshot, streaming, and disk-headroom capacity model.

Intermediate120–160 minutesDisk-headroom + compaction labApache Cassandra 5.0.9 · cqlsh/nodetool · SSTable tools on copies · RF=3 · UCS · BIG default / BTI optionalLast reviewed: September 2026

Learning outcomes

AtlasMart stores 1.2 TiB of live Cassandra data on a node with a 1.5 TiB volume and concludes that 300 GiB free is “25% headroom, so enough.” That percentage is meaningless without knowing SSTable size distribution, compaction strategy, pending compactions, snapshot/backup retention, streaming plans, restore staging, and failure mode. This lesson turns file-level evidence into a capacity decision.

01

Connect SSTable count/size distributions to read amplification, write amplification and compaction work.

02

Distinguish live logical data size, compressed SSTable bytes, snapshot true size, and temporary rewrite/streaming space.

03

Build a transparent headroom worksheet from measured node/table data instead of a universal free-space percentage.

04

Use flush/compaction in the disposable lab to observe before/after SSTable count, disk bytes and read metrics.

05

Create production acceptance/rollback criteria before compaction, topology, repair, backup or restore operations consume headroom.

Chapter 11 lab baseline

The mandatory labs continue the disposable AtlasMart environment used by earlier chapters: pinned cassandra:5.0.9, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 11 uses keyspace atlasmart_storage with NetworkTopologyStrategy, replication factor (RF) 3, normally LOCAL_QUORUM, and tables that explicitly use UnifiedCompactionStrategy (UCS). Cassandra's current default SSTable format is BIG unless sstable.selected_format is changed; Cassandra 5.0 also supports BTI trie-indexed SSTables. Authentication, client TLS, internode TLS, and remote JMX stay disabled only inside this isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional; mandatory storage evidence uses cqlsh, nodetool, Docker/Linux filesystem tools, and Cassandra's bundled SSTable utilities. Re-check nodetool version, java -version, actual cassandra.yaml, disk free space, and selected SSTable format before interpreting output.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Storage terms before touching the filesystem

An SSTable (Sorted String Table) is Cassandra's immutable on-disk representation produced by a memtable flush, compaction, streaming, bulk load, or related storage workflow. A component is one file belonging to an SSTable generation; component sets and names depend on SSTable format/version. A partition is the rows sharing a partition key and is sorted with other partitions by token order inside SSTable data; rows inside a partition follow clustering order. Compression chunks are independently compressed blocks of the Data component, letting Cassandra read/decompress only relevant chunks instead of the full file. Compaction reads SSTables and writes replacement SSTables, then retires old ones when safe. Streaming transfers replica data between nodes for bootstrap, rebuild, repair, replacement, and topology movement. Disk headroom is free capacity reserved not only for live data but for temporary overlap during these operations. Offline SSTable tools inspect or transform SSTables outside normal CQL/native-protocol execution; many explicitly require Cassandra to be stopped and therefore must never be pointed casually at live production paths.

1. SSTable count is a symptom; overlap and workload determine cost

More SSTables can increase read amplification because a read may need to consider data/tombstones from several immutable generations. But “20 SSTables is bad” is not a universal rule: compaction strategy, overlap, query key presence, Bloom filters, partition shape and repaired state matter. Likewise one huge SSTable can create different concerns around streaming/rewrite duration and large partitions. Measure SSTables per Read, table latency, compaction backlog, file sizes and partition distributions together.

bash · verify the disposable three-node course cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 sh -lc "grep -n -A8 -B2 '^sstable:' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && df -i /var/lib/cassandra"
CQL · create the bounded AtlasMart storage fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_storageWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_storage.orders_by_customer_day (    customer_id text,    order_day date,    order_time timestamp,    order_id uuid,    status text,    total decimal,    note text,    PRIMARY KEY ((customer_id, order_day), order_time, order_id)) WITH CLUSTERING ORDER BY (order_time DESC, order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND compression = {'class':'LZ4Compressor','chunk_length_in_kb':'16'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'PAID',129.90,'first');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'PACKING',89.50,'second');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-77','2026-09-07','2026-09-07T18:02:00Z',00000000-0000-0000-0000-000000000003,'CREATED',44.00,'third');DESCRIBE TABLE atlasmart_storage.orders_by_customer_day;SELECT * FROM atlasmart_storage.orders_by_customer_dayWHERE customer_id='cust-42' AND order_day='2026-09-07';
bash · create several small immutable generations intentionally
docker exec atlasmart-cass-1 nodetool flush atlasmart_storage orders_by_customer_dayfor i in 1 2 3; do  docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_storage.orders_by_customer_day SET note='generation-$i' WHERE customer_id='cust-42' AND order_day='2026-09-07' AND order_time='2026-09-07T18:00:00Z' AND order_id=00000000-0000-0000-0000-000000000001;"  docker exec atlasmart-cass-1 nodetool flush atlasmart_storage orders_by_customer_daydonedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_storage.orders_by_customer_daydocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_storage orders_by_customer_daydocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_storage -type f -path '*orders_by_customer_day*' -name '*Data.db' -printf '%s %f\n' | sort -n"

2. Disk accounting has multiple meanings

tablestats reports table/SSTable metrics; du reports filesystem blocks reachable from paths; snapshots often use hard links so “snapshot size” and “additional physical bytes” are not the same; nodetool listsnapshots distinguishes size-on-disk and true disk-space concepts. During compaction, old SSTables coexist with newly written replacements until the operation commits. During streaming or restore, incoming/staged files can coexist with existing data. Therefore a capacity model must state which byte concept it uses.

Bucket Example evidence Why it can coexist
Current live SSTables tablestats + table-directory Data/component bytes serving current replica data
Compaction output in progress compactionstats + free-space trend new SSTables written before old ones retire
Snapshots listsnapshots + snapshots directories hard links retain old SSTables after live compaction/delete
Incremental backup files backups directories/statusbackup retains new SSTables for backup pipeline
Incoming streaming/repair netstats + free-space trend receiver materializes replica data/repair outputs
Restore staging/copies restore runbook staging path copied backup set may coexist before load/refresh
Filesystem/system margin df -h/-i, logs/commitlog/saved caches node needs more than data-table bytes

3. A transparent headroom worksheet

Use a scenario model rather than a magic “keep X% free” rule. Define measured current bytes, plausible simultaneous operations, and a safety/uncertainty reserve. Avoid summing impossible worst cases if they are operationally mutually exclusive; conversely, do not omit combinations your runbook actually permits.

text · example planning worksheet (illustrative numbers, not a recommendation)
Measured current filesystem used on data volume:          1,200 GiBMeasured free capacity:                                      600 GiBPlanned operation envelope (example only):  largest eligible compaction rewrite still in progress:     180 GiB  snapshot-retained old SSTables during window:               90 GiB  expected bootstrap/repair incoming data on this node:       120 GiB  restore/backup staging allowed concurrently:                  0 GiB  (runbook forbids overlap)  logs/commitlog/other growth allowance during window:         25 GiB  uncertainty / measurement error reserve:                     60 GiB                                                           ---------Scenario temporary requirement:                              475 GiBRemaining free if all allowed overlap occurs:                125 GiBDecision is NOT “125 GiB is enough” by itself.Validate disk throughput, inode space, operation duration, abort/rollback behavior,compaction free-space constraints, and whether emergency growth can be added.
bash · capture the inputs on the disposable lab
docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra; df -i /var/lib/cassandra"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_storage.orders_by_customer_daydocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool listsnapshotsdocker exec atlasmart-cass-1 nodetool netstats -Hdocker exec atlasmart-cass-1 sh -lc "du -sh /var/lib/cassandra/data /var/lib/cassandra/commitlog /var/lib/cassandra/saved_caches 2>/dev/null"

4. Controlled before/after compaction evidence

Force compaction only on the disposable fixture to demonstrate temporary rewrite and resulting file consolidation. Capture disk/free space immediately before, during if possible, and after. Tiny data may complete too quickly to observe temporary overlap; report that limitation rather than inventing a spike.

bash · controlled lab compaction and re-measurement
docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_storage.orders_by_customer_day# Lab only; production manual compaction requires an explicit reason and headroom plan.docker exec atlasmart-cass-1 nodetool compact atlasmart_storage orders_by_customer_daydocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_storage.orders_by_customer_daydocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_storage orders_by_customer_day
Wrong capacity model: “dataset size × RF is disk requirement.”

RF explains logical replica multiplicity across nodes, not per-node temporary space during compaction, repair, bootstrap, snapshots, backup retention, or restore staging. Per-node ownership is also not necessarily equal under skew/topology. Capacity planning must use actual node/table metrics and the operations allowed to overlap.

5. Pre-operation acceptance and rollback gates

Gate Proceed only if... Rollback/abort signal
Free bytes/inodes scenario envelope + safety margin fits unexpected free-space decline / inode exhaustion
Foreground SLO p95/p99 and error rates have margin tail latency/errors exceed maintenance threshold
Compaction pending/throughput state is understood backlog or write amplification grows uncontrollably
Streaming/repair scope and throttles are explicit source/target disk/network saturation or repeated session failure
Snapshot/backup retention/off-host status known snapshots retain more old SSTables than modeled
Recovery abort path and restore source verified operation creates a state runbook cannot safely reverse

Check your understanding

  1. Why is raw SSTable count insufficient to judge read health?
  2. Why can compaction temporarily need extra disk even if it eventually frees space?
  3. Why can a snapshot keep disk usage high after compaction?
  4. What is wrong with dataset-bytes × RF as the only capacity formula?
  5. What should a capacity decision record besides free percentage?
Review the answers

1. Cost depends on overlap, query shape, Bloom decisions, partition/tombstone distribution, compaction state and SSTables actually touched per read.

2. It writes replacement SSTables while old SSTables still exist, and only retires old files after the new output is safely committed.

3. Hard-linked snapshot references can retain old SSTable files that the live table no longer needs.

4. It omits per-node ownership differences and temporary overlap from compaction, streaming, repair, snapshots/backups, restore staging and system overhead.

5. Measured bytes/inodes, operation overlap envelope, compaction/streaming state, workload SLOs, device throughput, duration, uncertainty reserve and rollback/expansion plan.

Production judgment

SSTable files expose valuable operational evidence, but they are not an application contract. Production decisions must combine logical data shape with physical storage state: partition rows/bytes, mutation and TTL/delete rates, RF/CL, read/write p95/p99, SSTables per read, compaction strategy and backlog, repaired/unrepaired state, compression ratio, CPU/decompression cost, disk throughput/latency, temporary compaction/streaming space, snapshot/backup retention, topology changes, repair cadence, and restore objectives. Include JVM/GC, page cache/off-heap use, driver timeouts/retries/idempotency, network bandwidth, tenant isolation, encryption-at-rest expectations, filesystem/device behavior, and managed-service restrictions.

Do not plan disk capacity as “live dataset bytes × RF” only. Flush, compaction, repair, streaming, snapshots, incremental backups, anti-compaction, restore staging, and operational safety margins can temporarily retain additional SSTables. Do not manually delete or edit component files to recover space. If space is critical, first stop unsafe automation, measure ownership/snapshots/compaction/streaming, and choose a documented recovery path with rollback. Chapter 12 now focuses on compaction itself: why merges exist, how UCS/STCS/LCS/TWCS differ, and how to measure read/write/space amplification without freezing old strategy folklore into defaults.

Summary and next bridge

SSTable count, size and temporary copies translate directly into latency and capacity risk only when combined with compaction, snapshots, streaming, repair and workload evidence. Chapter 12 takes the next step: choose and operate compaction strategies by measured amplification and workload fit.

Authoritative references

These are version-sensitive sources of truth. Re-check them when regenerating the lesson because SSTable formats, utilities, defaults, and topology procedures evolve.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.