Chapter 11 · SSTables and On-Disk Storage Internals

Immutable Sorted Storage, Partitions on Disk, Compression Chunks, and Sequential I / O

Follow immutable token/clustering order, compression chunks, and I/O granularity from flush through reads and compaction.

Intermediate110–150 minutesCompression + partition layout labApache Cassandra 5.0.9 · cqlsh/nodetool · SSTable tools on copies · RF=3 · UCS · BIG default / BTI optionalLast reviewed: September 2026

Learning outcomes

AtlasMart sees a 420 MiB Data.db file and assumes every point read must decompress hundreds of megabytes. Another engineer claims Cassandra performs purely sequential I/O because SSTables are sorted. Both statements miss the block-level and index-driven behavior. This lesson connects immutable ordering, partition locality, compression chunks, and actual I/O granularity.

01

Explain token ordering between partitions and clustering ordering within a partition on disk.

02

Explain why immutability enables append/flush/compaction workflows but creates version/read amplification.

03

Relate compression chunk size to read/decompression granularity and compression ratio/CPU tradeoffs.

04

Use metadata and partition-sampling tools on a safe copy to inspect token ranges, timestamps, compression, rows/bytes, and tombstones.

05

Distinguish sequential compaction/streaming advantages from point-read random I/O and avoid oversimplified “Cassandra is sequential” claims.

Chapter 11 lab baseline

The mandatory labs continue the disposable AtlasMart environment used by earlier chapters: pinned cassandra:5.0.9, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 11 uses keyspace atlasmart_storage with NetworkTopologyStrategy, replication factor (RF) 3, normally LOCAL_QUORUM, and tables that explicitly use UnifiedCompactionStrategy (UCS). Cassandra's current default SSTable format is BIG unless sstable.selected_format is changed; Cassandra 5.0 also supports BTI trie-indexed SSTables. Authentication, client TLS, internode TLS, and remote JMX stay disabled only inside this isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional; mandatory storage evidence uses cqlsh, nodetool, Docker/Linux filesystem tools, and Cassandra's bundled SSTable utilities. Re-check nodetool version, java -version, actual cassandra.yaml, disk free space, and selected SSTable format before interpreting output.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Storage terms before touching the filesystem

An SSTable (Sorted String Table) is Cassandra's immutable on-disk representation produced by a memtable flush, compaction, streaming, bulk load, or related storage workflow. A component is one file belonging to an SSTable generation; component sets and names depend on SSTable format/version. A partition is the rows sharing a partition key and is sorted with other partitions by token order inside SSTable data; rows inside a partition follow clustering order. Compression chunks are independently compressed blocks of the Data component, letting Cassandra read/decompress only relevant chunks instead of the full file. Compaction reads SSTables and writes replacement SSTables, then retires old ones when safe. Streaming transfers replica data between nodes for bootstrap, rebuild, repair, replacement, and topology movement. Disk headroom is free capacity reserved not only for live data but for temporary overlap during these operations. Offline SSTable tools inspect or transform SSTables outside normal CQL/native-protocol execution; many explicitly require Cassandra to be stopped and therefore must never be pointed casually at live production paths.

1. Sorted means ordered by storage keys—not “every workload is sequential”

With the default Murmur3 partitioner, partitions in an SSTable are arranged by decorated partition-key/token order. Within each partition, clustering rows follow the table's clustering order. This lets Cassandra use indexes to narrow where a partition or clustering slice lives and lets compaction merge already sorted streams efficiently. However, a point read can still require non-sequential storage access to Bloom/index/trie structures and one or more compressed data chunks across several SSTables. Sequential characteristics are strongest in flush, compaction, streaming, and bounded scans over adjacent storage regions—not a guarantee that arbitrary request workloads become sequential I/O.

Operation Typical storage pattern Why
Memtable flush write a new sorted SSTable sequentially memtable content is emitted as immutable ordered storage
Compaction read sorted SSTables + write sorted replacements merge operation can process ordered streams
Point partition read index/filter probes + selected data chunks locates one partition across candidate SSTables
Clustering slice selected chunks/blocks within a partition clustering order bounds the slice
Streaming sections or eligible entire SSTables moves ownership/repair data between nodes

2. Compression chunks: smaller disk footprint, bounded decompression unit

Cassandra's table compression divides Data.db into independently compressed chunks. CompressionInfo.db records chunk locations and lengths. A read locates the relevant compressed chunk, reads it, decompresses the whole chunk, and then continues row reconciliation. Thus a 16 KiB configured chunk does not mean “each row is 16 KiB”; it is an I/O/decompression unit. Larger chunks can improve compression ratio and metadata efficiency but may decompress more irrelevant bytes for small random reads. Smaller chunks can reduce that waste but increase metadata/fragmentation overhead. Benchmark representative data and workload instead of choosing a universal chunk size.

bash · verify the disposable three-node course cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 sh -lc "grep -n -A8 -B2 '^sstable:' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && df -i /var/lib/cassandra"
CQL · create the bounded AtlasMart storage fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_storageWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_storage.orders_by_customer_day (    customer_id text,    order_day date,    order_time timestamp,    order_id uuid,    status text,    total decimal,    note text,    PRIMARY KEY ((customer_id, order_day), order_time, order_id)) WITH CLUSTERING ORDER BY (order_time DESC, order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND compression = {'class':'LZ4Compressor','chunk_length_in_kb':'16'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'PAID',129.90,'first');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'PACKING',89.50,'second');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-77','2026-09-07','2026-09-07T18:02:00Z',00000000-0000-0000-0000-000000000003,'CREATED',44.00,'third');DESCRIBE TABLE atlasmart_storage.orders_by_customer_day;SELECT * FROM atlasmart_storage.orders_by_customer_dayWHERE customer_id='cust-42' AND order_day='2026-09-07';
bash · flush and capture compression/size evidence
docker exec atlasmart-cass-1 nodetool flush atlasmart_storage orders_by_customer_daydocker exec atlasmart-cass-1 nodetool tablestats atlasmart_storage.orders_by_customer_daydocker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_storage -type f -path '*orders_by_customer_day*' -printf '%f %s bytes\n' | sort"docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"

Record SSTable Compression Ratio from tablestats. A ratio near 0.5, for example, means compressed SSTable data occupies roughly half the uncompressed size represented by that metric; it does not mean reads are “50% faster.” CPU, chunk size, data compressibility, page cache, SSTable overlap, and query shape all matter.

3. Build a safe offline copy and inspect partition/metadata shape

Many bundled SSTable tools warn that Cassandra must be stopped. To make inspection reproducible without pointing them at a live mutable directory, this course uses a snapshot as a stable source, briefly stops one disposable replica, copies the snapshot out, restarts the replica, and then runs the utility in an ephemeral Cassandra 5.0.9 container against the copied files mounted read-only. The other replicas remain available during the short stop; still, use this only in the course cluster.

bash · snapshot, stop one disposable node, copy snapshot, restart
mkdir -p ch11-inspect# Flush+snapshot a single table on node 3. Snapshot defaults to a flush unless --skip-flush is used.docker exec atlasmart-cass-3 nodetool snapshot -t ch11-storage-copy --table orders_by_customer_day atlasmart_storagedocker exec atlasmart-cass-3 nodetool listsnapshotsdocker stop atlasmart-cass-3# docker cp works with a stopped container. Copy only the snapshot tree for this table.TABLE_DIR=$(docker run --rm --volumes-from atlasmart-cass-3 --entrypoint sh cassandra:5.0.9 -lc "find /var/lib/cassandra/data/atlasmart_storage -maxdepth 1 -type d -name 'orders_by_customer_day-*' | head -1")docker cp "atlasmart-cass-3:${TABLE_DIR}/snapshots/ch11-storage-copy/." ./ch11-inspect/docker start atlasmart-cass-3# Wait for node 3 to return UN before continuing.docker exec atlasmart-cass-1 nodetool statusls -lh ch11-inspect | head
Portability note

The TABLE_DIR=$(...) line is POSIX-shell syntax. On Windows PowerShell, obtain the printed table directory in one command, copy that path into a $TableDir variable, then run docker cp. The lesson's Cassandra semantics are identical; only host-shell syntax differs.

bash · inspect copied SSTable metadata and partition distribution read-only
docker run --rm -v "$PWD/ch11-inspect:/inspect:ro" --entrypoint sh cassandra:5.0.9 -lc '  f=$(find /inspect -name "*-Data.db" | head -1)  echo "Data component: $f"  /opt/cassandra/bin/sstablemetadata "$f" | head -120  /opt/cassandra/bin/sstablepartitions --csv "$f" | head -20' 

Expected evidence includes partitioner/token boundaries, timestamp/deletion ranges, compression details, estimated partition metrics, and possibly repaired/level metadata. Exact fields differ by SSTable format and patch; treat human-readable output as diagnostic evidence, not a parse-stable API.

4. Wrong shortcut: “large Data.db means large partition”

An SSTable can contain many small partitions or fewer large ones. File size alone cannot identify hot/wide partitions. Use tablehistograms/tablestats, sstablepartitions on a safe copy, and query/business-key knowledge. Conversely, one partition can be spread across several SSTables, so no single file's size tells you the full logical partition footprint.

Check your understanding

  1. How are partitions ordered inside an SSTable with Murmur3Partitioner?
  2. Why does a compressed Data.db not require full-file decompression for one point read?
  3. Does a larger compression chunk always improve reads?
  4. Why is compaction naturally suited to sequential I/O?
  5. Why can SSTable file size not diagnose partition size by itself?
Review the answers

1. By decorated partition key/token order; rows within each partition follow clustering order.

2. CompressionInfo metadata lets Cassandra locate and decompress only the relevant compression chunk(s).

3. No. It can improve ratio but may force more irrelevant bytes to be decompressed for small random reads; benchmark representative data.

4. It merges ordered immutable SSTable streams and writes new ordered SSTables.

5. Each SSTable may hold many partitions, and one logical partition can span multiple SSTables.

Production judgment

SSTable files expose valuable operational evidence, but they are not an application contract. Production decisions must combine logical data shape with physical storage state: partition rows/bytes, mutation and TTL/delete rates, RF/CL, read/write p95/p99, SSTables per read, compaction strategy and backlog, repaired/unrepaired state, compression ratio, CPU/decompression cost, disk throughput/latency, temporary compaction/streaming space, snapshot/backup retention, topology changes, repair cadence, and restore objectives. Include JVM/GC, page cache/off-heap use, driver timeouts/retries/idempotency, network bandwidth, tenant isolation, encryption-at-rest expectations, filesystem/device behavior, and managed-service restrictions.

Do not plan disk capacity as “live dataset bytes × RF” only. Flush, compaction, repair, streaming, snapshots, incremental backups, anti-compaction, restore staging, and operational safety margins can temporarily retain additional SSTables. Do not manually delete or edit component files to recover space. If space is critical, first stop unsafe automation, measure ownership/snapshots/compaction/streaming, and choose a documented recovery path with rollback. Lesson 3 turns the safe-copy workflow into a disciplined toolbox for sstablemetadata, sstabledump, sstablepartitions, and related utilities without mutating production files.

Summary and next bridge

Immutable sorted storage creates efficient flush/merge/stream paths while read I/O remains index- and chunk-driven. Compression trades disk/network bytes for CPU and chunk-level read granularity. Next, use offline tools safely and learn which outputs are diagnostic versus stable interfaces.

Authoritative references

These are version-sensitive sources of truth. Re-check them when regenerating the lesson because SSTable formats, utilities, defaults, and topology procedures evolve.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.