Chapter 11 · SSTables and On-Disk Storage Internals

SSTable Components: Data, Index, Summary, Filter, Statistics, Compression, and TOC

Map Cassandra 5.0 SSTable components across BIG and BTI formats and learn why component files are evidence, not a stable application API.

Intermediate100–140 minutesComponent inventory labApache Cassandra 5.0.9 · cqlsh/nodetool · SSTable tools on copies · RF=3 · UCS · BIG default / BTI optionalLast reviewed: September 2026

Learning outcomes

AtlasMart receives a storage alert showing thousands of files under one table directory. An operator proposes deleting “obvious index files” because Data.db looks like the only file containing real rows. This lesson replaces that dangerous assumption with a format-aware map of an SSTable component set and a read-only inventory workflow.

01

Identify Data, index/trie, summary, Bloom-filter, statistics, compression, checksum, TOC, and optional SAI components without treating filenames as a stable API.

02

Distinguish Cassandra 5.0 BIG and BTI SSTable formats and explain why component names differ.

03

Use nodetool datapaths/tablestats plus read-only filesystem listing to connect a logical table to its physical component sets.

04

Explain what each component proves and what it cannot prove about logical correctness or recoverability.

05

Reject manual component deletion/editing and replace it with documented compaction, cleanup, snapshot, restore, or capacity actions.

Chapter 11 lab baseline

The mandatory labs continue the disposable AtlasMart environment used by earlier chapters: pinned cassandra:5.0.9, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. Chapter 11 uses keyspace atlasmart_storage with NetworkTopologyStrategy, replication factor (RF) 3, normally LOCAL_QUORUM, and tables that explicitly use UnifiedCompactionStrategy (UCS). Cassandra's current default SSTable format is BIG unless sstable.selected_format is changed; Cassandra 5.0 also supports BTI trie-indexed SSTables. Authentication, client TLS, internode TLS, and remote JMX stay disabled only inside this isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional; mandatory storage evidence uses cqlsh, nodetool, Docker/Linux filesystem tools, and Cassandra's bundled SSTable utilities. Re-check nodetool version, java -version, actual cassandra.yaml, disk free space, and selected SSTable format before interpreting output.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Storage terms before touching the filesystem

An SSTable (Sorted String Table) is Cassandra's immutable on-disk representation produced by a memtable flush, compaction, streaming, bulk load, or related storage workflow. A component is one file belonging to an SSTable generation; component sets and names depend on SSTable format/version. A partition is the rows sharing a partition key and is sorted with other partitions by token order inside SSTable data; rows inside a partition follow clustering order. Compression chunks are independently compressed blocks of the Data component, letting Cassandra read/decompress only relevant chunks instead of the full file. Compaction reads SSTables and writes replacement SSTables, then retires old ones when safe. Streaming transfers replica data between nodes for bootstrap, rebuild, repair, replacement, and topology movement. Disk headroom is free capacity reserved not only for live data but for temporary overlap during these operations. Offline SSTable tools inspect or transform SSTables outside normal CQL/native-protocol execution; many explicitly require Cassandra to be stopped and therefore must never be pointed casually at live production paths.

1. One logical SSTable is a coordinated component set

The Data.db component contains serialized partition/row data, but Cassandra relies on companion components to locate, validate, decompress, summarize, and reason about it. In Cassandra 5.0 the exact set is format-specific. BIG remains the default SSTable format unless configured otherwise. BTI, introduced in Cassandra 5.0, uses trie-indexed structures and changes the classic index component layout. Consequently, a runbook that assumes every generation has identical Index.db/Summary.db files across versions can misdiagnose a healthy node.

Component/concept Role Format/version caveat
Data.db serialized partitions and rows core data component; interpretation is Cassandra-version-specific
Index.db classic BIG partition/row index into Data.db format-dependent; do not assume in BTI
Summary.db sampled classic index summary that narrows index search tied to formats using the classic index path
Partitions.db BTI partition trie/index Cassandra 5.0 BTI evidence
Rows.db BTI row-index data for sufficiently wide partitions may not contain entries for every partition
Filter.db Bloom filter for partition-key membership false positives cost work; not authoritative row data
CompressionInfo.db compressed chunk offsets/lengths present when compression is enabled
Statistics.db timestamps, tombstones, clustering, repair/compression/TTL metadata metadata, not a substitute for Data.db
Digest.crc32 / checksum components integrity evidence for data/components not a backup or replica-consistency proof
TOC.txt table of contents listing SSTable components useful inventory, not an invitation to edit files
SAI*.db Storage-Attached Index component files when SAI exists only relevant when indexes are configured

2. Create one generation, then map CQL → data directory → components

bash · verify the disposable three-node course cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 sh -lc "grep -n -A8 -B2 '^sstable:' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && df -i /var/lib/cassandra"
CQL · create the bounded AtlasMart storage fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_storageWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_storage.orders_by_customer_day (    customer_id text,    order_day date,    order_time timestamp,    order_id uuid,    status text,    total decimal,    note text,    PRIMARY KEY ((customer_id, order_day), order_time, order_id)) WITH CLUSTERING ORDER BY (order_time DESC, order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'}  AND compression = {'class':'LZ4Compressor','chunk_length_in_kb':'16'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'PAID',129.90,'first');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'PACKING',89.50,'second');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-77','2026-09-07','2026-09-07T18:02:00Z',00000000-0000-0000-0000-000000000003,'CREATED',44.00,'third');DESCRIBE TABLE atlasmart_storage.orders_by_customer_day;SELECT * FROM atlasmart_storage.orders_by_customer_dayWHERE customer_id='cust-42' AND order_day='2026-09-07';
bash · flush only in the disposable lab and locate the table directory
# Lab only: force a known memtable-to-SSTable transition.docker exec atlasmart-cass-1 nodetool flush atlasmart_storage orders_by_customer_day# Ask Cassandra for configured data paths first.docker exec atlasmart-cass-1 nodetool datapaths atlasmart_storage.orders_by_customer_day# Then inventory the matching table directory read-only.docker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_storage -maxdepth 1 -type d -name 'orders_by_customer_day-*' -print"docker exec atlasmart-cass-1 sh -lc "find /var/lib/cassandra/data/atlasmart_storage -maxdepth 2 -type f -name '*Data.db' -o -name '*Index.db' -o -name '*Summary.db' -o -name '*Partitions.db' -o -name '*Rows.db' -o -name '*Filter.db' -o -name '*Statistics.db' -o -name '*CompressionInfo.db' -o -name '*TOC.txt' | sort"

Expected shape: one SSTable generation normally produces several files sharing the same generation/version stem. The exact stem, format code, and component names depend on Cassandra patch, selected SSTable format, UUID-generation configuration, and whether compression/SAI are enabled. Record what your node produced instead of copying a sample filename.

bash · correlate component inventory with table metrics and disk usage
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_storage.orders_by_customer_daydocker exec atlasmart-cass-1 sh -lc "du -sh /var/lib/cassandra/data/atlasmart_storage/*orders_by_customer_day*"docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"

3. The dangerous shortcut: “delete metadata files; Cassandra can regenerate them”

Some components can theoretically be reconstructed from primary data by specialized tools, but that does not make arbitrary file deletion safe. Cassandra tracks SSTables as coordinated units while compaction, snapshots, streaming, and lifecycle transactions may be active. Deleting a component can turn an otherwise readable generation into a startup/read failure, break checksums/index lookups, or create a restore problem that is harder than the original disk alert.

Do not edit or delete SSTable component files by hand.

Use supported operations such as compaction, snapshot cleanup, nodetool cleanup after ownership changes, documented restore/reload workflows, or capacity expansion. If corruption is suspected, preserve evidence and work from copies. Component-level surgery on live production data is not a routine space-management technique.

bash · evidence-first response to a disk alert
docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool listsnapshotsdocker exec atlasmart-cass-1 nodetool netstats -Hdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra; du -sh /var/lib/cassandra/data/* 2>/dev/null | sort -h | tail"

4. Verification checklist

  • You can identify the actual selected SSTable format before interpreting component names.
  • You can map atlasmart_storage.orders_by_customer_day to its per-table data directory without guessing a UUID suffix.
  • You can explain why TOC.txt, Bloom/filter, statistics, compression, and index components are not optional junk.
  • You can distinguish a component listing from a logical row/replica consistency proof.
  • No command in the lab mutates or deletes SSTable files directly.

Check your understanding

  1. Why is Data.db not the only file that matters?
  2. Why can two Cassandra 5.0 nodes show different index component names?
  3. Does a present TOC.txt prove the SSTable is logically correct?
  4. What should you do before deleting files to recover disk space?
  5. Why is an SSTable component set implementation evidence rather than an application API?
Review the answers

1. Cassandra uses companion index/trie, Bloom, compression, statistics, checksum and TOC components to locate, validate and interpret the data efficiently and safely.

2. They may use different SSTable formats such as default BIG versus optional BTI, or different format/version settings.

3. No. It inventories expected components; correctness also depends on readable components, checksums/metadata, Cassandra version/format, and replica semantics.

4. Measure compaction, streaming, snapshots, ownership and disk consumers, then use documented Cassandra lifecycle operations or add capacity.

5. Applications interact through CQL/native protocol; on-disk format/components can change across Cassandra versions and configured SSTable formats.

Production judgment

SSTable files expose valuable operational evidence, but they are not an application contract. Production decisions must combine logical data shape with physical storage state: partition rows/bytes, mutation and TTL/delete rates, RF/CL, read/write p95/p99, SSTables per read, compaction strategy and backlog, repaired/unrepaired state, compression ratio, CPU/decompression cost, disk throughput/latency, temporary compaction/streaming space, snapshot/backup retention, topology changes, repair cadence, and restore objectives. Include JVM/GC, page cache/off-heap use, driver timeouts/retries/idempotency, network bandwidth, tenant isolation, encryption-at-rest expectations, filesystem/device behavior, and managed-service restrictions.

Do not plan disk capacity as “live dataset bytes × RF” only. Flush, compaction, repair, streaming, snapshots, incremental backups, anti-compaction, restore staging, and operational safety margins can temporarily retain additional SSTables. Do not manually delete or edit component files to recover space. If space is critical, first stop unsafe automation, measure ownership/snapshots/compaction/streaming, and choose a documented recovery path with rollback. Lesson 2 goes deeper into why immutable sorted storage and block compression produce these components and how token/clustering order shapes sequential and random I/O.

Summary and next bridge

An SSTable is a versioned component set, not “one data file plus disposable helpers.” The next lesson follows partitions inside immutable sorted storage and connects compression chunks to actual read and compaction I/O.

Authoritative references

These are version-sensitive sources of truth. Re-check them when regenerating the lesson because SSTable formats, utilities, defaults, and topology procedures evolve.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.