Chapter 11 · SSTables and On-Disk Storage Internals
SSTable Tools and Metadata Inspection Without Mutating Production Data
Inspect copied SSTables with metadata/dump/partition tools while preserving live production data and security boundaries.
Learning outcomes
During an AtlasMart incident, an engineer wants to run
sstabledump directly against files that Cassandra
is compacting. The intent—inspect a suspicious row—is
reasonable; the execution is unsafe. This lesson builds a
repeatable “copy first, inspect second” workflow and classifies
SSTable tools by read-only versus mutating risk.
Use snapshot/copy isolation before offline inspection and explain why many bundled tools require Cassandra to be stopped.
Use sstablemetadata, sstabledump, and sstablepartitions on copied SSTables to answer targeted operational questions.
Distinguish diagnostic tools from mutating/rewrite tools such as scrub, split, upgrade, relevel, or repaired-set operations.
Interpret metadata/timestamps/tombstones/partition samples without treating human-readable output as a permanent automation contract.
Preserve chain-of-custody evidence and avoid exposing sensitive application values unnecessarily during troubleshooting.
The mandatory labs continue the disposable AtlasMart
environment used by earlier chapters: pinned
cassandra:5.0.9, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. Chapter 11 uses keyspace
atlasmart_storage with
NetworkTopologyStrategy, replication factor (RF)
3, normally LOCAL_QUORUM, and tables that
explicitly use UnifiedCompactionStrategy (UCS). Cassandra's
current default SSTable format is BIG unless
sstable.selected_format is changed; Cassandra 5.0
also supports BTI trie-indexed SSTables. Authentication,
client TLS, internode TLS, and remote JMX stay disabled only
inside this isolated learning network. The Apache Cassandra
Java Driver 4.19.3 is optional; mandatory storage evidence
uses cqlsh, nodetool, Docker/Linux
filesystem tools, and Cassandra's bundled SSTable utilities.
Re-check nodetool version,
java -version, actual
cassandra.yaml, disk free space, and selected
SSTable format before interpreting output.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Storage terms before touching the filesystem
An SSTable (Sorted String Table) is Cassandra's immutable on-disk representation produced by a memtable flush, compaction, streaming, bulk load, or related storage workflow. A component is one file belonging to an SSTable generation; component sets and names depend on SSTable format/version. A partition is the rows sharing a partition key and is sorted with other partitions by token order inside SSTable data; rows inside a partition follow clustering order. Compression chunks are independently compressed blocks of the Data component, letting Cassandra read/decompress only relevant chunks instead of the full file. Compaction reads SSTables and writes replacement SSTables, then retires old ones when safe. Streaming transfers replica data between nodes for bootstrap, rebuild, repair, replacement, and topology movement. Disk headroom is free capacity reserved not only for live data but for temporary overlap during these operations. Offline SSTable tools inspect or transform SSTables outside normal CQL/native-protocol execution; many explicitly require Cassandra to be stopped and therefore must never be pointed casually at live production paths.
1. Offline tool rule: the filesystem is not a second control plane
CQL and nodetool are supported online
control/observation interfaces. SSTable utilities work beneath
that layer. The official tool documentation repeatedly warns
that Cassandra must be stopped for many utilities and does not
automatically verify this prerequisite. Even a read-only-looking
tool can observe a component set while live compaction or
deletion changes the directory around it. For production
analysis, prefer snapshots/backup copies or a stopped node and
preserve the original evidence.
| Tool | Primary use | Course safety rule |
|---|---|---|
sstablemetadata |
timestamps, tokens, tombstones, compression/repair metadata | run against copied SSTables; docs require Cassandra stopped |
sstabledump |
JSON/internal dump or key enumeration | copy first; sensitive row values may be exposed |
sstablepartitions |
large-partition rows/bytes/cells/tombstones statistics | prefer copied data; CSV for automation |
sstableutil |
find SSTables associated with table | use as discovery, then work from copies |
sstableverify |
verify SSTable integrity | maintenance/offline semantics; do not casually run on live paths |
sstablescrub, sstablesplit,
sstableupgrade
|
rewrite/repair/transform files | mutating operational tools; require explicit runbook, stop state, headroom, backup/rollback |
2. Prepare a stable inspection set
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 sh -lc "grep -n -A8 -B2 '^sstable:' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && df -i /var/lib/cassandra"
CREATE KEYSPACE IF NOT EXISTS atlasmart_storageWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_storage.orders_by_customer_day ( customer_id text, order_day date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((customer_id, order_day), order_time, order_id)) WITH CLUSTERING ORDER BY (order_time DESC, order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND compression = {'class':'LZ4Compressor','chunk_length_in_kb':'16'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'PAID',129.90,'first');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-07','2026-09-07T18:01:00Z',00000000-0000-0000-0000-000000000002,'PACKING',89.50,'second');INSERT INTO atlasmart_storage.orders_by_customer_day(customer_id,order_day,order_time,order_id,status,total,note)VALUES ('cust-77','2026-09-07','2026-09-07T18:02:00Z',00000000-0000-0000-0000-000000000003,'CREATED',44.00,'third');DESCRIBE TABLE atlasmart_storage.orders_by_customer_day;SELECT * FROM atlasmart_storage.orders_by_customer_dayWHERE customer_id='cust-42' AND order_day='2026-09-07';
# Produce more than one generation so metadata inspection is meaningful.docker exec atlasmart-cass-3 nodetool flush atlasmart_storage orders_by_customer_daydocker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_storage.orders_by_customer_day SET status='SHIPPED' WHERE customer_id='cust-42' AND order_day='2026-09-07' AND order_time='2026-09-07T18:00:00Z' AND order_id=00000000-0000-0000-0000-000000000001;"docker exec atlasmart-cass-3 nodetool flush atlasmart_storage orders_by_customer_daydocker exec atlasmart-cass-3 nodetool snapshot -t ch11-tools-copy --table orders_by_customer_day atlasmart_storagedocker exec atlasmart-cass-3 nodetool listsnapshots
rm -rf ch11-tools-inspect && mkdir ch11-tools-inspect# Discover the concrete table directory while the node is still running.TABLE_DIR=$(docker exec atlasmart-cass-3 sh -lc "find /var/lib/cassandra/data/atlasmart_storage -maxdepth 1 -type d -name 'orders_by_customer_day-*' | head -1")docker stop atlasmart-cass-3docker cp "atlasmart-cass-3:${TABLE_DIR}/snapshots/ch11-tools-copy/." ./ch11-tools-inspect/docker start atlasmart-cass-3docker exec atlasmart-cass-1 nodetool statusfind ch11-tools-inspect -maxdepth 1 -type f -printf '%f %s bytes\n' | sort
3. Answer three questions from the copy
Question A: what time/token/compression/repair range does
this SSTable cover?
Use sstablemetadata.
Question B: which partition keys or cells are physically
present?
Use a narrow sstabledump mode such as key
enumeration or a specific key before dumping an entire table.
Question C: are there unexpectedly large/tombstone-heavy
partitions?
Use sstablepartitions, preferably CSV for scripts.
docker run --rm -v "$PWD/ch11-tools-inspect:/inspect:ro" --entrypoint sh cassandra:5.0.9 -lc ' f=$(find /inspect -name "*-Data.db" | sort | head -1) /opt/cassandra/bin/sstablemetadata "$f" | head -160'
docker run --rm -v "$PWD/ch11-tools-inspect:/inspect:ro" --entrypoint sh cassandra:5.0.9 -lc ' f=$(find /inspect -name "*-Data.db" | sort | head -1) /opt/cassandra/bin/sstabledump -e "$f" | head -80'
docker run --rm -v "$PWD/ch11-tools-inspect:/inspect:ro" --entrypoint sh cassandra:5.0.9 -lc ' /opt/cassandra/bin/sstablepartitions --csv /inspect | head -40'
sstabledump can expose application values that
normal authorization controls would protect at the CQL layer.
Treat copied SSTables and dumps as sensitive data. Restrict
filesystem access, avoid uploading raw dumps to tickets/chat
systems, redact tenant/customer values, and delete the local
inspection copy when the incident or lab is finished.
4. What not to do
Do not use sed, a hex editor, rm, or
file-copy replacement to “fix” live SSTables. Do not invoke
mutating offline utilities until the node state,
backup/rollback, compatible Cassandra version, available disk
headroom, and post-operation verification are explicit. If the
node is healthy enough for CQL/nodetool, prefer supported online
operations. If corruption prevents startup, preserve originals
and follow a documented recovery plan rather than improvising
component edits.
docker exec atlasmart-cass-3 nodetool clearsnapshot -t ch11-tools-copy atlasmart_storagerm -rf ch11-tools-inspect# Confirm node and table still serve data normally.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT customer_id,order_day,order_time,status FROM atlasmart_storage.orders_by_customer_day WHERE customer_id='cust-42' AND order_day='2026-09-07';"
Check your understanding
- Why not run sstabledump directly on a live production table directory?
- Why enumerate keys before dumping the whole SSTable?
- Which output form is better for automation from sstablepartitions?
- Is sstablesplit a harmless inspection tool?
- What is the safest first response to suspected on-disk corruption?
Review the answers
1. The tool documentation requires Cassandra to be stopped, and live compaction/lifecycle activity can change component sets while you inspect them; work from a stable copy.
2. It minimizes sensitive-data exposure and output volume while helping target the partition of interest.
3. Its documented CSV mode, because human-readable output can change formatting across versions.
4. No. It transforms SSTables and requires Cassandra to be stopped; it needs an explicit maintenance/rollback plan.
5. Preserve evidence/copies, identify scope and replica health, then follow supported verification/recovery procedures rather than manually editing files.
Production judgment
SSTable files expose valuable operational evidence, but they are not an application contract. Production decisions must combine logical data shape with physical storage state: partition rows/bytes, mutation and TTL/delete rates, RF/CL, read/write p95/p99, SSTables per read, compaction strategy and backlog, repaired/unrepaired state, compression ratio, CPU/decompression cost, disk throughput/latency, temporary compaction/streaming space, snapshot/backup retention, topology changes, repair cadence, and restore objectives. Include JVM/GC, page cache/off-heap use, driver timeouts/retries/idempotency, network bandwidth, tenant isolation, encryption-at-rest expectations, filesystem/device behavior, and managed-service restrictions.
Do not plan disk capacity as “live dataset bytes × RF” only. Flush, compaction, repair, streaming, snapshots, incremental backups, anti-compaction, restore staging, and operational safety margins can temporarily retain additional SSTables. Do not manually delete or edit component files to recover space. If space is critical, first stop unsafe automation, measure ownership/snapshots/compaction/streaming, and choose a documented recovery path with rollback. Lesson 4 follows immutable SSTable data as Cassandra deliberately streams it during bootstrap, repair, rebuild, and topology changes, where network and disk headroom become coupled.
Summary and next bridge
Offline SSTable tools are powerful because they bypass normal query abstractions; that is also why they need disciplined isolation. Copy first, inspect narrowly, protect sensitive values, and separate diagnostic tools from file-rewriting tools. Next, connect SSTables to streaming and topology operations.
Authoritative references
These are version-sensitive sources of truth. Re-check them when regenerating the lesson because SSTable formats, utilities, defaults, and topology procedures evolve.
- Apache Cassandra downloads / current GA baseline
- Apache Cassandra storage engine and SSTable components
- Cassandra 5.0 cassandra.yaml SSTable format and streaming settings
- Apache Cassandra compression guidance
- Apache Cassandra SSTable tools safety overview
- sstablemetadata
- sstabledump
- sstablepartitions
- nodetool netstats
- Repair and streaming differences