Measure compression as a storage/CPU/cache-density tradeoff rather than assuming the smallest file is always the fastest workload.
Compression for Collections and Indexes: Space, CPU, Cache Density, and Workload Tradeoffs
AtlasMart can save disk with Snappy, zstd, or zlib, but storage savings can change CPU cost, cache density, and I/O. The right compressor is therefore workload evidence, not a universal ranking.
Learning objectives
Compare Snappy, zstd, zlib, and no block compression using the same AtlasMart documents.
Inspect the actual collection creation string and storage statistics rather than assuming a compressor was applied.
Explain why collection data is uncompressed in the WiredTiger internal cache even when disk files are compressed.
Observe index prefix compression and separate collection compression from index behavior.
Evaluate storage savings together with CPU, read/write latency, and cache/I/O effects.
This lesson pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where a driver workload is useful.
Topology: disposable standalone. The host publishes only
127.0.0.1:27188. Authentication and TLS are
disabled only for this isolated disposable lab; production
security remains the Chapter 22 prerequisite. Default read/write
concern and primary read preference are used unless a comparison
says otherwise.
FCV is observed and never changed.
Atlas/Search/Vector Search/KMS/Enterprise capabilities are not
required. WiredTiger internals are treated as version-sensitive
implementation details; use supported MongoDB commands and
metrics instead of editing .wt files or
undocumented knobs. All compressor comparisons use separate
disposable collections in the same process so server version and
host conditions stay as constant as practical. The workload
still is not a substitute for production benchmarking. Product
runtime labs were not executed in the generation environment, so
cache ratios, checkpoint durations, journal sync times, disk
bytes, and latency percentiles must be measured locally rather
than copied as invented values.
1. Compression changes both capacity and the amount of work around I/O
WiredTiger uses Snappy block compression by default for most
ordinary collections and prefix compression for indexes. zstd
and zlib can often compress more, while none avoids
block-compression CPU. Time-series collections have a different
default (zstd), but this lesson uses ordinary AtlasMart event
documents. The critical distinction is that compressed files can
improve disk capacity and filesystem-cache density while
collection data in the WiredTiger internal cache is represented
uncompressed for manipulation.
| Choice | Typical motivation | Tradeoff to measure |
|---|---|---|
| Snappy | Default balance for ordinary collections | Moderate compression with relatively low CPU |
| zstd | Higher compression ratio / less I/O | Additional compression/decompression CPU; level is configurable and version-sensitive |
| zlib | Alternative high compression | CPU cost may be greater for the workload |
| none | CPU-sensitive or already-incompressible payload experiments | More disk bytes and potentially more I/O/filesystem-cache pressure |
| Index prefix compression | Default index optimization | Affects index representation; not the same mechanism as collection block compression |
2. Create one collection per compressor and verify the creation strings
docker rm -f atlasmart-ch24-l3 2>/dev/null || truedocker volume rm atlasmart-ch24-l3-db 2>/dev/null || truedocker run -d --name atlasmart-ch24-l3 \ -p 127.0.0.1:27188:27017 \ -v atlasmart-ch24-l3-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27188/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27188/admin" --quiet --eval ' printjson(db.adminCommand({buildInfo:1})); printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1})); printjson(db.serverStatus().storageEngine);'
const d=db.getSiblingDB("atlasmart");for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) d[n].drop();d.createCollection("cmp_none", {storageEngine:{wiredTiger:{configString:"block_compressor=none"}}});d.createCollection("cmp_snappy",{storageEngine:{wiredTiger:{configString:"block_compressor=snappy"}}});d.createCollection("cmp_zstd", {storageEngine:{wiredTiger:{configString:"block_compressor=zstd"}}});d.createCollection("cmp_zlib", {storageEngine:{wiredTiger:{configString:"block_compressor=zlib"}}});for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) { const s=d[n].stats(); print(n, (s.wiredTiger?.creationString||"").split(",").filter(x=>x.startsWith("block_compressor="))[0]);}
3. Load identical compressible AtlasMart documents and record storage plus write cost
python -m venv /tmp/atlasmart-ch24-compress-venv. /tmp/atlasmart-ch24-compress-venv/bin/activatepython -m pip install --disable-pip-version-check "pymongo==4.17.0"python - <<'PY'from statistics import medianfrom time import perf_counterfrom pymongo import MongoClientclient=MongoClient("mongodb://127.0.0.1:27188")db=client.atlasmartnames=["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]for name in names: c=db[name]; samples=[] for batch in range(50): docs=[] for j in range(500): i=batch*500+j docs.append({ "_id":i, "tenantId":f"tenant-{i%8}", "kind":"inventory-snapshot", "sku":f"SKU-{i%2000:05d}", "payload":"sensor="+("A"*700)+";warehouse="+("W"*(i%5+20))+";status=healthy;", "counter":i%100 }) t=perf_counter(); c.insert_many(docs,ordered=True); samples.append((perf_counter()-t)*1000) q=sorted(samples) st=db.command("collStats",name) print(name,{ "logical_size":st["size"], "storage_size":st["storageSize"], "p50_batch_ms":median(q), "p95_batch_ms":q[int(.95*(len(q)-1))] })PY
Do not rank compressors from one run. The generated payload is intentionally compressible so the storage effect is visible; real payloads may already be compressed/encrypted or have very different entropy. Re-run with representative BSON shapes and report CPU and storage-device context.
4. Add the same index and inspect prefix-compression evidence separately
const d=db.getSiblingDB("atlasmart");for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) { d[n].createIndex({tenantId:1,sku:1,counter:1},{name:"idx_tenant_sku_counter"}); const s=d[n].stats({indexDetails:true,indexDetailsName:"idx_tenant_sku_counter"}); const creation=s.indexDetails?.idx_tenant_sku_counter?.creationString || ""; printjson({name:n,indexBytes:s.indexSizes?.idx_tenant_sku_counter, prefixSetting:creation.split(",").filter(x=>x.startsWith("prefix_compression="))[0]});}
Index prefix compression is distinct from the collection block compressor. Also note that index bytes in the WiredTiger internal cache use an internal representation that can still benefit from prefix compression, while collection values are uncompressed there.
5. Read-path experiment: storage savings do not automatically mean lower query latency
const d=db.getSiblingDB("atlasmart");for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) { const t=Date.now(); const docs=d[n].find({tenantId:"tenant-3",counter:{$gte:40,$lt:60}}).hint("idx_tenant_sku_counter").limit(1000).toArray(); const ms=Date.now()-t; printjson({name:n,returned:docs.length,wallClockMs:ms,stats:d[n].stats().storageSize});}
“zstd made the collection smallest, therefore zstd is the fastest” is not a valid inference. Compression can reduce physical reads and improve filesystem-cache density while increasing CPU. The result depends on data entropy, working set, CPU headroom, storage latency, cache state, read/write mix, and compression level.
6. Production judgment
Start with MongoDB defaults unless representative evidence justifies a change. Compare logical bytes, storage bytes, index bytes, CPU, read/write distributions, cache/page I/O, and host disk metrics. If the application encrypts or pre-compresses payloads, database compression may produce little space savings. If the workload is I/O-bound, stronger compression can improve total latency despite added CPU; if CPU-bound, the opposite can occur. Existing collections keep the compressor chosen at creation, so changing the server default affects new collections rather than rewriting old data. The next lesson deliberately constrains cache to make eviction and admission signals observable.
Check your understanding
- Why can stronger on-disk compression improve read latency even though decompression uses CPU?
- Is collection block compression the same as index prefix compression?
- Why must the creation string be inspected in the lab?
-
Does a smaller
storageSizemean fewer WiredTiger internal cache bytes for the same values? - What must accompany a compression recommendation?
Review the answers
1. It can reduce physical I/O and increase filesystem-cache density enough to outweigh decompression CPU on an I/O-bound workload.
2. No. They are separate mechanisms with separate configuration and representations.
3. It proves the collection actually uses the intended compressor instead of relying on an assumption.
4. Not directly; collection values in the WiredTiger internal cache use an uncompressed internal representation.
5. Representative data entropy, CPU, storage, cache state, read/write mix, latency distributions, and measured storage/index sizes.
Authoritative references
WiredTiger metrics and internal field names are implementation- and version-sensitive. The lesson uses documented MongoDB interfaces for evidence and requires re-checking the current server manual before relying on exact metric names or defaults in a later release.
- WiredTiger Storage Engine
- Storage FAQ
- serverStatus
- Self-Managed Diagnostics FAQ
- Journaling
- Configure Journaling
- Write Concern
- Write Operation Performance
- Configuration File Options
- Server Parameters
- mongod Options
- db.collection.stats()
- $collStats
- dbStats
- db.createCollection() Storage Engine Options
- Create Indexes and Storage Engine Options
- Performance Tuning
- Production Notes
- Log Messages
- MongoDB 8.3 Release Notes
- mongosh Release Notes
- PyMongo Release Notes