Measure compression as a storage/CPU/cache-density tradeoff rather than assuming the smallest file is always the fastest workload.

Compression for Collections and Indexes: Space, CPU, Cache Density, and Workload Tradeoffs

AtlasMart can save disk with Snappy, zstd, or zlib, but storage savings can change CPU cost, cache density, and I/O. The right compressor is therefore workload evidence, not a universal ranking.

Advanced120–200 minutesCompression comparison labMongoDB 8.3.8 · WiredTiger · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Compare Snappy, zstd, zlib, and no block compression using the same AtlasMart documents.

02

Inspect the actual collection creation string and storage statistics rather than assuming a compressor was applied.

03

Explain why collection data is uncompressed in the WiredTiger internal cache even when disk files are compressed.

04

Observe index prefix compression and separate collection compression from index behavior.

05

Evaluate storage savings together with CPU, read/write latency, and cache/I/O effects.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where a driver workload is useful. Topology: disposable standalone. The host publishes only 127.0.0.1:27188. Authentication and TLS are disabled only for this isolated disposable lab; production security remains the Chapter 22 prerequisite. Default read/write concern and primary read preference are used unless a comparison says otherwise. FCV is observed and never changed. Atlas/Search/Vector Search/KMS/Enterprise capabilities are not required. WiredTiger internals are treated as version-sensitive implementation details; use supported MongoDB commands and metrics instead of editing .wt files or undocumented knobs. All compressor comparisons use separate disposable collections in the same process so server version and host conditions stay as constant as practical. The workload still is not a substitute for production benchmarking. Product runtime labs were not executed in the generation environment, so cache ratios, checkpoint durations, journal sync times, disk bytes, and latency percentiles must be measured locally rather than copied as invented values.

1. Compression changes both capacity and the amount of work around I/O

WiredTiger uses Snappy block compression by default for most ordinary collections and prefix compression for indexes. zstd and zlib can often compress more, while none avoids block-compression CPU. Time-series collections have a different default (zstd), but this lesson uses ordinary AtlasMart event documents. The critical distinction is that compressed files can improve disk capacity and filesystem-cache density while collection data in the WiredTiger internal cache is represented uncompressed for manipulation.

Choice Typical motivation Tradeoff to measure
Snappy Default balance for ordinary collections Moderate compression with relatively low CPU
zstd Higher compression ratio / less I/O Additional compression/decompression CPU; level is configurable and version-sensitive
zlib Alternative high compression CPU cost may be greater for the workload
none CPU-sensitive or already-incompressible payload experiments More disk bytes and potentially more I/O/filesystem-cache pressure
Index prefix compression Default index optimization Affects index representation; not the same mechanism as collection block compression

2. Create one collection per compressor and verify the creation strings

start the comparison server
docker rm -f atlasmart-ch24-l3 2>/dev/null || truedocker volume rm atlasmart-ch24-l3-db 2>/dev/null || truedocker run -d --name atlasmart-ch24-l3 \  -p 127.0.0.1:27188:27017 \  -v atlasmart-ch24-l3-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27188/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27188/admin" --quiet --eval '  printjson(db.adminCommand({buildInfo:1}));  printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}));  printjson(db.serverStatus().storageEngine);'
create collections with explicit per-collection compressors
const d=db.getSiblingDB("atlasmart");for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) d[n].drop();d.createCollection("cmp_none",  {storageEngine:{wiredTiger:{configString:"block_compressor=none"}}});d.createCollection("cmp_snappy",{storageEngine:{wiredTiger:{configString:"block_compressor=snappy"}}});d.createCollection("cmp_zstd",  {storageEngine:{wiredTiger:{configString:"block_compressor=zstd"}}});d.createCollection("cmp_zlib",  {storageEngine:{wiredTiger:{configString:"block_compressor=zlib"}}});for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) {  const s=d[n].stats();  print(n, (s.wiredTiger?.creationString||"").split(",").filter(x=>x.startsWith("block_compressor="))[0]);}

3. Load identical compressible AtlasMart documents and record storage plus write cost

PyMongo deterministic load with batch latency samples
python -m venv /tmp/atlasmart-ch24-compress-venv. /tmp/atlasmart-ch24-compress-venv/bin/activatepython -m pip install --disable-pip-version-check "pymongo==4.17.0"python - <<'PY'from statistics import medianfrom time import perf_counterfrom pymongo import MongoClientclient=MongoClient("mongodb://127.0.0.1:27188")db=client.atlasmartnames=["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]for name in names:    c=db[name]; samples=[]    for batch in range(50):        docs=[]        for j in range(500):            i=batch*500+j            docs.append({                "_id":i,                "tenantId":f"tenant-{i%8}",                "kind":"inventory-snapshot",                "sku":f"SKU-{i%2000:05d}",                "payload":"sensor="+("A"*700)+";warehouse="+("W"*(i%5+20))+";status=healthy;",                "counter":i%100            })        t=perf_counter(); c.insert_many(docs,ordered=True); samples.append((perf_counter()-t)*1000)    q=sorted(samples)    st=db.command("collStats",name)    print(name,{        "logical_size":st["size"],        "storage_size":st["storageSize"],        "p50_batch_ms":median(q),        "p95_batch_ms":q[int(.95*(len(q)-1))]    })PY

Do not rank compressors from one run. The generated payload is intentionally compressible so the storage effect is visible; real payloads may already be compressed/encrypted or have very different entropy. Re-run with representative BSON shapes and report CPU and storage-device context.

4. Add the same index and inspect prefix-compression evidence separately

create identical indexes and inspect index details
const d=db.getSiblingDB("atlasmart");for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) {  d[n].createIndex({tenantId:1,sku:1,counter:1},{name:"idx_tenant_sku_counter"});  const s=d[n].stats({indexDetails:true,indexDetailsName:"idx_tenant_sku_counter"});  const creation=s.indexDetails?.idx_tenant_sku_counter?.creationString || "";  printjson({name:n,indexBytes:s.indexSizes?.idx_tenant_sku_counter,    prefixSetting:creation.split(",").filter(x=>x.startsWith("prefix_compression="))[0]});}

Index prefix compression is distinct from the collection block compressor. Also note that index bytes in the WiredTiger internal cache use an internal representation that can still benefit from prefix compression, while collection values are uncompressed there.

5. Read-path experiment: storage savings do not automatically mean lower query latency

same query shape across all four collections
const d=db.getSiblingDB("atlasmart");for (const n of ["cmp_none","cmp_snappy","cmp_zstd","cmp_zlib"]) {  const t=Date.now();  const docs=d[n].find({tenantId:"tenant-3",counter:{$gte:40,$lt:60}}).hint("idx_tenant_sku_counter").limit(1000).toArray();  const ms=Date.now()-t;  printjson({name:n,returned:docs.length,wallClockMs:ms,stats:d[n].stats().storageSize});}
Deliberately wrong conclusion

“zstd made the collection smallest, therefore zstd is the fastest” is not a valid inference. Compression can reduce physical reads and improve filesystem-cache density while increasing CPU. The result depends on data entropy, working set, CPU headroom, storage latency, cache state, read/write mix, and compression level.

6. Production judgment

Start with MongoDB defaults unless representative evidence justifies a change. Compare logical bytes, storage bytes, index bytes, CPU, read/write distributions, cache/page I/O, and host disk metrics. If the application encrypts or pre-compresses payloads, database compression may produce little space savings. If the workload is I/O-bound, stronger compression can improve total latency despite added CPU; if CPU-bound, the opposite can occur. Existing collections keep the compressor chosen at creation, so changing the server default affects new collections rather than rewriting old data. The next lesson deliberately constrains cache to make eviction and admission signals observable.

Check your understanding

  1. Why can stronger on-disk compression improve read latency even though decompression uses CPU?
  2. Is collection block compression the same as index prefix compression?
  3. Why must the creation string be inspected in the lab?
  4. Does a smaller storageSize mean fewer WiredTiger internal cache bytes for the same values?
  5. What must accompany a compression recommendation?
Review the answers

1. It can reduce physical I/O and increase filesystem-cache density enough to outweigh decompression CPU on an I/O-bound workload.

2. No. They are separate mechanisms with separate configuration and representations.

3. It proves the collection actually uses the intended compressor instead of relying on an assumption.

4. Not directly; collection values in the WiredTiger internal cache use an uncompressed internal representation.

5. Representative data entropy, CPU, storage, cache state, read/write mix, latency distributions, and measured storage/index sizes.

Authoritative references

WiredTiger metrics and internal field names are implementation- and version-sensitive. The lesson uses documented MongoDB interfaces for evidence and requires re-checking the current server manual before relying on exact metric names or defaults in a later release.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.