Chapter 32Lesson 01~155 minutes

Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting: Concepts, Architecture, and Mental Model

Build a causal storage model that separates image/content blobs, snapshots or classic layers, container writable state, volumes, BuildKit cache, logs, metadata, filesystem capacity, inode pressure, and reclaimable estimates before any deletion.

StorageCopy-on-writeoverlay2SnapshottersDisk evidence

Learning objectives

  • Explain where Docker disk usage can originate and why one aggregate “Docker size” number is insufficient for diagnosis.
  • Distinguish Engine 29 containerd snapshotter storage from classic overlay2 while preserving the same copy-on-write mental model.
  • Separate immutable image/content storage, container writable layers, volumes, BuildKit cache, logs, and metadata.
  • Interpret bytes, shared bytes, unique bytes, reclaimable estimates, free space, and inode availability without double-counting shared layers.
  • Define a safe evidence-first path from pressure observation to exact object selection and bounded cleanup.

1. The practical problem: “Docker uses too much disk” is not a diagnosis

Disk-pressure incidents often begin with a host-volume alert or a Docker Desktop disk image growing unexpectedly. The tempting response is deletion. The correct response is classification. Docker can consume storage through OCI image blobs, unpacked snapshots or classic graph-driver layers, each container’s writable layer, named and anonymous volumes, BuildKit cache, local log storage, metadata, and—on Desktop—the Linux VM disk image that contains these resources.

Those categories have different ownership and recovery rules. An unused image might be rebuildable. A volume might contain the only copy of application data. Build cache may be disposable but expensive to regenerate. A container log can grow without changing docker ps --size. Therefore this chapter treats deletion as the last step after the object and its dependencies are proven.

2. Causal storage model

Storage pressure is the sum of independently owned storage classes, observed before any cleanup
flowchart TD
  A[Registry/image content blobs] --> F[Docker storage domains]
  B[Snapshot or classic image layers] --> F
  C[Container writable layer] --> F
  D[Volumes] --> F
  E[BuildKit cache and logs] --> F
  F --> G[Backing filesystem or Desktop disk image]
  G --> H[Bytes + inodes + I/O pressure]
  H --> I[Read-only inventory]
  I --> J[Exact retention decision]
  J --> K[Targeted deletion or no action]
            

The top nodes are intentionally separate. Registry/image content gives Docker immutable bytes. A snapshotter or storage driver materializes filesystem state. Containers add copy-on-write changes. Volumes live outside the writable layer. Builders retain intermediate cache. Logging drivers can retain stdout/stderr independently. All of that eventually competes for capacity, inodes, and I/O on one or more filesystems.

The causal order for operations is therefore: capacity signal → identify storage domain → inventory exact objects → establish references/retention → predict reclaimed state → remove only owned targets → remeasure. Skipping the inventory step turns cleanup into guesswork.

3. Engine 29 has two storage architectures you may encounter

Observed architecture What stores image/container filesystem state Evidence to capture Operational rule
Engine 29 fresh install / migrated containerd image store containerd content store + snapshotter, normally overlayfs on Linux docker info, especially DriverStatus showing io.containerd.snapshotter.v1 Do not use old overlay2 directory assumptions as the management interface.
Upgraded/classic daemon Classic graph driver, commonly overlay2 docker info storage-driver and backing-filesystem fields Classic OverlayFS terminology is directly applicable, but internal directories still are not an operator API.
Docker Desktop Engine storage inside a managed Linux VM/disk image Docker Desktop settings + Docker CLI evidence Host-visible free space and VM disk-image allocation are separate observations.

4. Copy-on-write: why starting many containers is cheaper than copying images

An image filesystem is read-only from the container’s perspective. Containers share those image layers and receive a writable top layer or writable snapshot. Reads can come from lower image content; the first modification of a file can require materializing changed data into the writable layer. This is the essence of copy-on-write (CoW).

CoW explains two common surprises. First, virtual size is not additive across containers because image bytes are shared. Second, write-heavy application data in a container writable layer can create extra copy-up work and makes lifecycle coupling worse. Durable or high-write data normally belongs in volumes or external storage, not the ephemeral writable layer.

5. What docker ps --size proves—and what it omits

docker ps --size reports a container’s writable-layer size plus a “virtual” size that combines the image’s read-only data with that writable state. Docker’s documentation warns that shared image layers make virtual-size totals misleading.

It also does not account for everything attributable to a workload: volume bytes, bind-mounted host data, logging-driver files, swap, and some metadata are outside that writable-layer number. Use it to answer “how much did this container write into its own copy-on-write layer?”, not “how much disk does this application consume in total?”

6. docker system df is the Docker-wide starting inventory

docker system df
docker system df -v

The summary separates images, containers, local volumes, and build cache and reports active/total counts, size, and reclaimable estimates. The verbose form adds object-level detail, including shared and unique image bytes. “Reclaimable” is a candidate estimate based on references known to Docker; it is not permission to delete. Business retention, disaster recovery, future rollback, and preexisting lab ownership remain separate decisions.

7. Builder cache is its own lifecycle

docker buildx ls
docker buildx du
docker buildx du --format=pretty

BuildKit cache can be large even when few containers exist. docker buildx du reports the currently selected builder’s records, whether they are reclaimable, shared, mutable/private, and their sizes. Shared cache may refer to storage also needed by images; deleting cache metadata is not the same as deleting a uniquely owned blob.

BuildKit also has periodic garbage collection. Current defaults use ordered policies based on cache type, age, sharing, and space thresholds. Manual pruning should complement an understood cache budget, not replace one.

8. Volumes are persistent data, not “leftover container files”

Chapter 20 established the lifecycle separation: deleting a container does not imply deleting its named volume. Disk-pressure work must preserve that boundary. A volume can be unused right now and still be intentionally retained for restore, rollback, or a stopped service.

Record each candidate’s name, driver, labels, consuming containers, backup status, and retention owner before removal. Never infer safety solely from “not mounted by a running container.”

9. Logs can exhaust disk without growing the writable layer

Docker’s default json-file logging driver does not rotate by default. High-volume stdout/stderr can therefore consume substantial storage while docker ps --size remains small. Docker recommends the local driver for many non-Kubernetes use cases because it rotates automatically and stores logs efficiently.

Log files are daemon-managed internals. Inspect logging configuration through Docker; do not truncate or edit Docker’s internal log files behind the daemon. A safe prevention policy sets bounded rotation for newly created containers and treats existing containers explicitly because daemon logging defaults are not retroactive.

10. Capacity includes inodes, not only bytes

A filesystem can report free gigabytes and still reject file creation if its inode pool is exhausted. Workloads that create huge numbers of tiny files—package caches, temporary trees, extracted dependency stores, or application churn—can hit this failure mode.

On an authorized local Linux daemon, pair byte evidence (df -h) with inode evidence (df -i) for the filesystems that actually contain Docker and managed containerd state. With remote contexts or Docker Desktop, host-shell filesystem commands may observe the wrong machine, so label that evidence carefully.

11. Engine 29 data-root nuance

Docker’s daemon documentation now makes an important split explicit. General daemon data such as volumes/configs remains under Docker’s data root (by default /var/lib/docker on Linux). When the Engine 29 containerd image store is active, image contents and container snapshots are stored under managed containerd state, documented under /var/lib/containerd by default.

Therefore setting Docker’s data-root alone does not relocate all image/snapshot bytes in the containerd-image-store architecture. Capacity planning and migration plans must identify every active storage domain first. This chapter does not mutate those roots.

12. Safety contract for cleanup

Never begin with broad cleanup. The failure mode docker system prune -a --volumes can remove multiple resource classes and erase useful rollback/data state. Likewise, deleting files manually under Docker or containerd data directories bypasses metadata ownership and can corrupt the store.

The mandatory labs use exact container/image/volume/builder names, preflight ownership guards, and labels. Prune commands are studied as policy mechanisms, but broad destructive execution is not the learning path.

13. Read-only first-pass checklist

docker context show
docker version
docker info
docker system df
docker system df -v
docker ps -a --size
docker volume ls
docker buildx ls
docker buildx du
docker info --format 'root={{.DockerRootDir}} driver={{.Driver}} logging={{.LoggingDriver}}'
docker info --format '{{json .DriverStatus}}'

Capture this before touching anything. It establishes endpoint identity, Engine version, storage architecture, logging baseline, object inventory, and builder-cache state.

Knowledge check

Why can’t you sum the “virtual size” of every container to estimate disk use?

A container has a 20 KB writable layer but the host loses gigabytes. Name two storage classes to inspect next.

What does “reclaimable” mean in docker system df?

Why is data-root alone insufficient for Engine 29 containerd-image-store migration planning?

What extra filesystem resource should be checked when bytes are available but file creation returns ENOSPC?

Next lesson

Next: Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting: Guided Hands-On Workflow and Core Operations

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Verified baseline date:

2026-09-22. Docker Engine 29.8.1 is the current Engine 29 baseline used for version-sensitive discussion. Fresh Engine 29 installs use the containerd image store by default; upgraded daemons can remain on classic overlay2. Always record actual local Engine/Desktop architecture.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.