Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting: Diagnostics, Failure Modes, Security, and Performance
Diagnose capacity, inode, copy-on-write, logging, cache, migration, and storage-backend failures while preserving first-failure evidence and refusing unsupported edits to Docker-managed storage.
Learning objectives
- Diagnose storage incidents without deleting evidence or broad sets of resources.
- Differentiate byte exhaustion, inode exhaustion, writable-layer growth, volume growth, log growth, cache growth, and metadata/backend problems.
- Interpret copy-on-write symptoms and storage-backend evidence before changing architecture.
- Repair one intentionally broken storage pattern by moving durable data out of the writable layer without hiding the original cause.
- Define safe escalation for data-root/storage-driver migration and Docker Desktop disk-image issues.
1. Preserve first-failure evidence
Before cleanup, capture the error, affected container/image/volume
IDs, timestamps, context, Engine version,
docker system df -v, docker ps -a --size,
builder disk usage, logging configuration, and host/VM capacity
evidence. Deletion can make an incident appear “fixed” while
destroying the information needed to prevent recurrence.
Keep the original command error or application log exactly as observed. A recreated container with empty state is not proof of the original failure mechanism.
2. Evidence-first diagnostic sequence
- Preserve the first error and a UTC timestamp.
-
Confirm
docker context show, client/server versions, and host/Desktop architecture. - Confirm image-store/storage-driver/snapshotter and Docker root evidence.
-
Measure Docker-wide usage with
docker system df -v. -
Measure container writable layers with
docker ps -a --size. - Inspect volumes/mounts and their retention owners.
-
Inspect the selected builder with
docker buildx du. - Inspect logging driver/options and external sink health.
- Check filesystem free bytes and inodes on the daemon’s actual storage domains.
- Apply the least destructive correction and remeasure only the relevant scope.
3. Failure mode: broad system prune as first response
docker system prune -a --volumes. This widens one
storage symptom into deletion across stopped containers, networks,
images, build cache, and—when requested—volumes.
Repair the process, not just the disk: inventory first, identify exact resource classes, protect retained volumes/release images, then delete only proven disposable targets. If recurring pressure comes from build cache or logs, set an explicit budget/rotation policy so the incident does not repeat.
4. Failure mode: manually deleting Docker internal directories
Removing files beneath Docker or managed containerd data directories bypasses the metadata/control plane. The filesystem may show free space while Docker’s databases, references, snapshots, or content metadata now point at missing objects. This is corruption, not cleanup.
Use supported Docker/BuildKit lifecycle commands. If the store is already inconsistent, freeze mutation, collect daemon logs and architecture/version evidence, back up application data, and follow a version-specific upstream recovery procedure.
5. Failure mode: confusing volume data with writable-layer data
docker inspect <container> --format '{{json .Mounts}}'
docker ps -a --size --filter id=<container>
docker system df -v
A growing writable-layer number points toward writes outside intended mounts. A growing volume can leave writable-layer size nearly unchanged. Diagnose mount targets and application paths before moving or deleting data.
6. Intentionally broken pattern: durable state written into the container layer
The following disposable demonstration intentionally writes “state” to the writable layer. It is broken by design because replacing the container loses that state and CoW growth is tied to container lifecycle.
docker run -d \
--name dca32-broken-state \
--label devops-academy.lab=chapter32-diagnostic alpine:3.22.1 sh -c 'dd if=/dev/zero of=/var/lib/app-state.bin bs=1M count=8; sleep 300'
docker ps --size --filter name=dca32-broken-state
docker exec dca32-broken-state ls -lh /var/lib/app-state.bin
docker inspect dca32-broken-state --format '{{json .Mounts}}'
Interpretation: the writable size grows and the mount list has no durable mount for that path. Preserve those outputs as the causal evidence; do not simply remove/recreate the container yet.
7. Repair the broken pattern with a bounded volume
docker volume create --label devops-academy.lab=chapter32-diagnostic dca32-diagnostic-data
docker run \
--rm \
--volumes-from dca32-broken-state \
--mount type=volume,src=dca32-diagnostic-data,dst=/recovered alpine:3.22.1 sh -c 'cp /var/lib/app-state.bin /recovered/app-state.bin 2>/dev/null || true'
The command above demonstrates why recovery design must be
deliberate: --volumes-from only brings mounts, not
arbitrary writable-layer paths from another container, so this
attempted shortcut will not copy the broken container’s private
writable file. Preserve that result—it is useful failure evidence.
The supported correction is application-specific export/copy before replacement. For this synthetic lab, use Docker’s copy API, then load the file into the volume:
mkdir -p dca32-recovery
docker cp dca32-broken-state:/var/lib/app-state.bin dca32-recovery/app-state.bin
docker run \
--rm \
--mount type=bind,src="$(pwd)/dca32-recovery",dst=/src,readonly \
--mount type=volume,src=dca32-diagnostic-data,dst=/data alpine:3.22.1 sh -c 'cp /src/app-state.bin /data/app-state.bin; sha256sum /data/app-state.bin'
docker run --rm --mount type=volume,src=dca32-diagnostic-data,dst=/var/lib,readonly alpine:3.22.1 ls -lh /var/lib/app-state.bin
Now the evidence distinguishes recovery from redesign: first extract the only copy of data, then place future durable state in a volume.
8. Cleanup this diagnostic example exactly
docker rm -f dca32-broken-state
docker volume rm dca32-diagnostic-data
rm -rf dca32-recovery
Only the explicitly named diagnostic resources are deleted. No prune operation is necessary.
9. Failure mode: inode exhaustion
Symptoms can include no space left on device even when
df -h reports free capacity. On the daemon host,
inspect df -i for the actual Docker/containerd
filesystems. Then identify which workload or cache creates many
small files; deleting a few large images will not repair inode
pressure if the source is millions of tiny cache files.
On Docker Desktop, filesystem internals live inside the managed VM.
Use Desktop-supported diagnostics/settings instead of assuming the
host’s C:, macOS APFS, or Linux root filesystem
directly represents the Engine filesystem.
10. Failure mode: unbounded local logs
If writable-layer sizes are small but disk growth correlates with
stdout/stderr volume, inspect .HostConfig.LogConfig and
the daemon’s logging default. The default
json-file driver does not rotate unless configured. Do
not truncate daemon-managed log files manually. Apply bounded
per-container settings in a disposable test, then plan a controlled
recreation/default-policy rollout.
11. Failure mode: data-root move without complete architecture plan
Stopping Docker and copying only DockerRootDir can be incomplete on Engine 29’s containerd image store because managed image/snapshot state is a separate storage domain. Conversely, copying directories while daemons are writing to them risks inconsistency. A migration requires downtime/consistency, backups, architecture-specific official steps, permission/filesystem validation, and rollback.
12. Failure mode: changing storage backend on existing state
Switching storage drivers or image-store architecture can hide existing local images/containers or require migration. Preserve immutable images in a registry/export when appropriate, back up volumes separately, record current daemon configuration and version, and validate the new backend on disposable hosts first.
A snapshotter or driver is a materialization mechanism, not image identity. The same immutable manifest digest can be materialized through different supported backends.
13. Triage table
| Symptom | Most discriminating evidence | Likely owning layer | Least destructive next action |
|---|---|---|---|
| Host disk full; writable layers small | Log config + volume/cache inventory | Logs, volumes, builder cache, image content | Measure exact category; rotate/retire exact owner |
| ENOSPC with free GB | Filesystem inode usage | Backing filesystem/inode allocation | Find high-file-count source; target that lifecycle |
| One container writable layer grows rapidly | docker ps --size + mounts |
Application write path / CoW layer | Move intended durable/high-write path to volume |
| Build host fills after many builds | docker buildx du |
BuildKit cache | Tune GC budget or clean known dedicated builder |
| After storage switch, images “disappear” | DriverStatus + config/history | Store architecture/migration state | Stop further mutation; verify old store still exists and follow official migration/rollback |
Knowledge check
Why should first-failure evidence be captured before cleanup?
Cleanup can destroy the object/log/cache state that proves what consumed space or why the operation failed.
What does the failed --volumes-from recovery
attempt teach?
Volumes-from exposes another container’s mounts, not its private writable layer; the storage ownership boundary matters.
A filesystem has 30 GB free but Docker gets ENOSPC. What should you inspect?
Inode availability on the actual daemon storage filesystem, plus the source of high file counts.
Why is manual deletion under Docker’s data directories unsafe?
It bypasses Docker/containerd metadata ownership and can leave references pointing to missing state, corrupting the store.
A container’s writable layer is tiny but disk usage keeps climbing. Which chapter-23 evidence is immediately relevant?
Logging driver/options and log retention, because local logs are not included in writable-layer size.
Official references and version notes
Diagnostic rule: every correction in this lesson preserves the distinction between artifact identity, materialized filesystem state, persistent application data, build cache, and daemon-managed logs. Internal Docker/containerd files are evidence, not a supported cleanup interface.
- Docker Docs — containerd image store with Docker Engine — Engine 29 fresh-install default, snapshotter reporting, switching, and migration boundaries.
- Docker Docs — Storage drivers — image layers, container writable layers, copy-on-write, and container-size accounting caveats.
-
Docker Docs — OverlayFS storage driver
— classic
overlay2, backing-filesystem requirements, lower/upper/merged/workdir semantics, and migration cautions. - Docker Docs — Select a storage driver — current containerd snapshotter versus classic-driver guidance and platform distinctions.
-
Docker CLI —
docker system df— daemon disk-usage summary and verbose object-level accounting. -
Docker CLI —
docker system prune— deletion scope and supported filters; referenced here to understand risk, not as the lab cleanup path. -
Docker Buildx —
buildx du— builder-cache usage, shared/private/reclaimable records, and detailed evidence. -
Docker Buildx —
buildx prune— cache filters and space-budget controls. - Docker Build — Build garbage collection — periodic BuildKit GC, default policy shape, and builder-specific configuration.
- Docker Docs — Docker daemon configuration overview — Docker data-root and the separate managed containerd root used by the Engine 29 containerd image store.
-
Docker Docs — Configure logging drivers
—
json-filegrowth risk, rotation guidance, and the rotatinglocaldriver. - Docker Engine 29 release notes — current Engine 29 behavior and storage/image-store fixes.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.