Chapter 32Lesson 04~175 minutes

Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting: Diagnostics, Failure Modes, Security, and Performance

Diagnose capacity, inode, copy-on-write, logging, cache, migration, and storage-backend failures while preserving first-failure evidence and refusing unsupported edits to Docker-managed storage.

ENOSPCInodesMigrationDiagnosticsRecovery

Learning objectives

  • Diagnose storage incidents without deleting evidence or broad sets of resources.
  • Differentiate byte exhaustion, inode exhaustion, writable-layer growth, volume growth, log growth, cache growth, and metadata/backend problems.
  • Interpret copy-on-write symptoms and storage-backend evidence before changing architecture.
  • Repair one intentionally broken storage pattern by moving durable data out of the writable layer without hiding the original cause.
  • Define safe escalation for data-root/storage-driver migration and Docker Desktop disk-image issues.

1. Preserve first-failure evidence

Before cleanup, capture the error, affected container/image/volume IDs, timestamps, context, Engine version, docker system df -v, docker ps -a --size, builder disk usage, logging configuration, and host/VM capacity evidence. Deletion can make an incident appear “fixed” while destroying the information needed to prevent recurrence.

Keep the original command error or application log exactly as observed. A recreated container with empty state is not proof of the original failure mechanism.

2. Evidence-first diagnostic sequence

  1. Preserve the first error and a UTC timestamp.
  2. Confirm docker context show, client/server versions, and host/Desktop architecture.
  3. Confirm image-store/storage-driver/snapshotter and Docker root evidence.
  4. Measure Docker-wide usage with docker system df -v.
  5. Measure container writable layers with docker ps -a --size.
  6. Inspect volumes/mounts and their retention owners.
  7. Inspect the selected builder with docker buildx du.
  8. Inspect logging driver/options and external sink health.
  9. Check filesystem free bytes and inodes on the daemon’s actual storage domains.
  10. Apply the least destructive correction and remeasure only the relevant scope.

3. Failure mode: broad system prune as first response

Unsafe pattern: reaching immediately for docker system prune -a --volumes. This widens one storage symptom into deletion across stopped containers, networks, images, build cache, and—when requested—volumes.

Repair the process, not just the disk: inventory first, identify exact resource classes, protect retained volumes/release images, then delete only proven disposable targets. If recurring pressure comes from build cache or logs, set an explicit budget/rotation policy so the incident does not repeat.

4. Failure mode: manually deleting Docker internal directories

Removing files beneath Docker or managed containerd data directories bypasses the metadata/control plane. The filesystem may show free space while Docker’s databases, references, snapshots, or content metadata now point at missing objects. This is corruption, not cleanup.

Use supported Docker/BuildKit lifecycle commands. If the store is already inconsistent, freeze mutation, collect daemon logs and architecture/version evidence, back up application data, and follow a version-specific upstream recovery procedure.

5. Failure mode: confusing volume data with writable-layer data

docker inspect <container> --format '{{json .Mounts}}'
docker ps -a --size --filter id=<container>
docker system df -v

A growing writable-layer number points toward writes outside intended mounts. A growing volume can leave writable-layer size nearly unchanged. Diagnose mount targets and application paths before moving or deleting data.

6. Intentionally broken pattern: durable state written into the container layer

The following disposable demonstration intentionally writes “state” to the writable layer. It is broken by design because replacing the container loses that state and CoW growth is tied to container lifecycle.

docker run -d \
  --name dca32-broken-state \
  --label devops-academy.lab=chapter32-diagnostic   alpine:3.22.1   sh -c 'dd if=/dev/zero of=/var/lib/app-state.bin bs=1M count=8; sleep 300'

docker ps --size --filter name=dca32-broken-state
docker exec dca32-broken-state ls -lh /var/lib/app-state.bin
docker inspect dca32-broken-state --format '{{json .Mounts}}'

Interpretation: the writable size grows and the mount list has no durable mount for that path. Preserve those outputs as the causal evidence; do not simply remove/recreate the container yet.

7. Repair the broken pattern with a bounded volume

docker volume create   --label devops-academy.lab=chapter32-diagnostic   dca32-diagnostic-data

docker run \
  --rm \
  --volumes-from dca32-broken-state \
  --mount type=volume,src=dca32-diagnostic-data,dst=/recovered   alpine:3.22.1   sh -c 'cp /var/lib/app-state.bin /recovered/app-state.bin 2>/dev/null || true'

The command above demonstrates why recovery design must be deliberate: --volumes-from only brings mounts, not arbitrary writable-layer paths from another container, so this attempted shortcut will not copy the broken container’s private writable file. Preserve that result—it is useful failure evidence.

The supported correction is application-specific export/copy before replacement. For this synthetic lab, use Docker’s copy API, then load the file into the volume:

mkdir -p dca32-recovery
docker cp dca32-broken-state:/var/lib/app-state.bin dca32-recovery/app-state.bin

docker run \
  --rm \
  --mount type=bind,src="$(pwd)/dca32-recovery",dst=/src,readonly \
  --mount type=volume,src=dca32-diagnostic-data,dst=/data   alpine:3.22.1   sh -c 'cp /src/app-state.bin /data/app-state.bin; sha256sum /data/app-state.bin'

docker run --rm   --mount type=volume,src=dca32-diagnostic-data,dst=/var/lib,readonly   alpine:3.22.1   ls -lh /var/lib/app-state.bin

Now the evidence distinguishes recovery from redesign: first extract the only copy of data, then place future durable state in a volume.

8. Cleanup this diagnostic example exactly

docker rm -f dca32-broken-state
docker volume rm dca32-diagnostic-data
rm -rf dca32-recovery

Only the explicitly named diagnostic resources are deleted. No prune operation is necessary.

9. Failure mode: inode exhaustion

Symptoms can include no space left on device even when df -h reports free capacity. On the daemon host, inspect df -i for the actual Docker/containerd filesystems. Then identify which workload or cache creates many small files; deleting a few large images will not repair inode pressure if the source is millions of tiny cache files.

On Docker Desktop, filesystem internals live inside the managed VM. Use Desktop-supported diagnostics/settings instead of assuming the host’s C:, macOS APFS, or Linux root filesystem directly represents the Engine filesystem.

10. Failure mode: unbounded local logs

If writable-layer sizes are small but disk growth correlates with stdout/stderr volume, inspect .HostConfig.LogConfig and the daemon’s logging default. The default json-file driver does not rotate unless configured. Do not truncate daemon-managed log files manually. Apply bounded per-container settings in a disposable test, then plan a controlled recreation/default-policy rollout.

11. Failure mode: data-root move without complete architecture plan

Stopping Docker and copying only DockerRootDir can be incomplete on Engine 29’s containerd image store because managed image/snapshot state is a separate storage domain. Conversely, copying directories while daemons are writing to them risks inconsistency. A migration requires downtime/consistency, backups, architecture-specific official steps, permission/filesystem validation, and rollback.

12. Failure mode: changing storage backend on existing state

Switching storage drivers or image-store architecture can hide existing local images/containers or require migration. Preserve immutable images in a registry/export when appropriate, back up volumes separately, record current daemon configuration and version, and validate the new backend on disposable hosts first.

A snapshotter or driver is a materialization mechanism, not image identity. The same immutable manifest digest can be materialized through different supported backends.

13. Triage table

Symptom Most discriminating evidence Likely owning layer Least destructive next action
Host disk full; writable layers small Log config + volume/cache inventory Logs, volumes, builder cache, image content Measure exact category; rotate/retire exact owner
ENOSPC with free GB Filesystem inode usage Backing filesystem/inode allocation Find high-file-count source; target that lifecycle
One container writable layer grows rapidly docker ps --size + mounts Application write path / CoW layer Move intended durable/high-write path to volume
Build host fills after many builds docker buildx du BuildKit cache Tune GC budget or clean known dedicated builder
After storage switch, images “disappear” DriverStatus + config/history Store architecture/migration state Stop further mutation; verify old store still exists and follow official migration/rollback

Knowledge check

Why should first-failure evidence be captured before cleanup?

What does the failed --volumes-from recovery attempt teach?

A filesystem has 30 GB free but Docker gets ENOSPC. What should you inspect?

Why is manual deletion under Docker’s data directories unsafe?

A container’s writable layer is tiny but disk usage keeps climbing. Which chapter-23 evidence is immediately relevant?

Next lesson

Next: Checkpoint Lab — Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Diagnostic rule: every correction in this lesson preserves the distinction between artifact identity, materialized filesystem state, persistent application data, build cache, and daemon-managed logs. Internal Docker/containerd files are evidence, not a supported cleanup interface.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.