Chapter 32Lesson 03~165 minutes

Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting: Configuration, Design Choices, and Tradeoffs

Choose storage backend, filesystem placement, retention policy, logging rotation, cleanup cadence, and BuildKit cache budgets from workload evidence rather than blanket prune habits.

RetentionData rootLoggingGCTradeoffs

Learning objectives

  • Choose between classic overlay2 and containerd snapshotter terminology based on observed Engine architecture.
  • Design disk placement and monitoring that covers Docker data, managed containerd state, logs, volumes, and builders.
  • Set retention and cache budgets from rebuild cost, rollback needs, and capacity SLOs rather than arbitrary age alone.
  • Compare automatic BuildKit GC with manual targeted cleanup and explain when each belongs in operations.
  • Document migration prerequisites and rollback boundaries before changing storage drivers, snapshotters, or data locations.

1. Design from storage classes, not one “Docker partition” assumption

A production storage design begins by listing actual state classes: image/content bytes, unpacked root filesystems, writable layers, volumes, build cache, local logs, daemon metadata, and external bind-mounted application data. Map each class to the filesystem or VM disk that holds it, then assign capacity, performance, backup, retention, and alerting expectations.

Engine 29 makes this mapping more important because the containerd image store can place image/snapshot state under managed containerd storage while Docker’s general data root still holds other data.

2. Classic overlay2 versus containerd overlayfs snapshotter

Question Classic overlay2 Engine 29 containerd image store
Abstraction Docker graph/storage driver containerd content store + snapshotter
Common Linux backend overlay2 using kernel OverlayFS containerd overlayfs snapshotter
Fresh Engine 29 default No Yes, except documented unsupported modes such as current userns-remap limitation
Local multi-platform/attestation support Classic store has limitations Containerd store enables richer multi-platform/attestation storage
Operational inspection Docker CLI/info; do not edit overlay2 dirs Docker CLI/info; do not mutate managed containerd content/snapshots

3. OverlayFS copy-up has workload consequences

OverlayFS presents lower read-only layers and an upper writable layer. A write to a file originating in a lower layer can cause copy-up into upper state. For many normal application patterns this is efficient, but large frequently modified files or write-heavy databases are poor candidates for the container writable layer. Volumes bypass much of this lifecycle coupling and are the standard persistence path.

Do not “optimize” by editing Docker’s upper/lower directories. Optimize by changing application storage placement, image composition, cache behavior, and filesystem capacity with supported interfaces.

4. Separate Docker data filesystem: benefits and boundaries

Benefit Why it helps Boundary
Capacity isolation Docker growth does not consume the host root filesystem as quickly Containerd image-store bytes may be on a separate managed containerd root
Monitoring clarity Dedicated alerts can target Docker-specific bytes/inodes/I/O You must still monitor volumes/logs/builders and any other storage domains
I/O planning Choose SSD/filesystem suited to workload Do not place multiple independent daemons on the same data directory
Maintenance Backup/migration scope becomes explicit Daemon stop, backup, consistency, and supported migration steps remain required

5. Retention policy: classify by rebuildability and business value

State Typical rebuildability Retention question Preferred control
Pulled/release image Often re-pullable if registry retains immutable digest Needed for rollback/offline recovery? Release retention policy + digest records
Container writable layer Usually ephemeral Is application writing data in wrong place? Keep small; move durable data to volume
Named volume Potentially irreplaceable Backup/RPO/owner/expiry? Application data lifecycle; never generic prune first
Build cache Re-creatable but may be expensive How much feedback-time value per GB? BuildKit GC budget + dedicated builder policy
Logs Operational evidence, sometimes regulated Retention/search/RTO/sensitivity? Rotation + external aggregation where appropriate

6. BuildKit automatic GC versus manual cleanup

Current BuildKit garbage collection runs periodically and evaluates ordered policies. Defaults preferentially reclaim easily regenerated or stale cache and use space thresholds. For most developer systems, Docker states the defaults are sufficient. Large builders or constrained disks may justify explicit reservedSpace, maxUsedSpace, minFreeSpace, or custom policies.

Manual buildx prune is an immediate intervention. It supports filters and space targets, but it is still a deletion operation. A mature setup first budgets cache, then monitors buildx du, and uses manual pruning only when a known builder and policy justify it.

7. Automatic versus manual cleanup

Approach Strength Risk Good fit
Exact object deletion Maximum auditability and narrow blast radius Operational effort Incidents, training, protected hosts
Label/age-filtered prune Policy can scale to many disposable objects Filter semantics/labels must be governed consistently CI workers with enforced labeling
BuildKit GC Continuous cache-budget enforcement Mis-sized budget can hurt cache hit rate or disk headroom Long-lived builders
Ephemeral worker recreation Very strong lifecycle boundary Requires workflow designed for disposability CI/build fleet nodes

8. Cache budget is a performance decision, not only a storage decision

Build cache trades disk for feedback latency and network/compute savings. Too little cache increases rebuild time and dependency downloads; too much can starve the host. A sensible budget uses empirical data: cache hit rate, rebuild time, average/peak cache size, host free-space SLO, and cost of re-fetching dependencies.

For a dedicated builder, record docker buildx du over time and set a budget that preserves the most valuable working set while maintaining emergency free space.

9. Logs require a retention policy of their own

The default json-file driver exists for compatibility and does not rotate unless configured. Docker’s local driver rotates by default (current defaults are five files of 20 MB each before compression, about 100 MB per container maximum before compression). Production policy must decide whether local retention is enough or logs should be shipped to a managed sink.

Changing daemon logging defaults only affects newly created containers. A rollout plan therefore includes recreation timing and validation, not only a daemon.json edit.

10. Data-root migration: plan both Docker and containerd domains

A storage relocation is an administrative maintenance event, not a live copy exercise. Preserve an inventory, stop/coordinate workloads, back up persistent application data, verify free space/inodes/permissions/filesystem support at the target, and use the current official procedure for the active architecture.

For the classic store, Docker’s OverlayFS documentation shows stop/copy/configure/start verification patterns. With Engine 29’s containerd image store, Docker explicitly documents that Docker data-root does not move managed containerd image/snapshot state. A plan that copies only /var/lib/docker can therefore be incomplete.

11. Storage-driver or snapshotter changes are migrations, not tuning knobs

Changing a storage backend can make existing local images and containers inaccessible until you switch back or migrate them. This is not equivalent to changing a harmless daemon preference. Export/push immutable images, back up application data, document current backend and versions, test rollback, and schedule downtime when required.

Docker Desktop manages these internals differently; editing Linux Engine storage-driver configuration is not a supported Desktop method. Use Desktop’s supported disk-image controls instead.

12. Worked decision scenario

Scenario Choice Prerequisites/evidence Why
Developer workstation, frequent builds, disk occasionally tight Keep BuildKit GC defaults; monitor buildx du; bounded local logs Observed cache trend + disk headroom Preserves fast feedback without routine broad prune
CI runner dedicated to one pipeline and recreated daily Ephemeral builder/worker lifecycle Jobs have no durable local state; artifacts pushed immutably Lifecycle itself bounds cache and stale objects
Stateful database container growing writable layer Move database data to managed volume; verify backup/RPO Mount inspection + writable-layer growth + data backup plan Separates durable data from container lifecycle and CoW layer
Engine 29 fresh host needs larger storage Map DockerRootDir and managed containerd storage before migration DriverStatus + filesystem capacity/inodes + backups Avoids moving only half of the active storage architecture

13. Decision record template

Storage architecture observed:
Storage domains/filesystems:
Capacity + inode baseline:
Largest categories/objects:
Protected resources and owners:
Rebuildable resources:
Log retention:
Build-cache budget:
Deletion/migration mechanism:
Expected reclaimed bytes / performance effect:
Rollback and verification:
Evidence timestamp + Engine/Desktop version:

Knowledge check

Why is overlay2 not the right label for every Engine 29 filesystem path?

Why can aggressive cache cleanup increase cost even when it frees space?

What is the strongest reason not to prune an unused named volume automatically?

Does changing data-root relocate Engine 29 managed containerd image/snapshot data?

Why should logging retention be part of disk governance?

Next lesson

Next: Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting: Diagnostics, Failure Modes, Security, and Performance

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Design baseline: Docker currently recommends understanding actual storage architecture first. BuildKit GC is periodic and policy-driven; manual prune is immediate and should be scoped to an understood builder. Docker’s default json-file logging driver remains unrotated unless configured, while local rotates by default.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.