Storage Drivers, overlay2, Copy-on-Write Behavior, Disk Usage, Pruning, and Storage Troubleshooting: Configuration, Design Choices, and Tradeoffs
Choose storage backend, filesystem placement, retention policy, logging rotation, cleanup cadence, and BuildKit cache budgets from workload evidence rather than blanket prune habits.
Learning objectives
-
Choose between classic
overlay2and containerd snapshotter terminology based on observed Engine architecture. - Design disk placement and monitoring that covers Docker data, managed containerd state, logs, volumes, and builders.
- Set retention and cache budgets from rebuild cost, rollback needs, and capacity SLOs rather than arbitrary age alone.
- Compare automatic BuildKit GC with manual targeted cleanup and explain when each belongs in operations.
- Document migration prerequisites and rollback boundaries before changing storage drivers, snapshotters, or data locations.
1. Design from storage classes, not one “Docker partition” assumption
A production storage design begins by listing actual state classes: image/content bytes, unpacked root filesystems, writable layers, volumes, build cache, local logs, daemon metadata, and external bind-mounted application data. Map each class to the filesystem or VM disk that holds it, then assign capacity, performance, backup, retention, and alerting expectations.
Engine 29 makes this mapping more important because the containerd image store can place image/snapshot state under managed containerd storage while Docker’s general data root still holds other data.
2. Classic overlay2 versus containerd overlayfs snapshotter
| Question | Classic overlay2 |
Engine 29 containerd image store |
|---|---|---|
| Abstraction | Docker graph/storage driver | containerd content store + snapshotter |
| Common Linux backend | overlay2 using kernel OverlayFS |
containerd overlayfs snapshotter |
| Fresh Engine 29 default | No | Yes, except documented unsupported modes such as current userns-remap limitation |
| Local multi-platform/attestation support | Classic store has limitations | Containerd store enables richer multi-platform/attestation storage |
| Operational inspection |
Docker CLI/info; do not edit overlay2 dirs
|
Docker CLI/info; do not mutate managed containerd content/snapshots |
3. OverlayFS copy-up has workload consequences
OverlayFS presents lower read-only layers and an upper writable layer. A write to a file originating in a lower layer can cause copy-up into upper state. For many normal application patterns this is efficient, but large frequently modified files or write-heavy databases are poor candidates for the container writable layer. Volumes bypass much of this lifecycle coupling and are the standard persistence path.
Do not “optimize” by editing Docker’s upper/lower directories. Optimize by changing application storage placement, image composition, cache behavior, and filesystem capacity with supported interfaces.
4. Separate Docker data filesystem: benefits and boundaries
| Benefit | Why it helps | Boundary |
|---|---|---|
| Capacity isolation | Docker growth does not consume the host root filesystem as quickly | Containerd image-store bytes may be on a separate managed containerd root |
| Monitoring clarity | Dedicated alerts can target Docker-specific bytes/inodes/I/O | You must still monitor volumes/logs/builders and any other storage domains |
| I/O planning | Choose SSD/filesystem suited to workload | Do not place multiple independent daemons on the same data directory |
| Maintenance | Backup/migration scope becomes explicit | Daemon stop, backup, consistency, and supported migration steps remain required |
5. Retention policy: classify by rebuildability and business value
| State | Typical rebuildability | Retention question | Preferred control |
|---|---|---|---|
| Pulled/release image | Often re-pullable if registry retains immutable digest | Needed for rollback/offline recovery? | Release retention policy + digest records |
| Container writable layer | Usually ephemeral | Is application writing data in wrong place? | Keep small; move durable data to volume |
| Named volume | Potentially irreplaceable | Backup/RPO/owner/expiry? | Application data lifecycle; never generic prune first |
| Build cache | Re-creatable but may be expensive | How much feedback-time value per GB? | BuildKit GC budget + dedicated builder policy |
| Logs | Operational evidence, sometimes regulated | Retention/search/RTO/sensitivity? | Rotation + external aggregation where appropriate |
6. BuildKit automatic GC versus manual cleanup
Current BuildKit garbage collection runs periodically and evaluates
ordered policies. Defaults preferentially reclaim easily regenerated
or stale cache and use space thresholds. For most developer systems,
Docker states the defaults are sufficient. Large builders or
constrained disks may justify explicit reservedSpace,
maxUsedSpace, minFreeSpace, or custom
policies.
Manual buildx prune is an immediate intervention. It
supports filters and space targets, but it is still a deletion
operation. A mature setup first budgets cache, then monitors
buildx du, and uses manual pruning only when a known
builder and policy justify it.
7. Automatic versus manual cleanup
| Approach | Strength | Risk | Good fit |
|---|---|---|---|
| Exact object deletion | Maximum auditability and narrow blast radius | Operational effort | Incidents, training, protected hosts |
| Label/age-filtered prune | Policy can scale to many disposable objects | Filter semantics/labels must be governed consistently | CI workers with enforced labeling |
| BuildKit GC | Continuous cache-budget enforcement | Mis-sized budget can hurt cache hit rate or disk headroom | Long-lived builders |
| Ephemeral worker recreation | Very strong lifecycle boundary | Requires workflow designed for disposability | CI/build fleet nodes |
8. Cache budget is a performance decision, not only a storage decision
Build cache trades disk for feedback latency and network/compute savings. Too little cache increases rebuild time and dependency downloads; too much can starve the host. A sensible budget uses empirical data: cache hit rate, rebuild time, average/peak cache size, host free-space SLO, and cost of re-fetching dependencies.
For a dedicated builder, record docker buildx du over
time and set a budget that preserves the most valuable working set
while maintaining emergency free space.
9. Logs require a retention policy of their own
The default json-file driver exists for compatibility
and does not rotate unless configured. Docker’s
local driver rotates by default (current defaults are
five files of 20 MB each before compression, about 100 MB per
container maximum before compression). Production policy must decide
whether local retention is enough or logs should be shipped to a
managed sink.
Changing daemon logging defaults only affects newly created
containers. A rollout plan therefore includes recreation timing and
validation, not only a daemon.json edit.
10. Data-root migration: plan both Docker and containerd domains
A storage relocation is an administrative maintenance event, not a live copy exercise. Preserve an inventory, stop/coordinate workloads, back up persistent application data, verify free space/inodes/permissions/filesystem support at the target, and use the current official procedure for the active architecture.
For the classic store, Docker’s OverlayFS documentation shows
stop/copy/configure/start verification patterns. With Engine 29’s
containerd image store, Docker explicitly documents that Docker
data-root does not move managed containerd
image/snapshot state. A plan that copies only
/var/lib/docker can therefore be incomplete.
11. Storage-driver or snapshotter changes are migrations, not tuning knobs
Changing a storage backend can make existing local images and containers inaccessible until you switch back or migrate them. This is not equivalent to changing a harmless daemon preference. Export/push immutable images, back up application data, document current backend and versions, test rollback, and schedule downtime when required.
Docker Desktop manages these internals differently; editing Linux Engine storage-driver configuration is not a supported Desktop method. Use Desktop’s supported disk-image controls instead.
12. Worked decision scenario
| Scenario | Choice | Prerequisites/evidence | Why |
|---|---|---|---|
| Developer workstation, frequent builds, disk occasionally tight |
Keep BuildKit GC defaults; monitor buildx du;
bounded local logs
|
Observed cache trend + disk headroom | Preserves fast feedback without routine broad prune |
| CI runner dedicated to one pipeline and recreated daily | Ephemeral builder/worker lifecycle | Jobs have no durable local state; artifacts pushed immutably | Lifecycle itself bounds cache and stale objects |
| Stateful database container growing writable layer | Move database data to managed volume; verify backup/RPO | Mount inspection + writable-layer growth + data backup plan | Separates durable data from container lifecycle and CoW layer |
| Engine 29 fresh host needs larger storage | Map DockerRootDir and managed containerd storage before migration | DriverStatus + filesystem capacity/inodes + backups | Avoids moving only half of the active storage architecture |
13. Decision record template
Storage architecture observed:
Storage domains/filesystems:
Capacity + inode baseline:
Largest categories/objects:
Protected resources and owners:
Rebuildable resources:
Log retention:
Build-cache budget:
Deletion/migration mechanism:
Expected reclaimed bytes / performance effect:
Rollback and verification:
Evidence timestamp + Engine/Desktop version:
Knowledge check
Why is overlay2 not the right label for every
Engine 29 filesystem path?
Fresh Engine 29 uses the containerd image store and normally the containerd overlayfs snapshotter; upgraded hosts can still use classic overlay2.
Why can aggressive cache cleanup increase cost even when it frees space?
It lowers cache hit rates and forces rebuild, network download, and compute work that the cache was avoiding.
What is the strongest reason not to prune an unused named volume automatically?
“Unused now” does not prove disposable; it may contain retained application data with backup/rollback obligations.
Does changing data-root relocate Engine 29 managed
containerd image/snapshot data?
No. Current Docker docs explicitly separate Docker data-root from managed containerd root when the containerd image store is active.
Why should logging retention be part of disk governance?
Local logging can consume significant disk independently of container writable-layer size, especially unrotated json-file logs.
Official references and version notes
Design baseline: Docker currently recommends
understanding actual storage architecture first. BuildKit GC is
periodic and policy-driven; manual prune is immediate and should be
scoped to an understood builder. Docker’s default
json-file logging driver remains unrotated unless
configured, while local rotates by default.
- Docker Docs — containerd image store with Docker Engine — Engine 29 fresh-install default, snapshotter reporting, switching, and migration boundaries.
- Docker Docs — Storage drivers — image layers, container writable layers, copy-on-write, and container-size accounting caveats.
-
Docker Docs — OverlayFS storage driver
— classic
overlay2, backing-filesystem requirements, lower/upper/merged/workdir semantics, and migration cautions. - Docker Docs — Select a storage driver — current containerd snapshotter versus classic-driver guidance and platform distinctions.
-
Docker CLI —
docker system df— daemon disk-usage summary and verbose object-level accounting. -
Docker CLI —
docker system prune— deletion scope and supported filters; referenced here to understand risk, not as the lab cleanup path. -
Docker Buildx —
buildx du— builder-cache usage, shared/private/reclaimable records, and detailed evidence. -
Docker Buildx —
buildx prune— cache filters and space-budget controls. - Docker Build — Build garbage collection — periodic BuildKit GC, default policy shape, and builder-specific configuration.
- Docker Docs — Docker daemon configuration overview — Docker data-root and the separate managed containerd root used by the Engine 29 containerd image store.
-
Docker Docs — Configure logging drivers
—
json-filegrowth risk, rotation guidance, and the rotatinglocaldriver. - Docker Engine 29 release notes — current Engine 29 behavior and storage/image-store fixes.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.