Containerd Image Store, Snapshotters, Runtime Internals, OCI Runtime Specs, and Docker Engine Evolution: Diagnostics, Failure Modes, Security, and Performance
Diagnose image-store and runtime failures by preserving evidence, separating digest identity from unpacked snapshots, and refusing destructive direct edits to Docker-managed containerd state.
Learning objectives
- Apply an evidence-first sequence to image-store, snapshotter, task/shim, and OCI runtime failures.
- Recognize destructive anti-patterns such as deleting Docker-managed content with ctr or editing data-root internals.
- Diagnose wrong-platform and wrong-storage-assumption failures without erasing the first error.
- Plan downgrade/recovery around explicit storage migration state instead of version-number optimism.
- Separate immutable image identity from local snapshot/runtime state during incident analysis.
1. Preserve first-failure evidence before touching runtime state
Low-level runtime incidents are exactly where “try a restart” destroys the most useful clues. Before restarting Docker, switching storage backends, deleting images, or removing containers, preserve: UTC time, context, Engine version/info, image reference/digest/platform, container inspect state, events/logs, disk/cgroup/security evidence, and the exact first error.
If the host is unhealthy, copy evidence to a separate filesystem. Do not store the only incident log inside a container writable layer that you are about to recreate.
2. Diagnostic sequence by owning layer
| Layer | Evidence | Typical question |
|---|---|---|
| Client/context |
docker context show, client/server versions
|
Am I talking to the intended daemon? |
| Engine/store mode | full docker info, DriverStatus |
containerd image store or classic driver? |
| Registry/content | imagetools/index/digests, pull error | Does the requested platform/content exist? |
| Snapshot/unpack | daemon logs, disk/filesystem support, store mode | Could content be materialized into a rootfs? |
| Container/task | docker inspect state, runtime, events | Was the object created? Did start reach runtime? |
| Shim/runc/kernel | daemon/runtime error, PID/process evidence | Did OCI execution fail? |
| Application | container logs/health/exit code | Did the launched process itself fail? |
3. Failure mode: using ctr to delete Docker-managed content
The tempting anti-pattern is to see a suspicious blob/snapshot
through ctr and remove it directly. That bypasses
Docker’s metadata/reference accounting. Even if the low-level
deletion succeeds, Engine may still retain references or assumptions
about that object.
4. Failure mode: looking only under /var/lib/docker
An Engine 29 containerd-image-store host can store image content and
snapshots under /var/lib/containerd while other Docker
data remains under /var/lib/docker. A capacity monitor
that watches only the old tree can miss the actual growth.
Conversely, an operator may see a small
/var/lib/docker/overlay2 and incorrectly conclude that
images disappeared.
Repair the monitoring model: confirm store mode, record
both supported data-root locations, use
docker system df, and monitor the actual filesystem
hosting containerd data. Do not move directories manually while
services are running.
5. Intentionally broken example: request a nonexistent platform
This failure is safe because it should stop at manifest selection before any container runtime state exists.
set +e
docker pull --platform=linux/definitely-not-a-real-arch alpine:3.22.1 > dca31-platform-error.txt 2>&1
RC=$?
set -e
printf 'exit=%s
' "$RC"
cat dca31-platform-error.txt
Interpret the evidence: a “no matching manifest” style error is an index/platform selection failure. It is not a snapshotter corruption, runc crash, firewall problem, or application error. Confirm available platforms:
docker buildx imagetools inspect alpine:3.22.1
Then choose a platform actually present and compatible with your host/emulation policy. The least-destructive fix is a correct reference/platform, not a daemon restart.
6. Failure mode: legacy graph-driver field assumptions
Old troubleshooting guides often start with “check Storage Driver:
overlay2” and then reason entirely from
/var/lib/docker/overlay2. On modern Engine this can be
incomplete. Preserve full docker info and inspect
DriverStatus; the current containerd-backed path
exposes snapshotter-specific status.
Use the field as evidence of architecture, not as a cue to open internal directories and modify them.
7. Failure mode: treating snapshot state as image identity
A snapshot is local materialization. Its keys/parents/mounts can differ across hosts and snapshotters even when both hosts run the exact same OCI manifest digest. Therefore snapshot identifiers should never replace registry/image digests in promotion, SBOM, signature, or incident correlation records.
If a snapshot is lost but registry content and application data are intact, the image can often be pulled/unpacked again. If the immutable digest changes, you are no longer discussing the same artifact.
8. Failure mode: downgrade after a storage migration
“The old Engine binary starts” is not a rollback test. Storage metadata, snapshotter behavior, API expectations, and migration state can cross compatibility boundaries. Before upgrading or migrating, document the currently supported downgrade/rollback path and keep recoverable image/application data outside assumptions about internal stores.
For upgraded classic-store hosts, switching back can reveal the previously hidden classic objects because Docker keeps stores separate. That is different from downgrading an Engine package after an experimental migration. Test the exact path on a disposable clone.
9. Failure mode: separate Docker data-root but full root filesystem
A team moves data-root to a large disk and later
enables the containerd image store. The root partition fills because
containerd still uses its own default root unless configured
separately. Evidence: Engine store mode, filesystem usage,
DockerRootDir, containerd root configuration, and
docker system df.
The correction is a planned storage-location/capacity change with backup and service-management procedure—not a broad prune or manual copy while daemons are active.
10. Security-sensitive diagnostics
Access to Docker’s Engine socket or managed containerd socket is
powerful. Do not expose either remotely merely to make debugging
easier. Do not mount them into untrusted containers. Do not use
--privileged, seccomp=unconfined, or
disabled LSM/firewall protections to “see if runtime internals
work.” Those actions change the security boundary and contaminate
the experiment.
11. Performance questions need layer-specific measurements
If startup is slow, determine whether time is spent resolving registry metadata, downloading compressed blobs, unpacking through the snapshotter, creating the runtime task, or initializing the application. Snapshotter benchmarks that ignore network cache state or application startup are not enough. Likewise, disk usage should separate content blobs, unpacked snapshots, writable layers, BuildKit cache, volumes, and logs.
12. Least-destructive recovery pattern
- Freeze and copy first-failure evidence.
- Confirm context/host and Engine version/store mode.
- Confirm immutable image reference/platform.
- Check daemon/runtime logs and disk/filesystem health.
- Reproduce with the smallest disposable image/container.
- Correct only the owning layer.
- Verify digest/container/process state after the fix.
- Only then remove exact failed lab objects.
Knowledge check
A pull fails with “no matching manifest.” Which layer owns the first failure?
Registry/index/platform selection, before snapshot/task/application layers.
Why is deleting a suspicious snapshot with ctr a poor first repair?
It mutates Docker-managed state out of band, can create metadata divergence, and destroys evidence without proving the root cause.
What does a snapshot ID prove about the registry image digest?
Very little by itself. Snapshot state is local materialization; immutable OCI digest is the artifact identity.
Why can DockerRootDir monitoring miss capacity pressure after switching stores?
containerd image content/snapshots can live in a separate containerd root such as /var/lib/containerd; Docker data-root does not automatically move it.
What is a valid downgrade plan?
A tested, version-specific procedure with preserved images/application data, storage-mode/migration evidence, backups, and a known path to restore the prior state—not simply reinstalling an older binary.
Official references and version notes
Diagnostics baseline: Docker’s current Engine 29 documentation explicitly warns that managed/embedded containerd endpoints are for debugging and that changes from other clients can conflict with the daemon. Treat all low-level state as evidence unless an upstream-supported recovery procedure says otherwise.
- Docker Docs — containerd image store with Docker Engine — Engine 29 fresh-install defaults, snapshotters, disk layout, switching, and experimental migration guidance.
- Docker Docs — Storage drivers — distinction between classic graph drivers and the Engine 29 containerd image store.
- Docker Docs — Select a storage driver — current storage-backend matrix and platform notes.
-
Docker Docs — Docker daemon configuration overview
—
/var/lib/dockerversus/var/lib/containerdand data-root implications. - Docker Docs — Run containerd in the Docker daemon — Engine 29.7+ experimental embedded-containerd mode and debugging endpoint warning.
- Docker Engine 29 release notes — Engine 29.8.1 baseline, containerd 2.3.5 static-binary packaging, BuildKit 0.33.0 and runc 1.5.1 updates.
- containerd 2.3 — Features — namespaces, images, root filesystems, snapshots, containers, tasks, and OCI runtime integration.
-
containerd — Snapshotters
— core snapshotter behavior and the
overlayfsnaming used by containerd. - OCI Image Specification 1.1.1 — image indexes, manifests, configs, layers, and content-addressable descriptors.
-
OCI Runtime Specification 1.3.0
— runtime bundle,
config.json, execution environment, and lifecycle model.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.