Chapter 31Lesson 04~170 minutes

Containerd Image Store, Snapshotters, Runtime Internals, OCI Runtime Specs, and Docker Engine Evolution: Diagnostics, Failure Modes, Security, and Performance

Diagnose image-store and runtime failures by preserving evidence, separating digest identity from unpacked snapshots, and refusing destructive direct edits to Docker-managed containerd state.

DiagnosticsState ownershipDowngrade riskEvidenceRecovery

Learning objectives

  • Apply an evidence-first sequence to image-store, snapshotter, task/shim, and OCI runtime failures.
  • Recognize destructive anti-patterns such as deleting Docker-managed content with ctr or editing data-root internals.
  • Diagnose wrong-platform and wrong-storage-assumption failures without erasing the first error.
  • Plan downgrade/recovery around explicit storage migration state instead of version-number optimism.
  • Separate immutable image identity from local snapshot/runtime state during incident analysis.

1. Preserve first-failure evidence before touching runtime state

Low-level runtime incidents are exactly where “try a restart” destroys the most useful clues. Before restarting Docker, switching storage backends, deleting images, or removing containers, preserve: UTC time, context, Engine version/info, image reference/digest/platform, container inspect state, events/logs, disk/cgroup/security evidence, and the exact first error.

If the host is unhealthy, copy evidence to a separate filesystem. Do not store the only incident log inside a container writable layer that you are about to recreate.

2. Diagnostic sequence by owning layer

Layer Evidence Typical question
Client/context docker context show, client/server versions Am I talking to the intended daemon?
Engine/store mode full docker info, DriverStatus containerd image store or classic driver?
Registry/content imagetools/index/digests, pull error Does the requested platform/content exist?
Snapshot/unpack daemon logs, disk/filesystem support, store mode Could content be materialized into a rootfs?
Container/task docker inspect state, runtime, events Was the object created? Did start reach runtime?
Shim/runc/kernel daemon/runtime error, PID/process evidence Did OCI execution fail?
Application container logs/health/exit code Did the launched process itself fail?

3. Failure mode: using ctr to delete Docker-managed content

The tempting anti-pattern is to see a suspicious blob/snapshot through ctr and remove it directly. That bypasses Docker’s metadata/reference accounting. Even if the low-level deletion succeeds, Engine may still retain references or assumptions about that object.

Do not “repair” managed state with direct mutation. Preserve Docker and optional read-only containerd listings. Then use supported Docker lifecycle operations or restore/migrate from documented backups. If corruption requires deeper recovery, follow vendor/upstream support guidance on a copy/disposable host rather than improvising deletes.

4. Failure mode: looking only under /var/lib/docker

An Engine 29 containerd-image-store host can store image content and snapshots under /var/lib/containerd while other Docker data remains under /var/lib/docker. A capacity monitor that watches only the old tree can miss the actual growth. Conversely, an operator may see a small /var/lib/docker/overlay2 and incorrectly conclude that images disappeared.

Repair the monitoring model: confirm store mode, record both supported data-root locations, use docker system df, and monitor the actual filesystem hosting containerd data. Do not move directories manually while services are running.

5. Intentionally broken example: request a nonexistent platform

This failure is safe because it should stop at manifest selection before any container runtime state exists.

set +e
docker pull --platform=linux/definitely-not-a-real-arch alpine:3.22.1   > dca31-platform-error.txt 2>&1
RC=$?
set -e
printf 'exit=%s
' "$RC"
cat dca31-platform-error.txt

Interpret the evidence: a “no matching manifest” style error is an index/platform selection failure. It is not a snapshotter corruption, runc crash, firewall problem, or application error. Confirm available platforms:

docker buildx imagetools inspect alpine:3.22.1

Then choose a platform actually present and compatible with your host/emulation policy. The least-destructive fix is a correct reference/platform, not a daemon restart.

6. Failure mode: legacy graph-driver field assumptions

Old troubleshooting guides often start with “check Storage Driver: overlay2” and then reason entirely from /var/lib/docker/overlay2. On modern Engine this can be incomplete. Preserve full docker info and inspect DriverStatus; the current containerd-backed path exposes snapshotter-specific status.

Use the field as evidence of architecture, not as a cue to open internal directories and modify them.

7. Failure mode: treating snapshot state as image identity

A snapshot is local materialization. Its keys/parents/mounts can differ across hosts and snapshotters even when both hosts run the exact same OCI manifest digest. Therefore snapshot identifiers should never replace registry/image digests in promotion, SBOM, signature, or incident correlation records.

If a snapshot is lost but registry content and application data are intact, the image can often be pulled/unpacked again. If the immutable digest changes, you are no longer discussing the same artifact.

8. Failure mode: downgrade after a storage migration

“The old Engine binary starts” is not a rollback test. Storage metadata, snapshotter behavior, API expectations, and migration state can cross compatibility boundaries. Before upgrading or migrating, document the currently supported downgrade/rollback path and keep recoverable image/application data outside assumptions about internal stores.

For upgraded classic-store hosts, switching back can reveal the previously hidden classic objects because Docker keeps stores separate. That is different from downgrading an Engine package after an experimental migration. Test the exact path on a disposable clone.

9. Failure mode: separate Docker data-root but full root filesystem

A team moves data-root to a large disk and later enables the containerd image store. The root partition fills because containerd still uses its own default root unless configured separately. Evidence: Engine store mode, filesystem usage, DockerRootDir, containerd root configuration, and docker system df.

The correction is a planned storage-location/capacity change with backup and service-management procedure—not a broad prune or manual copy while daemons are active.

10. Security-sensitive diagnostics

Access to Docker’s Engine socket or managed containerd socket is powerful. Do not expose either remotely merely to make debugging easier. Do not mount them into untrusted containers. Do not use --privileged, seccomp=unconfined, or disabled LSM/firewall protections to “see if runtime internals work.” Those actions change the security boundary and contaminate the experiment.

11. Performance questions need layer-specific measurements

If startup is slow, determine whether time is spent resolving registry metadata, downloading compressed blobs, unpacking through the snapshotter, creating the runtime task, or initializing the application. Snapshotter benchmarks that ignore network cache state or application startup are not enough. Likewise, disk usage should separate content blobs, unpacked snapshots, writable layers, BuildKit cache, volumes, and logs.

12. Least-destructive recovery pattern

  1. Freeze and copy first-failure evidence.
  2. Confirm context/host and Engine version/store mode.
  3. Confirm immutable image reference/platform.
  4. Check daemon/runtime logs and disk/filesystem health.
  5. Reproduce with the smallest disposable image/container.
  6. Correct only the owning layer.
  7. Verify digest/container/process state after the fix.
  8. Only then remove exact failed lab objects.

Knowledge check

A pull fails with “no matching manifest.” Which layer owns the first failure?

Why is deleting a suspicious snapshot with ctr a poor first repair?

What does a snapshot ID prove about the registry image digest?

Why can DockerRootDir monitoring miss capacity pressure after switching stores?

What is a valid downgrade plan?

Next lesson

Next: Checkpoint Lab — Containerd Image Store, Snapshotters, Runtime Internals, OCI Runtime Specs, and Docker Engine Evolution

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Diagnostics baseline: Docker’s current Engine 29 documentation explicitly warns that managed/embedded containerd endpoints are for debugging and that changes from other clients can conflict with the daemon. Treat all low-level state as evidence unless an upstream-supported recovery procedure says otherwise.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.