Chapter 04Lesson 04~105 minutes

Images, Layers, Content-Addressable Storage, Manifests, Config Objects, and Image Identity: Diagnostics, Failure Modes, Security, and Performance

Image incidents are often identity incidents. This lesson diagnoses tag drift, digest-type confusion, platform mismatch, misleading history/size assumptions, and unsafe cleanup by preserving the original manifest, config, layer, platform, and local-store evidence before changing anything.

DiagnosticsTag driftPlatform mismatchSafe cleanupEvidence

Learning objectives

  • Localize failures to name resolution, registry/index/manifest selection, config identity, layer transfer/unpack, local store, or runtime platform selection.
  • Diagnose the common mistake of using a config/Image ID where a registry manifest digest is required.
  • Explain why naive layer-size arithmetic and docker history alone can misrepresent image storage and provenance.
  • Preserve first-failure evidence before retagging, repulling, deleting references, or pruning content.
  • Repair failures with the smallest reversible action and verify the corrected digest/platform path independently.
Chapter 04 evidence baseline — verified 2026-09-21. This chapter uses a free/local/disposable Docker path. Version-sensitive image behavior is inspected from the active daemon and registry rather than assumed. The standards baseline is OCI Image Specification v1.1.1; current Docker Engine 29 fresh installs default to the containerd image store, while upgraded hosts can differ. The lab image is a small Docker Official Image, and all immutable references are captured at run time rather than copied from prose.

1. Preserve identity before “fixing” the image

Image troubleshooting is unusually vulnerable to evidence destruction. A repull can move a tag to new content. A retag can hide what name originally resolved. A prune can remove the exact blobs needed to reproduce a failure. Therefore the first action is read-only capture: context, Engine version, storage backend, tag/reference, RepoDigests, local Image ID, platform, registry index/manifest output, container image fields, and the exact error.

Do not begin with: docker system prune, deleting Docker data-root directories, retagging over evidence, or “pull latest again.” Those actions can change the state you are trying to explain.

2. Evidence-first diagnostic sequence

  1. Confirm host/platform, active context, client/server versions, and local image store behavior.
  2. Record the exact input reference and registry-side top-level digest/media type.
  3. Record the selected platform manifest and local config/Image ID.
  4. Inspect layer descriptors/DiffIDs and pull/unpack errors if transfer failed.
  5. Inspect the container object's .Config.Image and .Image if runtime behavior is involved.
  6. Only then test one hypothesis with the least destructive correction.
  7. Re-run the exact failed operation and compare identities before/after.

3. Intentionally broken example: using the Image ID as a registry digest

This is a safe identity error to reproduce with the Chapter 04 lab image. First ensure the versioned image exists locally and capture two different values:

SOURCE='busybox:1.37.0'
docker pull "$SOURCE"

IMAGE_ID=$(docker image inspect "$SOURCE" --format '{{.Id}}')
REPO_DIGEST=$(docker image inspect "$SOURCE" --format '{{index .RepoDigests 0}}')

printf 'Image ID:    %s\n' "$IMAGE_ID"
printf 'RepoDigest:  %s\n' "$REPO_DIGEST"

Now derive the repository name and deliberately attempt to use the config/Image ID as though it were a registry manifest digest:

REPO='busybox'
WRONG_DIGEST=${IMAGE_ID#sha256:}

docker pull "${REPO}@sha256:${WRONG_DIGEST}"

On a normal registry this should fail because the config blob digest is not the digest of a pullable image manifest/index. Preserve the exact registry error. The correction is not “try random hashes”; it is to use the repository-qualified RepoDigest captured from the registry relationship:

docker pull "$REPO_DIGEST"

The diagnosis is semantic: both values are SHA-256 digests, but they name different object types.

4. Failure mode: the tag moved

Symptom: yesterday and today both say app:stable, yet hosts run different bytes. Do not compare tag strings. Compare the digest resolved at each timestamp. If historical digest evidence is missing, the ambiguity cannot be repaired after the fact by inspecting today's tag.

Correction: record the digest during build/promotion/deploy, and pin production to an approved digest or at minimum verify that the resolved digest equals the approved one before rollout.

5. Failure mode: correct release, wrong or missing platform

Symptom: the tag resolves, but the host reports no matching manifest for its platform, or an emulated/incorrect variant is selected. Inspect the image index and host platform rather than blaming the container process.

docker info --format 'OS={{.OSType}} Arch={{.Architecture}}'
docker buildx imagetools inspect busybox:1.37.0

Check whether the index actually contains the required OS/architecture/variant descriptor. A manifest existing for amd64 does not prove arm64 support.

6. Failure mode: treating history or sizes as exact blob accounting

docker image history is useful narrative evidence, but history rows and filesystem layers are not one-to-one. Metadata-only steps can report 0B and have no rootfs layer. Docker's displayed image size is generally uncompressed logical image data, while transfer sizes in manifests are compressed blob sizes. Shared blobs also mean summing image sizes across tags overstates unique disk consumption.

Correction: use manifest descriptors for transferred blob sizes, Docker disk-accounting commands for local usage, and history for build-step narrative. State which quantity you are measuring.

7. Failure mode: deleting shared content to “fix corruption”

Deleting files under Docker's data root is not targeted image repair. Content, snapshots, metadata, containers, and images can share state. Manual deletion can transform one understandable image failure into daemon-wide corruption.

Correction: preserve daemon logs and object IDs, verify the exact affected reference, use Docker-supported remove/pull operations on a disposable scope when appropriate, and escalate persistent store corruption with a backup/recovery plan. Never teach rm -rf /var/lib/docker as troubleshooting.

8. Failure mode: assuming every Engine 29 host has the same store

A fresh Engine 29 installation normally uses the containerd image store. An upgraded older host may still use classic storage-driver state, and userns-remap changes the supported path. If one host can retain a multi-platform image locally and another cannot, inspect their store/configuration history before calling one “broken.”

9. Security: digest integrity is necessary but not sufficient

Successfully pulling by digest verifies that the fetched content matches that digest. It does not establish who authorized the digest, whether the publisher's process was trustworthy, or whether the image contains vulnerabilities. Later chapters bind SBOMs, scans, provenance, and signatures to exact digests. Do not weaken TLS or trust settings to make a digest pull “work.”

10. Performance: shared layers, transfer, and local reuse

When multiple images reference the same content-addressed layers, Docker can reuse already-present blobs. That reduces transfer and disk duplication. Performance diagnosis should therefore distinguish network transfer bytes, compressed layer sizes, uncompressed image size, snapshot usage, and total Docker data-root consumption. Optimize the measured bottleneck rather than deleting caches reflexively.

11. Repair checklist

Before changing state
  • Save the exact reference and registry response.
  • Save index/manifest/config identities and platform.
  • Save local Image ID/RepoDigests and relevant container image fields.
  • Save the first error and daemon/build/runtime logs.
Then
  • Correct one identity/platform/reference assumption.
  • Repeat only the smallest failed operation.
  • Confirm the new evidence chain and document what changed.
Next lesson

Next: Checkpoint Lab — Image Identity Dossier

Combine registry, local-store, platform, config, layer, and runtime evidence into one reproducible dossier and prove a digest-pinned pull/run.

Knowledge check

A registry says “manifest unknown” after you use the value from .Id. Which layer should you investigate first?

Two hosts use the same tag but different digests. Is Docker necessarily inconsistent?

Why is docker history insufficient to prove exact layer blobs?

What should you do before pruning after a suspected image corruption?

Official references and version notes

Current baseline, not a frozen requirement

Verified 2026-09-21: Docker Engine 29.8.1 remains the current Engine 29 patch line used by this course baseline. Fresh Docker Engine 29 installations use the containerd image store by default, while upgraded older daemons can retain the classic storage-driver store; userns-remap is a documented exception. OCI Image Specification v1.1.1 is the current stable release. The mandatory lab uses the Docker Official Image busybox:1.37.0 only as a small human-readable starting reference, then captures and uses the registry-provided digest at run time. Always record the actual daemon, storage backend, platform, Buildx version, media types, and resolved digests observed on the learner's system.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.