Chapter 40Lesson 04~195 minutes

Docker Performance Engineering: Build Speed, Image Size, Runtime Overhead, Network, Storage, and Resource Profiling: Diagnostics, Failure Modes, Security, and Performance

Diagnose slow Docker workloads without folklore: preserve evidence, separate host saturation from container limits, identify cache or I/O causes, and avoid security-for-speed shortcuts.

DiagnosticscgroupsHost saturationVarianceSecurity

Learning objectives

  • Preserve first-failure performance evidence before changing caches, resources, storage, or networking.
  • Diagnose one-run conclusions, cache-erasure bias, layer-count myths, over-parallelization, Desktop-to-Linux extrapolation, and security-for-speed shortcuts.
  • Interpret CPU throttling with inspect/cgroup/stats evidence instead of blaming the image.
  • Separate build, image, runtime, storage, network, host, and external dependency causes.
  • Apply the least destructive correction and re-run the smallest benchmark scope.

1. Performance incidents still need incident discipline

A slow build or request is an incident when it threatens delivery or SLOs. The same rule applies: preserve evidence before “fixing” it. Cache deletion, container recreation, host restarts, and configuration changes can remove the exact state needed to explain the regression.

2. Evidence-first diagnostic sequence

  1. preserve the first slow build/request log and timestamps;
  2. confirm host/platform, Docker version, context, builder, and source/image digest;
  3. confirm daemon/build/runtime state and exact resource limits;
  4. inspect cache vertices or process/health evidence;
  5. inspect mounts, storage, network/DNS, and external dependencies;
  6. check host/VM saturation;
  7. form one causal hypothesis;
  8. change one variable;
  9. re-run only the smallest comparable benchmark;
  10. keep both before/after packets.

3. Failure mode: optimizing from one run

One run captures noise: host scheduling, page cache, background antivirus/indexing, registry/CDN path, DNS cache, thermal state, Desktop VM activity, and clock granularity. Repeat the same workload. Report a median plus range or percentile distribution appropriate to the workload. If the change is smaller than normal variance, you do not yet have evidence of improvement.

4. Failure mode: clearing caches before every comparison

Cache removal can be valid for a deliberately cold test, but doing it before every run makes warm/incremental performance impossible to study and adds destructive noise. Preserve the cache state when the real workflow preserves it. Use an isolated builder if you need controlled cold/warm populations rather than deleting unrelated global cache.

5. Failure mode: “fewer layers means faster and smaller”

Layer count alone does not determine build speed or final bytes. Combining unrelated steps can make cache invalidation worse; splitting stable and volatile work can create more reusable vertices. Likewise, deleting a file in a later layer does not remove bytes from an earlier image layer. Inspect the actual graph, history, descriptors, and result.

6. Failure mode: parallelism beyond capacity

More parallel build vertices, test workers, or service replicas can reduce wall time until CPU, memory, disk, or network saturates. Beyond that point they can increase queueing, context switching, OOM risk, and cache contention. Pair application throughput with host utilization and tail latency instead of tuning worker counts by intuition.

7. Failure mode: generalizing Desktop results to native Linux

Docker Desktop introduces a managed Linux VM and host file-sharing boundary. CPU/memory limits may apply to that VM, and bind-mount performance can differ substantially from native Linux. Desktop measurements are valid for Desktop development; production conclusions require production-like Linux measurements.

8. Intentionally constrained example: apparent “slow image”

Create two identical compute workloads, but quietly constrain one to a quarter CPU. The slow result is caused by runtime governance, not image construction.

docker run -d \
  --name ch40-free \
  --label devops.academy.lab=ch40   alpine:3.22 sh -c 'i=0; while [ "$i" -lt 250000 ]; do echo "$i" | sha256sum >/dev/null; i=$((i+1)); done; sleep 30'

docker run -d \
  --name ch40-capped \
  --label devops.academy.lab=ch40 \
  --cpus=0.25   alpine:3.22 sh -c 'i=0; while [ "$i" -lt 250000 ]; do echo "$i" | sha256sum >/dev/null; i=$((i+1)); done; sleep 30'

docker inspect ch40-free ch40-capped   --format 'name={{.Name}} image={{.Image}} NanoCpus={{.HostConfig.NanoCpus}} Memory={{.HostConfig.Memory}}'
docker stats --no-stream ch40-free ch40-capped

Before changing anything, keep inspect and stats evidence. On Linux/cgroup v2, inspect the container’s own cgroup view if available:

docker exec ch40-capped sh -c 'printf "cpu.max="; cat /sys/fs/cgroup/cpu.max 2>/dev/null || true; printf "cpu.stat
"; cat /sys/fs/cgroup/cpu.stat 2>/dev/null || true' 

If cpu.stat shows throttling and inspect shows NanoCpus=250000000, the causal hypothesis is strong. The image is identical; the resource envelope differs.

9. Least-destructive correction and re-measure

docker update --cpus=1.0 ch40-capped
docker inspect ch40-capped --format 'NanoCpus={{.HostConfig.NanoCpus}}'
docker stats --no-stream ch40-capped
# Preserve post-change evidence, then clean only the lab containers.
docker rm -f ch40-free ch40-capped

The correction changes only the disposable container’s CPU limit. It does not disable cgroups or change daemon-wide policy. For a real workload, choose the new limit from measured SLO/capacity evidence rather than simply setting it to 1 CPU.

10. Security is not a benchmark variable to disable casually

Do not remove seccomp, AppArmor/SELinux, namespace isolation, TLS, firewall policy, or least privilege just to produce a faster number. If a security mechanism is suspected of measurable overhead, reproduce the claim in an isolated authorized environment, quantify it, and evaluate the risk tradeoff with a security owner. The default troubleshooting path keeps protections enabled.

11. Causal map for common symptoms

Symptom First evidence Likely layer(s) Do not jump to
Incremental build reruns dependency step BuildKit plain progress + Dockerfile inputs Build graph/cache Deleting all cache
Pull is slow, local start is fast Registry layer transfer timing/descriptors Registry/network/compression Changing CPU limit
High CPU + throttle counters Inspect + cpu.stat + host saturation Resource governance/capacity Rebuilding image blindly
Low CPU, high latency Block/network I/O, external timing, app logs I/O/network/external dependency Adding CPUs immediately
Desktop bind source is slow Mount type + Desktop VM/file-sharing evidence Host-coupled storage path Assuming production overlay is same
Build faster with more workers until it regresses Host CPU/memory/disk + wall time Capacity/parallelism Infinite concurrency

12. Incident packet

timestamp=<UTC time>
source=<commit or synthetic version>
image/base=<digests>
context=<docker context>
engine/buildx/buildkit=<versions>
platform=<native Linux / Desktop / remote / emulated>
metric=<exact start/stop boundary>
runs=<all samples>
host_state=<CPU/memory/disk/VM resources>
container_limits=<inspect values>
cache_state=<cold/warm/incremental + build log>
storage_network_path=<mount/driver/hops>
first_failure=<preserved logs/stats/cgroup evidence>
hypothesis=<one causal statement>
change=<one variable>
post_result=<same benchmark>
limits=<variance/external validity>

13. Smallest safe scope

If a build cache key is wrong, rebuild the relevant graph; do not restart the daemon. If a single container is CPU-throttled, inspect/update that container; do not disable cgroups. If a bind mount is slow on Desktop, profile that path; do not rewrite every production storage choice. The scope of the correction should match the scope of the evidence.

Knowledge check

Why is deleting all caches a poor default response to a slow incremental build?

Two containers use the same image; one is much slower. What evidence should precede an image rebuild?

Why can low CPU coincide with high request latency?

What is wrong with disabling seccomp to see whether a workload becomes faster?

What makes a performance correction “least destructive”?

Next lesson

Next: Checkpoint Lab — Docker Performance Engineering: Build Speed, Image Size, Runtime Overhead, Network, Storage, and Resource Profiling

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Current upstream baseline used for version-sensitive explanations: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, and BuildKit 0.33.0. Executable labs record the learner’s actual Docker, Buildx, BuildKit, kernel/cgroup, image digest, and host/VM environment before interpreting any timing.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.