Chapter 40Lesson 02~220 minutes

Docker Performance Engineering: Build Speed, Image Size, Runtime Overhead, Network, Storage, and Resource Profiling: Guided Hands-On Workflow and Core Operations

Run a bounded performance workflow: establish a repeatable baseline, compare cold/warm builds, inspect cache/image/runtime evidence, profile one I/O path, change one variable, and re-measure.

Hands-onCold/warm buildsdocker statsStorage I/OEvidence

Learning objectives

  • Create a bounded synthetic lab with stable names, labels, a recorded base-image digest, and a dedicated disposable builder.
  • Compare deliberately cold, unchanged warm, and realistic incremental builds.
  • Inspect BuildKit cache usage, local image identity/size, container startup, and runtime stats.
  • Measure one storage path repeatedly without turning a microbenchmark into a universal filesystem claim.
  • Change one Docker-specific variable and re-run the same measurement.

1. Lab contract and safety guard

The mandatory path uses only a public Alpine base image, synthetic files, one dedicated Buildx builder, exact ch40-* objects, and no daemon-wide configuration. It does not clear global caches, alter Docker Desktop settings, publish ports, or touch production registries.

Safety boundary. Run this on a disposable/local Docker context. Confirm docker context show before creating anything. If the active context points at a shared or production daemon, stop and switch to an authorized local context.

2. Preflight and immutable base identity

set -eu
printf 'context=%s
' "$(docker context show)"
docker version
docker info --format 'Server={{.ServerVersion}} Driver={{.Driver}} Cgroup={{.CgroupVersion}} CPUs={{.NCPU}} Memory={{.MemTotal}}'
docker buildx version

docker pull alpine:3.22
BASE_REF="$(docker image inspect alpine:3.22 --format '{{index .RepoDigests 0}}')"
printf 'base=%s
' "$BASE_REF"

The human-readable tag is convenient for discovery; the recorded RepoDigest is the evidence identity used by the Dockerfiles. If RepoDigests is unavailable in an unusual local/mirror setup, record the exact docker image inspect output and explain the limitation instead of inventing a digest.

3. Create the synthetic context

rm -rf ch40-perf-lab
mkdir -p ch40-perf-lab/ignored
cd ch40-perf-lab
printf 'dep-v1
' > dependency.lock
printf 'app-v1
' > app.txt
for i in $(seq 1 300); do printf 'ignored-%04d-%0200d
' "$i" 0 > "ignored/$i.txt"; done
printf 'ignored/
*.log
' > .dockerignore

cat > Dockerfile.baseline <<'EOF'
# syntax=docker/dockerfile:1
ARG BASE
FROM ${BASE} AS build
WORKDIR /work
COPY . .
RUN i=0; while [ "$i" -lt 3500 ]; do       printf '%s-%s
' "$(cat dependency.lock)" "$i" | sha256sum >/dev/null;       i=$((i+1));     done  && mkdir -p /out  && cp dependency.lock app.txt /out/
FROM ${BASE}
WORKDIR /app
COPY --from=build /out/ ./
CMD ["sh","-c","cat dependency.lock app.txt >/dev/null"]
EOF

cat > Dockerfile.optimized <<'EOF'
# syntax=docker/dockerfile:1
ARG BASE
FROM ${BASE} AS build
WORKDIR /work
COPY dependency.lock .
RUN i=0; while [ "$i" -lt 3500 ]; do       printf '%s-%s
' "$(cat dependency.lock)" "$i" | sha256sum >/dev/null;       i=$((i+1));     done
COPY app.txt .
RUN mkdir -p /out && cp dependency.lock app.txt /out/
FROM ${BASE}
WORKDIR /app
COPY --from=build /out/ ./
CMD ["sh","-c","cat dependency.lock app.txt >/dev/null"]
EOF

The two Dockerfiles produce the same intended runtime files. The single experimental variable is dependency-step placement relative to the frequently changing app.txt.

4. Create an isolated builder and record its identity

BUILDER=ch40-perf-builder
if docker buildx inspect "$BUILDER" >/dev/null 2>&1; then
  echo "Builder $BUILDER already exists; clean the previous lab first." >&2
  exit 1
fi
docker buildx create --name "$BUILDER" --driver docker-container --use
docker buildx inspect "$BUILDER" --bootstrap
docker buildx du --builder "$BUILDER"

Bootstrapping happens before timed builds so “starting the BuildKit service” is not accidentally counted as application build time. The builder remains a labeled-by-name disposable boundary you can remove exactly later.

5. Deliberately cold baseline build

BASE_REF="$(docker image inspect alpine:3.22 --format '{{index .RepoDigests 0}}')"
START=$SECONDS
docker buildx build \
  --builder ch40-perf-builder \
  --progress=plain \
  --no-cache \
  --load \
  --build-arg BASE="$BASE_REF" \
  --metadata-file baseline-cold.json   -f Dockerfile.baseline   -t ch40-perf:baseline . 2>&1 | tee baseline-cold.log
printf 'baseline_cold_seconds=%s
' "$((SECONDS-START))" | tee baseline-cold.time

--no-cache asks BuildKit not to reuse previous result cache for this build. It does not delete global Docker data. The completed build can still populate cache for the next run. Preserve the plain progress log because it shows context transfer and whether vertices execute.

6. Unchanged warm build

START=$SECONDS
docker buildx build \
  --builder ch40-perf-builder \
  --progress=plain \
  --load \
  --build-arg BASE="$BASE_REF"   -f Dockerfile.baseline   -t ch40-perf:baseline . 2>&1 | tee baseline-warm.log
printf 'baseline_warm_seconds=%s
' "$((SECONDS-START))" | tee baseline-warm.time

docker buildx du --builder ch40-perf-builder

Do not judge the cache only by wall time. Confirm that expensive vertices are marked cached. A busy host can make a cached build appear slower than an earlier run even though cache behavior is correct.

7. Warm both layouts, then make one source edit

docker buildx build --builder ch40-perf-builder --progress=plain --load   --build-arg BASE="$BASE_REF" -f Dockerfile.optimized -t ch40-perf:optimized .   2>&1 | tee optimized-warmup.log

printf 'app-v2
' > app.txt

START=$SECONDS
docker buildx build --builder ch40-perf-builder --progress=plain --load   --build-arg BASE="$BASE_REF" -f Dockerfile.baseline -t ch40-perf:baseline .   2>&1 | tee baseline-incremental.log
printf 'baseline_incremental_seconds=%s
' "$((SECONDS-START))" | tee baseline-incremental.time

START=$SECONDS
docker buildx build --builder ch40-perf-builder --progress=plain --load   --build-arg BASE="$BASE_REF" -f Dockerfile.optimized -t ch40-perf:optimized .   2>&1 | tee optimized-incremental.log
printf 'optimized_incremental_seconds=%s
' "$((SECONDS-START))" | tee optimized-incremental.time

Interpret the graph: the baseline COPY . . is invalidated by app.txt, so its expensive loop executes again. The optimized graph copies the stable dependency file before the expensive step; an app-only change should let that step stay cached. That is a causal explanation, not a claim that “more Dockerfile lines are faster.”

8. Inspect image identity and size

docker image inspect ch40-perf:baseline ch40-perf:optimized   --format 'ref={{join .RepoTags ","}} id={{.Id}} size={{.Size}}'
docker history --no-trunc ch40-perf:optimized

The final images are intentionally small and similar. This experiment changes build invalidation behavior, not the application payload. That separation is useful: you improved inner-loop performance without pretending that runtime size also changed.

9. Measure startup as a distribution, not a single trophy number

for i in 1 2 3 4 5; do
  START=$SECONDS
  docker run --rm --name "ch40-start-$i" ch40-perf:optimized
  printf 'startup_run=%s elapsed_seconds=%s
' "$i" "$((SECONDS-START))"
done

SECONDS has coarse one-second resolution, so very fast starts can all show zero. That is a measurement limitation, not proof of identical microsecond performance. Use a higher-resolution host timer if startup latency is the real optimization target, and state the timer/tool version in the evidence packet.

10. Runtime stats under a bounded synthetic load

docker run -d \
  --name ch40-load \
  --label devops.academy.lab=ch40 \
  --memory=96m \
  --cpus=0.50   alpine:3.22 sh -c 'i=0; while [ "$i" -lt 120000 ]; do echo "$i" | sha256sum >/dev/null; i=$((i+1)); done; sleep 20'

docker inspect ch40-load --format 'NanoCpus={{.HostConfig.NanoCpus}} Memory={{.HostConfig.Memory}} PidsLimit={{.HostConfig.PidsLimit}}'
docker stats --no-stream --format 'table {{.Name}}	{{.CPUPerc}}	{{.MemUsage}}	{{.BlockIO}}	{{.NetIO}}	{{.PIDs}}' ch40-load
docker wait ch40-load
docker inspect ch40-load --format 'Exit={{.State.ExitCode}} OOMKilled={{.State.OOMKilled}}' 

The CPU and memory limits are part of the test fixture. Do not compare this result with an unlimited run unless the purpose is specifically to measure the effect of the resource envelope.

11. Profile one storage path: named volume

docker volume create --label devops.academy.lab=ch40 ch40-perf-data
for i in 1 2 3 4 5; do
  echo "run=$i"
  docker run \
  --rm \
  --name "ch40-io-$i" \
  --mount type=volume,src=ch40-perf-data,dst=/data     alpine:3.22 sh -c 'time dd if=/dev/zero of=/data/blob bs=1M count=32 conv=fsync >/dev/null 2>&1; rm -f /data/blob'
done
docker volume inspect ch40-perf-data

This is a synthetic synchronous-write probe, not a database benchmark. It tells you about this daemon, this volume backend, this image/tool, and this moment. On Desktop, compare a named volume with a disposable bind mount only if cross-OS source I/O is the actual question; do not generalize either result to production Linux without re-measuring there.

12. Small challenge: identify the owning layer

A warm build is fast, but after changing one source file the dependency-install step runs again. Which layer should you investigate first: registry, runtime CPU, Docker DNS, or BuildKit graph/cache inputs? Preserve the plain build log, name the invalidated vertex, and propose the smallest Dockerfile change that could isolate stable dependency inputs from volatile source.

13. Bounded cleanup

docker rm -f ch40-load 2>/dev/null || true
docker volume rm ch40-perf-data 2>/dev/null || true
docker image rm ch40-perf:baseline ch40-perf:optimized 2>/dev/null || true
docker buildx rm ch40-perf-builder
cd ..
rm -rf ch40-perf-lab

Every cleanup command names an object created by this lab. It does not delete unrelated images, volumes, containers, or builder caches.

Knowledge check

Why did the lab bootstrap the builder before timing builds?

What does the baseline incremental log prove if the expensive RUN executes again after only app.txt changed?

Why is image size not the success metric for the COPY-order experiment?

What is wrong with reporting only the fastest of five storage runs?

A stats sample shows low container CPU but requests are slow. What layer should you inspect next?

Next lesson

Next: Docker Performance Engineering: Build Speed, Image Size, Runtime Overhead, Network, Storage, and Resource Profiling: Configuration, Design Choices, and Tradeoffs

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Current upstream baseline used for version-sensitive explanations: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, and BuildKit 0.33.0. Executable labs record the learner’s actual Docker, Buildx, BuildKit, kernel/cgroup, image digest, and host/VM environment before interpreting any timing.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.