Multi-Stage Builds, Minimal Runtime Images, Distroless Patterns, Debug Stages, and Build Separation: Concepts, Architecture, and Mental Model
A production image should contain the runtime contract, not an accidental copy of the build workstation. This lesson models multi-stage builds as an explicit artifact boundary: toolchains and tests live in build stages, selected artifacts cross with COPY --from, and the final runtime stage contains only the files, identity, configuration, and dependencies the application actually needs.
Learning objectives
- Trace source and dependency inputs through build/test stages, an explicit artifact-copy boundary, a minimal runtime stage, an optional debug target, and the final image identity.
- Explain why each FROM starts a stage with its own filesystem/configuration and why named stages are safer and clearer than positional stage numbers.
- Distinguish build-time toolchains and caches from runtime files, libraries, certificates, user identity, entrypoint, and other release dependencies.
- Compare scratch and Distroless mental models without assuming that a smaller image is automatically more secure or operationally complete.
- Define evidence that proves which artifact crossed stages, which user runs the process, which dependencies remain, and which image/config digest identifies the result.
1. The problem: build environments accidentally become production environments
A beginner can make a container image by starting from a compiler image, copying the source tree, running the compiler, and leaving everything in place. The result may run, but the release now contains the compiler, package manager, source code, temporary build output, dependency caches, shell tools, and any other state the build needed. That is not just a size problem. It expands what must be patched, reviewed, scanned, trusted, transferred, and explained during an incident.
Multi-stage builds solve this by making the artifact boundary explicit. Build/test stages may be intentionally rich. The runtime stage begins from a fresh base and receives only the files that are deliberately copied into it. The question changes from “what can we delete after building?” to “what must cross into production?”
2. Mental model: stages are separate roots connected by declared edges
flowchart TD
A[Source + dependency inputs] --> B[build stage\ncompiler + tests]
B --> C[explicit artifact\nCOPY --from]
C --> D[minimal runtime stage\napp + required runtime deps]
B --> E[optional debug stage\nshell + inspection tools]
D --> F[release image identity]
E --> G[diagnostic image identity]
Each FROM starts a new stage. A stage has its own root
filesystem and image configuration. A later stage does not inherit
previous-stage files merely because it appears later in the
Dockerfile. Files cross only when the Dockerfile explicitly bases
one stage on another or copies/mounts content from another stage,
image, or named context.
This means a compiler can exist in build and be absent
from runtime. It also means a shell can exist in
debug while the release image intentionally has none.
These are separate image states, not runtime modes of one mutable
image.
3. Name stages by responsibility, not position
# syntax=docker/dockerfile:1
FROM golang:1.26 AS build
# compile and test here
FROM busybox:1.37.0 AS debug
COPY --from=build /out/app /app
FROM scratch AS runtime
COPY --from=build /out/app /app
ENTRYPOINT ["/app"]
Stage names such as build, debug, and
runtime preserve meaning if the Dockerfile is
reordered. Numeric references such as --from=0 couple
the copy to source-file position and make later maintenance easier
to get wrong.
Current Dockerfile syntax also allows COPY --from to
read from an image or named context. That makes the source boundary
powerful, but it also makes source identity part of the supply
chain: an external image tag is a build input and should be reviewed
and pinned by digest when reproducibility matters.
4. The artifact boundary is narrower than the workspace
The strongest pattern is to decide what the runtime stage needs
before writing the copy. For a statically compiled service it may be
one executable plus a configuration schema. For a dynamically linked
application it may include a runtime, native libraries, CA roots,
timezone data, locale data, templates, or other assets. Copying
/workspace wholesale defeats the design even if the
Dockerfile technically has multiple stages.
| State | Usually build-only | Potential runtime dependency |
|---|---|---|
| Compilers/linkers | Yes | Rarely |
| Source and test fixtures | Yes | Only deliberate runtime assets |
| Package-manager caches | Yes | No |
| Application binary/runtime | No | Yes |
| CA certificates | Not necessarily | Yes for TLS clients |
| Native libraries/loader | Not necessarily | Yes for dynamically linked programs |
| Shell/debug tools | Debug-only by default | Only when the operational contract requires them |
5. scratch, Distroless, and minimal distributions solve different problems
scratch is Docker's reserved empty base. It has no
shell, package manager, CA bundle, users database, timezone data, or
C library unless you explicitly add those files. It is ideal for
artifacts that truly carry or do not need their runtime
dependencies, such as a carefully built static executable.
Distroless images are not empty. Their purpose is to contain the
application plus a deliberately small runtime dependency set while
omitting ordinary distribution utilities such as package managers
and shells. Current Distroless Debian 13 families include
static/base/cc and language runtimes, with separate
nonroot, debug, and
debug-nonroot variants. The project explicitly
documents debug variants as the shell-equipped path rather than
asking you to modify the production image.
A minimal distribution such as Alpine or a slim Debian/Ubuntu variant provides more familiar operational tooling and package-management behavior. That can be the correct choice when the application requires native libraries, dynamic package patching, or an on-image diagnostic contract. The design goal is purpose-built, not “always scratch.”
6. Runtime process semantics survive the stage boundary
The final stage owns runtime USER,
WORKDIR, ENV, ENTRYPOINT, and
CMD. Build-stage configuration does not automatically
become final-stage configuration. This is a frequent source of
surprises: a non-root user created in the builder is not present in
a scratch stage unless its identity files are copied, yet a numeric
USER 65532:65532 can still instruct the kernel to
launch the process with that UID/GID.
Images without a shell require exec/vector command form. Distroless
documentation explicitly calls this out, and scratch has the same
practical constraint: there is no /bin/sh -c available
to interpret shell-form commands.
7. Inspect before rebuilding
docker version
docker buildx version
docker buildx ls
docker buildx inspect --bootstrap
docker image inspect example:runtime \
--format 'ID={{.Id}} Size={{.Size}} User={{.Config.User}} Entrypoint={{json .Config.Entrypoint}}'
docker image history --no-trunc example:runtime
For a local image, .Id identifies the local image
configuration object; it is not automatically a registry repository
digest. When a Buildx metadata file reports a result digest,
preserve that separately. Chapter 04 established why these
identities must not be collapsed.
8. Evidence map for a multi-stage release
Source revision/hash, Dockerfile hash, build arguments that are not secrets, builder/frontend versions, target platform, base references and resolved digests.
Named source stage, copied paths, artifact checksum, file ownership/mode, and any external image/context identity.
Final base, libraries/certificates/assets, USER, workdir, entrypoint/CMD, image size, and expected shell/tool availability.
Build record/log, metadata-file digest, local image ID, runtime result, and separate debug-target identity.
9. Small challenge: what should cross the boundary?
A service compiles to one static binary but its build stage also contains Git, source code, unit-test fixtures, an SSH client, compiler cache, and a package manager. The service only writes logs to stdout and makes no TLS calls. Design the narrowest runtime copy and explain which evidence would prove the extra build tools did not cross into the release target.
Knowledge check
Does a second FROM automatically inherit files
from the first stage?
No. A new stage begins from its own base unless it explicitly uses a prior stage as its base or copies/mounts content from that stage.
Why is COPY --from=build /workspace / usually a
weak artifact boundary?
It can move source, caches, credentials left by bad build practices, test files, and tools into runtime. Prefer an explicitly named artifact path.
What does scratch provide?
An empty starting filesystem. Everything the application needs must come from copied artifacts or the kernel/runtime outside the image.
Why can a Distroless production image be paired with a debug variant?
It preserves a minimal production contract while providing a separate shell/tooling path for diagnosis instead of modifying the release image.
Does smaller image size prove fewer exploitable application vulnerabilities?
No. Size can reduce unnecessary packages and transfer cost, but application flaws, configuration, privileges, dependency vulnerabilities, and provenance still need independent evidence.
Official references and version notes
-
Docker multi-stage builds
— named stages,
COPY --from, stage reuse, and stopping at a specific target. -
Dockerfile reference
—
FROM, stage naming,COPY --from,USER,ENTRYPOINT, and current frontend semantics. -
Base images
— choosing base images and creating minimal images with the
reserved
scratchbase. - Docker build best practices — small trusted bases, rebuilding, pinning, decoupling applications, and non-root considerations.
-
docker buildx build—--target,--metadata-file, progress, output, and build-result evidence. -
docker image inspect— size and runtime configuration evidence. -
docker image history— image-history evidence and its limits. - GoogleContainerTools/distroless — current Distroless image families, Debian 13 tags, nonroot/debug variants, no-shell behavior, and signature guidance.
- Distroless base image contents — static/base/base-nossl runtime contents such as CA certificates, tzdata, glibc, and libssl.
- Distroless support policy — current Debian-family support timelines.
- Docker Engine 29 release notes — current Engine baseline and bundled component updates.
- Buildx releases — current Buildx release history.
- BuildKit releases — current BuildKit and built-in Dockerfile frontend history.
Version-sensitive statements were rechecked against primary
sources on 2026-09-21. Docker Engine
29.8.1 is the current Engine 29 patch baseline
(released 2026-09-15); Buildx 0.37.1 is the
current upstream Buildx release (2026-09-11); BuildKit
0.33.0 is the current upstream BuildKit release
(2026-09-02) and ships built-in Dockerfile frontend
1.27.0. Current Distroless documentation lists
Debian 13 families with latest, nonroot,
debug, and debug-nonroot variants and
warns that images intentionally lack a shell. Always record the
actual local versions, builder/frontend, base-image digests,
target platform, and external-image identity because installation
bundles and image tags evolve independently.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.