Dockerfile Fundamentals: FROM, RUN, COPY, ADD, WORKDIR, ENV, ARG, USER, CMD, and ENTRYPOINT: Concepts, Architecture, and Mental Model
A Dockerfile is a build definition, not a shell transcript. This lesson models how a Dockerfile frontend turns ordered instructions plus a bounded build context into filesystem changes, image configuration, cache records, and ultimately runtime defaults such as user, working directory, environment, entrypoint, and command.
Learning objectives
- Trace Dockerfile text and build context through the frontend, BuildKit execution graph, image layers/config, and runtime defaults.
- Explain what FROM, RUN, COPY, ADD, WORKDIR, ENV, ARG, USER, CMD, and ENTRYPOINT change and which effects persist.
- Distinguish build-time variables and transformations from image configuration and runtime container state.
- Relate parser/frontend syntax and cache keys to reproducibility without treating cache hits as artifact identity.
- Use digest-level base and output evidence when reproducibility matters.
# syntax=docker/dockerfile:1 for most users
who want the latest stable Dockerfile 1.x frontend, documents
ARG/ENV as inappropriate secret channels,
and supports docker build --check for Dockerfile build
checks. The labs resolve the base tag to a repository digest before
building and never require a registry push, paid service, privileged
mode, Docker socket mount, or broad prune.
1. The problem: a Dockerfile can be readable yet operationally ambiguous
A Dockerfile looks deceptively simple: one instruction per line, from a base image to a final command. But those lines operate in different phases. Some choose build inputs, some execute build-time processes, some copy filesystem content, and some write configuration that is only interpreted when a container is later created. If those phases are collapsed into “Docker ran these commands,” it becomes difficult to reason about cache, secrets, image identity, runtime users, or shutdown behavior.
Chapter 06 showed why an interactive fix in a running container is not an image change. Chapter 07 now moves the durable side of that lesson into source control: if a file, package, user, working directory, environment default, or startup command belongs to every replacement container, the Dockerfile should describe that intent clearly enough for a builder and a reviewer to reproduce it.
2. Mental model: source text becomes build state, then image state, then runtime defaults
flowchart TD
A[Dockerfile + build context] --> B[Dockerfile frontend]
B --> C[BuildKit build graph]
C --> D[RUN / COPY / ADD filesystem results]
C --> E[Image config: ENV USER WORKDIR ENTRYPOINT CMD]
D --> F[Image layers + config]
E --> F
F --> G[Image identity]
G --> H[Container creation]
H --> I[Runtime process + writable layer]
The Dockerfile frontend parses the file and converts it into a build graph. BuildKit executes the graph, reuses cache where inputs match, and produces image content and configuration. Only later does the Docker daemon use that image configuration to create a container. A successful build therefore does not mean a container was run, a registry was updated, or an application became healthy.
3. Parser/frontend syntax is part of the build input
Current Docker documentation recommends the parser directive
# syntax=docker/dockerfile:1 for most users who want
the latest stable 1.x Dockerfile frontend. If no syntax directive is
present, BuildKit uses a bundled frontend. The directive is parsed
before ordinary Dockerfile instructions, so it is not a
RUN command and does not create an image layer.
# syntax=docker/dockerfile:1
FROM busybox:1.37.0
WORKDIR /app
COPY app.sh .
CMD ["./app.sh"]
The :1 frontend reference is intentionally convenient
and moving. A high-assurance build may record or pin a more exact
frontend identity; this beginner chapter records the actual
builder/tool versions and treats the base-image digest and final
image identity as the critical executable evidence.
4. FROM establishes the root filesystem and configuration ancestry
FROM starts a build stage from a base image. A tag such
as busybox:1.37.0 is a human reference that can be
resolved to immutable content. When reproducibility matters, record
or use the repository digest that was actually selected. A digest
identifies content; a tag remains a name that a registry owner can
move.
An ARG declared before the first FROM can
parameterize the base reference. This is useful for a lab that
resolves a digest first and then passes it in explicitly. It does
not make the value a secret, and it does not make other build inputs
deterministic automatically.
5. RUN, COPY, and ADD change build-time filesystem state differently
RUN executes a command during the build and commits its
filesystem result into build state. COPY copies files
or directories from an allowed build source into the image.
ADD can do everything COPY can, but also has special
source behaviors such as local tar extraction and supported
remote/Git sources. Because those extra behaviors can be surprising,
COPY is the clearer default when simple copying is all you need.
Neither instruction means “copy from anywhere on the client machine.” Chapter 08 will define the build context boundary in detail. For now, remember that the builder can only use declared/available build inputs, and copied sensitive data can remain recoverable from image layers even if a later instruction deletes the visible file.
6. WORKDIR, ENV, USER, CMD, and ENTRYPOINT primarily shape image configuration
WORKDIR sets the working directory used by subsequent
instructions and by the runtime defaults. Use an absolute path so a
base-image change cannot silently reinterpret a relative working
directory. ENV records environment values in image
configuration and those values normally appear in containers created
from the image.
USER selects the default user/group for later build
instructions and for runtime
ENTRYPOINT/CMD. A common pattern is to
perform necessary package or filesystem setup as root, create or
prepare application-owned paths, and switch to a non-root UID before
the final runtime configuration.
7. ARG and ENV solve different lifecycle problems
ARG is a build-time parameter. It can control
instructions and is available to build steps after declaration, but
it is not automatically a runtime environment variable.
ENV becomes image configuration and therefore persists
for containers unless overridden. Both can leak sensitive
information through metadata, history, provenance, logs, or
resulting files. Docker's current guidance is explicit: do not pass
secrets with ARG or ENV; use BuildKit secret or SSH mounts instead.
not-a-secret when demonstrating checks.
Never replace them with a real token to “see what happens.” A lesson
about leakage should not create an actual leak.
8. ENTRYPOINT and CMD are image defaults, not build steps
ENTRYPOINT defines the executable that normally runs
when a container starts. CMD can define the default
executable when no ENTRYPOINT exists, or default arguments when
ENTRYPOINT is present. In exec/JSON form, Docker passes arguments
directly without an implicit shell; this makes process identity and
signal delivery more predictable.
ENTRYPOINT ["/app/app.sh"]
CMD ["academy"]
Running the image with docker run IMAGE learner keeps
the ENTRYPOINT and replaces the default CMD argument with
learner. Overriding ENTRYPOINT is a separate, explicit
operation. This separation lets an image define a stable executable
with user-adjustable defaults.
9. Cache is execution reuse, not release identity
BuildKit computes cache keys from the instruction and relevant inputs. Reordering independent instructions can improve reuse; changing an early copied file can invalidate later work. But a cache hit proves only that BuildKit considers the build step reusable under the current inputs. It does not prove that a registry tag points to the desired release or that a deployed container is healthy.
For package managers, combine update and install operations in a well-scoped instruction where the package manager requires those operations to stay coherent. Avoid opaque “one giant RUN” lines whose only goal is minimizing layers; readability, deterministic inputs, cleanup semantics, and cache behavior all matter.
10. What to inspect before trusting a built image
Dockerfile bytes, frontend syntax, context identity, base repository digest, build arguments, and builder version.
Image ID/config digest, history, filesystem layers, labels/metadata, and BuildKit result metadata where captured.
User, working directory, environment, entrypoint, command, and any declared ports/health metadata.
Container ID, actual process UID/PID, command line, environment overrides, exit state, logs, and application behavior.
11. Small challenge: classify the instruction
For each requirement, choose the instruction family and explain whether the result belongs to build-time filesystem state or image configuration: install a package, copy a binary, set a default working directory, pass a non-secret build variant, provide a runtime mode default, switch to UID 65532, and define a stable executable with an overridable argument.
Knowledge check
Does a successful RUN instruction prove a
container created from the final image is healthy?
No. RUN executes during the build. Runtime container creation, process execution, health, networking, and external service state are separate.
Why is COPY normally preferred over
ADD for ordinary local files?
COPY states the simple intent directly. ADD has extra behaviors such as local tar extraction and remote/Git sources that should be chosen only when needed.
Which normally persists into container runtime: ARG or ENV?
ENV is image configuration and normally becomes container environment. ARG is build-time, although its value can still leak through build metadata/history/provenance or resulting files.
Why does the chapter prefer exec-form ENTRYPOINT/CMD?
It avoids an implicit shell, gives predictable argument boundaries, and makes PID 1 and signal behavior easier to reason about.
What identity should you record when a base tag is used for human convenience?
The resolved repository digest/platform used by the build, plus the human-readable tag for traceability.
Official references and version notes
- Dockerfile reference — current syntax and semantics for FROM, RUN, COPY, ADD, WORKDIR, ENV, ARG, USER, CMD, ENTRYPOINT, parser directives, and instruction options.
- Building best practices — guidance on reusable images, cache-aware instruction order, package installation, pinning, and non-root design.
- Build variables — lifecycle and scope of ARG and ENV.
- Build secrets — secret and SSH mounts for sensitive build-time input.
- Build checks reference — Dockerfile checks including SecretsUsedInArgOrEnv, JSONArgsRecommended, and WorkdirRelativePath.
-
Checking build configuration
— running checks with
docker build --checkand interpreting results. - Docker Engine 29 release notes — current Engine 29 behavior and packaging updates.
- BuildKit releases — current builder/frontend changes that can affect Dockerfile features.
Version-sensitive statements were rechecked against primary
documentation on 2026-09-21. Docker Engine 29.8.1
is the current Engine 29 patch release at this checkpoint. Docker
currently recommends # syntax=docker/dockerfile:1 for
most users when an external stable Dockerfile frontend is desired;
that reference follows the latest stable 1.x frontend rather than
being an immutable artifact. Executable labs therefore record the
learner's actual Engine/CLI/Buildx/BuildKit state and resolve the
application base image to a repository digest before building.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.