Dockerfile Fundamentals: FROM, RUN, COPY, ADD, WORKDIR, ENV, ARG, USER, CMD, and ENTRYPOINT: Configuration, Design Choices, and Tradeoffs
Dockerfile instructions often have similar-looking alternatives with different lifecycle and security consequences. This lesson compares COPY with ADD, ARG with ENV, shell with exec form, ENTRYPOINT plus CMD with CMD alone, root with non-root runtime users, and convenient mutable inputs with reproducible identities.
Learning objectives
- Choose COPY or ADD based on actual source behavior rather than habit.
- Choose ARG or ENV according to build-time versus runtime lifecycle and never use either for secrets.
- Choose shell or exec form with explicit awareness of variable expansion, process trees, and signals.
- Design ENTRYPOINT/CMD defaults that are predictable yet overridable.
- Place privileged build steps before USER and keep the final runtime identity least-privileged.
# syntax=docker/dockerfile:1 for most users
who want the latest stable Dockerfile 1.x frontend, documents
ARG/ENV as inappropriate secret channels,
and supports docker build --check for Dockerfile build
checks. The labs resolve the base tag to a repository digest before
building and never require a registry push, paid service, privileged
mode, Docker socket mount, or broad prune.
1. Design goal: make the Dockerfile explain intent
A production Dockerfile is reviewed as both source code and a build policy. The best choice is not always the instruction that produces the fewest lines; it is the choice whose lifecycle, security, cache, and runtime consequences are explicit enough to audit and reproduce.
2. COPY versus ADD
| Question | COPY | ADD |
|---|---|---|
| Ordinary local files/directories | Preferred default. Intent is explicit. | Works, but communicates extra behavior you may not need. |
| Local tar extraction | Copies archive unchanged. | Can unpack recognized local archive formats. |
| Remote URL/Git source | Use a build context/named context or fetch in a controlled RUN when appropriate. | Supports remote/Git source features; version/frontend requirements apply. |
| Supply-chain reasoning | Simple source boundary. | Extra resolution behavior requires stronger source/checksum review. |
Use ADD because you need an ADD capability, not because it is “more powerful.” For remote sources, prefer immutable revisions/checksums where supported and preserve evidence of what was actually fetched.
3. ARG versus ENV
Use ARG for non-secret build parameters: base selection, feature switches, build metadata, or other values consumed while building. Use ENV for a default that should become image/runtime configuration. An ENV can be overridden at container creation; an ARG cannot be expected to appear at runtime unless you intentionally copy its value into ENV or a file.
4. Shell form versus exec form
Shell form is convenient when a shell is the point—for example
pipelines, variable expansion, conditional logic, or package-manager
command composition in a RUN. Exec form is usually
better for ENTRYPOINT/CMD because
arguments are explicit and the intended application can be PID 1
directly.
# Shell features are intentional during build
RUN set -eux; mkdir -p /app; printf '%s\n' "$APP_VERSION" > /app/version
# Runtime process boundaries are explicit
ENTRYPOINT ["/app/server"]
CMD ["--foreground"]
Exec form does not perform shell expansion by itself. If you need
expansion, invoke a shell explicitly and understand that the shell
becomes part of the process model unless it execs the
final application.
5. ENTRYPOINT plus CMD versus CMD alone
Use CMD alone when the entire default command is meant to be easy to
replace. Use ENTRYPOINT plus CMD when the image represents a stable
executable and CMD supplies default arguments. That design makes
docker run IMAGE --help or
docker run IMAGE alternate-arg intuitive while
preserving the executable.
A rigid ENTRYPOINT can be frustrating if users frequently need a diagnostic shell. A totally replaceable CMD can be too loose for purpose-built images. The decision should follow the image's contract, not a universal rule.
6. Root versus non-root runtime user
Many build operations need elevated privileges inside the build environment, but that does not imply the final application process should run as root. Perform package installation and filesystem ownership setup before the final USER instruction, then switch to the least-privileged UID/GID that can read/execute required files and write only intended paths.
Numeric UIDs are portable across images that may not contain a name-service entry, but named users make intent readable when the image defines them. Either way, verify effective UID in a running container and test required write paths rather than assuming the Dockerfile is correct.
7. Absolute WORKDIR versus inherited/relative paths
An explicit absolute WORKDIR /app makes later
COPY/RUN/CMD behavior easier to review and protects the build from
base-image working-directory changes. Docker's current build checks
include a warning for relative WORKDIR patterns that can become
surprising when a base image changes.
8. One well-scoped RUN versus opaque command chains
Package managers often require update/install/cleanup actions to be
coherent in one build step so cached metadata is not separated from
the install it governs. But “combine everything into one RUN” is not
a general readability rule. Group operations that share
lifecycle/cache semantics, use shell safeguards such as
set -eu where appropriate, and keep unrelated concerns
separate enough to diagnose.
Instruction order matters: copy stable dependency manifests before frequently changing source when that allows expensive dependency work to stay cached. Chapter 11 will go deeper into cache mounts and cache import/export.
9. Tag convenience versus digest pinning
A versioned tag is useful for human maintenance because it
communicates the upstream release line. A digest provides immutable
content identity. A common pattern is to record both:
busybox:1.37.0 as the upstream name and the resolved
@sha256:... as the executable build input. Pinning
means you must also own an update process; immutability without
maintenance can leave known vulnerabilities frozen indefinitely.
10. Worked decision table
| Requirement | Preferred choice | Evidence to retain |
|---|---|---|
| Copy local application binary | COPY | Source hash/context identity + resulting image ID |
| Unpack a controlled local tar | ADD, with the behavior documented | Archive hash + build log + resulting filesystem proof |
| Temporary private package credential | BuildKit secret mount | Secret ID/policy and redacted build evidence; never value |
| Runtime mode default | ENV | Image Config.Env + container override evidence |
| Build variant | ARG | Non-secret value/source + resulting build metadata |
| Purpose-built CLI image | ENTRYPOINT + CMD defaults | Config.Entrypoint/Cmd + override tests |
| Application does not need root | Final USER non-root | Config.User + runtime id output |
11. Scenario: choose for reproducibility, not just convenience
A team wants a base tag that auto-updates, an API token in ARG, source copied before dependency installation, and a shell-form CMD. The faster short-term design creates four separate risks: mutable source identity, secret exposure, unnecessary cache invalidation, and ambiguous PID 1/signal behavior. A better design records a base digest plus readable tag, moves the token to a secret mount, orders stable dependency inputs before volatile source when valid, and uses exec-form runtime configuration.
Knowledge check
When should ADD be preferred to COPY?
When you intentionally need an ADD-specific capability such as local tar extraction or supported remote/Git source behavior and have reviewed its source/version implications.
Why is ARG inappropriate for a token even though it is not a runtime ENV by default?
Build arguments can be exposed through history, provenance, logs, metadata, or their effects. Use a BuildKit secret/SSH mount instead.
What is the operational advantage of ENTRYPOINT plus CMD for a purpose-built CLI image?
The executable stays stable while CMD provides user-overridable default arguments.
Why can a digest-pinned base still be operationally unsafe over time?
It is immutable, not automatically maintained. You still need an explicit process to review and adopt patched upstream images.
Official references and version notes
- Dockerfile reference — current syntax and semantics for FROM, RUN, COPY, ADD, WORKDIR, ENV, ARG, USER, CMD, ENTRYPOINT, parser directives, and instruction options.
- Building best practices — guidance on reusable images, cache-aware instruction order, package installation, pinning, and non-root design.
- Build variables — lifecycle and scope of ARG and ENV.
- Build secrets — secret and SSH mounts for sensitive build-time input.
- Build checks reference — Dockerfile checks including SecretsUsedInArgOrEnv, JSONArgsRecommended, and WorkdirRelativePath.
-
Checking build configuration
— running checks with
docker build --checkand interpreting results. - Docker Engine 29 release notes — current Engine 29 behavior and packaging updates.
- BuildKit releases — current builder/frontend changes that can affect Dockerfile features.
Version-sensitive statements were rechecked against primary
documentation on 2026-09-21. Docker Engine 29.8.1
is the current Engine 29 patch release at this checkpoint. Docker
currently recommends # syntax=docker/dockerfile:1 for
most users when an external stable Dockerfile frontend is desired;
that reference follows the latest stable 1.x frontend rather than
being an immutable artifact. Executable labs therefore record the
learner's actual Engine/CLI/Buildx/BuildKit state and resolve the
application base image to a repository digest before building.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.