Chapter 05Lesson 03~100 minutes

Running Containers: create, run, start, stop, restart, rm, exec, attach, inspect, and Lifecycle State: Configuration, Design Choices, and Tradeoffs

Container operation involves deliberate choices: create-plus-start versus run, foreground versus detached, exec versus a separate diagnostic container, automatic removal versus retained failure evidence, and graceful-stop timing versus faster termination. This lesson evaluates those choices using reproducibility, safety, debugging, and rollback evidence.

Trade-offsForeground/detached--rmShutdownDiagnostics

Learning objectives

  • Choose between docker run and docker create plus docker start according to inspection and orchestration needs.
  • Choose foreground or detached execution according to ownership of terminal streams, operator interaction, and automation requirements.
  • Decide when docker exec is appropriate and when a separate diagnostic container or image-level change is safer and more reproducible.
  • Balance --rm convenience against the need to retain exit-state and filesystem evidence during failures.
  • Select stop timeouts and restart behavior based on application shutdown guarantees instead of arbitrary speed.
Chapter 05 evidence baseline — verified 2026-09-21. This chapter uses a free/local/disposable Docker path and captures the exact image digest and container ID before lifecycle changes. Version-sensitive behavior is verified against the active daemon instead of assumed. Docker Engine 29.8.1 is the current Engine 29 patch baseline at verification time; Engine 29.7 introduced a daemon-level default-stop-timeout option. Current Docker documentation distinguishes graceful stop from kill/force-removal, exec from attach, and explicit restart from restart policy. Labs target only named/labeled chapter-owned containers and never use broad prune or production workloads.

1. Design choices are lifecycle contracts

Container commands encode operational assumptions. A one-line docker run --rm ... may be perfect for a disposable compiler job and a poor choice for an intermittently failing service where post-exit evidence matters. The correct decision depends on state ownership, observability, restart responsibility, data durability, trust, and recovery requirements.

This lesson uses the Chapter 05 state model as the comparison framework: what object is created, what process exists, what survives exit, what evidence remains, and who is allowed to mutate or remove the state.

2. run versus create + start

Choice Strength Cost/risk Prefer when
docker run Concise create-and-start flow; easy for disposable jobs Creation and execution evidence happen together; repeated run creates new objects The desired configuration is already reviewed and immediate start is correct
docker create + start Lets you inspect final object configuration before execution More steps and object-management responsibility Teaching, debugging, pre-start verification, or automation that deliberately stages objects

In orchestrated production systems you usually do not manually stage containers this way; the orchestrator owns desired state. The mental model still matters because it explains why “object exists” and “process is running” are different observations.

3. Foreground versus detached

Foreground mode binds the operator/automation session closely to the container's standard streams and exit status. It is ideal for short jobs where the caller should wait and react to the process result. Detached mode returns control immediately and makes logs/inspect/events the normal observation surfaces.

Question Foreground Detached
Who waits for exit? Calling CLI/session Daemon continues after CLI returns
How are logs observed? Directly on attached streams docker logs or logging system
Signal risk Terminal signals can reach primary process Less accidental terminal coupling
Typical use Build/test/one-shot command, interactive tool Long-lived local service or background workload

Do not infer application health from detached mode. -d says the CLI detached; it does not say the application initialized successfully.

4. Exec versus a separate diagnostic container

docker exec is appropriate when you need a bounded command inside the exact running container namespaces/filesystem and the image actually contains the required tool. It is fast and context-rich, but it mutates runtime state if you change files and creates a privileged diagnostic path if you run it as root.

A separate diagnostic container can be better when you want tooling that should not be installed in the production image, when you want a clean ephemeral environment, or when you need to inspect a network endpoint without entering the target container. It still requires careful namespace/network/mount choices and should not be given broad host access.

Neither choice justifies --privileged or Docker-socket mounting as a shortcut. Diagnose the actual missing permission or observation boundary.

5. --rm versus retained failure evidence

--rm automatically removes a container when it exits. That is excellent for known-disposable successful commands because it prevents object accumulation. It is weaker for intermittent failures because the Docker object and its inspectable post-exit state disappear immediately.

A practical pattern is to use automatic removal for deterministic disposable tasks whose logs/results are exported elsewhere, and retain objects when post-mortem inspection is part of the diagnostic contract. In CI, this decision should be intentional and documented rather than inherited from a copied command.

6. Stop timeout is an application requirement

A stop timeout is not “how impatient Docker should be.” It is a contract between the runtime and the application. The app may need time to stop accepting work, drain in-flight requests, flush buffers, close transactions, persist checkpoints, or release leases. A timeout shorter than the app's safe shutdown path converts normal deployment into forced termination.

Conversely, an infinite timeout can make recovery hang forever when a process is wedged. Choose a bounded value based on measured shutdown behavior and upstream orchestration deadlines. Record the configured stop signal and timeout as part of deployment evidence.

7. Explicit restart versus restart policy

An explicit docker restart is an operator/API action. A restart policy is daemon configuration on the container that controls automatic behavior after exit/daemon restart. Current Docker policy options include no, on-failure[:max-retries], always, and unless-stopped, each with different semantics.

Automatic restart can improve availability for transient process failure, but it can also turn a deterministic crash into a fast loop that repeatedly consumes resources and obscures the first error. Logs/events, backoff behavior, restart count, and application dependency state must remain observable.

8. Name convenience versus immutable incident identity

Names are excellent operator interfaces, but they are not immutable evidence. An incident timeline should capture both the name and the full container ID. If api is removed and a new container is created as api, name-based screenshots can look continuous even though the object changed completely.

The image input should likewise be recorded by digest rather than tag alone. This gives an evidence chain: source/release → image digest → container ID → primary PID/timestamps → external service state.

9. Decision table

Scenario Recommended pattern Evidence to retain Why
One-shot formatter/test Foreground; --rm only if logs/results are already retained Image digest, command, exit code, output Caller owns completion
Local long-lived service Detached; explicit name/label; logs + inspect ID, PID, image digest, ports/mounts, stop policy CLI should not own lifetime
Intermittent crash diagnosis Retain object; avoid --rm; conservative restart First logs/events, exit code, timestamps, restart count Post-exit evidence is valuable
Need one in-container query Bounded exec as least-privileged user Command, user, timestamp, target ID Uses exact target context
Need rich troubleshooting tools Separate diagnostic image/container when practical Tool image digest, network/mount scope Keeps production image minimal and changes explicit

10. Worked scenario: payment worker shutdown

A payment worker receives jobs from a queue. It normally needs up to 18 seconds to finish an in-flight job and acknowledge it. The current container uses the default stop behavior and is sometimes force-killed during rapid deployments, causing duplicate work after restart.

  1. Measure real shutdown latency and verify the process handles the configured stop signal.
  2. Set a stop timeout safely above the measured drain window while respecting the deployment system's outer timeout.
  3. Record container ID, image digest, stop configuration, queue depth, and application shutdown log.
  4. Test termination on a disposable environment and verify no new work is accepted after the signal.
  5. Do not “solve” duplicates only by extending Docker timeout; the application still needs idempotency and correct queue acknowledgement semantics.

The Docker choice is only one layer in a reliable shutdown design.

11. Platform and version prerequisites

Native Linux, Docker Desktop, and Windows containers have different process-isolation boundaries and default stop timing. Current Docker documentation lists 10 seconds as the daemon fallback for Linux containers and 30 seconds for Windows containers when no explicit container/default timeout is configured. Engine 29.7 also added default-stop-timeout at the daemon level.

Therefore a runbook should say “inspect the configured stop timeout and platform,” not “Docker always waits 10 seconds.” The same principle applies to PID visibility, pause support, signal semantics, and TTY behavior.

12. Design challenge

You own two workloads: a disposable schema-linter that should leave no local object after success, and a flaky local API service whose first crash must be investigated. Choose foreground/detached mode, --rm or retention, restart policy, and diagnostic method for each. For every choice, list the evidence you expect to remain after the process exits and how you prevent cleanup from touching an unrelated container with the same application role.

Next lesson

Next: Diagnostics, Failure Modes, Security, and Performance

Apply these design choices to intentionally broken lifecycle cases while preserving the evidence that explains the first failure.

Knowledge check

Why might docker run be the wrong command when you need pre-start inspection?

When is --rm a strong choice?

Why can a very short stop timeout create application-level failures?

Why should incident notes record both a container name and ID?

Official references and version notes

  • docker container command group — current container-management surface, including create, run, start, stop, restart, kill, rm, exec, attach, inspect, pause, and wait.
  • docker container create — creates a container object without starting its primary process and records runtime configuration such as restart policy and stop timeout.
  • docker container run — create-and-start convenience behavior, foreground/detached operation, automatic removal, signal proxying, and stop configuration.
  • docker container start — starts an existing stopped container and optionally attaches standard streams.
  • docker container stop — graceful stop signal, timeout, and eventual SIGKILL escalation semantics.
  • docker container kill — immediate/default SIGKILL behavior and explicit signal selection.
  • docker container restart — stop-then-start semantics, configurable signal, and timeout behavior.
  • docker container rm — exact container removal; force-removing a running container uses SIGKILL and -v affects anonymous volumes.
  • docker container exec — starts an additional command only while the container's primary PID 1 is running; exec commands are not automatically restarted with the container.
  • docker container attach — attaches local standard streams to the existing ENTRYPOINT/CMD process, signal-proxy implications, detach keys, and throughput caveats.
  • docker container inspect — low-level container configuration and state evidence, including PID, exit code, restart count, timestamps, mounts, and networks.
  • Start containers automatically — restart-policy behavior, successful-start monitoring, manual-stop interaction, and distinction from live restore.
  • Docker Engine 29 release notes — current Engine 29 baseline; Engine 29.7 added the daemon-level default-stop-timeout option and 29.8.1 is the current patch release at chapter verification time.
Current baseline, not a frozen requirement

Verified 2026-09-21: Docker Engine 29.8.1 is the current Engine 29 patch release. Docker's current CLI documentation defines docker exec as a command that exists only while the container's primary process is running; docker stop sends the configured stop signal and escalates to SIGKILL after the timeout; docker kill defaults to SIGKILL; and docker rm --force kills a running container before removing the object. Linux and Windows defaults and process-isolation behavior differ, so every executable lab records the actual client/server, OS type, architecture, stop configuration, image digest, and container state observed on the learner's environment.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.