Chapter 05Lesson 04~110 minutes

Running Containers: create, run, start, stop, restart, rm, exec, attach, inspect, and Lifecycle State: Diagnostics, Failure Modes, Security, and Performance

Lifecycle failures are often made worse by the attempted fix. This lesson preserves first-failure state before restart or removal, diagnoses stopped-container exec errors, unsafe attachment, forced deletion, missing evidence from --rm, restart loops, and name/identity confusion, then applies the smallest correction.

DiagnosticsFirst-failure evidenceSignalsRestart loopsSafe repair

Learning objectives

  • Localize failures to client/context, container configuration, primary process, signal handling, exec/attach interaction, restart policy, storage/network dependency, or object cleanup.
  • Preserve inspect output, logs, events, IDs, image digest, exit code, and timestamps before restarting or deleting a failed container.
  • Diagnose why docker exec fails against a stopped container and why forced removal can destroy useful process evidence.
  • Recognize how restart loops and automatic removal can erase the stable inspection window needed for diagnosis.
  • Repair lifecycle failures with the least destructive change and independently verify the resulting state.
Chapter 05 evidence baseline — verified 2026-09-21. This chapter uses a free/local/disposable Docker path and captures the exact image digest and container ID before lifecycle changes. Version-sensitive behavior is verified against the active daemon instead of assumed. Docker Engine 29.8.1 is the current Engine 29 patch baseline at verification time; Engine 29.7 introduced a daemon-level default-stop-timeout option. Current Docker documentation distinguishes graceful stop from kill/force-removal, exec from attach, and explicit restart from restart policy. Labs target only named/labeled chapter-owned containers and never use broad prune or production workloads.

1. Diagnostic rule: preserve state before changing state

Lifecycle troubleshooting is unusually vulnerable to evidence destruction because the obvious commands—restart, remove, rerun—change the very object you need to inspect. Before any repair, capture the active context, daemon version, container ID, image identity, command, state, PID, exit code, restart count, timestamps, logs, events, mounts, networks, and labels.

Do not begin with: docker rm -f, broad prune, a blind daemon restart, or repeated docker run. Those actions may change identity, kill processes, or delete the post-failure object before you know the causal layer.

2. Evidence-first diagnostic sequence

  1. Confirm docker context show, client/server versions, OS type, and daemon reachability.
  2. Capture the exact container ID/name and image digest.
  3. Inspect .State, restart policy/count, configured command, stop signal/timeout, labels, mounts, and networks.
  4. Capture docker logs --timestamps and a bounded docker events window if available.
  5. Decide whether the failure is configuration, primary process, signal handling, exec/attach interaction, restart policy, storage/network dependency, or cleanup.
  6. Apply one least-destructive correction and verify the resulting state.

3. Broken example: exec against a stopped container

Create a disposable stopped object that exits immediately:

SOURCE='busybox:1.37.0'
PINNED=$(docker image inspect "$SOURCE" --format '{{index .RepoDigests 0}}' 2>/dev/null || true)
if [ -z "$PINNED" ]; then docker pull "$SOURCE"; PINNED=$(docker image inspect "$SOURCE" --format '{{index .RepoDigests 0}}'); fi

BROKEN='devops-academy-ch05-broken-exec'
docker run --name "$BROKEN" --label devops-academy.lab=ch05 "$PINNED" sh -c 'echo primary-finished; exit 17' || true

docker inspect "$BROKEN" --format 'ID={{.Id}} Status={{.State.Status}} Exit={{.State.ExitCode}} Finished={{.State.FinishedAt}}'
docker logs "$BROKEN"

docker exec "$BROKEN" sh -c 'echo should-not-run' 2> ch05-exec-error.txt || true
cat ch05-exec-error.txt

The repair is not to force an exec mechanism. The error is causal evidence: exec requires the primary process to be running. If you need the stopped object's files, choose an appropriate read-only/copy/export technique later in the course; if you need a new process from the same image, create a new disposable diagnostic container with an explicit purpose.

4. Broken example: trying to remove a running object

RUNNING='devops-academy-ch05-broken-rm'
docker run -d --name "$RUNNING" --label devops-academy.lab=ch05 "$PINNED" sh -c 'while :; do sleep 1; done'

docker inspect "$RUNNING" --format 'ID={{.Id}} Status={{.State.Status}} Pid={{.State.Pid}}'
docker rm "$RUNNING" 2> ch05-rm-error.txt || true
cat ch05-rm-error.txt

The refusal is protective. Do not immediately add -f. First decide whether graceful shutdown is required, preserve logs/state, then:

docker stop --timeout 4 "$RUNNING"
docker inspect "$RUNNING" --format 'Status={{.State.Status}} Exit={{.State.ExitCode}} Finished={{.State.FinishedAt}}'
docker rm "$RUNNING"

Force removal would have sent SIGKILL and deleted the object in one path. That is appropriate only when the exact disposable object must be forcibly terminated and the evidence/risk trade-off is explicit.

5. Failure mode: using attach when exec/logs are intended

If the goal is “open a troubleshooting shell,” docker attach is the wrong abstraction: it connects to the existing primary process. A Ctrl-c or forwarded signal can terminate that workload. If the goal is only to read output, docker logs avoids coupling your terminal to the primary stream. If the goal is a bounded command, docker exec is more explicit.

Attach also has throughput implications: Docker documents a buffer in the attach path and recommends logs rather than slow attached clients for performance-critical high-output processes.

6. Failure mode: automatic removal erased the object you needed

A team runs an intermittent job with --rm. The job fails, prints one vague line, and disappears. The team now lacks Docker's retained exit state, timestamps, command/configuration, and writable-layer context.

The correction is not “never use --rm.” Use it where post-exit evidence is unnecessary or exported elsewhere. For unstable workloads, retain the container until the diagnostic packet is captured, then perform exact cleanup.

7. Failure mode: restart policy hides the first failure

A container with an automatic restart policy crashes repeatedly. docker ps may show Restarting or a recently restarted process, while the original failure scrolls away. Capture .RestartCount, restart policy, timestamped logs, and events before manually restarting anything.

docker inspect CONTAINER --format \
'Status={{.State.Status}} RestartCount={{.RestartCount}} Policy={{json .HostConfig.RestartPolicy}} Started={{.State.StartedAt}} Finished={{.State.FinishedAt}}'

docker logs --timestamps --tail 200 CONTAINER

Do not disable a restart policy on a production workload merely to make diagnosis easier unless change ownership, impact, and rollback are authorized. Reproduce the failure in a disposable environment when possible.

8. Failure mode: same name, different object

An operator screenshots docker ps -a --filter name=api, removes the container, deploys a replacement also called api, and assumes later observations belong to the same object. The cure is simple: incident evidence must record the full container ID and image digest at every transition.

Names are operational handles. IDs and digests are identity evidence.

9. Failure mode: killing PID 1 before capturing state

docker kill defaults to SIGKILL. It is useful when a process will not terminate and a hard stop is justified, but it does not provide graceful cleanup. If the workload is disposable and a kill is necessary, capture logs, inspect state, external dependency state, and any storage evidence first.

Chapter 06 goes deeper into PID 1 and signal forwarding. For now, remember that shell-form entrypoints and applications that do not handle the configured stop signal can make graceful shutdown behave differently than expected.

10. Security-sensitive troubleshooting boundaries

  • Do not add --privileged to make an exec command work.
  • Do not mount the Docker socket into a diagnostic container as a shortcut.
  • Do not use chmod 777 on host paths to bypass a permission diagnosis.
  • Do not disable seccomp, AppArmor/SELinux, TLS, or host firewall controls to “prove Docker works.”
  • Do not print environment variables wholesale when they may contain credentials.
  • Do not delete runtime/data-root directories to recover a stuck container.

Each shortcut changes the trust boundary and may create a larger incident than the original lifecycle failure.

11. Performance implications of lifecycle choices

Repeated container creation has different costs from process restart inside an existing object, and excessive restart loops can consume CPU, log I/O, network connections, and dependency capacity. Attaching a slow client to a high-output process can also affect output flow. Measure before optimizing: startup latency, restart frequency, log volume, memory/CPU, dependency reconnects, and image/local-store behavior.

Do not keep a broken service alive merely to avoid startup cost. Correctness and recoverability come first; performance tuning follows measured evidence.

12. Cleanup the intentionally broken objects

docker ps -a --filter 'label=devops-academy.lab=ch05' --format 'ID={{.ID}} Name={{.Names}} Status={{.Status}}'

docker rm "$BROKEN" 2>/dev/null || true
rm -f ch05-exec-error.txt ch05-rm-error.txt

The running-removal example was already stopped and removed in its repair path. Before removing anything else, verify the label/name belongs to this lab. Do not translate this cleanup into a broad prune command.

13. Diagnostic challenge

A service has restarted six times. The current process is healthy, but users report a 30-second outage every hour. Your teammate wants to run docker restart “to clear it.” Write the evidence sequence you would collect before changing state. Your answer must distinguish current PID from container ID, include restart policy/count and timestamped logs/events, preserve image digest and dependency/network/storage context, and explain why another restart could obscure the periodic trigger.

Next lesson

Next: Checkpoint Lab

Build one complete lifecycle timeline and prove every state transition with inspect evidence before and after the command.

Knowledge check

Why does docker exec fail after the primary process exits?

What is the safe interpretation of a failed docker rm on a running container?

Why can a restart loop be an observability problem even if the service eventually comes back?

What immutable or stable identities belong in lifecycle incident evidence?

Official references and version notes

  • docker container command group — current container-management surface, including create, run, start, stop, restart, kill, rm, exec, attach, inspect, pause, and wait.
  • docker container create — creates a container object without starting its primary process and records runtime configuration such as restart policy and stop timeout.
  • docker container run — create-and-start convenience behavior, foreground/detached operation, automatic removal, signal proxying, and stop configuration.
  • docker container start — starts an existing stopped container and optionally attaches standard streams.
  • docker container stop — graceful stop signal, timeout, and eventual SIGKILL escalation semantics.
  • docker container kill — immediate/default SIGKILL behavior and explicit signal selection.
  • docker container restart — stop-then-start semantics, configurable signal, and timeout behavior.
  • docker container rm — exact container removal; force-removing a running container uses SIGKILL and -v affects anonymous volumes.
  • docker container exec — starts an additional command only while the container's primary PID 1 is running; exec commands are not automatically restarted with the container.
  • docker container attach — attaches local standard streams to the existing ENTRYPOINT/CMD process, signal-proxy implications, detach keys, and throughput caveats.
  • docker container inspect — low-level container configuration and state evidence, including PID, exit code, restart count, timestamps, mounts, and networks.
  • Start containers automatically — restart-policy behavior, successful-start monitoring, manual-stop interaction, and distinction from live restore.
  • Docker Engine 29 release notes — current Engine 29 baseline; Engine 29.7 added the daemon-level default-stop-timeout option and 29.8.1 is the current patch release at chapter verification time.
Current baseline, not a frozen requirement

Verified 2026-09-21: Docker Engine 29.8.1 is the current Engine 29 patch release. Docker's current CLI documentation defines docker exec as a command that exists only while the container's primary process is running; docker stop sends the configured stop signal and escalates to SIGKILL after the timeout; docker kill defaults to SIGKILL; and docker rm --force kills a running container before removing the object. Linux and Windows defaults and process-isolation behavior differ, so every executable lab records the actual client/server, OS type, architecture, stop configuration, image digest, and container state observed on the learner's environment.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.