Chapter 24Lesson 01~120 minutes

Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Concepts, Architecture, and Mental Model

Separate Docker process state, health status, restart policy, dependency readiness, graceful shutdown, and external availability into one causal reliability model.

Health checksRestart policiesGraceful shutdownDependency healthEvidence-first

Learning objectives

  • Separate container process state, Docker health state, dependency readiness, restart policy, graceful-stop behavior, and externally observed availability.
  • Explain why an unhealthy container is still running and why Engine restart policies react to process exit rather than generic health failure.
  • Interpret healthcheck timing, failing streak, probe output, restart count, exit code, and stop signal/timeout as distinct evidence.
  • Connect these mechanisms to a layered self-healing model without treating a local probe as an SLA.
Chapter 24 principle. “Running,” “healthy,” “ready for a dependency,” and “reachable by an external client” are four different claims. Docker health adds an application-oriented signal to process state; restart policies react to container exit; Compose can gate creation on dependency health; and external systems still need their own reachability/SLA evidence.

1. The practical problem: a running process can still be unusable

A container can be running because PID 1 has not exited while the application is deadlocked, serving only errors, or still initializing. Conversely, a healthcheck can fail because the probe itself is broken even while the application serves real traffic correctly. A recovery design must therefore name exactly which layer observed failure and which mechanism is allowed to react.

Chapter 23 gave you time-correlated logs, events, stats, and inspect evidence. Chapter 24 uses those signals to reason about health and recovery rather than guessing from a single status word.

2. Mental model: state, health, and recovery are separate transitions

The container first enters a process lifecycle state. If a healthcheck exists, Docker independently runs a probe inside the container and records starting, healthy, or unhealthy. A restart policy observes process termination and may start the container again. Compose can delay a dependent service until a dependency reports healthy. An external monitor or load balancer may still see a different result because it tests a different path.

Health, restart, and dependency lifecycle
  flowchart TD
    A[Container object] --> B[PID 1 starts]
    B --> C[Application initializes]
    C --> D[Healthcheck executions]
    D -->|exit 0| E[healthy]
    D -->|consecutive exit 1| F[unhealthy]
    B -->|process exits| G[Exited + exit code]
    G --> H{Restart policy?}
    H -->|eligible| B
    H -->|not eligible| I[remains exited]
    E --> J[Compose dependency can proceed]
    E --> K[External check may still fail]
    F --> L[Engine records health; does not generically restart]
            

3. Process state and health state are orthogonal

docker inspect exposes normal runtime state such as .State.Status, .State.Running, .State.ExitCode, .State.OOMKilled, and .RestartCount. With a healthcheck, .State.Health adds health status, failing streak, and recent probe records. An unhealthy container normally remains running until its main process exits or an operator/orchestrator takes another action.

docker inspect NAME \
  --format 'status={{.State.Status}} running={{.State.Running}} exit={{.State.ExitCode}} restarts={{.RestartCount}} health={{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}}'

4. HEALTHCHECK timing is a state machine, not a cron line

The Dockerfile HEALTHCHECK supports --interval, --timeout, --start-period, --start-interval, and --retries. During the start period, failures do not count toward the retry threshold; a success during that period ends startup grace and subsequent failures count normally. The probe command returns 0 for success and 1 for unhealthy; exit code 2 is reserved.

Current Docker also stores a bounded amount of probe output in inspect state, which is useful for diagnosis but must not contain secrets.

Field Operational meaning Evidence to preserve
interval Cadence after normal startup behavior begins. Configured value + timestamps of recent probe records.
timeout Maximum runtime for one probe before failure. Probe duration/output and any timeout evidence.
start_period Initialization grace in which failures do not count. Expected application startup time and observed first success.
start_interval Probe frequency during start period; Engine 25+. Actual Engine version and configured value.
retries Consecutive failures required for unhealthy. Failing streak and last probe records.

5. Restart policies react to exits, not “unhealthy”

Docker restart policies are no, on-failure[:max-retries], always, and unless-stopped. They evaluate container termination semantics. They do not mean “restart when the health status becomes unhealthy.” Docker also applies a restart backoff and treats a container as successfully started only after it has stayed up long enough for restart-policy monitoring.

Policy Restarts after non-zero exit? Daemon restart behavior Manual-stop nuance
no No No automatic restart Remains stopped.
on-failure[:N] Yes, optionally bounded Does not restart merely because daemon restarts Manual stop suppresses policy until manually started again.
always Yes, regardless of exit code Starts again with daemon unless manual-stop suppression applies After manual stop, policy is ignored until daemon restart or manual start.
unless-stopped Yes, regardless of exit code Starts after daemon restart only if it had not been intentionally stopped Preserves intentional stopped state across daemon restart.

6. Graceful shutdown is signal delivery plus a deadline

docker stop sends the container’s configured stop signal, defaulting to SIGTERM if none is configured. Docker waits for the stop timeout, then escalates to SIGKILL. On Linux the default timeout is 10 seconds unless the container has a different configured default; Windows containers use a different default. The application must therefore arrange for PID 1 to receive and handle the signal and complete cleanup before the deadline.

docker inspect NAME --format 'stop_signal={{.Config.StopSignal}} stop_timeout={{.HostConfig.StopTimeout}}'
# Read-only: do not infer graceful handling from these fields alone; confirm with logs/events.

7. Compose dependency health is startup gating, not resilience

Compose starts dependencies in dependency order, but plain startup ordering does not mean readiness. With long-form depends_on and condition: service_healthy, Compose waits for the dependency healthcheck before creating the dependent service. That is useful startup coordination, but the dependent application still needs retry/reconnect logic for failures that happen later.

8. External availability is another boundary

A local healthcheck can pass while DNS, routing, firewall rules, TLS termination, reverse proxy configuration, or an upstream dependency makes the service unavailable to users. Conversely, an external probe can fail because its own path is broken while the container is locally healthy. Keep “container health” and “user-visible availability” separate in dashboards and incident notes.

9. DevOps connection: self-healing needs causal evidence

Automatic recovery is useful only when its trigger is understood. Restarting a crashed process is different from removing an unhealthy endpoint from traffic, delaying a dependent startup, or retrying an external database connection. Preserve exact image identity, health history, exit code, restart count, signal/timeout configuration, and external-path evidence so recovery does not erase the root cause.

Knowledge check

A container shows status=running and health=unhealthy. Must the Docker Engine restart it?

What does start_period protect against?

Why is service_healthy not a substitute for application retry logic?

What happens after docker stop sends the stop signal and the timeout expires?

Why can an external availability check disagree with Docker health?

Next lesson

Next: Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Guided Hands-On Workflow and Core Operations

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used for compatibility notes. Current Dockerfile HEALTHCHECK supports interval, timeout, start_period, start_interval (Engine 25+), and retries. Engine 29 includes a fix for health checks delayed too long when start_interval exceeds start_period. The labs record the actual Engine/CLI/Compose versions before relying on these semantics.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.