Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Concepts, Architecture, and Mental Model
Separate Docker process state, health status, restart policy, dependency readiness, graceful shutdown, and external availability into one causal reliability model.
Learning objectives
- Separate container process state, Docker health state, dependency readiness, restart policy, graceful-stop behavior, and externally observed availability.
-
Explain why an
unhealthycontainer is still running and why Engine restart policies react to process exit rather than generic health failure. - Interpret healthcheck timing, failing streak, probe output, restart count, exit code, and stop signal/timeout as distinct evidence.
- Connect these mechanisms to a layered self-healing model without treating a local probe as an SLA.
1. The practical problem: a running process can still be unusable
A container can be running because PID 1 has not exited
while the application is deadlocked, serving only errors, or still
initializing. Conversely, a healthcheck can fail because the probe
itself is broken even while the application serves real traffic
correctly. A recovery design must therefore name exactly which layer
observed failure and which mechanism is allowed to react.
Chapter 23 gave you time-correlated logs, events, stats, and inspect evidence. Chapter 24 uses those signals to reason about health and recovery rather than guessing from a single status word.
2. Mental model: state, health, and recovery are separate transitions
The container first enters a process lifecycle state. If a
healthcheck exists, Docker independently runs a probe inside the
container and records starting, healthy,
or unhealthy. A restart policy observes process
termination and may start the container again. Compose can delay a
dependent service until a dependency reports healthy. An external
monitor or load balancer may still see a different result because it
tests a different path.
flowchart TD
A[Container object] --> B[PID 1 starts]
B --> C[Application initializes]
C --> D[Healthcheck executions]
D -->|exit 0| E[healthy]
D -->|consecutive exit 1| F[unhealthy]
B -->|process exits| G[Exited + exit code]
G --> H{Restart policy?}
H -->|eligible| B
H -->|not eligible| I[remains exited]
E --> J[Compose dependency can proceed]
E --> K[External check may still fail]
F --> L[Engine records health; does not generically restart]
3. Process state and health state are orthogonal
docker inspect exposes normal runtime state such as
.State.Status, .State.Running,
.State.ExitCode, .State.OOMKilled, and
.RestartCount. With a healthcheck,
.State.Health adds health status, failing streak, and
recent probe records. An unhealthy container normally remains
running until its main process exits or an operator/orchestrator
takes another action.
docker inspect NAME \
--format 'status={{.State.Status}} running={{.State.Running}} exit={{.State.ExitCode}} restarts={{.RestartCount}} health={{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}}'
4. HEALTHCHECK timing is a state machine, not a cron line
The Dockerfile HEALTHCHECK supports
--interval, --timeout,
--start-period, --start-interval, and
--retries. During the start period, failures do not
count toward the retry threshold; a success during that period ends
startup grace and subsequent failures count normally. The probe
command returns 0 for success and 1 for unhealthy; exit code 2 is
reserved.
Current Docker also stores a bounded amount of probe output in inspect state, which is useful for diagnosis but must not contain secrets.
| Field | Operational meaning | Evidence to preserve |
|---|---|---|
interval |
Cadence after normal startup behavior begins. | Configured value + timestamps of recent probe records. |
timeout |
Maximum runtime for one probe before failure. | Probe duration/output and any timeout evidence. |
start_period |
Initialization grace in which failures do not count. | Expected application startup time and observed first success. |
start_interval |
Probe frequency during start period; Engine 25+. | Actual Engine version and configured value. |
retries |
Consecutive failures required for unhealthy.
|
Failing streak and last probe records. |
5. Restart policies react to exits, not “unhealthy”
Docker restart policies are no,
on-failure[:max-retries], always, and
unless-stopped. They evaluate container termination
semantics. They do not mean “restart when the health status becomes
unhealthy.” Docker also applies a restart backoff and treats a
container as successfully started only after it has stayed up long
enough for restart-policy monitoring.
| Policy | Restarts after non-zero exit? | Daemon restart behavior | Manual-stop nuance |
|---|---|---|---|
no |
No | No automatic restart | Remains stopped. |
on-failure[:N] |
Yes, optionally bounded | Does not restart merely because daemon restarts | Manual stop suppresses policy until manually started again. |
always |
Yes, regardless of exit code | Starts again with daemon unless manual-stop suppression applies | After manual stop, policy is ignored until daemon restart or manual start. |
unless-stopped |
Yes, regardless of exit code | Starts after daemon restart only if it had not been intentionally stopped | Preserves intentional stopped state across daemon restart. |
6. Graceful shutdown is signal delivery plus a deadline
docker stop sends the container’s configured stop
signal, defaulting to SIGTERM if none is configured.
Docker waits for the stop timeout, then escalates to
SIGKILL. On Linux the default timeout is 10 seconds
unless the container has a different configured default; Windows
containers use a different default. The application must therefore
arrange for PID 1 to receive and handle the signal and complete
cleanup before the deadline.
docker inspect NAME --format 'stop_signal={{.Config.StopSignal}} stop_timeout={{.HostConfig.StopTimeout}}'
# Read-only: do not infer graceful handling from these fields alone; confirm with logs/events.
7. Compose dependency health is startup gating, not resilience
Compose starts dependencies in dependency order, but plain startup
ordering does not mean readiness. With long-form
depends_on and condition: service_healthy,
Compose waits for the dependency healthcheck before creating the
dependent service. That is useful startup coordination, but the
dependent application still needs retry/reconnect logic for failures
that happen later.
8. External availability is another boundary
A local healthcheck can pass while DNS, routing, firewall rules, TLS termination, reverse proxy configuration, or an upstream dependency makes the service unavailable to users. Conversely, an external probe can fail because its own path is broken while the container is locally healthy. Keep “container health” and “user-visible availability” separate in dashboards and incident notes.
9. DevOps connection: self-healing needs causal evidence
Automatic recovery is useful only when its trigger is understood. Restarting a crashed process is different from removing an unhealthy endpoint from traffic, delaying a dependent startup, or retrying an external database connection. Preserve exact image identity, health history, exit code, restart count, signal/timeout configuration, and external-path evidence so recovery does not erase the root cause.
Knowledge check
A container shows status=running and health=unhealthy. Must the Docker Engine restart it?
No. A healthcheck adds health state; Engine restart policies react to process exit, not generic unhealthy status.
What does start_period protect against?
It gives the application initialization time during which failed probes do not count toward the unhealthy retry threshold; a successful probe ends that grace behavior.
Why is service_healthy not a substitute for application retry logic?
It gates startup on initial dependency health, but dependencies can fail or restart later. The client must remain resilient after startup.
What happens after docker stop sends the stop signal and the timeout expires?
Docker forcibly terminates the container with SIGKILL if it has not exited.
Why can an external availability check disagree with Docker health?
They traverse different paths and dependencies. Docker health runs inside the container; external checks may include DNS, routing, TLS, proxies, firewalls, and upstream systems.
Official references and version notes
- Dockerfile reference — HEALTHCHECK — health states, timing, exit codes, output retention, and health-status events.
- Dockerfile reference — STOPSIGNAL — image-defined graceful-stop signal.
- Docker CLI — docker container stop — stop signal and timeout/escalation behavior.
-
Docker Docs — Restart policies
—
on-failure,always,unless-stopped, successful-start threshold, and manual-stop behavior. - Docker CLI — restart policies — backoff and retry semantics.
-
Docker Compose — startup/shutdown order
—
service_started,service_healthy, andservice_completed_successfully. - Compose services — healthcheck — Compose healthcheck fields and override behavior.
- Compose services — depends_on — dependency conditions and explicit Compose restart propagation.
- Docker Engine 29 release notes — current Engine baseline and healthcheck fixes.
Docker
Engine 29.8.1 is the current Engine baseline used for compatibility
notes. Current Dockerfile HEALTHCHECK supports
interval, timeout,
start_period, start_interval (Engine 25+),
and retries. Engine 29 includes a fix for health checks
delayed too long when start_interval exceeds
start_period. The labs record the actual
Engine/CLI/Compose versions before relying on these semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.