Checkpoint Lab — Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns
Engineer startup delay, unhealthy state, process crash, and graceful shutdown, then assemble a health/restart/signal evidence packet.
Learning objectives
- Engineer startup delay, unhealthy-but-running behavior, and a process crash in one disposable scenario.
- Capture a correlated packet of health history, Engine events, restart evidence, signal handling, and dependency gating.
- Predict which mechanisms will react to each failure before triggering it, then compare prediction to evidence.
- State the limits of Docker-local self-healing and bridge the result into the security model of Chapter 25.
1. Create the checkpoint workspace and capture assumptions
rm -rf dkr24-checkpoint dkr24-capstone
mkdir -p dkr24-checkpoint dkr24-capstone
docker context show | tee dkr24-checkpoint/context.txt
docker version | tee dkr24-checkpoint/docker-version.txt
docker info | tee dkr24-checkpoint/docker-info.txt
docker compose version | tee dkr24-checkpoint/compose-version.txt 2>&1 || true
docker buildx version | tee dkr24-checkpoint/buildx-version.txt 2>&1 || true
docker image inspect alpine:3.22 --format '{{json .RepoDigests}}' | tee dkr24-checkpoint/image-digests.txt
2. Write predictions before execution
Record predictions in plain text before the run. At minimum include:
-
Startup delay: health should stay
startinguntil the readiness marker appears; the dependent service should wait. - Unhealthy condition: PID 1 should stay running, health should become unhealthy, and restart count should not change merely because of health.
-
Crash: a non-zero process exit under bounded
on-failureshould increment restart count. - Graceful stop: TERM should be logged and the process should exit before the stop timeout.
cat > dkr24-checkpoint/predictions.txt <<'EOF'
startup-delay: health=starting then healthy; worker waits for healthy
unhealthy-marker-removed: process stays running; restart_count unchanged
synthetic-crash: on-failure performs bounded restart attempt(s)
graceful-stop: SIGTERM handled before timeout; no forced-kill evidence expected
EOF
3. Define the disposable Compose scenario
The api service waits four seconds before declaring
itself ready. Its healthcheck uses only shell test. The
worker starts only after api is healthy.
The restart policy is on-failure:2, but health failure
alone will not trigger it.
cat > dkr24-capstone/compose.yaml <<'YAML'
name: dkr24-capstone
services:
api:
image: alpine:3.22
restart: on-failure:2
stop_signal: SIGTERM
stop_grace_period: 5s
command:
- sh
- -c
- |
trap 'echo TERM-received; exit 0' TERM INT
rm -f /tmp/ready
echo startup-begin
sleep 4
touch /tmp/ready
echo startup-ready
while :; do sleep 1; done
healthcheck:
test: ["CMD-SHELL", "test -f /tmp/ready"]
interval: 2s
timeout: 1s
retries: 2
start_period: 5s
start_interval: 1s
labels:
academy: docker-ch24
scenario: checkpoint
worker:
image: alpine:3.22
depends_on:
api:
condition: service_healthy
command: ["sh", "-c", "echo worker-started-after-api-health; sleep 30"]
labels:
academy: docker-ch24
scenario: checkpoint
YAML
(cd dkr24-capstone && docker compose config) | tee dkr24-checkpoint/compose-normalized.yaml
4. Observe startup delay and dependency gating
START_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)
(cd dkr24-capstone && docker compose up -d)
sleep 1
(cd dkr24-capstone && docker compose ps) | tee dkr24-checkpoint/ps-during-startup.txt
sleep 6
(cd dkr24-capstone && docker compose ps) | tee dkr24-checkpoint/ps-after-health.txt
(cd dkr24-capstone && docker compose logs --timestamps) | tee dkr24-checkpoint/startup-logs.txt
5. Create unhealthy-but-running state and preserve evidence
Remove the readiness marker without stopping PID 1. Compare restart count before and after health becomes unhealthy.
API_ID=$(cd dkr24-capstone && docker compose ps -q api)
BEFORE=$(docker inspect "$API_ID" --format '{{.RestartCount}}')
docker exec "$API_ID" rm -f /tmp/ready
sleep 6
docker inspect "$API_ID" \
--format 'id={{.Id}} status={{.State.Status}} running={{.State.Running}} health={{.State.Health.Status}} streak={{.State.Health.FailingStreak}} restarts={{.RestartCount}}' | tee dkr24-checkpoint/unhealthy-state.txt
docker inspect "$API_ID" --format '{{json .State.Health.Log}}' | tee dkr24-checkpoint/unhealthy-health-log.json
AFTER=$(docker inspect "$API_ID" --format '{{.RestartCount}}')
printf 'before=%s after=%s
' "$BEFORE" "$AFTER" | tee dkr24-checkpoint/unhealthy-restart-count.txt
6. Trigger a real process crash and observe bounded restart behavior
First restore health, then deliberately terminate PID 1 with a non-zero exit from inside the disposable container. The health transition and the process crash are intentionally separate events.
docker exec "$API_ID" touch /tmp/ready
sleep 3
# Synthetic crash: ask PID 1 shell to exit non-zero via TERM is not suitable,
# because the trap exits 0. Instead run a one-shot crash container under the same policy.
docker rm -f dkr24-crash 2>/dev/null || true
docker run -d \
--name dkr24-crash \
--label academy=docker-ch24 \
--label scenario=checkpoint-crash \
--restart on-failure:1 alpine:3.22 sh -c 'sleep 11; echo crash-42 >&2; exit 42'
sleep 25
docker inspect dkr24-crash \
--format 'status={{.State.Status}} exit={{.State.ExitCode}} restarts={{.RestartCount}} policy={{.HostConfig.RestartPolicy.Name}}' | tee dkr24-checkpoint/crash-state.txt
docker logs --timestamps dkr24-crash 2>&1 | tee dkr24-checkpoint/crash-logs.txt
7. Verify graceful shutdown separately
Graceful shutdown is not the same as health or restart behavior. Stop only the API service and preserve its TERM log and final state.
(cd dkr24-capstone && docker compose stop -t 5 api)
(cd dkr24-capstone && docker compose logs --timestamps api) | tee dkr24-checkpoint/graceful-api-logs.txt
docker inspect "$API_ID" \
--format 'exit={{.State.ExitCode}} status={{.State.Status}} finished={{.State.FinishedAt}}' | tee dkr24-checkpoint/graceful-api-state.txt
8. Capture the Engine timeline and final inventories
END_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)
docker events \
--since "$START_UTC" \
--until "$END_UTC" \
--filter type=container \
--format '{{json .}}' | grep -E 'dkr24-capstone|dkr24-crash|docker-ch24' > dkr24-checkpoint/engine-events.jsonl || true
(cd dkr24-capstone && docker compose ps -a) | tee dkr24-checkpoint/final-compose-ps.txt
docker inspect dkr24-crash > dkr24-checkpoint/crash-inspect.json
9. Interpret mechanism by mechanism
| Event | Expected mechanism | What should not be inferred |
|---|---|---|
| Startup delay |
Health remains starting; Compose worker waits
for service_healthy.
|
Not a crash and not necessarily an error. |
| Readiness marker removed |
Health becomes unhealthy; API process remains
running.
|
No generic Engine restart should be inferred from unhealthy alone. |
| Crash container exits 42 |
on-failure:1 performs bounded restart behavior.
|
Does not prove the application became healthy after restart. |
| Compose stop API | SIGTERM + five-second grace path should allow trap to exit cleanly. | A clean stop is not the same as a successful healthcheck. |
10. Required evidence packet and limitations note
The packet should contain context/version evidence, normalized
Compose model, image digest, startup logs, health history, restart
counts, exit codes, Engine events, and graceful-stop logs. Add an
assumptions.txt stating that the lab is
local/disposable, uses Alpine 3.22, tests Docker-local behavior
only, and does not claim external SLA or production orchestration
behavior.
cat > dkr24-checkpoint/assumptions.txt <<'EOF'
Local disposable lab only.
No production endpoint or real secret.
Docker health is container-local evidence, not an SLA.
Compose service_healthy gates startup but does not replace runtime retry/reconnect logic.
Restart policy responds to process exit; unhealthy alone is not a generic Engine restart trigger.
Engine 29.8.1 is the documentation baseline; actual local versions are recorded in this packet.
EOF
11. Cleanup and rollback
Remove only the checkpoint resources after evidence is complete. Do not use volume/image/system prune.
(cd dkr24-capstone && docker compose down --remove-orphans)
docker rm -f dkr24-crash 2>/dev/null || true
rm -rf dkr24-capstone
tar -czf dkr24-checkpoint.tgz dkr24-checkpoint/
12. What Chapter 24 adds to the production Docker model
You can now separate process lifecycle, container health, dependency readiness, restart behavior, graceful shutdown, and external availability—and attach the correct evidence to each. That prevents restart loops from masquerading as resilience and healthchecks from masquerading as user availability. Chapter 25 now moves from reliability boundaries to security boundaries: namespaces, Linux capabilities, seccomp, AppArmor, SELinux, devices, and privilege containment.
Knowledge check
In the checkpoint, why should unhealthy-state restart count remain unchanged?
Because the API process is still running. Health failure alone does not trigger an Engine restart policy.
What mechanism delays the worker until the API is ready?
Compose depends_on with condition: service_healthy, driven by the API healthcheck.
Why is the synthetic crash tested in a separate container?
The API TERM trap exits cleanly by design. A separate bounded crash scenario isolates non-zero process-exit restart behavior from graceful-stop behavior.
What does a successful graceful stop prove?
That the configured stop signal reached the process and it exited within the grace period; it does not prove user-visible availability or health.
What is the most important limitation statement for this lab?
Docker-local health/restart evidence is not an external SLA and does not replace application retry logic or production orchestration/traffic-management checks.
Official references and version notes
- Dockerfile reference — HEALTHCHECK — health states, timing, exit codes, output retention, and health-status events.
- Dockerfile reference — STOPSIGNAL — image-defined graceful-stop signal.
- Docker CLI — docker container stop — stop signal and timeout/escalation behavior.
-
Docker Docs — Restart policies
—
on-failure,always,unless-stopped, successful-start threshold, and manual-stop behavior. - Docker CLI — restart policies — backoff and retry semantics.
-
Docker Compose — startup/shutdown order
—
service_started,service_healthy, andservice_completed_successfully. - Compose services — healthcheck — Compose healthcheck fields and override behavior.
- Compose services — depends_on — dependency conditions and explicit Compose restart propagation.
- Docker Engine 29 release notes — current Engine baseline and healthcheck fixes.
Docker
Engine 29.8.1 is the current Engine baseline used for compatibility
notes. Current Dockerfile HEALTHCHECK supports
interval, timeout,
start_period, start_interval (Engine 25+),
and retries. Engine 29 includes a fix for health checks
delayed too long when start_interval exceeds
start_period. The labs record the actual
Engine/CLI/Compose versions before relying on these semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.