Chapter 24Lesson 05~120 minutes

Checkpoint Lab — Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns

Engineer startup delay, unhealthy state, process crash, and graceful shutdown, then assemble a health/restart/signal evidence packet.

Health checksRestart policiesGraceful shutdownDependency healthEvidence-first

Learning objectives

  • Engineer startup delay, unhealthy-but-running behavior, and a process crash in one disposable scenario.
  • Capture a correlated packet of health history, Engine events, restart evidence, signal handling, and dependency gating.
  • Predict which mechanisms will react to each failure before triggering it, then compare prediction to evidence.
  • State the limits of Docker-local self-healing and bridge the result into the security model of Chapter 25.
Checkpoint objective. Build one evidence packet that distinguishes three incidents: startup delay, unhealthy-but-running behavior, and a real process crash. Before triggering each, write down which mechanism you expect to react: health state, restart policy, Compose dependency gating, graceful stop handling, or no automatic action.

1. Create the checkpoint workspace and capture assumptions

rm -rf dkr24-checkpoint dkr24-capstone
mkdir -p dkr24-checkpoint dkr24-capstone

docker context show | tee dkr24-checkpoint/context.txt
docker version | tee dkr24-checkpoint/docker-version.txt
docker info | tee dkr24-checkpoint/docker-info.txt
docker compose version | tee dkr24-checkpoint/compose-version.txt 2>&1 || true
docker buildx version | tee dkr24-checkpoint/buildx-version.txt 2>&1 || true
docker image inspect alpine:3.22 --format '{{json .RepoDigests}}'   | tee dkr24-checkpoint/image-digests.txt

2. Write predictions before execution

Record predictions in plain text before the run. At minimum include:

  • Startup delay: health should stay starting until the readiness marker appears; the dependent service should wait.
  • Unhealthy condition: PID 1 should stay running, health should become unhealthy, and restart count should not change merely because of health.
  • Crash: a non-zero process exit under bounded on-failure should increment restart count.
  • Graceful stop: TERM should be logged and the process should exit before the stop timeout.
cat > dkr24-checkpoint/predictions.txt <<'EOF'
startup-delay: health=starting then healthy; worker waits for healthy
unhealthy-marker-removed: process stays running; restart_count unchanged
synthetic-crash: on-failure performs bounded restart attempt(s)
graceful-stop: SIGTERM handled before timeout; no forced-kill evidence expected
EOF

3. Define the disposable Compose scenario

The api service waits four seconds before declaring itself ready. Its healthcheck uses only shell test. The worker starts only after api is healthy. The restart policy is on-failure:2, but health failure alone will not trigger it.

cat > dkr24-capstone/compose.yaml <<'YAML'
name: dkr24-capstone
services:
  api:
    image: alpine:3.22
    restart: on-failure:2
    stop_signal: SIGTERM
    stop_grace_period: 5s
    command:
      - sh
      - -c
      - |
        trap 'echo TERM-received; exit 0' TERM INT
        rm -f /tmp/ready
        echo startup-begin
        sleep 4
        touch /tmp/ready
        echo startup-ready
        while :; do sleep 1; done
    healthcheck:
      test: ["CMD-SHELL", "test -f /tmp/ready"]
      interval: 2s
      timeout: 1s
      retries: 2
      start_period: 5s
      start_interval: 1s
    labels:
      academy: docker-ch24
      scenario: checkpoint
  worker:
    image: alpine:3.22
    depends_on:
      api:
        condition: service_healthy
    command: ["sh", "-c", "echo worker-started-after-api-health; sleep 30"]
    labels:
      academy: docker-ch24
      scenario: checkpoint
YAML

(cd dkr24-capstone && docker compose config)   | tee dkr24-checkpoint/compose-normalized.yaml

4. Observe startup delay and dependency gating

START_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)
(cd dkr24-capstone && docker compose up -d)

sleep 1
(cd dkr24-capstone && docker compose ps)   | tee dkr24-checkpoint/ps-during-startup.txt
sleep 6
(cd dkr24-capstone && docker compose ps)   | tee dkr24-checkpoint/ps-after-health.txt
(cd dkr24-capstone && docker compose logs --timestamps)   | tee dkr24-checkpoint/startup-logs.txt

5. Create unhealthy-but-running state and preserve evidence

Remove the readiness marker without stopping PID 1. Compare restart count before and after health becomes unhealthy.

API_ID=$(cd dkr24-capstone && docker compose ps -q api)
BEFORE=$(docker inspect "$API_ID" --format '{{.RestartCount}}')
docker exec "$API_ID" rm -f /tmp/ready
sleep 6

docker inspect "$API_ID" \
  --format 'id={{.Id}} status={{.State.Status}} running={{.State.Running}} health={{.State.Health.Status}} streak={{.State.Health.FailingStreak}} restarts={{.RestartCount}}'   | tee dkr24-checkpoint/unhealthy-state.txt
docker inspect "$API_ID" --format '{{json .State.Health.Log}}'   | tee dkr24-checkpoint/unhealthy-health-log.json
AFTER=$(docker inspect "$API_ID" --format '{{.RestartCount}}')
printf 'before=%s after=%s
' "$BEFORE" "$AFTER"   | tee dkr24-checkpoint/unhealthy-restart-count.txt

6. Trigger a real process crash and observe bounded restart behavior

First restore health, then deliberately terminate PID 1 with a non-zero exit from inside the disposable container. The health transition and the process crash are intentionally separate events.

docker exec "$API_ID" touch /tmp/ready
sleep 3

# Synthetic crash: ask PID 1 shell to exit non-zero via TERM is not suitable,
# because the trap exits 0. Instead run a one-shot crash container under the same policy.
docker rm -f dkr24-crash 2>/dev/null || true
docker run -d \
  --name dkr24-crash \
  --label academy=docker-ch24 \
  --label scenario=checkpoint-crash \
  --restart on-failure:1   alpine:3.22 sh -c 'sleep 11; echo crash-42 >&2; exit 42'

sleep 25
docker inspect dkr24-crash \
  --format 'status={{.State.Status}} exit={{.State.ExitCode}} restarts={{.RestartCount}} policy={{.HostConfig.RestartPolicy.Name}}'   | tee dkr24-checkpoint/crash-state.txt
docker logs --timestamps dkr24-crash 2>&1   | tee dkr24-checkpoint/crash-logs.txt

7. Verify graceful shutdown separately

Graceful shutdown is not the same as health or restart behavior. Stop only the API service and preserve its TERM log and final state.

(cd dkr24-capstone && docker compose stop -t 5 api)
(cd dkr24-capstone && docker compose logs --timestamps api)   | tee dkr24-checkpoint/graceful-api-logs.txt
docker inspect "$API_ID" \
  --format 'exit={{.State.ExitCode}} status={{.State.Status}} finished={{.State.FinishedAt}}'   | tee dkr24-checkpoint/graceful-api-state.txt

8. Capture the Engine timeline and final inventories

END_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)
docker events \
  --since "$START_UTC" \
  --until "$END_UTC" \
  --filter type=container \
  --format '{{json .}}'   | grep -E 'dkr24-capstone|dkr24-crash|docker-ch24'   > dkr24-checkpoint/engine-events.jsonl || true

(cd dkr24-capstone && docker compose ps -a)   | tee dkr24-checkpoint/final-compose-ps.txt
docker inspect dkr24-crash > dkr24-checkpoint/crash-inspect.json

9. Interpret mechanism by mechanism

Event Expected mechanism What should not be inferred
Startup delay Health remains starting; Compose worker waits for service_healthy. Not a crash and not necessarily an error.
Readiness marker removed Health becomes unhealthy; API process remains running. No generic Engine restart should be inferred from unhealthy alone.
Crash container exits 42 on-failure:1 performs bounded restart behavior. Does not prove the application became healthy after restart.
Compose stop API SIGTERM + five-second grace path should allow trap to exit cleanly. A clean stop is not the same as a successful healthcheck.

10. Required evidence packet and limitations note

The packet should contain context/version evidence, normalized Compose model, image digest, startup logs, health history, restart counts, exit codes, Engine events, and graceful-stop logs. Add an assumptions.txt stating that the lab is local/disposable, uses Alpine 3.22, tests Docker-local behavior only, and does not claim external SLA or production orchestration behavior.

cat > dkr24-checkpoint/assumptions.txt <<'EOF'
Local disposable lab only.
No production endpoint or real secret.
Docker health is container-local evidence, not an SLA.
Compose service_healthy gates startup but does not replace runtime retry/reconnect logic.
Restart policy responds to process exit; unhealthy alone is not a generic Engine restart trigger.
Engine 29.8.1 is the documentation baseline; actual local versions are recorded in this packet.
EOF

11. Cleanup and rollback

Remove only the checkpoint resources after evidence is complete. Do not use volume/image/system prune.

(cd dkr24-capstone && docker compose down --remove-orphans)
docker rm -f dkr24-crash 2>/dev/null || true
rm -rf dkr24-capstone

tar -czf dkr24-checkpoint.tgz dkr24-checkpoint/

12. What Chapter 24 adds to the production Docker model

You can now separate process lifecycle, container health, dependency readiness, restart behavior, graceful shutdown, and external availability—and attach the correct evidence to each. That prevents restart loops from masquerading as resilience and healthchecks from masquerading as user availability. Chapter 25 now moves from reliability boundaries to security boundaries: namespaces, Linux capabilities, seccomp, AppArmor, SELinux, devices, and privilege containment.

Knowledge check

In the checkpoint, why should unhealthy-state restart count remain unchanged?

What mechanism delays the worker until the API is ready?

Why is the synthetic crash tested in a separate container?

What does a successful graceful stop prove?

What is the most important limitation statement for this lab?

Next lesson

Next: Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Concepts, Architecture, and Mental Model

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used for compatibility notes. Current Dockerfile HEALTHCHECK supports interval, timeout, start_period, start_interval (Engine 25+), and retries. Engine 29 includes a fix for health checks delayed too long when start_interval exceeds start_period. The labs record the actual Engine/CLI/Compose versions before relying on these semantics.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.