Checkpoint Lab — Production Container Patterns, Immutable Delivery, Configuration, Statelessness, Sidecars, and Operational Contracts
Define, run, replace, fail, recover, and audit a complete local production contract while proving digest identity, state continuity, health, bounded resources/logs, graceful shutdown, and exact cleanup.
Learning objectives
- Execute a complete local production contract from digest-pinned release through replacement and recovery.
- Predict and verify image, container, storage, network, security, resource, logging, health, and restart state changes.
- Prove durable state survives application replacement while the application container identity changes.
- Trigger one controlled process failure and distinguish Engine restart recovery from health state and external request success.
- Produce an evidence packet and exact rollback/cleanup record suitable for review.
1. Checkpoint scenario and success criteria
The checkpoint uses the Chapter 38 synthetic app. Success means you can show: one immutable v1 digest, one immutable v2 digest, a replacement to v2 without data loss, non-root/read-only runtime, bounded resources/logs, health plus external request evidence, graceful-stop evidence, one controlled process crash recovered by restart policy, companion-service behavior, a synthetic state backup/checksum, and exact cleanup. A human-readable release tag alone is not accepted as release identity.
2. Preflight and current assumptions
set -eu
LAB=dca38-checkpoint
mkdir -p "$LAB"/{src,config,secrets,evidence}
date -u +%Y-%m-%dT%H:%M:%SZ | tee "$LAB/evidence/time-start.txt"
docker context show | tee "$LAB/evidence/context.txt"
docker version | tee "$LAB/evidence/docker-version.txt"
docker compose version | tee "$LAB/evidence/compose-version.txt"
docker buildx version | tee "$LAB/evidence/buildx-version.txt"
docker info | tee "$LAB/evidence/docker-info-full.txt"
Expected course baseline at authoring time: Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. Do not fail the checkpoint merely because packaged component versions differ; record the actual environment and note compatibility assumptions.
3. Create source, config, fake secret, and image
# dca38-checkpoint/src/app.py
from http.server import BaseHTTPRequestHandler, HTTPServer
from pathlib import Path
import json, os, signal
DATA=Path('/data/state.json')
VERSION=os.getenv('APP_VERSION','v1')
state={'starts':0}
if DATA.exists(): state=json.loads(DATA.read_text())
state['starts']+=1; DATA.write_text(json.dumps(state)+'
')
message=Path('/etc/dca38/app.conf').read_text().strip()
assert Path('/run/secrets/app_token').read_text().startswith('FAKE-')
print(f'start version={VERSION} starts={state["starts"]}', flush=True)
class H(BaseHTTPRequestHandler):
def log_message(self, fmt, *args): print('request '+fmt%args, flush=True)
def do_GET(self):
if self.path=='/health': body=b'ok
'
elif self.path=='/state': body=(json.dumps({'version':VERSION,**state})+'
').encode()
else: body=f'{VERSION} {message}
'.encode()
self.send_response(200); self.end_headers(); self.wfile.write(body)
server=HTTPServer(('0.0.0.0',8080),H)
def stop(sig, frame): print(f'shutdown signal={sig}', flush=True); raise SystemExit(0)
signal.signal(signal.SIGTERM,stop)
server.serve_forever()
# dca38-checkpoint/src/observer.py
import time, urllib.request
while True:
try:
print('observer '+urllib.request.urlopen('http://app:8080/health',timeout=2).read().decode().strip(), flush=True)
except Exception as e:
print('observer unavailable='+type(e).__name__, flush=True)
time.sleep(5)
# dca38-checkpoint/Dockerfile
# syntax=docker/dockerfile:1
FROM python:3.13-alpine
RUN addgroup -g 10001 app && adduser -D -u 10001 -G app app && mkdir -p /app /data && chown -R app:app /app /data
WORKDIR /app
COPY --chown=app:app src/ /app/
ENV PYTHONDONTWRITEBYTECODE=1 PYTHONUNBUFFERED=1
USER app
CMD ["python","/app/app.py"]
printf 'message=checkpoint-production
' > dca38-checkpoint/config/app.conf
printf 'FAKE-DCA38-CHECKPOINT-TOKEN
' > dca38-checkpoint/secrets/app_token.txt
sha256sum dca38-checkpoint/src/* dca38-checkpoint/config/app.conf | tee dca38-checkpoint/evidence/source-config.sha256
4. Start disposable local registry and capture its identity
docker pull registry:3.1.1
docker image inspect registry:3.1.1 --format 'id={{.Id}} repoDigests={{json .RepoDigests}}' | tee dca38-checkpoint/evidence/registry-image.txt
docker run -d \
--name dca38-checkpoint-registry \
--label academy.lab=dca38-checkpoint -p 127.0.0.1:50039:5000 registry:3.1.1 | tee dca38-checkpoint/evidence/registry-container-id.txt
5. Build/push v1 and record the immutable subject
docker build \
--label academy.lab=dca38-checkpoint \
--label academy.release=v1 -t localhost:50039/dca38/checkpoint:v1 dca38-checkpoint | tee dca38-checkpoint/evidence/build-v1.log
docker push localhost:50039/dca38/checkpoint:v1 | tee dca38-checkpoint/evidence/push-v1.log
V1=$(docker image inspect localhost:50039/dca38/checkpoint:v1 --format '{{index .RepoDigests 0}}')
printf '%s
' "$V1" | tee dca38-checkpoint/evidence/release-v1.txt
6. Production contract
# dca38-checkpoint/compose.yaml
services:
app:
image: ${APP_IMAGE:?digest reference required}
environment:
APP_VERSION: ${APP_VERSION:-v1}
user: "10001:10001"
read_only: true
cap_drop: ["ALL"]
security_opt: ["no-new-privileges:true"]
tmpfs: ["/tmp:rw,noexec,nosuid,size=16m"]
configs:
- source: app_config
target: /etc/dca38/app.conf
secrets: [app_token]
volumes: ["dca38-checkpoint-data:/data"]
networks: [front]
ports: ["127.0.0.1:18039:8080"]
healthcheck:
test: ["CMD","python","-c","import urllib.request; urllib.request.urlopen('http://127.0.0.1:8080/health',timeout=1).read()"]
interval: 4s
timeout: 2s
retries: 5
start_period: 4s
restart: unless-stopped
stop_grace_period: 10s
cpus: 0.50
mem_limit: 128m
pids_limit: 64
logging:
driver: local
options: {max-size: "10m", max-file: "3"}
labels: {academy.lab: dca38-checkpoint}
observer:
image: ${OBSERVER_IMAGE:?digest reference required}
command: ["python","/app/observer.py"]
user: "10001:10001"
read_only: true
cap_drop: ["ALL"]
security_opt: ["no-new-privileges:true"]
tmpfs: ["/tmp:rw,noexec,nosuid,size=8m"]
networks: [front]
depends_on:
app: {condition: service_healthy}
restart: unless-stopped
cpus: 0.20
mem_limit: 64m
pids_limit: 32
logging:
driver: local
options: {max-size: "5m", max-file: "2"}
labels: {academy.lab: dca38-checkpoint}
configs:
app_config: {file: ./config/app.conf}
secrets:
app_token: {file: ./secrets/app_token.txt}
volumes:
dca38-checkpoint-data:
name: dca38-checkpoint-data
labels: {academy.lab: dca38-checkpoint}
networks:
front:
name: dca38-checkpoint-front
labels: {academy.lab: dca38-checkpoint}
7. Predict before execution
Prediction A: two containers will be created from the same v1 digest, but only app receives the secret, data volume, and host port.
Prediction B: app rootfs is read-only; /data persists and /tmp is ephemeral.
Prediction C: app external health succeeds only on 127.0.0.1:18039, while observer reaches app by service DNS on dca38-checkpoint-front.
Prediction D: replacing app with v2 changes app container ID and configured digest while the volume name remains dca38-checkpoint-data.
Prediction E: SIGKILL of the app process causes restartCount to rise on the same container because restart policy handles process exit.
8. Deploy v1 and capture contract evidence
printf 'APP_IMAGE=%s
OBSERVER_IMAGE=%s
APP_VERSION=v1
' "$V1" "$V1" > dca38-checkpoint/.env
docker compose \
--env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml config | tee dca38-checkpoint/evidence/compose-v1.rendered.yaml
docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml up -d
CID1=$(docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml ps -q app)
printf '%s
' "$CID1" | tee dca38-checkpoint/evidence/app-cid-v1.txt
docker inspect "$CID1" \
--format 'configured={{.Config.Image}} imageId={{.Image}} user={{.Config.User}} readonly={{.HostConfig.ReadonlyRootfs}} restart={{.HostConfig.RestartPolicy.Name}} memory={{.HostConfig.Memory}} nanoCpus={{.HostConfig.NanoCpus}} pids={{.HostConfig.PidsLimit}} mounts={{json .Mounts}} log={{json .HostConfig.LogConfig}}' | tee dca38-checkpoint/evidence/app-contract-v1.txt
docker volume inspect dca38-checkpoint-data | tee dca38-checkpoint/evidence/volume.txt
docker network inspect dca38-checkpoint-front | tee dca38-checkpoint/evidence/network.txt
9. Verify health, traffic, state, logs, resources
docker inspect "$CID1" --format 'health={{.State.Health.Status}} restartCount={{.RestartCount}}' | tee dca38-checkpoint/evidence/health-v1.txt
curl -fsS http://127.0.0.1:18039/health | tee dca38-checkpoint/evidence/external-health-v1.txt
curl -fsS http://127.0.0.1:18039/state | tee dca38-checkpoint/evidence/state-v1.txt
docker logs --timestamps --tail 30 "$CID1" | tee dca38-checkpoint/evidence/app-logs-v1.txt
OCID=$(docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml ps -q observer)
docker logs --timestamps --tail 20 "$OCID" | tee dca38-checkpoint/evidence/observer-logs-v1.txt
docker stats --no-stream "$CID1" "$OCID" | tee dca38-checkpoint/evidence/stats-v1.txt
10. Verify graceful shutdown without confusing it with crash recovery
docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml stop app
docker logs --timestamps --tail 20 "$CID1" | tee dca38-checkpoint/evidence/graceful-stop-v1.txt
docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml start app
curl -fsS http://127.0.0.1:18039/state | tee dca38-checkpoint/evidence/state-after-graceful-start.txt
A manual Compose stop suppresses ordinary restart-policy behavior until restarted; this test is for signal/grace-period handling, not self-healing.
11. Build/push v2 and replace only app
cp dca38-checkpoint/src/app.py dca38-checkpoint/src/app.py.v1
# The image bytes change through a harmless release label file.
printf 'release=v2
' > dca38-checkpoint/src/release.txt
docker build --label academy.lab=dca38-checkpoint --label academy.release=v2 -t localhost:50039/dca38/checkpoint:v2 dca38-checkpoint | tee dca38-checkpoint/evidence/build-v2.log
docker push localhost:50039/dca38/checkpoint:v2 | tee dca38-checkpoint/evidence/push-v2.log
V2=$(docker image inspect localhost:50039/dca38/checkpoint:v2 --format '{{index .RepoDigests 0}}')
printf '%s
' "$V2" | tee dca38-checkpoint/evidence/release-v2.txt
printf 'APP_IMAGE=%s
OBSERVER_IMAGE=%s
APP_VERSION=v2
' "$V2" "$V1" > dca38-checkpoint/.env
docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml up -d --no-deps app
CID2=$(docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml ps -q app)
printf '%s
' "$CID2" | tee dca38-checkpoint/evidence/app-cid-v2.txt
printf 'old=%s
new=%s
' "$CID1" "$CID2" | tee dca38-checkpoint/evidence/replacement-ids.txt
curl -fsS http://127.0.0.1:18039/state | tee dca38-checkpoint/evidence/state-v2.txt
The observer remains pinned to v1 intentionally; its independent service identity demonstrates that replacing one service need not silently change every companion. A real compatibility contract decides whether the observer must move with the app.
12. Capture backup/checksum and prove volume continuity
docker cp "$CID2:/data/state.json" dca38-checkpoint/evidence/state-backup.json
sha256sum dca38-checkpoint/evidence/state-backup.json | tee dca38-checkpoint/evidence/state-backup.sha256
docker volume inspect dca38-checkpoint-data \
--format 'name={{.Name}} driver={{.Driver}} labels={{json .Labels}}' | tee dca38-checkpoint/evidence/volume-after-replacement.txt
13. Controlled process failure and restart recovery
BEFORE=$(docker inspect "$CID2" --format '{{.RestartCount}}')
printf '%s
' "$BEFORE" | tee dca38-checkpoint/evidence/restart-count-before.txt
docker kill --signal=KILL "$CID2" | tee dca38-checkpoint/evidence/controlled-crash.txt
sleep 3
AFTER=$(docker inspect "$CID2" --format '{{.RestartCount}}')
printf '%s
' "$AFTER" | tee dca38-checkpoint/evidence/restart-count-after.txt
docker inspect "$CID2" \
--format 'status={{.State.Status}} health={{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}} exit={{.State.ExitCode}} oom={{.State.OOMKilled}} restartCount={{.RestartCount}}' | tee dca38-checkpoint/evidence/post-crash-inspect.txt
curl -fsS http://127.0.0.1:18039/state | tee dca38-checkpoint/evidence/state-after-crash-recovery.txt
docker logs --timestamps --tail 40 "$CID2" | tee dca38-checkpoint/evidence/app-logs-after-crash.txt
The same container ID is expected because Engine restart policy restarts the container’s process. A replacement deployment, by contrast, creates a new container ID. This distinction is part of the checkpoint.
14. Verify companion and contract after recovery
OCID=$(docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml ps -q observer)
docker logs --timestamps --tail 40 "$OCID" | tee dca38-checkpoint/evidence/observer-after-recovery.txt
docker inspect "$CID2" \
--format 'configured={{.Config.Image}} readonly={{.HostConfig.ReadonlyRootfs}} security={{json .HostConfig.SecurityOpt}} capDrop={{json .HostConfig.CapDrop}}' | tee dca38-checkpoint/evidence/security-final.txt
docker stats --no-stream "$CID2" "$OCID" | tee dca38-checkpoint/evidence/stats-final.txt
15. Verification checklist
| Contract assertion | Required proof |
|---|---|
| Immutable v1/v2 releases | two distinct registry digest references |
| Replacement not mutation | CID1 != CID2; configured image ref changes to v2 digest |
| State survives replacement | same named volume + increasing/consistent state + backup checksum |
| Non-root/read-only | runtime user + ReadonlyRootfs + narrow mounts/tmpfs |
| Health and reachability | Docker health state + host loopback request |
| Resource envelope | inspect memory/CPU/PID settings + stats sample |
| Bounded logging | per-service local driver options |
| Graceful shutdown | SIGTERM/shutdown log from controlled stop |
| Process recovery | restart count increases after controlled SIGKILL |
| Companion least privilege | no host port, secret, or state volume granted |
16. Limitations and recovery note
Checkpoint limitations:
- Single Docker host only: no cross-host rescheduling or HA claim.
- Local loopback registry: no production TLS/auth/authorization/retention policy tested.
- File-backed fake secret: no external secret-manager encryption/rotation/audit tested.
- Synthetic JSON state: not a database-consistent backup exercise.
- Health endpoint is local readiness evidence, not an SLA or full dependency proof.
- CPU/memory/PID limits are illustrative; production values require workload measurement.
17. Exact cleanup and rollback
docker compose --env-file dca38-checkpoint/.env -p dca38-checkpoint -f dca38-checkpoint/compose.yaml down
docker volume rm dca38-checkpoint-data 2>/dev/null || true
docker network rm dca38-checkpoint-front 2>/dev/null || true
docker image rm localhost:50039/dca38/checkpoint:v1 localhost:50039/dca38/checkpoint:v2 2>/dev/null || true
docker rm -f dca38-checkpoint-registry 2>/dev/null || true
printf 'Retain dca38-checkpoint/evidence for review.
'
Rollback during the live checkpoint means restore the prior digest/config reference and recreate only the app after confirming data compatibility. Cleanup is separate and deletes the intentionally disposable state volume only after evidence review.
18. Required evidence packet
| Family | Required artifact |
|---|---|
| Versions/context | Docker, Compose, Buildx, host/context info |
| Source/config | source/config checksums; fake-secret source identity without value in logs |
| Release | v1 and v2 pushed digest references; build/push logs |
| Runtime | old/new app container IDs; configured image refs; runtime user/security |
| Storage | volume identity before/after; state endpoints; backup/checksum |
| Network | loopback publication and internal network membership |
| Health | health state/history plus external request |
| Resources/logging | HostConfig limits/log driver + stats samples |
| Shutdown/recovery | graceful-stop log; restart counts and crash recovery evidence |
| Companion | observer ID/logs and narrow grants |
| Limitations | explicit single-host/registry/secret/backup/health caveats |
| Cleanup | exact project/volume/network/images/registry removed |
19. What Chapter 38 adds to the production operating model
The course now has a complete single-host production container contract: immutable release identity, external runtime configuration, narrow secrets, durable state outside the writable layer, least-privilege runtime, health plus external reachability, bounded resources/logs, graceful shutdown, process restart behavior, companion-service boundaries, replacement, backup evidence, rollback identity, and exact cleanup.
Knowledge check
What proves a release replacement occurred rather than a process restart?
The app container ID changes and the configured image digest changes; a restart policy restart keeps the same container ID and increments RestartCount.
What proves the state boundary worked?
The named volume identity remains constant across replacement and the state/backup evidence remains available while the app container ID changes.
Why capture both Docker health and a loopback HTTP request?
They prove different paths: the internal healthcheck and the host-published application path.
Does restart-policy recovery prove high availability?
No. It proves same-host process recovery while the Docker Engine and host remain available.
What should a rollback name?
The exact prior image digest, compatible config/secret state, data/schema assumptions, and the smallest service replacement needed.
Official references and version notes
Checkpoint baseline: 2026-09-22. Engine 29.8.1, Compose 5.5.1, Buildx 0.37.1 and BuildKit 0.33.0 are the current upstream baselines referenced. Distribution Registry 3.1.1 is used only as a disposable loopback registry. Learners must record their actual installed component versions and observed image digests.
- Docker Docs — Use Compose in production — single-host production use, production-specific overrides, and service recreation.
- Docker Docs — Why use Compose? — current single-host deployment boundary and application-model use cases.
- Docker Docs — Compose services reference — healthcheck, restart, read-only rootfs, resource limits, logging, configs, secrets, stop signal and grace period.
- Docker Docs — Compose Deploy Specification — resource limits/reservations and deployment-oriented service controls.
- Docker Docs — Start containers automatically — restart-policy semantics and the successful-start threshold.
- Docker Docs — Resource constraints — CPU/memory governance and OOM implications.
- Docker Docs — Volumes — persistent data lifecycle independent of containers.
- Docker Docs — Storage — writable-layer versus volume/tmpfs durability boundaries.
- Docker Docs — Manage secrets securely in Compose — file-mounted secret access and environment-variable risk.
- Docker Docs — Compose secrets reference — top-level secret sources and service grants.
- Docker Docs — Configure logging drivers — default json-file behavior and recommendation for bounded local logging.
- Docker Docs — Local logging driver — automatic rotation, max-size/max-file, and daemon-owned log files.
- Docker Engine 29 release notes — current Engine-era baseline; 29.8.1 released 2026-09-15.
- Docker Compose releases — current Compose 5.5.1 baseline used for version-sensitive examples.
- Docker Buildx releases — current Buildx 0.37.1 baseline.
- BuildKit releases — current BuildKit 0.33.0 baseline.
- Docker Official Image — registry — current Distribution Registry 3.1.1 local-registry image used in the optional digest-pinning lab path.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.