Production Capstone: Build, Secure, Publish, Operate, Observe, and Recover a Complete Dockerized Application: Failure Injection, Troubleshooting, and Recovery Drill
Inject controlled build, process/health, network/DNS, storage/permission, and recovery failures while preserving first-failure evidence and repairing only the causal layer.
Learning objectives
- Inject at least one build/input, process/health, network/DNS, storage/permission, and recovery failure without damaging unrelated Docker state.
- Preserve the exact failed IDs, logs, events, build traces, network/mount evidence, and timestamps before repair.
- Apply the Chapter 41 diagnostic order and repair only the causal layer.
- Measure recovery timing and validate recovered state by immutable image digest plus application/data evidence.
- Update the runbook after each drill with detection signal, cause, repair, verification, rollback, and prevention.
1. Incident drill rules
mkdir -p evidence/incidents
printf 'drill_start=%s
' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" > evidence/incidents/drill.txt
printf 'context=%s
subject=%s
' "$(docker context show)" "$SUBJECT" >> evidence/incidents/drill.txt
APP_REF="$SUBJECT" docker compose -p ch42cap ps > evidence/incidents/pre-ps.txt
docker inspect "$(APP_REF="$SUBJECT" docker compose -p ch42cap ps -q app)" > evidence/incidents/pre-app-inspect.json
2. Failure A — build/input error
Create a separate broken Dockerfile that refers to a missing input. This leaves the committed production Dockerfile untouched.
cat > Dockerfile.broken <<'EOF'
# syntax=docker/dockerfile:1.27
ARG BASE_IMAGE
FROM ${BASE_IMAGE}
COPY missing-capstone-input.txt /app/input.txt
USER 10001:10001
EOF
docker buildx build --builder ch42-builder --platform linux/amd64 --build-arg BASE_IMAGE="$BASE_REF" --progress=plain -f Dockerfile.broken . > evidence/incidents/build-fail.log 2>&1 || true
grep -E 'missing-capstone-input|not found|failed' evidence/incidents/build-fail.log | head -20 || true
Localization: source/build-context ownership. The registry, runtime and Compose deployment are unrelated. Repair by correcting the declared input; do not rebuild the release subject as part of incident recovery. Remove only the disposable broken Dockerfile after evidence capture.
rm Dockerfile.broken
printf 'build_incident=preserved_no_release_rebuild
' > evidence/incidents/build-fix.txt
3. Failure B — controlled unhealthy application
The app health endpoint deliberately returns 503 when
/data/unhealthy exists. Create that marker inside the
declared state volume, then observe health transitions without
killing the container first.
APP_ID="$(APP_REF="$SUBJECT" docker compose -p ch42cap ps -q app)"
docker exec "$APP_ID" sh -c 'touch /data/unhealthy'
sleep 12
docker inspect "$APP_ID" --format '{{json .State.Health}}' > evidence/incidents/unhealthy-health.json
APP_REF="$SUBJECT" docker compose -p ch42cap logs --no-color --timestamps --tail=100 app > evidence/incidents/unhealthy-app.log
END="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
docker events --since 5m --until "$END" --filter container="$APP_ID" > evidence/incidents/unhealthy-events.txt || true
curl -sS -i http://127.0.0.1:18080/health > evidence/incidents/unhealthy-http.txt || true
Interpretation: the process can remain running while health is unhealthy. Standard Engine restart policy does not generically restart a merely unhealthy container. Repair the application/data condition that the probe reports.
docker exec "$APP_ID" sh -c 'rm -f /data/unhealthy'
for i in 1 2 3 4 5 6; do curl -fsS http://127.0.0.1:18080/health && break || sleep 3; done
docker inspect "$APP_ID" --format '{{json .State.Health}}' > evidence/incidents/healthy-after-fix.json
4. Failure C — operations-network disconnect
Disconnect the app from the internal operations network only. The
loopback client path can remain up through frontend,
while the observer loses DNS/reachability to app. This
demonstrates that “one health path works” does not mean every path
works.
APP_ID="$(APP_REF="$SUBJECT" docker compose -p ch42cap ps -q app)"
docker network disconnect ch42cap_ops "$APP_ID"
sleep 12
docker network inspect ch42cap_ops > evidence/incidents/ops-network-fail.json
APP_REF="$SUBJECT" docker compose -p ch42cap logs --no-color --timestamps --tail=100 observer > evidence/incidents/observer-network-fail.log
curl -fsS http://127.0.0.1:18080/health > evidence/incidents/frontend-still-healthy.json
docker network connect --alias app ch42cap_ops "$APP_ID"
sleep 5
APP_REF="$SUBJECT" docker compose -p ch42cap logs --no-color --timestamps --tail=50 observer > evidence/incidents/observer-network-fixed.log
The repair reconnects one endpoint and restores its service alias. It does not touch host DNS, firewall rules, or the unrelated frontend network.
5. Failure D — storage write denied by mount contract
Do not corrupt the live volume. Mount the same volume read-only into a one-off container running as the app UID and attempt a write.
docker run \
--rm \
--name ch42-storage-probe \
--user 10001:10001 \
--mount type=volume,source=ch42cap_state,target=/data,readonly "$SUBJECT" python -c "open('/data/should-fail','w').write('x')" > evidence/incidents/storage-fail.log 2>&1 || true
docker create --name ch42-storage-inspect --user 10001:10001 --mount type=volume,source=ch42cap_state,target=/data,readonly "$SUBJECT" true >/dev/null
docker inspect ch42-storage-inspect --format '{{json .Mounts}}' > evidence/incidents/storage-mount.json
docker rm ch42-storage-inspect >/dev/null
# Minimal corrected contract in a separate one-off probe.
docker run \
--rm \
--name ch42-storage-probe-fixed \
--user 10001:10001 \
--mount type=volume,source=ch42cap_state,target=/data "$SUBJECT" python -c "open('/data/write-probe','w').write('ok')"
docker run --rm --mount type=volume,source=ch42cap_state,target=/data,readonly "$ALPINE_REF" cat /data/write-probe > evidence/incidents/storage-fixed.txt
The evidence differentiates mount authority from UID/GID ownership. If the mount were writable but UID 10001 still failed, then numeric ownership/LSM/user-namespace evidence would be the next layer.
6. Take an application-consistent backup
For this SQLite teaching workload, stop the app and observer so no process holds the database open, archive the named volume, hash the archive, then resume the original deployment. This creates a clear backup point for the recovery drill.
mkdir -p evidence/backups
BACKUP_START="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
APP_REF="$SUBJECT" docker compose -p ch42cap stop observer app
docker run --rm --mount type=volume,source=ch42cap_state,target=/data,readonly --mount type=bind,source="$PWD/evidence/backups",target=/backup "$ALPINE_REF" tar -czf /backup/state.tgz -C /data .
sha256sum evidence/backups/state.tgz | tee evidence/backups/state.tgz.sha256
printf 'backup_started=%s
subject=%s
' "$BACKUP_START" "$SUBJECT" > evidence/backups/metadata.txt
APP_REF="$SUBJECT" docker compose -p ch42cap start app observer
curl -fsS http://127.0.0.1:18080/ | tee evidence/backups/original-after-backup.json
7. Failure E — destroy only a disposable restore target, then recover
The original state volume remains intact. Recovery targets a brand-new volume so the exercise proves the archive, not luck from leftover data.
docker volume rm ch42_state_restore 2>/dev/null || true
docker volume create ch42_state_restore >/dev/null
RESTORE_START_EPOCH="$(date +%s)"
docker run --rm --mount type=volume,source=ch42_state_restore,target=/restore --mount type=bind,source="$PWD/evidence/backups",target=/backup,readonly "$ALPINE_REF" tar -xzf /backup/state.tgz -C /restore
cat > compose.recovery.yaml <<'EOF'
services:
app:
ports: !override
- "127.0.0.1:18081:8080"
observer:
profiles: [disabled-in-recovery]
EOF
STATE_VOLUME=ch42_state_restore APP_REF="$SUBJECT" docker compose -p ch42restore -f compose.yaml -f compose.recovery.yaml up -d app
for i in 1 2 3 4 5 6 7 8; do curl -fsS http://127.0.0.1:18081/ && break || sleep 2; done
RESTORE_END_EPOCH="$(date +%s)"
printf 'rto_seconds=%s
' "$((RESTORE_END_EPOCH-RESTORE_START_EPOCH))" | tee evidence/incidents/recovery-time.txt
curl -fsS http://127.0.0.1:18081/ | tee evidence/incidents/recovered-state.json
REC_ID="$(STATE_VOLUME=ch42_state_restore APP_REF="$SUBJECT" docker compose -p ch42restore -f compose.yaml -f compose.recovery.yaml ps -q app)"
docker inspect "$REC_ID" --format 'Image={{.Image}} Status={{.State.Status}} Health={{.State.Health.Status}}' | tee evidence/incidents/recovered-runtime.txt
Recovery succeeds only if the restored application returns the expected counter and the running container uses the intended release digest. A green container alone is not restore proof. The measured RTO belongs only to this host and dataset.
8. Verify RPO explicitly
Compare the counter in original-after-backup.json with
recovered-state.json. Any increments performed after
the backup point are intentionally outside the archive and therefore
outside the RPO. Write the observation into the incident record.
9. Runbook update template
Incident:
Detection signal:
UTC first-failure timestamp:
Context / host / Engine version:
Release subject digest:
Owning layer:
Original object IDs:
First-failure evidence files:
Root cause:
Minimal repair:
Verification repeated:
Rollback path:
Data/RPO impact:
Measured recovery time:
Residual risk:
Prevention/control update:
Owner:
Follow-up due date:
10. What the drills teach
| Failure | Owning layer | Repair that should not happen |
|---|---|---|
| Missing Dockerfile input | Source/build context | Restart Docker or prune cache first |
| 503 health marker | Application/data health contract | Assume restart policy will fix unhealthy state |
| Ops network disconnect | Container network endpoint/DNS alias | Change host DNS/firewall simultaneously |
| Read-only volume probe | Mount authority | chmod 777 or privileged container |
| Restore exercise | Backup/recovery lifecycle | Declare success because a blank container starts |
11. Preserve the recovery environment for final review
Leave ch42restore running until Lesson 5 has compared
original and recovered evidence. Do not remove the original volume,
registry storage, builder, signature public key, or backup archive
yet.
Knowledge check
Why is the storage failure injected with a separate read-only probe instead of changing the live app mount?
It demonstrates the same mount-authority failure without risking the working stateful service or contaminating recovery evidence.
Why can the observer fail while the loopback health endpoint still succeeds?
They use different network paths; the app can remain on frontend while its ops endpoint is disconnected.
What makes the backup application-consistent for this SQLite lab?
The app is stopped before the volume is archived, so no process is mutating/holding the database during the tar operation.
What two checks are required before calling restore successful?
The runtime must use the intended immutable image digest and the application must return the expected persisted data/state.
Why is measured RTO not a universal performance claim?
It depends on this host, dataset size, storage, registry/cache state, and lab conditions; it is evidence for this run only.
Official references and version notes
2026-09-22. The capstone records the learner’s actual versions before execution. Reference baselines used for compatibility discussion are Docker Engine/CLI 29.8.1, BuildKit 0.33.0, Buildx 0.37.1, Docker Compose 5.5.1, Trivy 0.74.0, and Cosign 3.1.3. Packaged Docker Desktop/Engine installations may expose different bundled containerd/runc versions, so the evidence packet records what the active daemon actually reports.
- Docker Engine 29 release notes — current Engine 29.8.1 behavior, fixes, known issues, and component changes.
- Docker Docs — Multi-platform builds — image indexes, native/emulated/cross-compilation strategies, and platform verification.
- Docker Docs — Build attestations — BuildKit SBOM and provenance metadata and image-store/registry requirements.
- Docker Docs — SBOM attestations — SPDX SBOM creation and inspection with Buildx imagetools.
- Docker Docs — Provenance attestations — SLSA provenance modes and inspection.
- docker buildx imagetools inspect — digest, platform, SBOM, and provenance inspection.
- Docker Docs — Use Compose in production — single-host production guidance and production-specific overrides.
- Docker Docs — Use secrets in Compose — runtime secret files and scope; local Compose secrets are not a substitute for an external encrypted secret manager.
- Docker Docs — Volumes — persistent data lifecycle, backup patterns, and container replacement semantics.
- Docker Docs — Resource constraints — CPU, memory, and runtime governance.
- Docker Docs — Configure logging drivers — bounded log retention and driver tradeoffs.
- Docker Docs — Seccomp security profiles — default syscall filtering and safer customization guidance.
- Docker Docs — Rootless mode — threat reduction and operational limits.
- dockerd reference — local registry trust behavior; 127.0.0.0/8 local registries are treated as insecure for testing, but production registries should use trusted TLS.
- Trivy releases — free local vulnerability scanner version identity used in the capstone.
- Cosign releases — signature tooling version identity used in the capstone.
- Open Container Initiative — image/runtime/distribution specifications underlying Docker artifact and runtime interoperability.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.