Chapter 42Lesson 04~300 minutes

Production Capstone: Build, Secure, Publish, Operate, Observe, and Recover a Complete Dockerized Application: Failure Injection, Troubleshooting, and Recovery Drill

Inject controlled build, process/health, network/DNS, storage/permission, and recovery failures while preserving first-failure evidence and repairing only the causal layer.

Failure injectionTroubleshootingRecoveryRTO/RPORunbook

Learning objectives

  • Inject at least one build/input, process/health, network/DNS, storage/permission, and recovery failure without damaging unrelated Docker state.
  • Preserve the exact failed IDs, logs, events, build traces, network/mount evidence, and timestamps before repair.
  • Apply the Chapter 41 diagnostic order and repair only the causal layer.
  • Measure recovery timing and validate recovered state by immutable image digest plus application/data evidence.
  • Update the runbook after each drill with detection signal, cause, repair, verification, rollback, and prevention.

1. Incident drill rules

Safety boundary. Do not restart Docker, change daemon security, mount the Docker socket, use --privileged, delete the failed object before evidence capture, or run any Docker-wide prune. All failures are scoped to ch42 resources.
mkdir -p evidence/incidents
printf 'drill_start=%s
' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" > evidence/incidents/drill.txt
printf 'context=%s
subject=%s
' "$(docker context show)" "$SUBJECT" >> evidence/incidents/drill.txt
APP_REF="$SUBJECT" docker compose -p ch42cap ps > evidence/incidents/pre-ps.txt
docker inspect "$(APP_REF="$SUBJECT" docker compose -p ch42cap ps -q app)" > evidence/incidents/pre-app-inspect.json

2. Failure A — build/input error

Create a separate broken Dockerfile that refers to a missing input. This leaves the committed production Dockerfile untouched.

cat > Dockerfile.broken <<'EOF'
# syntax=docker/dockerfile:1.27
ARG BASE_IMAGE
FROM ${BASE_IMAGE}
COPY missing-capstone-input.txt /app/input.txt
USER 10001:10001
EOF

docker buildx build --builder ch42-builder --platform linux/amd64   --build-arg BASE_IMAGE="$BASE_REF" --progress=plain   -f Dockerfile.broken . > evidence/incidents/build-fail.log 2>&1 || true

grep -E 'missing-capstone-input|not found|failed' evidence/incidents/build-fail.log | head -20 || true

Localization: source/build-context ownership. The registry, runtime and Compose deployment are unrelated. Repair by correcting the declared input; do not rebuild the release subject as part of incident recovery. Remove only the disposable broken Dockerfile after evidence capture.

rm Dockerfile.broken
printf 'build_incident=preserved_no_release_rebuild
' > evidence/incidents/build-fix.txt

3. Failure B — controlled unhealthy application

The app health endpoint deliberately returns 503 when /data/unhealthy exists. Create that marker inside the declared state volume, then observe health transitions without killing the container first.

APP_ID="$(APP_REF="$SUBJECT" docker compose -p ch42cap ps -q app)"
docker exec "$APP_ID" sh -c 'touch /data/unhealthy'
sleep 12
docker inspect "$APP_ID" --format '{{json .State.Health}}' > evidence/incidents/unhealthy-health.json
APP_REF="$SUBJECT" docker compose -p ch42cap logs --no-color --timestamps --tail=100 app > evidence/incidents/unhealthy-app.log
END="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
docker events --since 5m --until "$END" --filter container="$APP_ID" > evidence/incidents/unhealthy-events.txt || true
curl -sS -i http://127.0.0.1:18080/health > evidence/incidents/unhealthy-http.txt || true

Interpretation: the process can remain running while health is unhealthy. Standard Engine restart policy does not generically restart a merely unhealthy container. Repair the application/data condition that the probe reports.

docker exec "$APP_ID" sh -c 'rm -f /data/unhealthy'
for i in 1 2 3 4 5 6; do curl -fsS http://127.0.0.1:18080/health && break || sleep 3; done
docker inspect "$APP_ID" --format '{{json .State.Health}}' > evidence/incidents/healthy-after-fix.json

4. Failure C — operations-network disconnect

Disconnect the app from the internal operations network only. The loopback client path can remain up through frontend, while the observer loses DNS/reachability to app. This demonstrates that “one health path works” does not mean every path works.

APP_ID="$(APP_REF="$SUBJECT" docker compose -p ch42cap ps -q app)"
docker network disconnect ch42cap_ops "$APP_ID"
sleep 12
docker network inspect ch42cap_ops > evidence/incidents/ops-network-fail.json
APP_REF="$SUBJECT" docker compose -p ch42cap logs --no-color --timestamps --tail=100 observer > evidence/incidents/observer-network-fail.log
curl -fsS http://127.0.0.1:18080/health > evidence/incidents/frontend-still-healthy.json

docker network connect --alias app ch42cap_ops "$APP_ID"
sleep 5
APP_REF="$SUBJECT" docker compose -p ch42cap logs --no-color --timestamps --tail=50 observer > evidence/incidents/observer-network-fixed.log

The repair reconnects one endpoint and restores its service alias. It does not touch host DNS, firewall rules, or the unrelated frontend network.

5. Failure D — storage write denied by mount contract

Do not corrupt the live volume. Mount the same volume read-only into a one-off container running as the app UID and attempt a write.

docker run \
  --rm \
  --name ch42-storage-probe \
  --user 10001:10001 \
  --mount type=volume,source=ch42cap_state,target=/data,readonly   "$SUBJECT" python -c "open('/data/should-fail','w').write('x')"   > evidence/incidents/storage-fail.log 2>&1 || true

docker create --name ch42-storage-inspect   --user 10001:10001   --mount type=volume,source=ch42cap_state,target=/data,readonly   "$SUBJECT" true >/dev/null
docker inspect ch42-storage-inspect --format '{{json .Mounts}}' > evidence/incidents/storage-mount.json
docker rm ch42-storage-inspect >/dev/null

# Minimal corrected contract in a separate one-off probe.
docker run \
  --rm \
  --name ch42-storage-probe-fixed \
  --user 10001:10001 \
  --mount type=volume,source=ch42cap_state,target=/data   "$SUBJECT" python -c "open('/data/write-probe','w').write('ok')"

docker run --rm --mount type=volume,source=ch42cap_state,target=/data,readonly   "$ALPINE_REF" cat /data/write-probe > evidence/incidents/storage-fixed.txt

The evidence differentiates mount authority from UID/GID ownership. If the mount were writable but UID 10001 still failed, then numeric ownership/LSM/user-namespace evidence would be the next layer.

6. Take an application-consistent backup

For this SQLite teaching workload, stop the app and observer so no process holds the database open, archive the named volume, hash the archive, then resume the original deployment. This creates a clear backup point for the recovery drill.

mkdir -p evidence/backups
BACKUP_START="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
APP_REF="$SUBJECT" docker compose -p ch42cap stop observer app

docker run --rm   --mount type=volume,source=ch42cap_state,target=/data,readonly   --mount type=bind,source="$PWD/evidence/backups",target=/backup   "$ALPINE_REF" tar -czf /backup/state.tgz -C /data .
sha256sum evidence/backups/state.tgz | tee evidence/backups/state.tgz.sha256
printf 'backup_started=%s
subject=%s
' "$BACKUP_START" "$SUBJECT" > evidence/backups/metadata.txt
APP_REF="$SUBJECT" docker compose -p ch42cap start app observer
curl -fsS http://127.0.0.1:18080/ | tee evidence/backups/original-after-backup.json

7. Failure E — destroy only a disposable restore target, then recover

The original state volume remains intact. Recovery targets a brand-new volume so the exercise proves the archive, not luck from leftover data.

docker volume rm ch42_state_restore 2>/dev/null || true
docker volume create ch42_state_restore >/dev/null
RESTORE_START_EPOCH="$(date +%s)"

docker run --rm   --mount type=volume,source=ch42_state_restore,target=/restore   --mount type=bind,source="$PWD/evidence/backups",target=/backup,readonly   "$ALPINE_REF" tar -xzf /backup/state.tgz -C /restore

cat > compose.recovery.yaml <<'EOF'
services:
  app:
    ports: !override
      - "127.0.0.1:18081:8080"
  observer:
    profiles: [disabled-in-recovery]
EOF

STATE_VOLUME=ch42_state_restore APP_REF="$SUBJECT"   docker compose -p ch42restore -f compose.yaml -f compose.recovery.yaml up -d app
for i in 1 2 3 4 5 6 7 8; do curl -fsS http://127.0.0.1:18081/ && break || sleep 2; done
RESTORE_END_EPOCH="$(date +%s)"
printf 'rto_seconds=%s
' "$((RESTORE_END_EPOCH-RESTORE_START_EPOCH))" | tee evidence/incidents/recovery-time.txt
curl -fsS http://127.0.0.1:18081/ | tee evidence/incidents/recovered-state.json
REC_ID="$(STATE_VOLUME=ch42_state_restore APP_REF="$SUBJECT" docker compose -p ch42restore -f compose.yaml -f compose.recovery.yaml ps -q app)"
docker inspect "$REC_ID" --format 'Image={{.Image}} Status={{.State.Status}} Health={{.State.Health.Status}}' | tee evidence/incidents/recovered-runtime.txt

Recovery succeeds only if the restored application returns the expected counter and the running container uses the intended release digest. A green container alone is not restore proof. The measured RTO belongs only to this host and dataset.

8. Verify RPO explicitly

Compare the counter in original-after-backup.json with recovered-state.json. Any increments performed after the backup point are intentionally outside the archive and therefore outside the RPO. Write the observation into the incident record.

9. Runbook update template

Incident:
Detection signal:
UTC first-failure timestamp:
Context / host / Engine version:
Release subject digest:
Owning layer:
Original object IDs:
First-failure evidence files:
Root cause:
Minimal repair:
Verification repeated:
Rollback path:
Data/RPO impact:
Measured recovery time:
Residual risk:
Prevention/control update:
Owner:
Follow-up due date:

10. What the drills teach

Failure Owning layer Repair that should not happen
Missing Dockerfile input Source/build context Restart Docker or prune cache first
503 health marker Application/data health contract Assume restart policy will fix unhealthy state
Ops network disconnect Container network endpoint/DNS alias Change host DNS/firewall simultaneously
Read-only volume probe Mount authority chmod 777 or privileged container
Restore exercise Backup/recovery lifecycle Declare success because a blank container starts

11. Preserve the recovery environment for final review

Leave ch42restore running until Lesson 5 has compared original and recovered evidence. Do not remove the original volume, registry storage, builder, signature public key, or backup archive yet.

Knowledge check

Why is the storage failure injected with a separate read-only probe instead of changing the live app mount?

Why can the observer fail while the loopback health endpoint still succeeds?

What makes the backup application-consistent for this SQLite lab?

What two checks are required before calling restore successful?

Why is measured RTO not a universal performance claim?

Next lesson

Next: Production Capstone: Build, Secure, Publish, Operate, Observe, and Recover a Complete Dockerized Application: Final Operational Review and Handoff

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. The capstone records the learner’s actual versions before execution. Reference baselines used for compatibility discussion are Docker Engine/CLI 29.8.1, BuildKit 0.33.0, Buildx 0.37.1, Docker Compose 5.5.1, Trivy 0.74.0, and Cosign 3.1.3. Packaged Docker Desktop/Engine installations may expose different bundled containerd/runc versions, so the evidence packet records what the active daemon actually reports.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.