Chapter 06Lesson 04~115 minutes

Container Filesystems, Writable Layers, docker cp, exec Workflows, Processes, Signals, and PID 1: Diagnostics, Failure Modes, Security, and Performance

Filesystem and PID failures are easy to misdiagnose because runtime mutation can hide the real source state. This lesson preserves first-failure evidence, then diagnoses writable-layer drift, secret copying, mount masking, broken signal forwarding, unreaped children, and exec-based hot fixes without weakening isolation or deleting evidence.

DiagnosticsSignal forwardingFilesystem driftSecretsSafe repair

Learning objectives

  • Preserve container ID, image digest, mounts, docker diff output, process tree, logs, events, stop configuration, and copied-file provenance before repair.
  • Diagnose a shell-form signal-forwarding failure by proving which process is PID 1 and where SIGTERM is delivered.
  • Recognize that copying credentials into a container can leave sensitive runtime state and that docker cp should never be treated as a secret-management mechanism.
  • Diagnose content hidden by a mount separately from content deleted or changed in the writable layer.
  • Replace an exec-based hot fix with a declared image change while retaining evidence of the original failure.
Chapter 06 evidence baseline — verified 2026-09-21. This chapter uses free/local/disposable Docker resources, synthetic non-secret files, digest-identified base content, and exact labels/names. Docker Engine 29.8.1 is the current Engine 29 patch baseline at verification time, but executable steps record the learner's actual Engine/CLI/platform/storage state. Current documentation distinguishes immutable image layers, the unique container writable layer, explicit mounts, and process state; documents docker cp for running or stopped containers; documents docker exec as an additional process while PID 1 runs; and documents --init as a Tini-backed init option. No lab disables seccomp/LSM/firewall/TLS, mounts the Docker socket, uses privileged mode, or performs broad prune.

1. Diagnostic rule: preserve the filesystem/process boundary before changing it

When a container behaves differently from the image you expected, do not begin by copying files, opening an interactive shell, or restarting it. First capture context/version, container ID, image digest, mounts, diff output, configured command, PID 1/process tree, stop settings, logs/events, exit state, and the path you believe is wrong.

A hot fix can erase the original state. docker exec, docker cp, restart, forced removal, or mount changes can all invalidate the evidence that would have explained the failure.

2. Evidence-first diagnostic sequence

  1. Confirm context, Engine/CLI patch level, OS type, and daemon reachability.
  2. Capture exact image/container identity.
  3. Inspect .Mounts before interpreting filesystem visibility.
  4. Capture docker diff and relevant copied-out non-secret files.
  5. Capture docker top, /proc/1/cmdline where appropriate, logs, events, and stop configuration.
  6. Classify the failure: image content, writable drift, mounted content, permissions/identity, process launch, signal forwarding, child reaping, or external dependency.
  7. Test one hypothesis in a disposable reproduction, then encode the durable fix in declared source/configuration.

3. Failure mode: writable-layer edits became “configuration”

An operator enters a container, changes /etc/widget.conf, and reports that the deployment is fixed. The service now works, but the image digest and source remain unchanged. A later replacement reintroduces the fault.

Diagnosis:

docker inspect CONTAINER --format 'ID={{.Id}} Image={{.Image}} Status={{.State.Status}}'
docker inspect CONTAINER --format '{{json .Mounts}}'
docker diff CONTAINER | tee runtime-diff.txt
docker cp CONTAINER:/etc/widget.conf ./widget.conf.observed

Compare the observed file with the declared source in a safe workspace. The durable repair is a reviewed image/configuration change followed by a replacement test, not an attempt to snapshot the manually modified container as the new release.

4. Failure mode: a secret was copied into the container

docker cp api.key CONTAINER:/app/ may look convenient, but it creates sensitive runtime filesystem state outside a controlled secret lifecycle. It can also make incident evidence and backups unexpectedly sensitive. Do not reproduce this with a real key.

Safe correction: revoke/rotate any real credential that was exposed, preserve only sanitized evidence, identify where copies/logs/backups may exist, and migrate to an approved secret injection mechanism. Never teach “delete the file and move on” as proof that the secret is gone.

Engine 29.5.1 fixed several docker cp vulnerabilities involving archive/path handling. Keeping Docker patched is part of the security boundary, especially when copying from containers you do not fully trust.

5. Failure mode: mount masking misdiagnosed as image deletion

The image contains /app/defaults/policy.json, but the running container reports “file not found.” Before rebuilding, inspect mounts:

docker inspect CONTAINER --format '{{range .Mounts}}{{println .Type .Source "->" .Destination "RW=" .RW}}{{end}}'

If a mount targets /app/defaults, its filesystem can obscure the image content. Reproduce in a disposable container without the mount to prove the image file exists. Do not copy a replacement into the mounted path until you identify who owns that storage and whether writes are authorized.

6. Intentionally broken example: shell-form launch hides the application from SIGTERM

This lab uses a disposable script that logs its PID and handles TERM. Two images differ only in startup form. Resolve the same BusyBox digest as in Lesson 2, then create the files:

mkdir -p ch06-signal-broken && cd ch06-signal-broken
printf '%s\n' '#!/bin/sh' \
  'echo app-start pid=$$' \
  'trap '\''echo app-received-TERM pid=$$; exit 0'\'' TERM' \
  'while :; do sleep 1; done' > app.sh
chmod 755 app.sh

cat > Dockerfile.bad <<'EOF'
ARG BASE_IMAGE
FROM ${BASE_IMAGE}
COPY app.sh /app.sh
RUN chmod 0555 /app.sh
CMD /app.sh
EOF

cat > Dockerfile.good <<'EOF'
ARG BASE_IMAGE
FROM ${BASE_IMAGE}
COPY app.sh /app.sh
RUN chmod 0555 /app.sh
CMD ["/app.sh"]
EOF

Build with the same pinned BASE_IMAGE, start both with chapter labels, then inspect their process trees:

docker build -f Dockerfile.bad --build-arg BASE_IMAGE="$PINNED" -t devops-academy-ch06:bad .
docker build -f Dockerfile.good --build-arg BASE_IMAGE="$PINNED" -t devops-academy-ch06:good .

docker run -d --name devops-academy-ch06-bad --label devops-academy.lab=ch06 devops-academy-ch06:bad
docker run -d --name devops-academy-ch06-good --label devops-academy.lab=ch06 devops-academy-ch06:good

docker top devops-academy-ch06-bad
docker top devops-academy-ch06-good

In the bad image, shell form inserts /bin/sh -c as PID 1 and the application script is a child. In the good image, the application script is launched directly. Stop each with a short, explicit timeout and compare logs, duration, and exit state:

time docker stop --timeout 3 devops-academy-ch06-bad
docker logs devops-academy-ch06-bad
docker inspect devops-academy-ch06-bad --format 'Exit={{.State.ExitCode}} Finished={{.State.FinishedAt}}'

time docker stop --timeout 3 devops-academy-ch06-good
docker logs devops-academy-ch06-good
docker inspect devops-academy-ch06-good --format 'Exit={{.State.ExitCode}} Finished={{.State.FinishedAt}}'
Interpret evidence, not folklore. Exact exit codes/timing can vary by platform and shell implementation. The causal test is which process is PID 1, whether the application logs its TERM handler, and whether Docker had to escalate after the timeout.

7. Failure mode: child processes are not reaped

Long-running services that fork helper processes can accumulate zombies if PID 1 never reaps them. Do not diagnose this by blindly adding a full init system. First capture the process table repeatedly and prove defunct children exist. Then fix the application process model or use a minimal init such as Docker's --init when appropriate.

Evidence should show the before/after process table and the .HostConfig.Init setting. The improvement is “zombie processes are reaped,” not merely “we added --init.”

8. Failure mode: exec-based hot fix became deployment practice

A runbook says: “after deployment, exec into the container and change three files.” This creates an undeclared second deployment phase that cannot be reproduced from the image digest. It also makes rollback ambiguous because a restarted/replaced container may lose the changes.

Correction: move the intended files/commands into reviewed image source or an explicit runtime configuration mechanism. Use the exec procedure only as a documented diagnostic experiment with start/end timestamps and a follow-up immutable build.

9. Performance diagnosis: writable-layer I/O versus persistent mounts

Write-heavy workloads can suffer copy-on-write overhead and grow the container writable layer. Do not conclude “Docker storage is slow” from one measurement. Capture actual daemon storage mode, the path's mount ownership, write pattern, container size, and host I/O metrics. Compare the same workload with the correct persistent mount in a disposable benchmark before changing production.

Later storage chapters cover volume drivers and backup/recovery. Here, the diagnostic lesson is ownership first: a performance complaint at a mounted path is not a writable-layer problem.

10. Clean up only the intentional broken-example resources

docker ps -a --filter label=devops-academy.lab=ch06 --format 'table {{.ID}}\t{{.Names}}\t{{.Status}}'

docker rm devops-academy-ch06-bad devops-academy-ch06-good 2>/dev/null || true
docker image rm devops-academy-ch06:bad devops-academy-ch06:good 2>/dev/null || true

If the listing contains other Chapter 06 resources from another lab step, do not remove them by broad label without confirming ownership. Exact names are safer during diagnosis.

11. Incident challenge: the file is wrong and shutdown is slow

A service has a changed /app/config.json, a bind mount at /app/config, shell-form startup, and a 30-second forced shutdown. Design an evidence sequence that determines: whether the visible file comes from the mount or writable layer, what the image originally contained, which process is PID 1, whether the application receives TERM, and which correction should be image/configuration versus runtime. Do not propose chmod 777, privileged mode, socket mounts, or force deletion.

Next lesson

Next: Checkpoint Lab

Combine filesystem drift, copy evidence, process/signal proof, immutable rebuild, and exact cleanup into one end-to-end dossier.

Knowledge check

What should you inspect before concluding an image file was deleted?

Why is copying a secret into a container not fixed simply by deleting the file later?

What evidence distinguishes shell-form and exec-form signal behavior?

Why is an exec hot fix not a complete incident resolution?

Official references and version notes

  • Docker storage overview — explains the ephemeral per-container writable layer and why persistent data belongs in explicit mounts rather than the container layer.
  • Storage drivers and writable layers — copy-on-write concepts, writable-layer behavior, and the Engine 29 distinction between the containerd image store and classic storage-driver examples.
  • docker container diff — reports added, changed, and deleted paths in the container filesystem relative to its initial image-backed state.
  • docker container cp — current copy semantics, stopped-container support, destination ownership, archive mode, symlink handling, and path rules.
  • docker container exec — starts a new command only while the container's primary PID 1 is running and does not make image changes durable.
  • docker container run — current runtime options including --init, stop signals/timeouts, mounts, TTY behavior, and security controls.
  • docker container stop — graceful signal delivery and timeout-to-SIGKILL escalation.
  • docker container attach — documents special PID 1 signal behavior and terminal signal-proxy implications.
  • JSONArgsRecommended build check — why exec-form CMD/ENTRYPOINT avoids an unintended shell parent and improves signal handling.
  • Docker build best practices — entrypoint scripts should normally exec the final application so it becomes PID 1 and receives signals directly.
  • Bind mounts — mounting over an existing container path obscures image/writable-layer content until the container is recreated without that mount.
  • Docker Engine 29 release notes — Engine 29.8.1 is the current patch baseline at verification time; recent 29.x releases also include multiple docker cp security and compatibility fixes.
Current baseline, not a frozen requirement

Verified 2026-09-21: Docker Engine 29.8.1 is the current Engine 29 patch release. Current Docker documentation states that each container gets a unique writable layer above immutable image layers and that deleting the container deletes that layer; docker cp can copy to or from running or stopped containers; docker exec starts an additional process only while the primary process is running; and --init uses Docker's Tini-backed init to perform normal init duties such as reaping child processes. Because Engine 29.5.x–29.7.x contained important docker cp security and compatibility fixes, learners should record their actual Engine/CLI versions and current security status rather than treating file-copy behavior as version-independent.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.