Chapter 25Lesson 02~120 minutes

Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Guided Hands-On Workflow and Core Operations

Harden disposable Docker containers with non-root execution, capability reduction, no-new-privileges, read-only rootfs, tmpfs, and safe security inspection.

Container securityCapabilitiesSeccompAppArmor / SELinuxLeast privilege

Learning objectives

  • Run a disposable application as a non-root user and inspect its effective security configuration.
  • Apply cap-drop=ALL and add back only one capability when a concrete operation requires it.
  • Combine no-new-privileges, read-only rootfs, and targeted tmpfs mounts without disabling seccomp or LSM policy.
  • Inspect AppArmor/SELinux/security-option state safely and interpret a denied operation without escalating the entire container.
Lab safety. This lesson uses disposable Alpine containers and a temporary local directory only. It never disables seccomp/AppArmor/SELinux, never uses privileged execution, never exposes the Docker API/socket to a container, never joins host PID/network namespaces, and never maps a real host device. Where host-specific LSM or seccomp behavior cannot be demonstrated portably, the lab inspects and explains it rather than weakening policy.

1. Preflight: record host/daemon security context

Security behavior depends on kernel, daemon, context, Desktop VM, LSM, and rootless/user-namespace state. Capture those facts before creating the lab.

mkdir -p dkr25-evidence

docker context show | tee dkr25-evidence/context.txt
docker version | tee dkr25-evidence/docker-version.txt
docker info | tee dkr25-evidence/docker-info.txt
docker info --format '{{json .SecurityOptions}}' | tee dkr25-evidence/security-options.json

docker pull alpine:3.22
docker image inspect alpine:3.22 --format '{{json .RepoDigests}}' | tee dkr25-evidence/alpine-digests.json

2. Run the application as non-root and prove the identity

Alpine provides a nobody identity suitable for this synthetic lab. The container performs only a bounded identity/read operation and then stays alive for inspection.

docker rm -f dkr25-nonroot 2>/dev/null || true

docker run -d --name dkr25-nonroot   --label academy=docker-ch25   --user 65534:65534   alpine:3.22 sh -c 'id; umask; while :; do sleep 5; done'

docker exec dkr25-nonroot id | tee dkr25-evidence/nonroot-id.txt
docker inspect dkr25-nonroot \
  --format 'user={{json .Config.User}} privileged={{.HostConfig.Privileged}} readonly={{.HostConfig.ReadonlyRootfs}}'   | tee dkr25-evidence/nonroot-inspect.txt

3. Drop all capabilities and verify a normal workload still runs

Many application processes do not need Linux capabilities at all. Start with none and test the actual application behavior. Here the synthetic workload only reads a file and writes to stdout.

docker rm -f dkr25-dropall 2>/dev/null || true

docker run -d \
  --name dkr25-dropall \
  --label academy=docker-ch25 \
  --user 65534:65534 \
  --cap-drop ALL   alpine:3.22 sh -c 'cat /etc/os-release >/dev/null; echo minimal-authority-ok; while :; do sleep 5; done'

docker logs dkr25-dropall | tee dkr25-evidence/dropall-logs.txt
docker inspect dkr25-dropall --format 'cap_drop={{json .HostConfig.CapDrop}} cap_add={{json .HostConfig.CapAdd}}'   | tee dkr25-evidence/dropall-inspect.txt

4. Add back one capability only for a demonstrated need

A low-number TCP port is a clear, bounded example. Modern kernels may permit unprivileged low-port binds depending on net.ipv4.ip_unprivileged_port_start, so first inspect that host/container assumption. The lesson therefore treats NET_BIND_SERVICE as a conditional example rather than pretending every host requires it.

docker run \
  --rm \
  --cap-drop ALL alpine:3.22 sh -c   'cat /proc/sys/net/ipv4/ip_unprivileged_port_start 2>/dev/null || true'   | tee dkr25-evidence/unprivileged-port-start.txt

# Record the exact configuration of a one-capability container.
docker rm -f dkr25-onecap 2>/dev/null || true
docker run -d \
  --name dkr25-onecap \
  --label academy=docker-ch25 \
  --user 65534:65534 \
  --cap-drop ALL \
  --cap-add NET_BIND_SERVICE   alpine:3.22 sh -c 'echo capability-example-ready; while :; do sleep 5; done'

docker inspect dkr25-onecap --format 'cap_drop={{json .HostConfig.CapDrop}} cap_add={{json .HostConfig.CapAdd}}'   | tee dkr25-evidence/onecap-inspect.txt
Do not cargo-cult the capability. If the application does not need it on your platform, retain cap-drop=ALL. Evidence of need—not tradition—justifies an added capability.

5. Combine read-only rootfs with targeted writable tmpfs

The root filesystem becomes immutable while /tmp remains writable and ephemeral. This pattern narrows runtime mutation without pretending that every application can run with zero writable paths.

docker rm -f dkr25-readonly 2>/dev/null || true

docker run -d \
  --name dkr25-readonly \
  --label academy=docker-ch25 \
  --user 65534:65534 \
  --cap-drop ALL \
  --read-only \
  --tmpfs /tmp:rw,nosuid,nodev,size=16m   alpine:3.22 sh -c 'echo ok >/tmp/app.tmp; while :; do sleep 5; done'

docker exec dkr25-readonly sh -c 'cat /tmp/app.tmp; (echo blocked >/etc/should-fail) 2>&1 || true'   | tee dkr25-evidence/readonly-behavior.txt
docker inspect dkr25-readonly \
  --format 'readonly={{.HostConfig.ReadonlyRootfs}} tmpfs={{json .HostConfig.Tmpfs}} mounts={{json .Mounts}}'   | tee dkr25-evidence/readonly-inspect.txt

6. Add no-new-privileges and inspect the declared security option

no-new-privileges prevents a process and its children from gaining additional privilege through mechanisms such as set-user-ID/set-group-ID executables. It complements, rather than replaces, UID choice, capabilities, seccomp, and LSM policy.

docker rm -f dkr25-hardened 2>/dev/null || true

docker run -d \
  --name dkr25-hardened \
  --label academy=docker-ch25 \
  --user 65534:65534 \
  --cap-drop ALL \
  --security-opt no-new-privileges=true \
  --read-only \
  --tmpfs /tmp:rw,nosuid,nodev,size=16m   alpine:3.22 sh -c 'echo hardened-ready >/tmp/status; while :; do sleep 5; done'

docker inspect dkr25-hardened \
  --format 'security_opt={{json .HostConfig.SecurityOpt}} readonly={{.HostConfig.ReadonlyRootfs}} cap_drop={{json .HostConfig.CapDrop}}'   | tee dkr25-evidence/hardened-inspect.txt

7. Inspect seccomp safely; do not disable it to manufacture a demo

Seccomp-denied syscalls vary by kernel, architecture, capabilities, and application. A portable course must not rely on an unsafe or brittle syscall demonstration. Instead, confirm daemon security options and preserve any natural EPERM evidence from a real application. If an application requires a syscall blocked by the default profile, review the exact syscall and use a narrowly reviewed custom profile only when justified.

docker info --format '{{json .SecurityOptions}}'   | tee dkr25-evidence/seccomp-daemon-evidence.json

docker inspect dkr25-hardened --format 'security_opt={{json .HostConfig.SecurityOpt}}'   | tee dkr25-evidence/seccomp-container-evidence.txt

8. Inspect AppArmor/SELinux state without turning policy off

The following commands are read-only and deliberately tolerant of hosts that do not provide a given interface, including Docker Desktop VMs.

docker info | grep -Ei 'apparmor|selinux|security options' -A8   | tee dkr25-evidence/lsm-docker-info.txt || true

if command -v aa-status >/dev/null 2>&1; then
  aa-status | tee dkr25-evidence/apparmor-status.txt
fi
if command -v getenforce >/dev/null 2>&1; then
  getenforce | tee dkr25-evidence/selinux-enforce.txt
fi

9. Challenge: identify the failing security layer before changing it

Suppose a non-root process cannot write /etc/app.conf in the hardened container. Which layer should you inspect first? The declared configuration already says the root filesystem is read-only, so changing capabilities or seccomp would be the wrong layer. Decide whether the application should write elsewhere, receive a narrow writable volume/tmpfs, or move the configuration into the image/build/runtime configuration model.

10. Exact cleanup

docker rm -f dkr25-nonroot dkr25-dropall dkr25-onecap dkr25-readonly dkr25-hardened 2>/dev/null || true
# Keep dkr25-evidence/ if you want to compare later lessons.

Knowledge check

Why does this lab not disable seccomp to demonstrate a denial?

What does cap-drop=ALL prove?

Why pair read-only rootfs with tmpfs?

What should you do if NET_BIND_SERVICE is not actually needed on your host?

A write fails under read-only rootfs. Why is adding a capability usually the wrong first fix?

Next lesson

Next: Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Configuration, Design Choices, and Tradeoffs

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version/security baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline. The built-in seccomp policy remains a moderately protective allowlist-style default and currently blocks roughly 44 syscalls out of 300+. Docker’s AppArmor integration uses a generated docker-default profile where AppArmor is active; Engine 29.8 adds support for generating that profile from a custom daemon template. The 2026 CVE-2026-31431 hardening changed default seccomp/LSM handling for AF_ALG/socketcall; Docker explicitly warns against disabling seccomp as a workaround. Actual host LSM, rootless/userns, kernel, daemon, and Desktop/VM state must therefore be inspected rather than assumed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.