Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Diagnostics, Failure Modes, Security, and Performance
Diagnose capability, seccomp, AppArmor/SELinux, filesystem, namespace, device, mount, and daemon-authorization failures without broad privilege escalation.
Learning objectives
- Diagnose capability, seccomp, LSM, filesystem, device, namespace, and Docker API/socket failures as different security layers.
- Preserve first-failure evidence before any policy change and avoid broad authority escalation as a troubleshooting shortcut.
- Explain why host PID/network namespace sharing, writable host mounts, broad capabilities, or daemon/API exposure alter the threat boundary.
- Repair the smallest missing permission or design assumption and verify the corrected posture independently.
1. Evidence-first security diagnostic sequence
| Step | Evidence to preserve | Why it matters |
|---|---|---|
| 1. Context/version | docker context show, version, platform |
Rules out wrong daemon/platform assumptions. |
| 2. Exact image/container identity | image digest, container ID, command, user | Proves the workload/config being diagnosed. |
| 3. Security config | cap add/drop, security options, readonly, devices, namespace modes, mounts | Enumerates declared authority. |
| 4. First failure | stderr, exit code, health/log/event timestamp | Preserves the actual symptom before retries change it. |
| 5. Host policy evidence | AppArmor/SELinux/audit logs when authorized; daemon SecurityOptions | Identifies mandatory policy denials. |
| 6. Smallest correction | one permission/path/policy/application change | Avoids hiding root cause with broad escalation. |
| 7. Re-run same scope | same image/input/test | Proves causality. |
2. Intentionally broken example: writable-path assumption
This disposable container deliberately tries to write under
/var/lib/app while the root filesystem is read-only.
The failure is preserved. The repair is a narrow tmpfs at the
required path—not added capabilities, not disabled security policy.
docker rm -f dkr25-broken dkr25-fixed 2>/dev/null || true
docker run \
--name dkr25-broken \
--user 65534:65534 \
--cap-drop ALL \
--security-opt no-new-privileges=true \
--read-only alpine:3.22 sh -c 'mkdir -p /var/lib/app && echo state >/var/lib/app/state' >dkr25-broken.stdout 2>dkr25-broken.stderr || true
docker inspect dkr25-broken \
--format 'exit={{.State.ExitCode}} readonly={{.HostConfig.ReadonlyRootfs}} user={{json .Config.User}} cap_drop={{json .HostConfig.CapDrop}} security_opt={{json .HostConfig.SecurityOpt}}' | tee dkr25-broken.inspect.txt
cat dkr25-broken.stderr
# Repair only the required writable path.
docker run \
--rm \
--name dkr25-fixed \
--user 65534:65534 \
--cap-drop ALL \
--security-opt no-new-privileges=true \
--read-only \
--tmpfs /var/lib/app:rw,nosuid,nodev,size=8m,uid=65534,gid=65534 alpine:3.22 sh -c 'echo state >/var/lib/app/state && cat /var/lib/app/state'
3. Missing capability versus ordinary file permission
Do not infer a capability problem from every EPERM.
First inspect the target object’s ownership/mode, runtime UID/GID,
requested capabilities, and filesystem mount mode. Only after a
reproducer shows a privileged kernel operation should a capability
change be considered. If one is needed, add the narrow capability
and keep all others dropped.
4. Seccomp denial: identify the syscall before changing policy
Seccomp commonly surfaces as an operation denied by the kernel. Capture the application trace/log and, where available, kernel/audit evidence. Check whether the syscall is documented as blocked and whether the application can avoid it. A reviewed custom profile is a last-mile compatibility tool, not a generic fix. Current Docker security guidance explicitly warns against disabling seccomp to work around the 2026 socket hardening issue.
5. AppArmor/SELinux denial: use audit evidence
LSM denials often include policy/profile and target-object details in host audit logs. Preserve those records before changing labels or profiles. Correct the label/policy for the intended resource or redesign the access. Do not switch the host to an unconfined policy merely to make the error disappear.
6. Host namespace sharing changes the diagnosis and the threat boundary
If a workload requires seeing host processes or using the host network stack, first ask why. Host namespace sharing is not just “more connectivity” or “better debugging”; it removes an isolation layer. Prefer dedicated diagnostic tooling, read-only telemetry, or a separate privileged operations plane rather than granting ordinary application containers host-wide visibility.
7. Writable host mounts and daemon/API exposure are high-authority paths
A writable bind mount can let a container change host data in that path. Docker daemon/API access can let a client create containers and request host resources, making it an administrative boundary. These should not be used as convenient fixes for file access or build tooling. If automation truly requires daemon operations, isolate the automation identity and service from untrusted application workloads.
8. “Non-root” can still carry dangerous authority
A non-root UID can still be connected to powerful groups, devices, mounts, network namespaces, or APIs. Conversely, UID 0 inside a user namespace may map to an unprivileged host UID. Security review must therefore describe the full mapping and authority set rather than using a single “root/non-root” label.
9. Security and performance: measure, do not assume
Default namespace, capability, seccomp, and LSM enforcement usually costs far less than the risk of removing the boundary. If a workload shows measurable overhead, benchmark the exact operation and preserve profiling evidence. Do not turn off a protection based on generic claims. Stronger VM/sandbox isolation can have larger resource costs; that is an architectural tradeoff tied to the threat model.
10. Troubleshooting shortcuts that hide the cause
| Shortcut | Why it is unsafe/misleading | Correct direction |
|---|---|---|
| Broad privilege escalation | Changes many capabilities, devices, and policy boundaries at once | Identify exact denied operation and grant only what is necessary. |
| Disabling syscall confinement | Expands kernel attack surface and hides syscall compatibility evidence | Use default policy; narrow custom exception only when reviewed. |
| Disabling AppArmor/SELinux | Removes mandatory policy from unrelated resources/workloads | Fix label/profile/policy and preserve audit evidence. |
| Sharing host PID/network by default | Collapses isolation to simplify debugging | Use targeted diagnostics and explicit network/service design. |
| Writable host-root-style mounts | Delegates broad host filesystem authority | Mount only the narrow path and access mode required, or use managed storage. |
| Docker daemon/API access in app container | Delegates container/host resource control | Use isolated administration/build service with explicit authorization. |
11. Cleanup the broken-example resources
docker rm -f dkr25-broken dkr25-fixed 2>/dev/null || true
rm -f dkr25-broken.stdout dkr25-broken.stderr dkr25-broken.inspect.txt
Knowledge check
Why is “Operation not permitted” insufficient to diagnose a missing capability?
Because the denial could come from normal permissions, seccomp, an LSM, read-only mount policy, device rules, namespace scope, or another security layer.
What made the intentionally broken example fail?
The application tried to write under a read-only root filesystem. The narrow fix was to provide a dedicated writable tmpfs at the required path.
Why must LSM audit evidence be captured before changing policy?
It identifies the profile/label/object/action that was denied and preserves causality before a policy change hides the original failure.
Why can a Docker daemon/API connection be more powerful than the UID inside the container?
Because daemon operations can create or control containers and request host resources; authority comes from the API permission boundary, not only Unix UID inside the caller container.
What is the right response to suspected security-control performance overhead?
Measure the exact workload and control with profiling/benchmark evidence, then make a threat-model-aware decision; do not disable protection based on assumption.
Official references and version notes
- Docker Engine security — daemon trust, kernel isolation, user namespaces, and host-hardening context.
- Seccomp security profiles for Docker — built-in profile behavior, blocked syscall rationale, and custom profile mechanics.
-
AppArmor security profiles for Docker
— generated
docker-defaultprofile and custom-policy workflow. - docker container run reference — capabilities, devices, read-only rootfs, security options, user, namespaces, and resource/security controls.
-
Compose service capability controls
—
cap_addandcap_drop. -
Compose service security_opt
— including
no-new-privilegesand LSM labels. - Compose read_only — read-only root filesystem semantics.
- Docker tmpfs mounts — ephemeral writable memory-backed state.
- Docker Engine 29 release notes — current release baseline and 2026 security changes.
Docker Engine 29.8.1 is the current Engine baseline. The built-in
seccomp policy remains a moderately protective allowlist-style
default and currently blocks roughly 44 syscalls out of 300+.
Docker’s AppArmor integration uses a generated
docker-default profile where AppArmor is active; Engine
29.8 adds support for generating that profile from a custom daemon
template. The 2026 CVE-2026-31431 hardening changed default
seccomp/LSM handling for
AF_ALG/socketcall; Docker explicitly warns
against disabling seccomp as a workaround. Actual host LSM,
rootless/userns, kernel, daemon, and Desktop/VM state must therefore
be inspected rather than assumed.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.