Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Concepts, Architecture, and Mental Model
Model Docker container security as layered namespaces, identity, Linux capabilities, seccomp, LSM policy, devices, mounts, and privilege boundaries.
Learning objectives
- Explain how namespaces, UID/GID identity, capabilities, seccomp, LSM policy, devices, mounts, and network authority form layered controls rather than one “container sandbox.”
- Distinguish process identity inside a container from authority over host resources and explain why non-root is necessary but not sufficient.
-
Read Docker security configuration from
docker infoanddocker inspectbefore changing any security setting. - Identify authority-expanding controls such as added capabilities, host namespaces, devices, writable host mounts, or privileged execution as separate trust decisions.
1. The practical problem: isolation is not the same as trust
Docker gives a process a distinct view of resources through namespaces, constrains resources with cgroups, and applies additional kernel authorization controls. Those mechanisms reduce the process’s authority, but none of them turns an arbitrary workload into a separate physical machine. A workload that receives broad capabilities, host namespaces, powerful devices, writable host paths, or Docker daemon access crosses important trust boundaries even if its prompt still looks like a container shell.
Chapter 24 separated process state from health and recovery. This chapter applies the same evidence-first discipline to security: first determine what identity and authority the process already has, then remove only what is unnecessary.
2. Mental model: layered authority from kernel to application
The host kernel is the common enforcement point. Namespaces control which resources a process can see. UID/GID and filesystem ownership determine normal identity checks. Linux capabilities split traditional root powers into narrower privileges. Seccomp filters syscall entry. AppArmor or SELinux can constrain permitted objects/actions through Linux Security Module policy. Device access, mounts, network modes, and daemon/API access can then expand or narrow the practical attack surface.
flowchart TD
A[Host kernel] --> B[Namespaces + cgroups baseline]
B --> C[UID/GID identity]
C --> D[Linux capability set]
D --> E[Seccomp syscall filter]
E --> F[AppArmor or SELinux policy]
F --> G[Devices + mounts + network namespace]
G --> H[Application process]
H --> I[Host/external side effects allowed only by accumulated authority]
3. Namespaces change visibility, not kernel ownership
Typical Linux containers use separate PID, mount, IPC, UTS, and network namespaces, with additional namespace behavior depending on daemon/runtime configuration. A container PID namespace can make the application appear as PID 1 inside the container while the host sees a different PID. Sharing a host namespace deliberately collapses one of those isolation layers, so host PID/network namespace choices are security architecture decisions rather than debugging conveniences.
docker info --format '{{json .SecurityOptions}}'
docker inspect CONTAINER --format 'pid_mode={{.HostConfig.PidMode}} network_mode={{.HostConfig.NetworkMode}} ipc_mode={{.HostConfig.IpcMode}}'
docker inspect CONTAINER --format 'host_pid={{.State.Pid}} user={{json .Config.User}}'
4. Runtime user and host authority are different questions
Running as a non-root user reduces what the application can do through normal Unix permissions and often reduces the impact of compromise. But “UID 1000 inside” does not by itself answer how that UID maps to the host, whether user namespaces are enabled, which supplementary groups or capabilities are present, which host files are mounted, or whether a device/API grants separate authority. Chapter 26 explores rootless Docker and user namespaces in depth; here they are treated as additional layers, not substitutes for the rest.
5. Linux capabilities: split root authority into named privileges
Docker starts ordinary Linux containers with a default subset of
capabilities rather than every capability. Controls such as
--cap-drop and --cap-add change that set.
The safest design begins with the operation the application must
perform, removes unneeded capabilities, and adds back only the
narrow capability proven necessary. Broad powers such as
CAP_SYS_ADMIN cover many kernel operations and
therefore deserve exceptional scrutiny.
| Question | Evidence | Interpretation |
|---|---|---|
| What did Docker request? |
.HostConfig.CapDrop /
.HostConfig.CapAdd
|
Declarative container configuration. |
| What does the process actually have? |
/proc/1/status capability masks; optional
capsh when the image contains it
|
Runtime process evidence. |
| Did an operation fail? | Application stderr + kernel/LSM audit evidence where authorized | May indicate normal permissions, missing capability, seccomp, LSM, or another boundary. |
6. Seccomp filters syscall entry
Docker’s built-in seccomp profile is applied on supported Linux hosts unless configuration changes it. Current Docker documentation describes it as an allowlist-style policy that returns an error for blocked syscalls. A “permission denied” result can therefore be deliberate sandbox behavior rather than a filesystem bug. Do not leap from one denied syscall to disabling the filter; identify the syscall and why the application needs it.
7. AppArmor and SELinux add mandatory access control
On AppArmor systems Docker normally uses a generated
docker-default container profile. On SELinux systems,
Docker integration depends on the daemon/distribution policy being
enabled and correctly labeled. These policies can deny operations
even when Unix mode bits and capabilities appear permissive. Their
audit logs are therefore important first-failure evidence. Do not
disable an LSM merely to make a workload run.
8. Read-only rootfs, tmpfs, mounts, devices, and sockets shape real authority
A read-only root filesystem narrows mutation but applications still
need selected writable paths for temporary files, sockets, or
caches; targeted tmpfs mounts can provide those paths without making
the entire image filesystem writable. Device mappings expose
host-backed interfaces. Bind mounts can expose host data. Access to
the Docker API or socket effectively delegates control over the
daemon and must be treated as high authority. Each is a separate
design decision visible in docker inspect.
9. Read-only baseline inspection before hardening
docker version
docker context show
docker info --format '{{json .SecurityOptions}}'
docker inspect CONTAINER --format 'user={{json .Config.User}} readonly={{.HostConfig.ReadonlyRootfs}} privileged={{.HostConfig.Privileged}}'
docker inspect CONTAINER --format 'cap_add={{json .HostConfig.CapAdd}} cap_drop={{json .HostConfig.CapDrop}}'
docker inspect CONTAINER --format 'security_opt={{json .HostConfig.SecurityOpt}} devices={{json .HostConfig.Devices}}'
docker inspect CONTAINER --format '{{json .Mounts}}'
docker info,
platform/daemon configuration, and runtime evidence.
10. Authority-expanding settings are architectural exceptions
| Control | What it changes | Safer default reasoning |
|---|---|---|
| Added capability | Expands a specific class of kernel operations | Add only a proven required capability; prefer application redesign when practical. |
| Host namespace sharing | Removes isolation for PID/network/IPC scope | Keep private namespaces unless the use case explicitly requires host scope. |
| Writable host bind | Lets container mutate selected host data | Use read-only or managed volumes; narrow source path and ownership. |
| Device mapping | Delegates access to host device interface | Expose only the exact device and access mode required. |
| Docker daemon/API access | Delegates daemon-level operations | Avoid in ordinary application containers; isolate build/automation services. |
| Privileged execution | Expands capabilities/devices and relaxes security confinement broadly | Treat as exceptional infrastructure-level authority, not a troubleshooting step. |
11. DevOps connection: reproducibility includes the security boundary
Two containers running the same image digest are not equivalent if one has a different user, capability set, seccomp/LSM policy, namespace mode, device list, or host mount. Record the exact image identity and runtime security configuration together. That makes the boundary auditable and makes rollback or incident comparison meaningful.
Knowledge check
Why does “non-root inside the container” not completely describe host authority?
Because UID mapping, capabilities, supplementary groups, namespaces, devices, mounts, and daemon/API access can independently grant or restrict authority.
What is the correct first response to an operation denied by seccomp or an LSM?
Preserve the error and policy/audit evidence, identify the exact required operation, and change only the smallest justified policy or application behavior.
What is the purpose of a read-only root filesystem?
It prevents normal writes to the container root filesystem, reducing mutation. It does not automatically make mounted paths, tmpfs paths, devices, or external services read-only.
Why is CAP_SYS_ADMIN high risk?
It gates a very broad collection of kernel operations, so adding it often expands authority far beyond one narrow application need.
What must be recorded with an image digest to describe a deployed security posture?
At least runtime user, capabilities, security options/seccomp/LSM state, namespace modes, read-only/mount/device configuration, and relevant daemon/host security context.
Official references and version notes
- Docker Engine security — daemon trust, kernel isolation, user namespaces, and host-hardening context.
- Seccomp security profiles for Docker — built-in profile behavior, blocked syscall rationale, and custom profile mechanics.
-
AppArmor security profiles for Docker
— generated
docker-defaultprofile and custom-policy workflow. - docker container run reference — capabilities, devices, read-only rootfs, security options, user, namespaces, and resource/security controls.
-
Compose service capability controls
—
cap_addandcap_drop. -
Compose service security_opt
— including
no-new-privilegesand LSM labels. - Compose read_only — read-only root filesystem semantics.
- Docker tmpfs mounts — ephemeral writable memory-backed state.
- Docker Engine 29 release notes — current release baseline and 2026 security changes.
Docker Engine 29.8.1 is the current Engine baseline. The built-in
seccomp policy remains a moderately protective allowlist-style
default and currently blocks roughly 44 syscalls out of 300+.
Docker’s AppArmor integration uses a generated
docker-default profile where AppArmor is active; Engine
29.8 adds support for generating that profile from a custom daemon
template. The 2026 CVE-2026-31431 hardening changed default
seccomp/LSM handling for
AF_ALG/socketcall; Docker explicitly warns
against disabling seccomp as a workaround. Actual host LSM,
rootless/userns, kernel, daemon, and Desktop/VM state must therefore
be inspected rather than assumed.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.