Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Configuration, Design Choices, and Tradeoffs
Choose default or custom seccomp, capability allowlists, LSM policy, read-only filesystems, rootless/user namespaces, and stronger isolation using explicit tradeoffs.
Learning objectives
- Choose between default and custom seccomp policy, a capability allowlist, LSM policy, read-only filesystem design, rootless/user namespaces, and stronger VM/sandbox isolation.
- Connect each control to portability, least privilege, diagnosability, failure isolation, host cost, and multi-tenant threat assumptions.
- State explicit prerequisites and observable evidence for a security design rather than treating a flag as proof of safety.
- Recognize when hostile multi-tenancy requires a stronger boundary than a shared kernel container.
1. Default seccomp versus a custom profile
The built-in profile tracks Docker’s runtime expectations and current hardening. A custom profile can solve a legitimate compatibility requirement or enforce a narrower application policy, but it creates maintenance responsibility: architecture differences, new application versions, kernel behavior, and Docker security updates must all be reviewed. Prefer the default unless there is evidence for a deliberate change.
2. Capability allowlist versus default capability set
A practical hardening path is to test with
cap-drop=ALL and then add back the smallest capability
set required. This creates an explicit allowlist and makes authority
changes visible in code review. The cost is compatibility work for
software that silently assumes traditional root privileges. Treat
capability additions as documented requirements with a reproducer.
3. LSM policy: host/distribution context matters
AppArmor and SELinux solve related mandatory-access-control problems through different policy models and are integrated differently by distributions. Do not ship one host’s policy assumption as if it were portable everywhere. Instead, document the required host control, provide a safe baseline on unsupported platforms, and test enforcement where the production environment actually runs.
4. Read-only filesystem versus writable runtime state
A read-only root filesystem makes unauthorized or accidental mutation harder and exposes software that improperly writes into its installation tree. It works best when writable state is classified: ephemeral state goes to tmpfs, persistent state goes to an explicit managed volume, and configuration comes from declared runtime inputs. This also improves incident reasoning because writable surfaces are enumerable.
5. Rootless Docker and user namespaces are additional layers
User namespaces remap container identities; rootless Docker runs the daemon and containers without a rootful daemon. Both can reduce host impact from certain compromise paths, but each changes networking, storage, device, and privileged-operation assumptions. They do not make broad mounts, weak application authorization, or exposed credentials safe. Chapter 26 treats these tradeoffs in depth.
6. Shared kernel versus stronger VM/sandbox boundary
For mutually hostile or strongly isolated tenants, the shared host kernel can be the wrong trust boundary even when container controls are hardened. A microVM, VM, sandboxed runtime, or dedicated host may provide a more appropriate isolation layer. This costs density, startup time, operational complexity, or hardware resources, but it changes the failure domain rather than only tightening one container policy.
7. Decision table: choose the control from the failure boundary
| Need / threat | Preferred control | Prerequisites/evidence | Tradeoff |
|---|---|---|---|
| Ordinary app does not need kernel privileges | Non-root + drop capabilities; add back narrowly | Application tests; inspect user/caps | May expose hidden privilege assumptions. |
| Reduce syscall surface | Built-in seccomp first; custom reviewed profile only when justified | Kernel seccomp support; denied-syscall evidence | Custom policy maintenance burden. |
| Constrain filesystem/network/object access beyond Unix permissions | AppArmor or SELinux policy | Host LSM active; audit evidence | Distribution-specific policy work. |
| Prevent rootfs mutation | Read-only rootfs + explicit tmpfs/volumes | Writable-path inventory | Requires application path cleanup. |
| Reduce rootful daemon/container host impact | Rootless Docker or user namespaces | Supported kernel/userspace; mapped IDs; feature review | Networking/storage/device limitations. |
| Run hostile multi-tenant workloads | VM/microVM/sandbox/dedicated host boundary | Threat model and platform support | Higher resource/operational cost. |
8. Worked scenario: image-processing API
Assume an API reads an uploaded image, writes temporary conversion output, and returns the result. It needs no host devices, no host namespace, no persistent rootfs writes, and no Docker daemon access. A justified baseline is a non-root UID, dropped capabilities, default seccomp/LSM policy, read-only rootfs, and a bounded tmpfs for temporary conversion. If the conversion library later needs one blocked syscall, preserve the failing syscall evidence and review the narrowest policy change instead of expanding unrelated permissions.
| Prediction | Observable proof |
|---|---|
| Root filesystem mutations fail |
Attempted write to a non-tmpfs root path fails and inspect
shows ReadonlyRootfs=true.
|
| Temporary conversion succeeds | Write/read within the dedicated tmpfs and inspect shows that tmpfs target. |
| No Linux capabilities retained | Inspect requested cap-drop and runtime capability evidence where available. |
| Default security layers remain enabled | Docker info/host policy state and absence of a container override that disables them. |
9. Portability and Docker Desktop
Docker Desktop runs Linux containers inside a managed Linux VM. The effective kernel, cgroup, namespace, AppArmor/SELinux availability, device surface, and host filesystem path semantics therefore differ from a native Linux Engine. Record the daemon context and platform before claiming that a host security control is active.
Knowledge check
When should you prefer Docker’s built-in seccomp profile?
By default, because it tracks Docker runtime expectations and hardening. Use a custom profile only for an evidenced compatibility or stronger-policy requirement you can maintain.
What is the value of cap-drop=ALL plus narrow add-back?
It turns capability authority into an explicit allowlist tied to observed application requirements.
Why is read-only rootfs easier to operate when writable paths are classified?
Because temporary, persistent, and configuration state each receive an explicit storage mechanism, making mutation surfaces predictable and auditable.
Does rootless Docker replace application-level authorization?
No. It reduces certain host-privilege risks but does not fix weak application auth, secrets handling, dangerous external permissions, or overly broad data exposure.
When is a VM/sandbox boundary preferable to further container hardening?
When the threat model involves hostile multi-tenancy or requires stronger kernel isolation than a shared host kernel can provide.
Official references and version notes
- Docker Engine security — daemon trust, kernel isolation, user namespaces, and host-hardening context.
- Seccomp security profiles for Docker — built-in profile behavior, blocked syscall rationale, and custom profile mechanics.
-
AppArmor security profiles for Docker
— generated
docker-defaultprofile and custom-policy workflow. - docker container run reference — capabilities, devices, read-only rootfs, security options, user, namespaces, and resource/security controls.
-
Compose service capability controls
—
cap_addandcap_drop. -
Compose service security_opt
— including
no-new-privilegesand LSM labels. - Compose read_only — read-only root filesystem semantics.
- Docker tmpfs mounts — ephemeral writable memory-backed state.
- Docker Engine 29 release notes — current release baseline and 2026 security changes.
Docker Engine 29.8.1 is the current Engine baseline. The built-in
seccomp policy remains a moderately protective allowlist-style
default and currently blocks roughly 44 syscalls out of 300+.
Docker’s AppArmor integration uses a generated
docker-default profile where AppArmor is active; Engine
29.8 adds support for generating that profile from a custom daemon
template. The 2026 CVE-2026-31431 hardening changed default
seccomp/LSM handling for
AF_ALG/socketcall; Docker explicitly warns
against disabling seccomp as a workaround. Actual host LSM,
rootless/userns, kernel, daemon, and Desktop/VM state must therefore
be inspected rather than assumed.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.