Production Container Patterns, Immutable Delivery, Configuration, Statelessness, Sidecars, and Operational Contracts: Diagnostics, Failure Modes, Security, and Performance
Diagnose mutable-container drift, writable-layer data loss, secret leakage, unbounded logs/resources, misleading health, brittle companion coupling, and replacement/recovery failures without erasing first-failure evidence.
Learning objectives
- Diagnose production-container failures by preserving immutable identity, runtime state, storage, security, network, health, resource, log, and external evidence before repair.
- Recognize runtime mutation and writable-layer state as hidden drift rather than durable fixes.
- Differentiate a healthcheck failure from an externally unreachable or dependency-broken application.
- Diagnose resource/log pressure and companion-service coupling without disabling security controls or using broad cleanup.
- Repair an intentionally broken read-only filesystem example by declaring the correct writable state boundary.
1. Evidence-first incident sequence
Before replacing or restarting anything, record the context/host, exact container ID, configured image reference and image ID, health history, restart count, exit/OOM state, mounts, security settings, resources, logs/events, network publication, volume identity, and the external request result. Then change the smallest layer that explains the failure.
2. Failure: SSH/exec patching becomes the “fix”
An operator edits a Python file inside the running container and service returns to normal. The immediate symptom is gone, but the release artifact is unchanged and the patch will disappear on replacement. Preserve the diff/evidence, reproduce the correction in source, rebuild a new digest, test it, and replace the container. Use the running mutation only as incident evidence, not as release state.
3. Failure: database or business state in writable layer
Container recreation appears to “randomly” delete data. Inspect
.Mounts and application paths. If the data directory is
absent from declared persistent storage, the behavior is
deterministic: it lived in the writable layer. Fix the data
architecture and restore from valid backup; do not copy random
Engine storage directories by hand.
4. Failure: secret baked into image or environment
Finding a credential in image history, image config, environment inspection, logs, or source is a credential incident. Preserve evidence, revoke/rotate the secret, remove it from future build/runtime paths, and verify the replacement digest/config. Merely deleting a file in a later image layer does not erase it from prior layers.
5. Failure: unbounded logs or resources
Disk exhaustion and host OOM can masquerade as application
instability. Inspect logging driver/options, Docker disk use,
container resource limits, docker stats, OOM state, and
host filesystem capacity. Apply bounded per-service logging and
measured resource limits; do not start with global prune or
arbitrary restarts.
6. Failure: healthcheck treated as SLA proof
A container is healthy, yet users cannot connect. Preserve the health history and then inspect the external path: bound interface, published host port, firewall/load balancer, DNS/TLS, and upstream dependencies. Healthy proves only the configured probe path succeeded.
7. Failure: restart loop hides the root cause
A restart policy can make a failed process repeatedly disappear and reappear. Capture the first exit code, logs, events, restart count, and OOM status before tuning the policy. If the application exits because configuration is invalid, more restarts add noise rather than resilience.
8. Failure: companion lifecycle assumption
A companion expects the primary app to exist continuously and exits permanently during an ordinary app replacement. The main app returns healthy but the observer/forwarder remains absent. A robust companion should retry the app endpoint and have its own restart/health contract, or the function should move to infrastructure whose lifecycle matches the requirement.
9. Intentionally broken example: undeclared write on read-only rootfs
# BROKEN LAB EXAMPLE — safe local demonstration
services:
app:
image: python:3.13-alpine
user: "10001:10001"
read_only: true
command:
- python
- -c
- |
from pathlib import Path
Path('/app/state.txt').write_text('business-state')
Expected first failure (exact wording varies):
OSError: [Errno 30] Read-only file system: '/app/state.txt'
The correct diagnosis is not “read-only is broken.” The application attempted a persistent write to an undeclared immutable path. Decide whether the data is ephemeral (tmpfs), durable (named volume/external database), or should not be written at runtime. Keep the root filesystem read-only and declare only the required writable path.
10. Safe repair of the broken example
services:
app:
image: python:3.13-alpine
user: "10001:10001"
read_only: true
tmpfs:
- /tmp:rw,noexec,nosuid,size=8m
volumes:
- state:/data
command:
- python
- -c
- |
from pathlib import Path
Path('/data/state.txt').write_text('synthetic-state')
volumes:
state: {}
In a complete application, ensure the mounted data path has correct numeric ownership for the runtime user and add backup/retention rules. Do not weaken the whole filesystem to make one path writable.
11. Layer-by-layer diagnostic matrix
| Symptom | Likely layer | Evidence before correction |
|---|---|---|
| different code than release | image/runtime drift | configured image ref, image ID, file checksum |
| data missing after recreate | storage lifecycle | mounts, volume ID, backup record |
| permission denied on intended data | identity/storage/security | UID/GID, mount owner, LSM/security opts |
| healthy but unreachable | network/external path | health history, published port, request/DNS/TLS evidence |
| exit 137 / OOM | resource/kernel | OOMKilled, memory limit, stats, host pressure |
| disk fills quickly | logging/storage | log driver/options, Docker df, host free space/inodes |
| companion absent after app change | service lifecycle | container IDs, restart policy, logs/events |
12. Smallest-safe-correction rule
Fix the layer that violated the contract: new digest for code, config revision for configuration, ownership/mount for data, resource envelope for pressure, logging policy for retention, network rule for reachability, or companion retry/lifecycle for coupling. Do not change daemon security, host filesystem permissions, or unrelated resources to compensate for an application-level defect.
13. Preserve recovery evidence
After the correction, retain both before and after identities: old/new digest, old/new container ID, volume identity, health/restart state, external request, state checksum, logs/events, and the specific change that explains recovery. Recovery is not “it works now”; it is a causal, reproducible state transition.
Knowledge check
A container is healthy but users cannot connect. What should you do next?
Keep the health evidence and inspect publication, host/network/firewall/TLS/DNS and external dependency state; health alone is not availability.
Why is an exec-based production patch not an acceptable final correction?
It creates drift outside the immutable release artifact and disappears on replacement.
A read-only rootfs causes /app/state.txt writes to
fail. What is the right question?
Whether that data is ephemeral, durable, or unnecessary; then declare the narrow correct writable path rather than disabling read-only globally.
What evidence distinguishes resource failure from application exception?
OOMKilled/exit code, configured limits, stats/cgroup evidence and host pressure, alongside application logs.
Why capture the old container ID and digest before replacement?
They anchor the incident to exact runtime bytes/state and make before/after recovery evidence auditable.
Official references and version notes
Diagnostic baseline: 2026-09-22. The lesson deliberately avoids broad cleanup, disabling security controls, socket mounting, privileged mode, or daemon-wide mutations as troubleshooting shortcuts.
- Docker Docs — Use Compose in production — single-host production use, production-specific overrides, and service recreation.
- Docker Docs — Why use Compose? — current single-host deployment boundary and application-model use cases.
- Docker Docs — Compose services reference — healthcheck, restart, read-only rootfs, resource limits, logging, configs, secrets, stop signal and grace period.
- Docker Docs — Compose Deploy Specification — resource limits/reservations and deployment-oriented service controls.
- Docker Docs — Start containers automatically — restart-policy semantics and the successful-start threshold.
- Docker Docs — Resource constraints — CPU/memory governance and OOM implications.
- Docker Docs — Volumes — persistent data lifecycle independent of containers.
- Docker Docs — Storage — writable-layer versus volume/tmpfs durability boundaries.
- Docker Docs — Manage secrets securely in Compose — file-mounted secret access and environment-variable risk.
- Docker Docs — Compose secrets reference — top-level secret sources and service grants.
- Docker Docs — Configure logging drivers — default json-file behavior and recommendation for bounded local logging.
- Docker Docs — Local logging driver — automatic rotation, max-size/max-file, and daemon-owned log files.
- Docker Engine 29 release notes — current Engine-era baseline; 29.8.1 released 2026-09-15.
- Docker Compose releases — current Compose 5.5.1 baseline used for version-sensitive examples.
- Docker Buildx releases — current Buildx 0.37.1 baseline.
- BuildKit releases — current BuildKit 0.33.0 baseline.
- Docker Official Image — registry — current Distribution Registry 3.1.1 local-registry image used in the optional digest-pinning lab path.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.