Chapter 26Lesson 04200–280 min

Containers, Reproducible Runtimes, and Distributed Test Environments: Diagnostics, Failure Modes, and Production Practices

Diagnose container-specific Robot Framework failures without hiding evidence: ownership errors, missing dependencies, localhost mistakes, path coupling, secret leakage, resource pressure, and disappearing results.

DiagnosticsUID/GIDlocalhost boundarySecrets & layersCPU/memory

Learning objectives

  • Classify failures as source/build, image, runtime, filesystem, network, dependency, capacity, external-system, or result-retention problems.
  • Preserve stopped containers and first-failure artifacts before applying a correction.
  • Diagnose root-owned results, wrong localhost targets, missing OS dependencies, floating-image drift and resource-induced timeouts.
  • Recognize container security anti-patterns such as baked secrets, Docker socket mounts, public test services and writable source trees.
  • Repair the smallest failing layer without blanket retries, giant timeouts, chmod 777, or deleting evidence.

Current compatibility baseline — verified 2026-08-31. Robot Framework 7.4.2 is the stable course baseline and requires Python 3.8+. The mandatory container lab uses the Docker Official Image python:3.12.14-slim-bookworm and installs the Robot Framework 7.4.2 universal wheel by exact version and SHA-256 hash. Docker image tags are mutable, so the lesson records the resolved image digest and explains full digest pinning for controlled pipelines. Pabot 5.2.2 is discussed as the current stable optional parallel executor; 5.3.0b1 is prerelease and is not required. No browser library, external API/database/SSH service, CI provider, Kubernetes cluster, paid platform, real credential, or production target is required.

1. Diagnostic sequence: preserve before you “clean”

  1. Preserve first-failure artifacts. Keep the stopped container, console output, Robot result directory and any service logs.
  2. Record versions/provenance. Robot/Python, image ID/digest, dependency manifest, Docker client/server.
  3. Confirm what actually executed. image, command/entrypoint, tests path, variables/tags, working directory.
  4. Inspect mounts and user. Compare Robot --outputdir with mount destinations and UID/GID permissions.
  5. Inspect network topology. Identify caller namespace and target address before changing ports/network mode.
  6. Inspect dependency/system state. Required CA certificates, browser binaries, OS packages or library versions.
  7. Inspect CPU/memory and parallelism. Container limits, OOM state, Pabot/CI fan-out and external contention.
  8. Apply the least destructive correction. Fix the path, ownership, dependency pin or endpoint.
  9. Rerun the smallest controlled slice. Preserve the old evidence; do not overwrite the only failing output.

2. Intentionally broken container contract — diagnose it before fixing it

# DO NOT ADOPT — diagnostic example
FROM python:3.12-slim
WORKDIR /work
COPY . /work
RUN pip install robotframework
ENV API_TOKEN=token-FAKE_DO_NOT_USE
CMD ["robot", "--outputdir", "/work/results", "tests"]

This file is parser-valid but operationally weak. The base and Robot dependency float. COPY . can capture unrelated files. Execution defaults to root. The fake token illustrates that ENV becomes image configuration/layer-visible metadata; real secrets must never be baked this way. Results are inside the container unless /work/results is deliberately preserved.

The repair is not a single magic flag. Use an exact/digest-recorded base, hash-pinned Python dependencies, narrow COPY, non-root USER, runtime secret injection from an approved store when needed, and an explicit result volume/mount.

3. Failure: root-owned or unwritable result files

Symptom A: host artifacts exist but cannot be deleted/edited by the developer. On native Linux, a root container wrote through a bind mount as UID 0.

Symptom B: Robot cannot create output.xml in a bind-mounted result directory. A non-root container UID lacks write permission.

docker inspect rf26-run --format '{{.Config.User}}'
docker inspect rf26-run --format '{{json .Mounts}}'
# Native Linux host evidence:
ls -ln results-container

Repair ownership deliberately: choose a fixed runtime UID/GID and prepare the destination accordingly, or override to the current host UID/GID for a development bind mount. Do not use chmod -R 777 as a generic workaround; it destroys the permission boundary you need to understand.

4. Failure: container localhost points to itself

Symptom: a host-side fixture responds to curl http://127.0.0.1:8000 on the host, but the Robot container gets connection refused.

Inside the test container, 127.0.0.1 means the test container. If the fixture is another Compose service, use its service DNS name. If the fixture intentionally runs on Docker Desktop's host, use the documented host.docker.internal path. On Linux Engine outside Desktop, host reachability has different setup and should be configured explicitly rather than assumed.

Do not respond by publishing the fixture on every interface. A local integration dependency should remain private unless external reachability is required.

5. Failure: the Python image does not contain your domain/system dependencies

A slim Python image contains Python and a minimal operating-system userland; it does not automatically include browsers, browser drivers, CA bundles beyond the base distribution's contents, database client libraries, SSH servers, fonts, compilers or arbitrary command-line tools. A suite that worked on a developer workstation can fail because those implicit host dependencies were never declared.

Repair the image manifest, not the running container. If a Chapter 16 Browser/Selenium suite is containerized, pin the browser/library-compatible image/dependencies documented for that stack. If TLS fails because the organization uses a private CA, add the approved CA through a controlled build/runtime mechanism; do not disable certificate verification.

6. Failure: tests depend on a host-only path or writable source tree

Symptom: C:\Users\..., /home/alice/..., or ${EXECDIR}-relative assumptions work locally but fail in the image. The container filesystem has a different root and working directory.

Use repository-relative imports and Robot automatic path variables intentionally, and pass runtime fixture/output roots as explicit variables when they are environmental. Mount only the paths the container needs. A test should not discover the host home directory by accident.

7. Failure: secrets baked into image layers or leaked through test logs

Do not place real credentials in Dockerfile ENV, ARG, copied config files, image labels, shell history or test source. Image history and registries can outlive the run. Runtime secret injection must come from the environment/secret mechanism approved by the execution platform, and Robot/library logging must still avoid exposing the value.

Robot Framework 7.4 Secret values can reduce Robot-side representation leaks, but a container does not make a leaked secret safe, and Secret is not encryption. Keep credentials fake in this chapter.

8. Failure: CPU/memory limits look like flaky Robot timing

Container limits can cause slow keywords, browser crashes, killed Python processes or intermittent timeouts. Inspect runtime limits and termination state before increasing Robot timeouts.

docker inspect rf26-run --format '{{json .HostConfig.Memory}} {{json .HostConfig.NanoCpus}}'
docker inspect rf26-run --format '{{json .State}}'
docker stats --no-stream rf26-run

Docker documents that containers have no CPU/memory limits by default unless configured. In CI, however, the runner itself may be constrained. A Pabot run can multiply pressure inside the same cgroup. Fix under-provisioning or concurrency first; a larger Robot timeout only hides the causal layer.

9. Failure: results vanish when the container is removed

Symptom: the console showed a Robot failure, the command used docker run --rm, and there is no output.xml afterward.

The output lived only in the container writable layer. Reproduce once with a retained container or a correctly mounted result location, then preserve the evidence before cleanup. In CI, artifact upload must run on failure as well as success; container exit status and result retention are separate pipeline responsibilities.

10. Failure: a test-only Remote/browser/service endpoint is publicly reachable

Port publishing without a host address may bind broadly. Keep test-only dependencies on private container networks and publish only what the host must access. Never expose an unauthenticated Remote library server or browser control endpoint to a public interface as a debugging shortcut.

11. Failure: yesterday’s Dockerfile rebuild behaves differently

Floating base tags and unpinned pip installs are dependencies that changed outside the repository. Compare the previously recorded image/base digest and package manifest with the new build. Repair by pinning/recording dependencies and treating updates as reviewed changes. Do not keep rerunning until the registry happens to return something that passes.

12. Container troubleshooting shortcuts to reject

  • Do not switch to --privileged or mount /var/run/docker.sock to solve ordinary test access problems.
  • Do not make the entire repository writable merely because one fixture needs a scratch directory.
  • Do not disable TLS/SSH verification to make a container connect.
  • Do not publish all service ports or use host networking to paper over an address mistake.
  • Do not run as root or chmod 777 because result ownership is inconvenient.
  • Do not delete the failed container/result directory before inspection.
  • Do not add blanket retries or giant timeouts for CPU/memory contention.
  • Do not prune all Docker volumes/images as a “cleanup” step in a shared developer/CI host.

Knowledge check

A container run passed, but the mounted host result directory is empty. Is that necessarily a Robot failure?

Why is --privileged a poor fix for a result permission error?

What should you inspect before increasing a timeout after moving a suite into a constrained container?

Why is an unauthenticated Remote library on a published public port especially dangerous?

13. Summary and bridge

Container diagnostics becomes manageable when you classify failures by layer and preserve the stopped runtime before cleanup. The checkpoint lab now applies that discipline to a deliberately wrong result-volume destination and proves that a green Robot run can still violate the evidence contract.

Next lesson

Checkpoint Lab — Containers, Reproducible Runtimes, and Distributed Test Environments

Continue with Checkpoint Lab — Containers, Reproducible Runtimes, and Distributed Test Environments. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

References and version anchors

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.