Service Containers, Job Containers, Docker Builds, and Integration-Test Environments: Diagnostics, Failure Modes, and Production Practices
Most containerized CI failures are topology failures disguised as application failures. This lesson preserves the first run and separates DNS/port/network mistakes, health-check failures, runner/Docker capability, unsafe privilege, credential leakage, mutable images and lost ephemeral filesystem state before changing the workflow.
Learning objectives
- Diagnose localhost/DNS/port failures from the process network namespace rather than trial-and-error edits.
- Separate service health, application connectivity, runner Docker capability and image build failures.
- Recognize unnecessary port exposure, privileged containers and Docker-socket access as security decisions.
- Preserve failed run IDs, logs, service/runner evidence and external side effects before rerunning.
- Repair the smallest causal layer and avoid unsafe registry pushes from low-trust pull requests.
1. Evidence-first diagnostic sequence
Preserve the first failed run and attempt. Record the source SHA and workflow revision. Confirm the event/trust level and token permissions. Inspect job graph/queue state. Confirm the job runs on Linux and whether steps are on the host or in a job container. Inspect service health and the exact address the application used. Only then inspect Docker build output, images, artifacts or external registries.
Rerun the smallest equivalent scope only after a causal hypothesis exists. A rerun can pull a different mutable image tag or execute on a newer runner image, so “the rerun passed” is not itself diagnosis.
2. Intentionally broken example: localhost inside a job container
The following topology is valid YAML and Redis can become healthy,
but the client address is wrong. The Python process runs inside the
job container, where 127.0.0.1 points back to the
Python container itself.
jobs:
broken:
runs-on: ubuntu-24.04
container:
image: python:3.13-slim
services:
redis:
image: redis:7.4-alpine
options: >-
--health-cmd "redis-cli ping"
--health-interval 5s
--health-timeout 3s
--health-retries 10
steps:
- name: Broken network assumption
shell: sh
run: |
python - <<'PY'
import socket
socket.create_connection(('127.0.0.1', 6379), timeout=3)
PY
Expected evidence is a connection failure from the application step
while the service health setup succeeded. The repair is one field:
connect to redis:6379. Do not add host networking,
privileged mode or random port mappings.
3. Health-check failure is a different layer
If the service health command itself is wrong—for example, it probes a nonexistent command/port—the job may fail during service initialization before application steps run. Preserve the service-startup logs and inspect the configured health command. Repair the health check or image contract, not the application hostname.
Health checks should be bounded and representative. Extremely long retries hide startup regressions; overly aggressive timeouts create flaky infrastructure failures.
4. Unnecessary published ports increase attack surface and confusion
A job container does not need host port publication to reach sibling services. Publishing ports anyway creates an extra host exposure and can mislead future maintainers into using localhost. Expose only the ports required by processes outside the Docker network.
On persistent self-hosted runners, unintended published ports can also collide with local services or create network reachability beyond the job. Treat port publication as a security/configuration change.
5. Registry credentials belong outside layers and logs
Do not write registry tokens into Dockerfiles,
ENV instructions, build arguments, command history or
copied configuration files. Those values can survive in layers or
metadata. Prefer scoped registry login/federation and BuildKit
secret mechanisms for build-time secrets.
Pull requests from forks or other untrusted sources must not be allowed to use privileged registry write credentials to publish tags. Build/test locally and publish only from a trusted release path.
6. Privileged containers and Docker socket are host-adjacent capabilities
--privileged grants a broad Linux capability/device
surface. Mounting /var/run/docker.sock into a container
often gives effective control over the host Docker daemon. These are
not normal fixes for “Docker command failed.”
If a workflow truly requires Docker-in-Docker or host daemon access, isolate the runner, document the threat model and keep untrusted PR code away from that runner. Chapter 9 and Chapter 10's runner isolation rules still apply.
7. Container filesystems do not survive job boundaries
A file written inside a job container exists only in that job's workspace/container mount state. Another job normally starts on a fresh runner with a fresh filesystem. Use job outputs for small control values, artifacts for explicit cross-job files, and caches only for best-effort reusable acceleration state.
Do not “fix” a missing file in a later job by assuming Docker volumes from the earlier hosted runner remain attached.
8. Mutable tags can turn reruns into different experiments
If redis:7.4-alpine or
python:3.13-slim moves between attempts, the same
workflow text can execute different image content. In the beginner
lab we accept that convenience but record the observed identities.
Production workflows that require reproducibility should pin image
digests and review updates deliberately.
The same rule applies to executable actions: pin full SHAs rather than moving action tags.
9. Do not push an unverified latest image from a PR
A pull request tests proposed code; it is not a release
authorization. Publishing latest from PR context can
overwrite a trusted consumer reference, leak registry credentials or
create a supply-chain confusion event. Keep PR builds local or
publish into isolated, run-scoped, non-production namespaces only
when policy explicitly supports it.
10. Failure classification matrix
| Symptom | Likely layer | Evidence first | Smallest repair |
|---|---|---|---|
| Redis healthy, client connection refused in job container | network namespace/address | service health + client hostname + container topology | use service label, e.g. redis |
| Job fails before steps because service never healthy | service image/health | service startup/health logs | fix health command/image config |
docker command missing |
runner/job environment | runner OS, job container, PATH/toolchain | move build to suitable Linux host runner or install supported tooling |
| Image differs on rerun with same YAML | mutable image reference | observed image digest/ID per attempt | pin registry digest |
| Later job cannot find built file | job filesystem boundary | job/runner IDs, artifact/output declarations | publish explicit artifact or regenerate deterministically |
| PR attempts registry push | trust/permission policy | event, actor, token permissions, registry audit | remove write path from untrusted event; use trusted release workflow |
Knowledge check
Why should the broken localhost run be preserved?
It is first-failure evidence showing a healthy service plus a wrong client network assumption. Deleting or replacing it loses causal proof.
What is the smallest repair for localhost inside a job container?
Change the client host to the service label such as redis; no privileged/network-mode workaround is needed.
Why is mounting the Docker socket not a normal troubleshooting step?
It grants broad control over the host Docker daemon and expands the trust boundary dramatically.
How should a file move from one job to another?
Use an explicit artifact for file-shaped data or a job output for small values; do not assume container or runner filesystem persistence.
Why can a blind rerun be misleading with image tags?
A mutable tag can resolve to different image content on the rerun, changing the experiment while the workflow text remains the same.
Official references and version notes
- GitHub Docs — workflow syntax: job containers and services — Linux requirement, shells, volumes, options and service networking.
- GitHub Docs — Redis service containers — host versus job-container topology and health checks.
- GitHub Docs — run jobs in a container — container credentials, ports, volumes and default shell behavior.
-
docker/setup-buildx-action v4.1.0
— current optional Buildx setup action, production reference
d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5. -
docker/build-push-action v7.3.0
— current optional Docker build action, production reference
53b7df96c91f9c12dcc8a07bcb9ccacbed38856a. -
docker/login-action v4.6.0
— registry authentication action, production reference
dbcb813823bdd20940b903addbd779551569679f.
Version-sensitive behavior was rechecked on
2026-09-09. GitHub job containers, service
containers and Docker container actions require a Linux runner; on
GitHub-hosted runners that means Ubuntu. When a job runs on the
host, mapped service ports are reached through
localhost; when the job itself runs in a container,
service containers share a Docker network and are reached by their
service labels without host-port publication. The mandatory labs
target ubuntu-24.04, use public version-family images
redis:7.4-alpine and python:3.13-slim,
and record resolved image metadata rather than claiming those tags
are immutable. They use the Docker CLI already present on the
hosted Ubuntu runner and perform no registry login or push.
Optional production Buildx examples should pin the full SHAs
listed above.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.