Chapter 15Lesson 04~190 minutes

Service Containers, Job Containers, Docker Builds, and Integration-Test Environments: Diagnostics, Failure Modes, and Production Practices

Most containerized CI failures are topology failures disguised as application failures. This lesson preserves the first run and separates DNS/port/network mistakes, health-check failures, runner/Docker capability, unsafe privilege, credential leakage, mutable images and lost ephemeral filesystem state before changing the workflow.

Evidence-firstlocalhost failureHealth checksPrivilegeFirst failure

Learning objectives

  • Diagnose localhost/DNS/port failures from the process network namespace rather than trial-and-error edits.
  • Separate service health, application connectivity, runner Docker capability and image build failures.
  • Recognize unnecessary port exposure, privileged containers and Docker-socket access as security decisions.
  • Preserve failed run IDs, logs, service/runner evidence and external side effects before rerunning.
  • Repair the smallest causal layer and avoid unsafe registry pushes from low-trust pull requests.

1. Evidence-first diagnostic sequence

Preserve the first failed run and attempt. Record the source SHA and workflow revision. Confirm the event/trust level and token permissions. Inspect job graph/queue state. Confirm the job runs on Linux and whether steps are on the host or in a job container. Inspect service health and the exact address the application used. Only then inspect Docker build output, images, artifacts or external registries.

Rerun the smallest equivalent scope only after a causal hypothesis exists. A rerun can pull a different mutable image tag or execute on a newer runner image, so “the rerun passed” is not itself diagnosis.

2. Intentionally broken example: localhost inside a job container

The following topology is valid YAML and Redis can become healthy, but the client address is wrong. The Python process runs inside the job container, where 127.0.0.1 points back to the Python container itself.

jobs:
  broken:
    runs-on: ubuntu-24.04
    container:
      image: python:3.13-slim
    services:
      redis:
        image: redis:7.4-alpine
        options: >-
          --health-cmd "redis-cli ping"
          --health-interval 5s
          --health-timeout 3s
          --health-retries 10
    steps:
      - name: Broken network assumption
        shell: sh
        run: |
          python - <<'PY'
          import socket
          socket.create_connection(('127.0.0.1', 6379), timeout=3)
          PY

Expected evidence is a connection failure from the application step while the service health setup succeeded. The repair is one field: connect to redis:6379. Do not add host networking, privileged mode or random port mappings.

3. Health-check failure is a different layer

If the service health command itself is wrong—for example, it probes a nonexistent command/port—the job may fail during service initialization before application steps run. Preserve the service-startup logs and inspect the configured health command. Repair the health check or image contract, not the application hostname.

Health checks should be bounded and representative. Extremely long retries hide startup regressions; overly aggressive timeouts create flaky infrastructure failures.

4. Unnecessary published ports increase attack surface and confusion

A job container does not need host port publication to reach sibling services. Publishing ports anyway creates an extra host exposure and can mislead future maintainers into using localhost. Expose only the ports required by processes outside the Docker network.

On persistent self-hosted runners, unintended published ports can also collide with local services or create network reachability beyond the job. Treat port publication as a security/configuration change.

5. Registry credentials belong outside layers and logs

Do not write registry tokens into Dockerfiles, ENV instructions, build arguments, command history or copied configuration files. Those values can survive in layers or metadata. Prefer scoped registry login/federation and BuildKit secret mechanisms for build-time secrets.

Pull requests from forks or other untrusted sources must not be allowed to use privileged registry write credentials to publish tags. Build/test locally and publish only from a trusted release path.

6. Privileged containers and Docker socket are host-adjacent capabilities

--privileged grants a broad Linux capability/device surface. Mounting /var/run/docker.sock into a container often gives effective control over the host Docker daemon. These are not normal fixes for “Docker command failed.”

If a workflow truly requires Docker-in-Docker or host daemon access, isolate the runner, document the threat model and keep untrusted PR code away from that runner. Chapter 9 and Chapter 10's runner isolation rules still apply.

7. Container filesystems do not survive job boundaries

A file written inside a job container exists only in that job's workspace/container mount state. Another job normally starts on a fresh runner with a fresh filesystem. Use job outputs for small control values, artifacts for explicit cross-job files, and caches only for best-effort reusable acceleration state.

Do not “fix” a missing file in a later job by assuming Docker volumes from the earlier hosted runner remain attached.

8. Mutable tags can turn reruns into different experiments

If redis:7.4-alpine or python:3.13-slim moves between attempts, the same workflow text can execute different image content. In the beginner lab we accept that convenience but record the observed identities. Production workflows that require reproducibility should pin image digests and review updates deliberately.

The same rule applies to executable actions: pin full SHAs rather than moving action tags.

9. Do not push an unverified latest image from a PR

A pull request tests proposed code; it is not a release authorization. Publishing latest from PR context can overwrite a trusted consumer reference, leak registry credentials or create a supply-chain confusion event. Keep PR builds local or publish into isolated, run-scoped, non-production namespaces only when policy explicitly supports it.

10. Failure classification matrix

Symptom Likely layer Evidence first Smallest repair
Redis healthy, client connection refused in job container network namespace/address service health + client hostname + container topology use service label, e.g. redis
Job fails before steps because service never healthy service image/health service startup/health logs fix health command/image config
docker command missing runner/job environment runner OS, job container, PATH/toolchain move build to suitable Linux host runner or install supported tooling
Image differs on rerun with same YAML mutable image reference observed image digest/ID per attempt pin registry digest
Later job cannot find built file job filesystem boundary job/runner IDs, artifact/output declarations publish explicit artifact or regenerate deterministically
PR attempts registry push trust/permission policy event, actor, token permissions, registry audit remove write path from untrusted event; use trusted release workflow

Knowledge check

Why should the broken localhost run be preserved?

What is the smallest repair for localhost inside a job container?

Why is mounting the Docker socket not a normal troubleshooting step?

How should a file move from one job to another?

Why can a blind rerun be misleading with image tags?

Next lesson

Preserve failure, repair topology, then prove image identity

Lesson 5 combines the chapter into one bounded checkpoint with a deliberately wrong address and a local Docker build.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked on 2026-09-09. GitHub job containers, service containers and Docker container actions require a Linux runner; on GitHub-hosted runners that means Ubuntu. When a job runs on the host, mapped service ports are reached through localhost; when the job itself runs in a container, service containers share a Docker network and are reached by their service labels without host-port publication. The mandatory labs target ubuntu-24.04, use public version-family images redis:7.4-alpine and python:3.13-slim, and record resolved image metadata rather than claiming those tags are immutable. They use the Docker CLI already present on the hosted Ubuntu runner and perform no registry login or push. Optional production Buildx examples should pin the full SHAs listed above.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.