Chapter 19Lesson 04~145 minutes

Docker Pipeline, Containerized Build Steps, Docker Agents, Sidecars, Registries, and Image Workflows: Diagnostics, Failure Modes, Security, and Performance

Diagnose Docker/Jenkins failures from first evidence across queue, trusted agent, daemon context, image/container/network, credentials, registry, and promotion state. Repair the narrowest layer without granting untrusted daemon access, hiding secrets, broad-pruning state, or rebuilding a release image.

DiagnosticsDaemon privilegeTag driftSecret leakageOrphansEvidence

Learning objectives

  • Preserve first-failure Jenkins, agent, daemon, container, network, and registry evidence.
  • Diagnose untrusted daemon access as a trust-design failure.
  • Distinguish mutable tag drift from immutable digest identity.
  • Repair registry secret exposure and orphaned resources without broad destructive shortcuts.
  • Diagnose performance using Docker/registry measurements before changing Jenkins concurrency.

1. Evidence-first sequence

  1. Preserve item/build/source/cause and first-failure timestamp.
  2. Confirm Jenkins core/Java/Docker Pipeline plugin baseline.
  3. Confirm queue item, selected node/labels, workspace, and trust class.
  4. Confirm Docker client/server/context/security mode.
  5. Capture input image digest, container/sidecar/network IDs, and logs.
  6. Confirm registry endpoint and credential ID/reference without printing the value.
  7. Capture push/pull response and registry digest if created.
  8. Check promotion/external target separately from Jenkins step result.
  9. Apply one narrow repair and retry only an idempotent scope.

2. Causal failure layers

Causal failure layers
flowchart TD
  A[Build/source evidence] --> B{Correct trusted Docker agent?}
  B -->|no| C[Queue / label / trust repair]
  B -->|yes| D{Expected Docker context?}
  D -->|no| E[Endpoint/context repair]
  D -->|yes| F{Pinned container starts?}
  F -->|no| G[Pull/runtime/workspace evidence]
  F -->|yes| H{Sidecar/network ready?}
  H -->|no| I[Readiness/log/network repair]
  H -->|yes| J{Registry auth/push succeeds?}
  J -->|no| K[Credential/repository/network repair]
  J -->|yes| L[Verify digest + promotion identity]

3. Failure: untrusted code controls the Docker daemon

Symptom: a fork/PR Jenkinsfile runs on the same daemon-capable worker used for trusted image publishing.

Diagnosis: this is a scheduler/trust-boundary defect, not merely a container configuration issue. Preserve the build/source/agent-label evidence and stop further untrusted scheduling to that pool through an authorized change.

Repair: move untrusted builds to constrained workers/builders without privileged host-daemon or release-registry authority. Keep trusted build/push after review on the protected pool. Do not rely on masking, shell quoting, or “container isolation” to contain arbitrary daemon commands.

4. Failure: tag drift

Suppose build 88 recorded only python:3.13-alpine3.24. The tag later resolves to different content and a rerun changes behavior. Preserve the current tag resolution plus the old build record. If the old build never captured a digest, state that exact environment reproduction is incomplete rather than inventing a digest.

For release images where the original digest exists, use that digest. Do not rebuild simply because the tag changed.

5. Failure: registry secret appears in retained output

Credentials Binding masking is a log-safety aid, not data-loss prevention. Secrets can leak through transformations, shell tracing, Docker auth files, debug output, archives, or sidecars.

Real exposure response: preserve non-secret evidence, revoke/rotate through the provider, remove unsafe retained copies through an authorized incident process, fix the Pipeline, and verify the replacement credential. Deleting a Jenkins build alone does not prove the credential is safe.

Use --password-stdin or a Jenkins registry wrapper, short binding scope, and a private Docker client config where shared-worker persistence matters.

6. Failure: aborted build leaves sidecar/network

Capture the exact container/network IDs, names, owner build, start time, and logs. Remove only the resources owned by that build. Fix the Jenkinsfile with try/finally or post { always { ... } } cleanup using unique names.

Do not use broad prune commands on a shared worker. They can delete other jobs’ state and erase first-failure evidence.

7. Failure: “it was in a container” is mistaken for host isolation

Inspect the actual container user, mounts, capabilities, devices, network mode, and whether it can reach Docker/host-management APIs. Default container isolation can be useful, but host/daemon access or dangerous mounts can cross the intended boundary.

8. Failure: promotion rebuilds the release

CI tests digest D1. A release job later runs docker build again from the same source and deploys D2. That is a different artifact even if the Git SHA is identical. Preserve D1 and the original producer build, then repair promotion so it consumes repository@D1 without rebuilding.

9. Failure: remote Docker API works but inside cannot see workspace

Current Jenkins Docker documentation warns that inside expects compatible filesystem visibility between the Jenkins agent and Docker server. A remote daemon can be reachable while the mounted Jenkins workspace path is meaningless on the daemon host.

Preserve the Docker endpoint, workspace path, and durable-task error. Use a builder/execution design that deliberately transfers context, or pair the Docker Pipeline agent and daemon with shared filesystem semantics. Do not rewrite paths blindly.

10. Measure before tuning

Symptom Measure first Bad shortcut
Slow stage start Image pull/cache and registry latency Add executors
Slow build CPU, disk I/O, builder activity Assume Groovy is slow
Slow push Layer sizes, upload bandwidth, registry response Blind retry
Disk pressure Images/layers/volumes by owner Global prune
Queue growth Executor use plus Docker host saturation Unlimited executor count

11. Smallest-safe repairs

  • Wrong Docker context → select intended context, not restart Jenkins.
  • Tag drift → use recorded digest, not rebuild.
  • Registry denied → repair scoped credential/repository authorization, not grant admin.
  • Orphan sidecar → remove exact build-owned ID and fix cleanup, not global prune.
  • Untrusted daemon access → change worker trust/scheduling, not shell escaping.
  • Remote workspace mismatch → change execution architecture, not disable durable-task safeguards.
Next lesson

Checkpoint Lab

Build once, test with a pinned container and sidecar, push to an authenticated disposable registry, promote by digest in a simulation, remove transient state, and prove exact attribution survives.

Knowledge check

Answer before revealing the explanation.

1. A fork PR can control the trusted Docker daemon. What failed?

2. A tag now points to a different digest. Should release promotion rebuild?

3. A registry password appears in an archived file. Is masking sufficient?

4. A sidecar remains after abort. What is the safe repair?

5. Why can remote Docker work while inside fails?

Official references and version notes

Assumption timestamp: 2026-09-17. Recheck Jenkins LTS/Java, Docker Pipeline health/version/dependencies, Docker Engine security/release notes, and all image tags/digests before repeating later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.