Chapter 36Lesson 04~195 minutes

Docker in CI/CD: Reproducible Builds, Buildx, Registry Caching, Ephemeral Runners, and Release Promotion: Diagnostics, Failure Modes, Security, and Performance

Diagnose mutable-tag races, rebuild-at-deploy drift, cache poisoning, shared-runner privilege, build-secret leakage, and false success before changing pipeline controls.

DiagnosticsMutable tagsCache poisoningRunner isolationEvidence

Learning objectives

  • Preserve first-failure evidence before rerunning or rebuilding a CI release.
  • Diagnose mutable-tag races, rebuild-at-deploy drift, cache poisoning, shared-runner privilege, and secret leakage by state layer.
  • Interpret registry and BuildKit evidence instead of declaring success from a single job status.
  • Repair with the least destructive policy or pipeline change.
  • Use a repeatable diagnostic sequence that distinguishes builder, artifact, registry, promotion and runtime state.

1. Evidence-first incident sequence

  1. Freeze the failed run ID, source SHA, workflow definition and timestamps.
  2. Capture runner/context, Docker, Buildx and builder evidence.
  3. Capture build metadata, logs, cache refs and image/index digest.
  4. Inspect registry resolution for every tag/alias involved.
  5. Inspect attestation/signature/test records bound to the subject digest.
  6. Inspect deployment/runtime digest independently.
  7. Only then change the smallest relevant layer and rerun the narrowest safe scope.

2. Failure: publishing only latest

Symptom: two incident responders pull latest at different times and discuss different bytes. Cause: the pipeline discarded immutable identity. Repair: capture the digest immediately after publication, store it as a pipeline output/evidence artifact, and treat latest only as a mutable convenience alias.

3. Failure: rebuilding at deploy time

Symptom: staging passed, production failed, even though both jobs used the same Git tag. Evidence shows two different image digests. Cause: base tags, package repositories, timestamps or build tooling changed between builds. Repair: deploy/promote the staging-tested digest. If a rebuild is required, treat it as a new release candidate.

4. Failure: cache poisoning across forks

BROKEN POLICY EXAMPLE
trusted-main writes cache ref: org/app:buildcache
fork PR receives same registry token
fork PR writes same cache ref
trusted-main later imports cache from org/app:buildcache

Evidence to preserve:
- actor/fork identity
- token scope
- cache ref + timestamps
- BuildKit logs showing import/export
- resulting digest and source SHA

Repair by separating write scopes. A trusted job may optionally read a cache populated by trusted history, while untrusted jobs write nowhere or to an isolated disposable cache namespace.

5. Failure: privileged shared runners

Symptom: one repository can see or affect another repository’s Docker state or credentials. This is not a Buildx performance issue; it is a runner/daemon isolation failure. Chapter 35’s rule applies: untrusted jobs must not receive a reusable host daemon authority. Move them to ephemeral or isolated build workers.

6. Failure: secret in build arguments

Symptom: a credential appears in image history, provenance, logs, or build metadata. Preserve the leak location and rotate/revoke the credential before editing history. Repair the Dockerfile/workflow to use BuildKit secret or SSH mounts, then rebuild a new digest and rescan all evidence surfaces.

7. Failure: race-prone mutable tags

Two jobs push the same branch tag. A downstream job pulls after the second push but assumes it received the first job’s result. Repair by passing digests between jobs, not tags. Tags may be updated after digest verification, but consumers should verify the resolved digest against the expected pipeline output.

8. Failure: build success mistaken for registry success

A build can complete while push, attestation publication, promotion or downstream pull later fails. Treat each transition as separate evidence: build metadata → push response → registry digest inspection → alias inspection → runtime pull/deploy verification.

9. Intentionally broken digest check and repair

EXPECTED='sha256:1111111111111111111111111111111111111111111111111111111111111111'
ACTUAL=$(docker image inspect dca36-app:ci --format '{{.Id}}' 2>/dev/null || true)
printf 'expected=%s
actual=%s
' "$EXPECTED" "$ACTUAL"
[ "$EXPECTED" = "$ACTUAL" ] || printf 'FAIL: subject identity mismatch; do not promote.
'

The intended failure is the mismatch itself. Do not “repair” it by changing the expected value until you know which artifact the tests actually evaluated. Reconstruct the chain from source SHA and build metadata.

10. Diagnostic decision table

Evidence mismatch Most likely layer Next safe check
source SHA correct, digest unexpected build inputs/cache/base/tooling Dockerfile/context/base digests/build logs
digest correct locally, registry tag differs registry publication/tag race registry digest resolution and push timestamps
registry digest correct, production differs deployment/reference/pull deployment config and runtime image digest
same digest, behavior differs runtime config/platform/external state platform manifest, env/config, network/storage dependencies
attestation absent output/registry/BuildKit capability driver/image store, push mode, attestation flags

11. Least-destructive corrections

  • Pin expected digest in downstream jobs.
  • Separate cache write scopes.
  • Rotate leaked credentials and rebuild once.
  • Move untrusted jobs to isolated runners/builders.
  • Re-run only the failed verification stage when artifact identity is unchanged.
  • Do not clear all caches or prune the daemon as a generic diagnostic step.

12. Preserve the negative evidence

A failed digest comparison, missing attestation, denied registry push, or mismatched runtime digest is useful evidence. Store it with timestamps and run ID. “Fixed by rerun” without preserving the original discrepancy weakens root-cause analysis.

Knowledge check

A deployment built from the same Git SHA has a different digest. Is it promotion?

A trusted build unexpectedly imports cache written by a fork. Which layer failed?

What is the first action after discovering a real secret in provenance?

Why is “the build job passed” incomplete incident evidence?

Should a digest mismatch be fixed by retagging immediately?

Next lesson

Next: Checkpoint Lab — Docker in CI/CD: Reproducible Builds, Buildx, Registry Caching, Ephemeral Runners, and Release Promotion

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Diagnostic baseline: 2026-09-22. Current Buildx/BuildKit make digest, metadata and attestation evidence easier to export, but pipeline correctness still depends on preserving identities between stages and on provider-specific runner/credential isolation.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.