Chapter 29Lesson 04~145 minutes

Image Vulnerability Scanning, Dependency Hygiene, Base-Image Maintenance, and Remediation Workflows: Diagnostics, Failure Modes, Security, and Performance

Diagnose tag drift, stale advisory data, runtime hot-fixes, severity-only triage, unowned suppressions, and false confidence without destroying first-failure evidence.

DiagnosticsTag driftDatabase freshnessFalse positivesEvidence

Learning objectives

  • Diagnose identity drift between a mutable tag and the artifact actually scanned.
  • Recognize stale database, severity-only, hot-fix, and unowned-suppression failure modes.
  • Preserve first-failure evidence before rebuilding or changing scanner policy.
  • Apply the least destructive correction to the causal layer.

1. Diagnostic order: do not start by suppressing the CVE

When scan output looks wrong, first preserve the report and artifact identity. Then confirm Docker context/version, exact image ID/digest/platform, scanner version/database state, package inventory, advisory record, fix availability, and only then runtime exploitability/context. Policy changes or suppressions come last because they can hide the very evidence needed to diagnose the mismatch.

Layer Evidence Typical failure
Client/context docker context show, version Scanning/building against an unexpected daemon or local store
Artifact identity image ID, RepoDigest, platform Mutable tag moved; wrong architecture variant
Scanner/data tool version, DB timestamp/status Stale DB, changed matching logic, partial analyzer support
Inventory package/version/ecosystem Package misidentified or package absent from expected layer
Advisory CVE, status, fix version, source Severity/fix status changed upstream
Runtime context listener, feature use, privileges, exposure Severity interpreted without applicability/exposure
Remediation source Dockerfile/base/lockfile Hot-fix in container instead of source rebuild

2. Broken example: the tag moved after the scan

This safe example creates two local images and deliberately moves one human-friendly tag. It demonstrates why a report named only after the tag is ambiguous.

mkdir -p da-ch29-drift && cd da-ch29-drift
printf 'FROM alpine:3.21.0
CMD ["sh","-c","echo A"]
' > Dockerfile.a
printf 'FROM alpine:3.24.2
CMD ["sh","-c","echo B"]
' > Dockerfile.b

docker build -t da-ch29:a -f Dockerfile.a .
docker tag da-ch29:a da-ch29:candidate
A_ID=$(docker image inspect da-ch29:candidate --format '{{.Id}}')
trivy image --format json -o scan-candidate.json da-ch29:candidate

docker build -t da-ch29:b -f Dockerfile.b .
docker tag da-ch29:b da-ch29:candidate
B_ID=$(docker image inspect da-ch29:candidate --format '{{.Id}}')
printf 'scanned-id=%s
current-tag-id=%s
' "$A_ID" "$B_ID"

The report remains valid for the bytes it scanned, but the label “candidate” now points elsewhere. The correction is not to rescan blindly and overwrite the report. Preserve A_ID, preserve the original JSON, record that the tag moved, and scan/compare the intended immutable identity or rebuild artifact explicitly.

3. Failure mode: patching the running container

Running apk upgrade or apt upgrade interactively inside an existing container may make that one process environment look different, but the release image, Dockerfile, registry digest, CI evidence, and future replicas remain unchanged. It also destroys reproducibility because the repair exists only in a writable layer.

Correct the declared base/package/dependency source, rebuild a new image, run tests, rescan, then promote the new digest. Keep the old digest and report so the before/after evidence is reviewable.

4. Failure mode: severity-only decisions

Severity is useful sorting data, not a complete risk model. A common anti-pattern is “ignore everything below High; fail everything above High” with no fixability, exposure, or exception ownership. That can create noisy gates, encourage blanket suppression, and miss lower-severity issues with real exploit paths.

Use severity as one input. Preserve advisory status/fix version and add runtime context. If policy is intentionally simple, document that limitation rather than presenting the threshold as a universal measure of risk.

5. Failure mode: stale vulnerability data

If the scanner database is old, a zero-finding report can be dangerously misleading. Capture the scanner/database metadata before trusting the result. For Grype, grype db status exposes status/build age and the tool normally validates database age; for Trivy, the database is fetched/maintained automatically unless update behavior is changed.

trivy --version
# Optional comparison if Grype is installed:
grype version
grype db status
Do not “fix” freshness by disabling validation. In controlled offline environments, mirror/import a reviewed database artifact and record its identity/timestamp instead.

6. Failure mode: permanent suppressions

An ignore entry without an owner and expiry is a deferred vulnerability with no revisit mechanism. Every exception should record the exact digest/platform, scanner/advisory identity, package/version, rationale, mitigation, owner, approval, created date, and expiry/review date. If the image digest changes, re-evaluate rather than automatically copying the exception to new bytes.

7. Failure mode: “zero findings means secure”

A scanner can only match what it inventories against what its data knows. It does not prove secure configuration, correct authorization, safe application logic, absence of secrets, provenance, signature validity, runtime least privilege, or network exposure. Chapter 30 will add SBOM/provenance/signature evidence, but even those signals remain distinct.

8. Exact cleanup and evidence retention

Keep the JSON report and identity values long enough to finish diagnosis. Remove only the three lab tags after the comparison.

docker image rm da-ch29:candidate da-ch29:a da-ch29:b
# Remove only the da-ch29-drift directory after preserving any report you need.

If removal reports that a shared image reference still exists, inspect exact references rather than force-deleting unrelated content.

Knowledge check

A scan report says candidate, but the tag now points to a new image ID. Is the old report invalid?

Why is an interactive package upgrade inside the container not release remediation?

What should you do before changing an ignore policy?

What does a stale scanner database undermine?

Why must exceptions be re-evaluated when the digest changes?

Next lesson

Next: Checkpoint Lab — Image Vulnerability Scanning, Dependency Hygiene, Base-Image Maintenance, and Remediation Workflows

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version/platform baseline, verified 2026-09-22.

The intentionally broken tag-drift lab uses only disposable local tags and retains the original JSON instead of overwriting it. Current Trivy/Grype database behavior is treated as version-sensitive evidence. The chapter never uses runtime package mutation, blanket suppression, stale-data bypass, or broad Docker cleanup as a troubleshooting shortcut.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.