Image Vulnerability Scanning, Dependency Hygiene, Base-Image Maintenance, and Remediation Workflows: Diagnostics, Failure Modes, Security, and Performance
Diagnose tag drift, stale advisory data, runtime hot-fixes, severity-only triage, unowned suppressions, and false confidence without destroying first-failure evidence.
Learning objectives
- Diagnose identity drift between a mutable tag and the artifact actually scanned.
- Recognize stale database, severity-only, hot-fix, and unowned-suppression failure modes.
- Preserve first-failure evidence before rebuilding or changing scanner policy.
- Apply the least destructive correction to the causal layer.
1. Diagnostic order: do not start by suppressing the CVE
When scan output looks wrong, first preserve the report and artifact identity. Then confirm Docker context/version, exact image ID/digest/platform, scanner version/database state, package inventory, advisory record, fix availability, and only then runtime exploitability/context. Policy changes or suppressions come last because they can hide the very evidence needed to diagnose the mismatch.
| Layer | Evidence | Typical failure |
|---|---|---|
| Client/context | docker context show, version |
Scanning/building against an unexpected daemon or local store |
| Artifact identity | image ID, RepoDigest, platform | Mutable tag moved; wrong architecture variant |
| Scanner/data | tool version, DB timestamp/status | Stale DB, changed matching logic, partial analyzer support |
| Inventory | package/version/ecosystem | Package misidentified or package absent from expected layer |
| Advisory | CVE, status, fix version, source | Severity/fix status changed upstream |
| Runtime context | listener, feature use, privileges, exposure | Severity interpreted without applicability/exposure |
| Remediation source | Dockerfile/base/lockfile | Hot-fix in container instead of source rebuild |
2. Broken example: the tag moved after the scan
This safe example creates two local images and deliberately moves one human-friendly tag. It demonstrates why a report named only after the tag is ambiguous.
mkdir -p da-ch29-drift && cd da-ch29-drift
printf 'FROM alpine:3.21.0
CMD ["sh","-c","echo A"]
' > Dockerfile.a
printf 'FROM alpine:3.24.2
CMD ["sh","-c","echo B"]
' > Dockerfile.b
docker build -t da-ch29:a -f Dockerfile.a .
docker tag da-ch29:a da-ch29:candidate
A_ID=$(docker image inspect da-ch29:candidate --format '{{.Id}}')
trivy image --format json -o scan-candidate.json da-ch29:candidate
docker build -t da-ch29:b -f Dockerfile.b .
docker tag da-ch29:b da-ch29:candidate
B_ID=$(docker image inspect da-ch29:candidate --format '{{.Id}}')
printf 'scanned-id=%s
current-tag-id=%s
' "$A_ID" "$B_ID"
The report remains valid for the bytes it scanned, but the label
“candidate” now points elsewhere. The correction is not to rescan
blindly and overwrite the report. Preserve A_ID,
preserve the original JSON, record that the tag moved, and
scan/compare the intended immutable identity or rebuild artifact
explicitly.
3. Failure mode: patching the running container
Running apk upgrade or
apt upgrade interactively inside an existing container
may make that one process environment look different, but the
release image, Dockerfile, registry digest, CI evidence, and future
replicas remain unchanged. It also destroys reproducibility because
the repair exists only in a writable layer.
Correct the declared base/package/dependency source, rebuild a new image, run tests, rescan, then promote the new digest. Keep the old digest and report so the before/after evidence is reviewable.
4. Failure mode: severity-only decisions
Severity is useful sorting data, not a complete risk model. A common anti-pattern is “ignore everything below High; fail everything above High” with no fixability, exposure, or exception ownership. That can create noisy gates, encourage blanket suppression, and miss lower-severity issues with real exploit paths.
Use severity as one input. Preserve advisory status/fix version and add runtime context. If policy is intentionally simple, document that limitation rather than presenting the threshold as a universal measure of risk.
5. Failure mode: stale vulnerability data
If the scanner database is old, a zero-finding report can be
dangerously misleading. Capture the scanner/database metadata before
trusting the result. For Grype, grype db status exposes
status/build age and the tool normally validates database age; for
Trivy, the database is fetched/maintained automatically unless
update behavior is changed.
trivy --version
# Optional comparison if Grype is installed:
grype version
grype db status
6. Failure mode: permanent suppressions
An ignore entry without an owner and expiry is a deferred vulnerability with no revisit mechanism. Every exception should record the exact digest/platform, scanner/advisory identity, package/version, rationale, mitigation, owner, approval, created date, and expiry/review date. If the image digest changes, re-evaluate rather than automatically copying the exception to new bytes.
7. Failure mode: “zero findings means secure”
A scanner can only match what it inventories against what its data knows. It does not prove secure configuration, correct authorization, safe application logic, absence of secrets, provenance, signature validity, runtime least privilege, or network exposure. Chapter 30 will add SBOM/provenance/signature evidence, but even those signals remain distinct.
8. Exact cleanup and evidence retention
Keep the JSON report and identity values long enough to finish diagnosis. Remove only the three lab tags after the comparison.
docker image rm da-ch29:candidate da-ch29:a da-ch29:b
# Remove only the da-ch29-drift directory after preserving any report you need.
If removal reports that a shared image reference still exists, inspect exact references rather than force-deleting unrelated content.
Knowledge check
A scan report says candidate, but the tag now
points to a new image ID. Is the old report invalid?
Not necessarily. It is evidence for the artifact originally scanned, but the tag label is now ambiguous. Recover/preserve the scanned immutable identity and bind the report to it.
Why is an interactive package upgrade inside the container not release remediation?
It changes only that container writable layer. Source, image digest, future replicas, and CI evidence remain unchanged.
What should you do before changing an ignore policy?
Preserve the original report, image identity, scanner/database evidence, package/advisory details, and context. Then diagnose the causal mismatch.
What does a stale scanner database undermine?
The completeness/timeliness of advisory matching. It can make known vulnerabilities appear absent.
Why must exceptions be re-evaluated when the digest changes?
The package inventory and runtime content may differ, so applicability and residual risk are not automatically transferable to new bytes.
Official references and version notes
- Docker Engine 29 release notes — current Engine/CLI behavior; Engine 29.8.1 is the current release baseline used in this chapter.
-
Docker build best practices
— rebuild cadence,
--pull, digest pinning, and auditable base-image updates. - Docker Scout quickstart — package/vulnerability analysis and remediation workflow; used here as an optional Docker-native comparison path.
- docker scout cves reference — supported artifact types and local/registry resolution semantics.
- Trivy installation — current free local scanner installation; current documented release is 0.74.0.
- Trivy databases — vulnerability-database retrieval, update controls, and cache behavior.
- Trivy image reference — image scanning, JSON output, severity filters, and fixed/unfixed filtering.
- Grype vulnerability database — database freshness, update behavior, and age validation for the optional Grype path.
- Grype releases — current Grype release evidence; v0.119.0 was released 2026-09-17.
- Alpine release branches — support windows for the exact base-image family used in the disposable lab.
The intentionally broken tag-drift lab uses only disposable local tags and retains the original JSON instead of overwriting it. Current Trivy/Grype database behavior is treated as version-sensitive evidence. The chapter never uses runtime package mutation, blanket suppression, stale-data bypass, or broad Docker cleanup as a troubleshooting shortcut.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.