Chapter 23Lesson 04~320 minutes

Container Registry, OCI Images, Cleanup, Authentication, Vulnerability Scanning, and Retention: Diagnostics, Failure Modes, Security, and Performance

Diagnose mutable-tag drift, registry authorization failures, cleanup mistakes, scan-to-image mismatches, and Self-Managed garbage-collection/storage problems using preserved source, pipeline, digest, API, and runner evidence.

Diagnostics401/403Tag driftCleanupScan identityGarbage collection

Learning objectives

  • Preserve tag/digest, source/pipeline, credential-scope, scan, and cleanup evidence before repair.
  • Diagnose 401/403, tag drift, cleanup loss, and scan-image mismatch without widening privilege.
  • Treat credential leakage as revoke/rotate-first incident response.
  • Separate project cleanup from Self-Managed registry garbage collection/storage administration.
  • Measure registry/network/image causes before making performance changes.
Availability baseline (verified 2026-08-22 against current GitLab 19.3 documentation). The integrated Container Registry is Free/Premium/Ultimate on GitLab.com, Self-Managed, and Dedicated; Self-Managed administrators must enable and operate it. Registry authentication supports job-scoped CI credentials and token credentials with registry scopes. Cleanup policies are Free across all three offerings. Container scanning is Free/Premium/Ultimate, but some vulnerability-management presentation/governance capabilities vary by tier. Protected container repositories are Free across all three offerings. Protected container tags are Free on GitLab.com and Self-Managed with the current registry backend prerequisites. Immutable container tags are Ultimate-only on GitLab.com/Self-Managed. The mandatory labs therefore require neither Ultimate nor Self-Managed administration, privileged runners, cloud spend, nor a persistent registry credential.

1. Diagnostic sequence: preserve evidence before changing registry state

Registry incidents become harder once tags are moved or deleted. Use the chapter-wide sequence: preserve evidence → identify host/project/ref/pipeline/job/runner/repository/tag/digest/deployment/policy scope → inspect permission and metadata → choose the least destructive correction → independently verify.

git rev-parse HEAD
glab ci get --output json | jq '{id,status,ref,sha} // .'
glab container-registry repository list --output json
# Then inspect the exact repository/tag with --details before mutation.

2. Failure: deployment uses a mutable tag that now points to another digest

Symptom: production says api:stable, but behavior changed without a deployment record. Preserve the runtime image ID/digest, then compare with current registry tag metadata and the release evidence.

glab container-registry tag view <repository-id> stable --output json > evidence/stable-now.json
jq '{name,digest,created_at} // .' evidence/stable-now.json

Repair: redeploy the known good digest, then change the release process to capture and deploy by digest. Do not “fix” history by force-moving Git tags or deleting registry evidence first.

3. Failure: registry login/push returns 401/403

Separate authentication from authorization. Check: registry host, token lifetime, username/token pairing, read_registry/write_registry scopes, project role, protected repository/tag rules, and whether the token was already revoked/expired.

Observation Likely layer Least-destructive check
Login denied Credential/token realm/proxy. Use current registry host, token type, expiry, and proxy header behavior.
Pull works; push denied Write scope or project/protection policy. Confirm write_registry + read_registry and role/rules.
CI push stops after job Expected short-lived credential lifecycle. Do not persist CI_REGISTRY_PASSWORD; use a proper long-lived machine identity only if necessary.
Recent protection change seems ignored JWT propagation window. Request a new auth token/client operation; do not broadly disable protection.

4. Failure: registry credential has unnecessary write scope

A deployment host that only pulls should not possess write_registry. If exposure is suspected, revoke/rotate the persistent credential first, then replace it with a read-only identity. Cleanup of logs or variables comes after revocation because deleting traces does not invalidate a credential.

Incident order: revoke/rotate → contain access → inspect registry/audit/pipeline evidence → replace with least privilege → clean leaked copies. Never print the old credential while investigating.

5. Failure: cleanup removes the rollback tag

Preserve deployment digest and registry metadata. A missing tag does not necessarily mean the manifest is already physically unavailable, but you must not rely on that accident. If the known digest remains pullable, restore a safe release reference only after authorization review. Then fix the retention rule so deployed/rollback-relevant releases are excluded from cleanup.

If storage is already reclaimed, recover from a trusted immutable artifact/registry replica or reproducibly rebuild only when policy allows—and record that rebuilt content may produce a different digest unless the build is truly reproducible.

6. Failure: vulnerability scan is associated with the wrong image

Symptom: a scan report looks clean, but it was produced for api:main while production runs another digest. Compare CS_IMAGE, pipeline SHA, report metadata/checksum, registry tag digest at scan time, and deployment digest.

Repair: scan the exact release image, persist digest-to-report mapping, and make release checks validate that identity. Never silence findings merely to make the wrong scan look acceptable.

7. Intentionally broken example: “cleanup succeeded, so storage is reclaimed”

Assume a project cleanup policy log shows tags removed, but Self-Managed disk usage is unchanged. The wrong conclusion is “cleanup is broken.” Tag deletion and blob garbage collection are separate.

# Project evidence: tags may be gone
glab container-registry tag list <repository-id> --details --output json

# Self-Managed administrator evidence (read-only concepts):
# - which registry metadata backend is active?
# - online GC queue/metrics when metadata DB is enabled
# - object-store/filesystem usage and registry logs

On current Self-Managed with the metadata database, online garbage collection processes unreferenced data asynchronously. Legacy storage has different offline procedures. Do not run third-party/offline GC against a metadata-database registry; current docs explicitly warn this can delete data associated with tagged images.

8. Self-Managed storage/garbage collection is an administrator boundary

GitLab 19.x prefers the registry metadata database for new Self-Managed installations; it enables online garbage collection and richer registry metadata. The database does not replace object storage/filesystem image data. Backups, migrations, GC health, storage consistency, and registry nodes are platform-administration concerns, not project pipeline steps.

In a production incident, project maintainers should preserve repository/tag/digest evidence and escalate storage/GC diagnostics to administrators rather than editing gitlab.rb from a CI job.

9. Performance symptoms: distinguish registry, runner, network, and image design

A slow pull can come from image size/layer churn, runner network, registry/object storage latency, authentication round trips, or cache misses. Measure before rewriting Dockerfiles or scaling runners. Large mutable images also increase scanning and transfer cost. Use layer-aware image design and digest evidence; do not trade away security controls merely to shave a few seconds.

10. High-frequency publishing creates tag/storage pressure

One tag per pipeline can be useful for traceability but creates lifecycle work. Standardize a tag namespace such as commit/review/release patterns, then write cleanup rules around that taxonomy. Keep release and deployed digests longer than ephemeral branch images. If GitLab.com cleanup runs only partially due to execution limits, verify successive runs instead of assuming one pass is exhaustive.

11. Minimum incident evidence pack

  • GitLab host, project path/ID, registry repository ID/path;
  • tag and current digest, plus historical/deployment digest if different;
  • source commit, pipeline/job IDs, runner/executor context;
  • credential type/scope/expiry identifier, never secret value;
  • scan report checksum and image reference/digest;
  • cleanup/protection policy state;
  • for Self-Managed, registry backend/GC/storage observations supplied by administrators.

Knowledge check

A pull works but push is denied. Which layers do you inspect?

Why is deleting leaked registry credentials from logs not the first response?

Cleanup deleted tags but disk usage stayed flat. What is the likely conceptual mistake?

How do you prove a scan belongs to production?

What must a project maintainer avoid during a Self-Managed GC incident?

Summary

Registry diagnostics are identity-first. Preserve the tag and digest before mutation, separate auth from authorization, keep credential incident response revoke-first, bind scans to image identity, and respect the boundary between project cleanup and Self-Managed storage garbage collection.

Official references

Primary sources used for the current GitLab 19.3 behavior taught in this lesson:

Next lesson

Checkpoint Lab

Prove one complete image lifecycle and deliberately demonstrate why tag equality is insufficient release evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.