Chapter 24Lesson 03200–275 min

Docker Registry Operations, Garbage Collection, Retention, Search, Proxying, and Production Troubleshooting: Configuration, Design Choices, and Tradeoffs

Design Docker registry operations deliberately. Retention, digest pinning, proxy freshness, routing, repository grouping, cleanup frequency, and rollback are coupled choices; optimize them against real delivery and recovery requirements rather than one global “keep 30 days” rule.

Retention designDigest pinningTrust zonesProxy freshnessCapacity

Learning objectives

  • Choose between mutable tags and digest pinning at development, release, and deployment boundaries.
  • Balance retention/storage cost against rollback, forensic, reproducibility, and incident-response needs.
  • Design proxy freshness and availability without confusing client cache, Nexus cache, and upstream state.
  • Choose group/trust-zone and routing patterns that do not let internal and public namespaces silently shadow each other.
  • Create a maintenance calendar that accounts for upload concurrency, GC, compaction, backup, and peak traffic.
Version baseline (27 August 2026). The chapter uses Nexus Repository 3.95.2-01 with Java 21 as its dated reference line. Docker behavior has changed substantially across releases; re-check the exact task names, known issues, routing mode, and release notes on the instance you operate.
Current 3.94.0–3.95.2 Docker cleanup warning. Sonatype documents an open issue where Docker manifest HEAD requests do not refresh an asset's lastDownloaded timestamp. An actively used image can therefore look inactive. Do not use Last Downloaded–based Docker cleanup policies on an affected instance until Sonatype publishes a fixed version.
Safety boundary. All destructive steps must target explicitly named disposable repositories and synthetic images. Never delete production tags/manifests, edit blob-store files, remove database rows, disable TLS validation, or run garbage collection/compaction on valuable repositories merely to observe behavior.

1. Start from lifecycle classes, not a universal TTL

A production registry usually contains at least four lifecycle classes: ephemeral branch/PR images, integration candidates, released application images, and base/runtime images. Their rollback and evidence value differ. A single age threshold applied to all four classes either wastes storage or destroys needed recovery history.

Write lifecycle intent first: owner, namespace, mutability, deployment criticality, rollback horizon, evidence retention, maximum storage growth, and deletion approval. Then implement repository/task controls.

2. Tag retention versus digest pinning

Reference style Strength Risk Recommended use
:latest / moving environment tag Human-friendly, easy promotion pointer. Same text can resolve to different bytes over time. Convenience only where mutability is explicitly accepted.
Immutable release tag such as :1.7.3 Readable and stable if write policy/process prevents overwrite. Still depends on registry governance to remain immutable. Release catalog and operator UX.
@sha256:... Exact manifest byte identity. Harder for humans; GC can remove digest-only content if no tag retains it. Deployment pinning/evidence, paired with retention tag/reference.

A strong release process commonly records both the immutable human version and the resolved digest. Deployment manifests should prefer the digest where the platform supports it, while registry retention ensures that digest remains reachable for the intended rollback window.

3. Storage cost versus rollback/debug value

A ten-minute rollback requirement implies that prior release bytes must still exist and be authorized. If cleanup deletes the previous release after seven quiet days, a disaster can convert a simple rollback into a rebuild—possibly from changed dependencies or unavailable source tools. Retention is therefore a reliability control, not only a storage-control feature.

For each release class, record minimum rollback count and maximum age. If current Nexus Pro/PostgreSQL retain-last-N is available, it can express a protected floor; otherwise use repository segmentation and controlled deletion automation. Never emulate Pro internals through unsupported database edits.

4. Shared layers make “sum of image sizes” a bad capacity model

Registry capacity planning should not add every manifest's layer sizes as though each image were independent. Shared base layers reduce physical storage, while temporary/incomplete uploads, soft-deleted blobs, and proxy churn add storage that may not appear in a simple tag inventory.

Measure actual blob-store growth, publish rate, retention age, proxy cache growth, GC/compaction lag, and available free-space headroom. Maintain enough headroom for upgrades, database operations, task logs, and temporary uploads.

5. Proxy availability versus freshness

A proxy improves availability and repeatability because cached layers can survive some upstream disruption. It also introduces cache freshness decisions. A mutable upstream tag may move; Nexus may have cached the previous manifest; the Docker client may also have a local copy. Production policy should prefer versioned/digest-pinned references and define when mutable tags are acceptable.

Do not globally invalidate proxy caches on every mismatch. First compare tag→digest at the client, Nexus, and upstream. Cache invalidation increases upstream load and can amplify rate-limit incidents.

6. Current 3.95.2 design constraint: do not trust Last Downloaded for Docker cleanup

The open 3.94.0–3.95.2 issue is not a minor metric cosmetic problem. Container runtimes may use manifest HEAD requests, and those requests currently fail to update lastDownloaded. A production cleanup policy that assumes “last downloaded 60 days ago” means “unused for 60 days” can delete actively used content.

On affected versions, Sonatype's documented workaround is to disable Docker cleanup policies that use Last Downloaded. Use safer criteria and explicit release protection until a fixed patch is installed and verified.

7. One group versus separate trust zones

A single group endpoint is convenient, but member ordering and namespace collisions matter. An internal image named platform/base and a public upstream image with the same name create ambiguity if both can resolve through the same group. Separate internal/public namespaces and repositories deliberately, and use content selectors/privileges/routing rules where appropriate.

Pattern Benefit Risk
One broad group Simple client configuration. Harder ownership boundaries; member order/shadowing risk.
Internal-only + public-proxy groups Clear trust boundaries and namespace policy. More client configuration/governance.
Environment groups Can separate dev/test/prod consumption. Risk of duplicating policy and drifting membership.

8. Routing mode is an operational choice, not cosmetic syntax

Path-based routing is current Sonatype's preferred Docker approach and avoids large numbers of port connectors. Port connectors reserve resources and complicate firewall/load-balancer rules. Subdomain routing can simplify references but is a Pro feature and requires DNS/TLS design. Whichever mode you use, standardize it and test clients before migration; do not run mixed modes accidentally.

9. Reverse proxy limits must match registry behavior

Large image uploads use resumable blob-upload flows and long-lived HTTP requests. Reverse proxies can break them with small body-size limits, buffering, short read/send timeouts, missing forwarded headers, or incorrect path rewriting. The correct fix is not to disable TLS or point clients around the reverse proxy. Measure the failing request, adjust the proxy for the intended upload profile, and retest with the smallest reproducible image.

10. Maintenance cadence: cleanup frequency versus IO/latency

Docker GC, blob compaction, repository cleanup, backup, database maintenance, support-ZIP generation, and heavy proxy refetches can compete for disk/database/network resources. A production calendar should avoid stacking them in the same peak window.

00:30  repository cleanup policy evaluation
01:15  Docker GC for selected repositories
02:30  file blob-store compaction (only where needed)
03:30  backup verification / snapshot coordination
business hours  avoid intensive maintenance; monitor ingest/pull latency

This is an illustrative sequence, not a universal schedule. Measure your repository size and task duration before setting recurrence.

11. Quarantine, retirement, deletion, and reclamation are different decisions

A vulnerable or superseded image may first be blocked from normal consumption while evidence is preserved. Retiring a tag removes a reference. Docker GC removes data that has become unreferenced. Blob compaction reclaims physical storage. Incident evidence/backup retention may require keeping bytes even when normal pulls are denied. Do not collapse governance, retention, and capacity reclamation into one destructive task.

12. Worked design: three repositories

Repository Content Retention Reference rule Maintenance
docker-dev Branch/PR images Short age-based lifecycle; no Last Downloaded on affected 3.95.2. Mutable branch tags allowed. Frequent cleanup; off-peak GC.
docker-release Approved release images Keep rollback floor + long age window. Version tag plus recorded digest; no overwrite. Conservative cleanup; backup aligned.
docker-public-proxy Approved upstream cache Capacity-driven cache retention. CI pins known digests for critical builds. Monitor upstream 429/freshness; avoid indiscriminate invalidation.

13. Capacity equation as a planning approximation

required_headroom ≈
  live_unique_layer_bytes
+ manifest/config/metadata
+ proxy_growth_between_cleanup_runs
+ temporary/incomplete_uploads
+ soft_deleted_bytes_waiting_for_compaction
+ backup/upgrade/task_working_space
+ safety_margin

Do not pretend this is an exact forecast. Measure unique physical storage and growth empirically because layer sharing and cache behavior are workload-dependent.

14. Knowledge check

Why should a release record both a version tag and a digest?

What is the downside of aggressively invalidating a Docker proxy cache whenever a tag looks stale?

Why may a broad Docker group be a supply-chain risk?

When is disk compaction a poor first response to “we have a vulnerable image”?

What makes Last Downloaded cleanup unsafe on the current reference line?

15. Summary and next step

Production registry design balances mutable developer workflows with immutable release identity, capacity with rollback, proxy availability with freshness, convenience with trust-zone clarity, and maintenance with IO latency. Lesson 4 applies a disciplined diagnostic sequence to the failure modes that appear when these boundaries are misconfigured.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.