Docker Registry Operations, Garbage Collection, Retention, Search, Proxying, and Production Troubleshooting: Diagnostics, Failure Modes, Security, and Performance
Troubleshoot Docker registry failures by preserving the exact image reference and HTTP evidence, then separating client cache, routing/TLS, bearer-token authorization, Nexus repository metadata, proxy/upstream behavior, and blob/database/task state. Fix the smallest verified cause.
Learning objectives
- Interpret common Docker client HTTP failures without immediately changing credentials, caches, or TLS controls.
- Diagnose stale tags by comparing client, Nexus, and upstream digest evidence.
- Separate reverse-proxy upload limits from Nexus repository/blob/database failures.
- Explain why cleanup during active publishing or manual blob deletion can create inconsistency/corruption risk.
- Use task logs, request evidence, disk/blob metrics, and controlled pulls to verify the least-destructive correction.
HEAD requests do not refresh an asset's
lastDownloaded timestamp. An actively used image can
therefore look inactive. Do not use Last Downloaded–based Docker
cleanup policies on an affected instance until Sonatype publishes a
fixed version.
1. Diagnostic sequence for Docker incidents
- Preserve the exact command/image reference, time, client version, and concise error.
- Confirm Nexus version/edition/runtime/database/blob-store type.
- Resolve DNS/TLS/routing mode and reverse-proxy path.
- Inspect Docker Bearer Token Realm, authentication challenge, token acquisition, and repository privileges.
- Identify hosted/proxy/group and member order.
- Compare tag → manifest digest in client, Nexus, and upstream where applicable.
- Inspect component/asset state and Docker GC/cleanup task history.
- Inspect proxy cache/remote health/upstream rate-limit headers.
- Inspect database/blob/disk and reverse-proxy upload limits.
- Apply one least-destructive correction and repeat a controlled pull/push.
2. HTTP 401: challenge or actual denial?
The initial Registry V2 request commonly returns
401 Unauthorized with a bearer challenge. A functioning
Docker client follows that challenge to obtain a token.
Troubleshooting must distinguish “normal challenge” from “token
request failed” and “token obtained but repository action denied.”
HTTP/1.1 401 Unauthorized
WWW-Authenticate: Bearer realm="https://repo.example.invalid/service/rest/v1/security/token",service="repo.example.invalid",scope="repository:ch24-docker-hosted:pull"
Redact real token values. Check Docker Bearer Token Realm, service URL/reverse-proxy headers, credential store, user/role privileges, and content-selector path semantics. Do not solve a 401 by granting administrator rights to CI.
3. HTTP 404: missing name/tag, wrong route, or policy-visible absence
A 404 can mean the tag genuinely does not exist, the client addressed the wrong repository path, a group member order/privilege hides the expected component, the proxy has no such remote content, or a stale metadata/reference was removed during cleanup. Verify the repository path and search exact image name/tag in Nexus before deleting caches.
curl -fsS "$NEXUS_URL/service/rest/v1/search?repository=$REPO&format=docker&docker.imageName=$IMAGE&docker.imageTag=$TAG"
4. HTTP 429: locate the rate limiter
429 Too Many Requests can be returned by an upstream
public registry, a reverse proxy/WAF, or another service in the
request path. Preserve Retry-After and vendor
rate-limit headers when present. If Nexus has the requested digest
fully cached, a pull may still succeed without upstream access; an
uncached tag/layer may fail.
Corrective patterns include authenticated upstream access, approved mirrors, fewer mutable-tag lookups, digest pinning, avoiding redundant CI pulls, and sizing cache retention. Do not add random public credentials or bypass Nexus policy.
5. Stale proxy tag: compare digests before invalidating
When upstream/app:stable moved from digest A to B, four
states may exist:
- Upstream tag = B.
- Nexus proxy cached tag = A.
- Docker client local image = A.
- Deployment manifest may be pinned to A intentionally.
Collect each state. If production is intentionally pinned to A, the “stale” report may be wrong. If the mutable tag should refresh, investigate proxy cache metadata/remote request evidence. Cache invalidation is a targeted recovery tool, not the first command.
6. Manifest/layer mismatch: preserve the failing digest
Errors such as unknown blob, manifest invalid, or digest mismatch require exact digest evidence. Do not delete all blobs. Record the requested manifest digest, layer digest, repository, task history, and whether the content was recently cleaned/imported/migrated. Modern Nexus releases include many Docker integrity and GC fixes; old-version folklore is unsafe on a current database.
If a layer is missing but metadata references it, escalate through supported recovery/data-repair guidance for the exact version. Chapter 25 will cover backup/restore; Chapter 26 covers migration. Do not fabricate the missing blob or edit database references manually.
7. Large push fails: inspect reverse-proxy limits before Nexus internals
Typical symptoms include 413 Request Entity Too Large,
502/504, interrupted resumable uploads, or client retry
loops. Compare reverse-proxy body-size limits, request buffering,
read/send timeouts, TLS termination, and upstream connection
timeouts against the failing layer size/duration.
# Illustrative reverse-proxy concepts; adapt to your approved baseline.
client_max_body_size 0;
proxy_request_buffering off;
proxy_buffering off;
proxy_read_timeout 300s;
proxy_send_timeout 300s;
These values are examples, not a blind production configuration. Measure the workload and preserve security controls. Partial uploads can leave temporary blob data; use the supported Docker - Delete incomplete uploads task after the underlying transport problem is fixed.
8. Cleanup during active publish: concurrency matters
Garbage collection, repository cleanup, large pushes, and compaction all consume IO/database resources and reason about references at specific points in time. Running aggressive maintenance during peak publication increases latency and complicates incident analysis. Current Nexus includes safeguards/fixes, but an operator should still schedule intensive maintenance deliberately and monitor task status.
If a task is already running, do not repeatedly start duplicate copies. Sonatype has specifically hardened Docker GC concurrency in modern releases; respect task status and repository scope.
9. Intentionally broken example: “rm the layer file to save disk”
# WRONG — do not run.
# rm -f /nexus-data/blobs/default/content/vol-*/chap-*/sha256-some-layer.bytes
This bypasses repository metadata and reference tracking. A surviving manifest may still point to the deleted layer, creating future pull failures and difficult recovery. The supported path is retention/deletion → Docker GC → compact/reclaim, with backup/recovery context. The diagnosis for “disk is full” is not “delete the largest .bytes file.”
10. Intentionally broken example: Last Downloaded cleanup on 3.95.2
Suppose prod/base:2026.08 is pulled daily, but the
runtime checks the manifest using HEAD. On Nexus 3.95.2
the open issue can leave lastDownloaded stale. A 30-day
inactivity policy may mark the active image for deletion.
Repair: disable the affected Last Downloaded Docker cleanup policy as Sonatype directs, preserve the active image, use safer retention evidence, and plan upgrade/verification when a fixed patch becomes available. Do not “fix” the timestamp directly in the database.
11. Performance diagnosis: identify the constrained stage
| Symptom | Likely evidence | Do not assume |
|---|---|---|
| Slow first pull, fast second pull | Upstream/network/proxy population dominates first request. | That Nexus CPU is necessarily undersized. |
| All pulls slow including cached layers | Blob IO, network, DB metadata lookup, JVM pressure, reverse proxy. | That upstream registry is the bottleneck. |
| Push pauses on one large layer | Reverse proxy timeout/body handling, client network, blob IO. | That manifest processing is at fault. |
| GC/compaction causes latency spike | Task IO/DB contention. | That adding heap fixes it. |
Record baseline request rate, cache state, blob backend, database, task load, and client version before tuning. Blind heap increases can hide the actual IO bottleneck and create longer GC pauses.
12. Minimal evidence bundle for support/escalation
incident-ch24/
├── timeline.md
├── nexus-version-edition-runtime.txt
├── client-version.txt
├── image-reference-and-digest.txt
├── routing-dns-tls.txt
├── auth-challenge-redacted.txt
├── repository-search.json
├── task-history.txt
├── relevant-nexus-log-excerpt.txt
├── reverse-proxy-log-excerpt.txt
├── disk-blob-metrics.txt
└── verification-after-fix.md
Review support bundles/logs for secrets before sharing. Authorization headers, bearer tokens, passwords, private registry credentials, and internal-only hostnames may require redaction according to organizational policy.
13. Failure matrix
| Failure | First evidence | Least-destructive correction |
|---|---|---|
| 401 after token flow | Realm/token/privilege scope | Fix scoped identity/role or route; do not grant admin. |
| 404 wrong tag | Nexus exact search + route | Correct reference/member or republish intended disposable content. |
| 429 upstream | Headers + proxy log + cache state | Authenticate/plan/cache; retry per policy. |
| Stale mutable tag | Client/Nexus/upstream digests | Targeted cache/freshness correction after identity comparison. |
| Disk not reclaimed | References + task history + soft-delete state | Run supported GC/compaction in safe scope. |
| Missing layer | Manifest/layer digest + logs + recent tasks | Use supported recovery/restore; never edit blob/db manually. |
14. Knowledge check
What evidence should you preserve first for a stale-tag incident?
The exact image reference and the resolved manifest digest at the client/Nexus/upstream points, plus timestamps/cache state. Without identity evidence, cache invalidation is guesswork.
A 413 occurs only for large pushes. Which layer should you inspect first?
The reverse proxy/ingress request body and timeout settings, then Nexus/blob IO. A size-dependent HTTP 413 strongly implicates the HTTP gateway path.
Why is direct deletion of a .bytes blob unsafe?
Nexus database/manifest metadata may still reference it. Direct deletion bypasses supported reference tracking and can corrupt future pulls.
What is the documented workaround for the current Docker lastDownloaded issue?
Disable affected Docker cleanup policies that use Last Downloaded until Sonatype provides a fixed patch, rather than editing timestamps.
When a cached digest pulls successfully but a new tag returns 429, what does that suggest?
Cached content is locally available, but resolving/fetching uncached remote content is hitting an upstream or intermediary rate limit.
15. Summary and next step
Docker troubleshooting is evidence-driven identity work: preserve the reference/digest, locate the failing protocol layer, verify repository/cache/task/blob state, and correct only the cause you can prove. Lesson 5 combines retention, shared-layer storage, a proxy failure, and structured troubleshooting into one checkpoint runbook.
Official references and version notes
- Sonatype: Docker Registry — Docker repository routing, Registry API support, manifest lists, OCI image support, and current path-based routing guidance.
- Sonatype: Repository Manager Concepts — Docker component/tag/manifest/layer relationships and storage implications.
- Sonatype: Cleanup Policies — Docker cleanup sequence, policy criteria, soft deletion, Docker GC interaction, and blob-store compaction.
- Sonatype: Tasks — current Docker GC, incomplete-upload cleanup, repository cleanup, and compact blob-store task behavior.
- Sonatype: Searching Docker — Docker client search constraints and Nexus repository/group behavior.
- Sonatype: Searching for Components — current SQL-backed Nexus search and Docker-specific image/tag/layer criteria.
- Sonatype: Search API — supported paginated search endpoints and format-specific fields.
- Sonatype: Docker Authentication — Docker Bearer Token Realm and client authentication flow.
- Sonatype: Proxy Repository for Docker — Docker Hub/private ECR proxy behavior and current upstream-authentication options.
-
Sonatype: Nexus Repository 3.95.0–3.95.2 Release Notes
— dated release line and open Docker
HEAD/lastDownloadedcleanup issue. - Sonatype: Nexus Repository 3.91.x Release Notes — Docker manifest/tag integrity and garbage-collection correctness fixes relevant to modern behavior.
- Sonatype: OCI Repositories — native OCI hosted/proxy/group support introduced in the 3.94 line; separate it from the older Docker repository format.
- Sonatype nexus-public 3.95.2-01 release — dated patch baseline used in this chapter.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.