Chapter 24Lesson 04220–300 min

Docker Registry Operations, Garbage Collection, Retention, Search, Proxying, and Production Troubleshooting: Diagnostics, Failure Modes, Security, and Performance

Troubleshoot Docker registry failures by preserving the exact image reference and HTTP evidence, then separating client cache, routing/TLS, bearer-token authorization, Nexus repository metadata, proxy/upstream behavior, and blob/database/task state. Fix the smallest verified cause.

Diagnostics401/404/429Reverse proxyManifest/layer integrityPerformance

Learning objectives

  • Interpret common Docker client HTTP failures without immediately changing credentials, caches, or TLS controls.
  • Diagnose stale tags by comparing client, Nexus, and upstream digest evidence.
  • Separate reverse-proxy upload limits from Nexus repository/blob/database failures.
  • Explain why cleanup during active publishing or manual blob deletion can create inconsistency/corruption risk.
  • Use task logs, request evidence, disk/blob metrics, and controlled pulls to verify the least-destructive correction.
Version baseline (27 August 2026). The chapter uses Nexus Repository 3.95.2-01 with Java 21 as its dated reference line. Docker behavior has changed substantially across releases; re-check the exact task names, known issues, routing mode, and release notes on the instance you operate.
Current 3.94.0–3.95.2 Docker cleanup warning. Sonatype documents an open issue where Docker manifest HEAD requests do not refresh an asset's lastDownloaded timestamp. An actively used image can therefore look inactive. Do not use Last Downloaded–based Docker cleanup policies on an affected instance until Sonatype publishes a fixed version.
Safety boundary. All destructive steps must target explicitly named disposable repositories and synthetic images. Never delete production tags/manifests, edit blob-store files, remove database rows, disable TLS validation, or run garbage collection/compaction on valuable repositories merely to observe behavior.

1. Diagnostic sequence for Docker incidents

  1. Preserve the exact command/image reference, time, client version, and concise error.
  2. Confirm Nexus version/edition/runtime/database/blob-store type.
  3. Resolve DNS/TLS/routing mode and reverse-proxy path.
  4. Inspect Docker Bearer Token Realm, authentication challenge, token acquisition, and repository privileges.
  5. Identify hosted/proxy/group and member order.
  6. Compare tag → manifest digest in client, Nexus, and upstream where applicable.
  7. Inspect component/asset state and Docker GC/cleanup task history.
  8. Inspect proxy cache/remote health/upstream rate-limit headers.
  9. Inspect database/blob/disk and reverse-proxy upload limits.
  10. Apply one least-destructive correction and repeat a controlled pull/push.

2. HTTP 401: challenge or actual denial?

The initial Registry V2 request commonly returns 401 Unauthorized with a bearer challenge. A functioning Docker client follows that challenge to obtain a token. Troubleshooting must distinguish “normal challenge” from “token request failed” and “token obtained but repository action denied.”

HTTP/1.1 401 Unauthorized
WWW-Authenticate: Bearer realm="https://repo.example.invalid/service/rest/v1/security/token",service="repo.example.invalid",scope="repository:ch24-docker-hosted:pull"

Redact real token values. Check Docker Bearer Token Realm, service URL/reverse-proxy headers, credential store, user/role privileges, and content-selector path semantics. Do not solve a 401 by granting administrator rights to CI.

3. HTTP 404: missing name/tag, wrong route, or policy-visible absence

A 404 can mean the tag genuinely does not exist, the client addressed the wrong repository path, a group member order/privilege hides the expected component, the proxy has no such remote content, or a stale metadata/reference was removed during cleanup. Verify the repository path and search exact image name/tag in Nexus before deleting caches.

curl -fsS "$NEXUS_URL/service/rest/v1/search?repository=$REPO&format=docker&docker.imageName=$IMAGE&docker.imageTag=$TAG"

4. HTTP 429: locate the rate limiter

429 Too Many Requests can be returned by an upstream public registry, a reverse proxy/WAF, or another service in the request path. Preserve Retry-After and vendor rate-limit headers when present. If Nexus has the requested digest fully cached, a pull may still succeed without upstream access; an uncached tag/layer may fail.

Corrective patterns include authenticated upstream access, approved mirrors, fewer mutable-tag lookups, digest pinning, avoiding redundant CI pulls, and sizing cache retention. Do not add random public credentials or bypass Nexus policy.

5. Stale proxy tag: compare digests before invalidating

When upstream/app:stable moved from digest A to B, four states may exist:

  • Upstream tag = B.
  • Nexus proxy cached tag = A.
  • Docker client local image = A.
  • Deployment manifest may be pinned to A intentionally.

Collect each state. If production is intentionally pinned to A, the “stale” report may be wrong. If the mutable tag should refresh, investigate proxy cache metadata/remote request evidence. Cache invalidation is a targeted recovery tool, not the first command.

6. Manifest/layer mismatch: preserve the failing digest

Errors such as unknown blob, manifest invalid, or digest mismatch require exact digest evidence. Do not delete all blobs. Record the requested manifest digest, layer digest, repository, task history, and whether the content was recently cleaned/imported/migrated. Modern Nexus releases include many Docker integrity and GC fixes; old-version folklore is unsafe on a current database.

If a layer is missing but metadata references it, escalate through supported recovery/data-repair guidance for the exact version. Chapter 25 will cover backup/restore; Chapter 26 covers migration. Do not fabricate the missing blob or edit database references manually.

7. Large push fails: inspect reverse-proxy limits before Nexus internals

Typical symptoms include 413 Request Entity Too Large, 502/504, interrupted resumable uploads, or client retry loops. Compare reverse-proxy body-size limits, request buffering, read/send timeouts, TLS termination, and upstream connection timeouts against the failing layer size/duration.

# Illustrative reverse-proxy concepts; adapt to your approved baseline.
client_max_body_size 0;
proxy_request_buffering off;
proxy_buffering off;
proxy_read_timeout 300s;
proxy_send_timeout 300s;

These values are examples, not a blind production configuration. Measure the workload and preserve security controls. Partial uploads can leave temporary blob data; use the supported Docker - Delete incomplete uploads task after the underlying transport problem is fixed.

8. Cleanup during active publish: concurrency matters

Garbage collection, repository cleanup, large pushes, and compaction all consume IO/database resources and reason about references at specific points in time. Running aggressive maintenance during peak publication increases latency and complicates incident analysis. Current Nexus includes safeguards/fixes, but an operator should still schedule intensive maintenance deliberately and monitor task status.

If a task is already running, do not repeatedly start duplicate copies. Sonatype has specifically hardened Docker GC concurrency in modern releases; respect task status and repository scope.

9. Intentionally broken example: “rm the layer file to save disk”

# WRONG — do not run.
# rm -f /nexus-data/blobs/default/content/vol-*/chap-*/sha256-some-layer.bytes

This bypasses repository metadata and reference tracking. A surviving manifest may still point to the deleted layer, creating future pull failures and difficult recovery. The supported path is retention/deletion → Docker GC → compact/reclaim, with backup/recovery context. The diagnosis for “disk is full” is not “delete the largest .bytes file.”

10. Intentionally broken example: Last Downloaded cleanup on 3.95.2

Suppose prod/base:2026.08 is pulled daily, but the runtime checks the manifest using HEAD. On Nexus 3.95.2 the open issue can leave lastDownloaded stale. A 30-day inactivity policy may mark the active image for deletion.

Repair: disable the affected Last Downloaded Docker cleanup policy as Sonatype directs, preserve the active image, use safer retention evidence, and plan upgrade/verification when a fixed patch becomes available. Do not “fix” the timestamp directly in the database.

11. Performance diagnosis: identify the constrained stage

Symptom Likely evidence Do not assume
Slow first pull, fast second pull Upstream/network/proxy population dominates first request. That Nexus CPU is necessarily undersized.
All pulls slow including cached layers Blob IO, network, DB metadata lookup, JVM pressure, reverse proxy. That upstream registry is the bottleneck.
Push pauses on one large layer Reverse proxy timeout/body handling, client network, blob IO. That manifest processing is at fault.
GC/compaction causes latency spike Task IO/DB contention. That adding heap fixes it.

Record baseline request rate, cache state, blob backend, database, task load, and client version before tuning. Blind heap increases can hide the actual IO bottleneck and create longer GC pauses.

12. Minimal evidence bundle for support/escalation

incident-ch24/
├── timeline.md
├── nexus-version-edition-runtime.txt
├── client-version.txt
├── image-reference-and-digest.txt
├── routing-dns-tls.txt
├── auth-challenge-redacted.txt
├── repository-search.json
├── task-history.txt
├── relevant-nexus-log-excerpt.txt
├── reverse-proxy-log-excerpt.txt
├── disk-blob-metrics.txt
└── verification-after-fix.md

Review support bundles/logs for secrets before sharing. Authorization headers, bearer tokens, passwords, private registry credentials, and internal-only hostnames may require redaction according to organizational policy.

13. Failure matrix

Failure First evidence Least-destructive correction
401 after token flow Realm/token/privilege scope Fix scoped identity/role or route; do not grant admin.
404 wrong tag Nexus exact search + route Correct reference/member or republish intended disposable content.
429 upstream Headers + proxy log + cache state Authenticate/plan/cache; retry per policy.
Stale mutable tag Client/Nexus/upstream digests Targeted cache/freshness correction after identity comparison.
Disk not reclaimed References + task history + soft-delete state Run supported GC/compaction in safe scope.
Missing layer Manifest/layer digest + logs + recent tasks Use supported recovery/restore; never edit blob/db manually.

14. Knowledge check

What evidence should you preserve first for a stale-tag incident?

A 413 occurs only for large pushes. Which layer should you inspect first?

Why is direct deletion of a .bytes blob unsafe?

What is the documented workaround for the current Docker lastDownloaded issue?

When a cached digest pulls successfully but a new tag returns 429, what does that suggest?

15. Summary and next step

Docker troubleshooting is evidence-driven identity work: preserve the reference/digest, locate the failing protocol layer, verify repository/cache/task/blob state, and correct only the cause you can prove. Lesson 5 combines retention, shared-layer storage, a proxy failure, and structured troubleshooting into one checkpoint runbook.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.