Chapter 24Lesson 05270–360 min

Checkpoint Lab — Docker Registry Operations, Garbage Collection, Retention, Search, Proxying, and Production Troubleshooting

Operate a small registry as if it mattered: build a multi-tag synthetic image inventory, define a retention SLA, predict and verify logical versus physical reclamation, simulate one proxy outage without touching a public registry, and produce an evidence bundle another operator could use to reproduce the diagnosis.

Checkpoint labRetention SLAStorage evidenceProxy outageRunbook

Learning objectives

  • Build a three-release synthetic Docker inventory with a moving convenience tag and exact recorded digests.
  • Apply a documented retention decision and prove which tags/manifests/layers remain referenced.
  • Measure logical state separately from physical blob-store reclamation and explain any gap.
  • Simulate a safe proxy failure with a local fixture and distinguish cached from uncached behavior.
  • Deliver a complete troubleshooting evidence bundle, verification checklist, cleanup/rollback plan, and production handoff.
Version baseline (27 August 2026). The chapter uses Nexus Repository 3.95.2-01 with Java 21 as its dated reference line. Docker behavior has changed substantially across releases; re-check the exact task names, known issues, routing mode, and release notes on the instance you operate.
Current 3.94.0–3.95.2 Docker cleanup warning. Sonatype documents an open issue where Docker manifest HEAD requests do not refresh an asset's lastDownloaded timestamp. An actively used image can therefore look inactive. Do not use Last Downloaded–based Docker cleanup policies on an affected instance until Sonatype publishes a fixed version.
Safety boundary. All destructive steps must target explicitly named disposable repositories and synthetic images. Never delete production tags/manifests, edit blob-store files, remove database rows, disable TLS validation, or run garbage collection/compaction on valuable repositories merely to observe behavior.

1. Scenario: Learner Registry Service

You operate a disposable Nexus registry for learner-example/registry-service. The application has three releases: 1.0, 1.1, and 1.2. The convenience tag stable points to 1.2. Releases share a common base layer. Your synthetic retention SLA says:

  • Keep the two newest release tags (1.1 and 1.2) for rollback.
  • Keep stable only as a convenience alias to 1.2.
  • Retire 1.0.
  • Never use Last Downloaded for this Docker policy on the affected 3.95.2 line.
  • Reclaim physical blob storage only after verifying no retained manifest references the candidate layers.

2. Preflight

  • Nexus reference: 3.95.2-01, Java 21, Community, single disposable node.
  • Database/blob: H2 + dedicated file blob store for the lab, or deterministic simulation if a dedicated blob store is unavailable.
  • Routing: trusted TLS path-based routing from Chapter 17; do not globally disable Docker TLS checks.
  • Credentials: disposable scoped publisher/reader; no admin credentials in scripts.
  • Proxy outage: local disposable upstream fixture only; no deliberate outage/load against Docker Hub or another public service.
  • Backups: no valuable data exists in the disposable repository. Production deletion would require backup/restore context covered in Chapter 25.

3. Create the evidence packet before changes

chapter24-checkpoint/
├── 00-preflight.md
├── 01-predictions.md
├── 02-repository-config.json
├── 03-tags-before.json
├── 04-search-before.json
├── 05-digests-before.txt
├── 06-layer-graph.json
├── 07-retention-decision.md
├── 08-task-history.txt
├── 09-storage-before.txt
├── 10-tags-after.json
├── 11-search-after.json
├── 12-storage-after.txt
├── 13-proxy-failure-evidence.md
├── 14-verification.md
└── 15-runbook-handoff.md

4. Deterministic layer graph

Even if you perform the live Docker push, generate this fixture so the expected reference graph is auditable:

import hashlib, json
from pathlib import Path

out = Path('chapter24-checkpoint')
out.mkdir(exist_ok=True)
sha = lambda s: 'sha256:' + hashlib.sha256(s.encode()).hexdigest()
base = sha('base-layer-v1')
app10, app11, app12 = sha('app-1.0'), sha('app-1.1'), sha('app-1.2')
m10 = sha('manifest:'+base+app10)
m11 = sha('manifest:'+base+app11)
m12 = sha('manifest:'+base+app12)
state = {
  'tags': {'1.0':m10, '1.1':m11, '1.2':m12, 'stable':m12},
  'manifests': {
    m10:{'layers':[base,app10]},
    m11:{'layers':[base,app11]},
    m12:{'layers':[base,app12]},
  }
}
(out/'06-layer-graph.json').write_text(json.dumps(state, indent=2)+'\n')
print(json.dumps(state, indent=2))

The base layer is referenced by all three manifests. stable and 1.2 reference the same manifest. Retiring 1.0 can make only app10 and manifest m10 candidates for later cleanup; the base layer remains live.

5. Live publication path

Using the Chapter 17 trusted endpoint, build/push the three harmless synthetic versions. Reuse the Lesson 2 Dockerfile, changing only version.txt. Record the exact digest printed after every push.

# Pseudocode-style loop; adapt to your approved local engine.
for v in 1.0 1.1 1.2; do
  printf 'version=%s
' "$v" > version.txt
  docker build -t "registry-service:$v" .
  docker tag "registry-service:$v" "$REGISTRY_HOST/$HOSTED_PATH:$v"
  docker push "$REGISTRY_HOST/$HOSTED_PATH:$v"
done

docker tag registry-service:1.2 "$REGISTRY_HOST/$HOSTED_PATH:stable"
docker push "$REGISTRY_HOST/$HOSTED_PATH:stable"

If the live environment is not available, use the deterministic graph and clearly mark the checkpoint as simulated for publication while still completing retention reasoning, task design, and evidence analysis.

6. Capture the before-state

Record:

  • Exact Nexus repository config and blob store.
  • Tags 1.0, 1.1, 1.2, stable.
  • Manifest digest for each tag; prove 1.2 == stable.
  • Nexus search results for image/tag.
  • Blob-store physical usage/available space.
  • Task schedule/status for repository cleanup, Docker GC, incomplete-upload cleanup, and compact blob store.

7. Required predictions

Write at least these predictions before execution:

  1. After retiring 1.0, tags 1.1, 1.2, and stable remain.
  2. The shared base layer remains referenced and cannot be reclaimed.
  3. The 1.0-specific layer can become unreferenced after its manifest is no longer retained.
  4. Physical disk space may not fall until Docker GC plus blob compaction complete.
  5. During a local upstream outage, a fully cached proxy digest may remain pullable while an uncached tag/image fails.

8. Apply the retention decision

Use the supported Nexus UI/Components API to identify the 1.0 component in the disposable hosted repository. Verify repository, image name, tag, and component ID. Retire only that disposable component. If you are using a cleanup policy in an authorized environment, preview candidates first and ensure the policy does not use Last Downloaded on the affected version.

No broad delete. Do not select “all components,” delete the repository, or remove the blob store to satisfy the checkpoint. The point is to prove targeted lifecycle control.

9. Run supported cleanup/reclamation in stages

  1. Capture logical after-state immediately after retiring 1.0.
  2. Run Docker - Delete unused manifests and images for the disposable hosted repository.
  3. Capture task log/history and search/tag state.
  4. Run Admin - Compact blob store only if the blob store is dedicated/safe for the lab.
  5. Capture physical storage after-state separately.

At each stage, note whether the change is repository metadata/reference state, soft-deleted blob state, or physical reclaimed bytes.

10. Safe proxy-failure drill

Prepare a local upstream registry fixture containing two synthetic tags, cached and uncached. Pull only cached through the Nexus proxy so its manifest/layers are present in Nexus. Stop the local upstream service or block only the disposable loopback route. Then:

  1. Request the already-cached digest/tag through Nexus.
  2. Request uncached.
  3. Record Nexus client output, proxy remote-health/log evidence, and which request required upstream access.
  4. Restore the local upstream and verify recovery.

This demonstrates availability/cache behavior without intentionally disrupting a public registry.

11. No-local-registry fallback for proxy failure

If you cannot run a local upstream, use this deterministic decision fixture:

requests = [
  {'name':'cached','manifestCached':True,'allLayersCached':True,'upstream':'down'},
  {'name':'partial','manifestCached':True,'allLayersCached':False,'upstream':'down'},
  {'name':'uncached','manifestCached':False,'allLayersCached':False,'upstream':'down'},
]
for r in requests:
    if r['manifestCached'] and r['allLayersCached']:
        outcome='may-serve-from-cache'
    else:
        outcome='requires-upstream-and-fails'
    print(r['name'], outcome)

Label this as a conceptual fallback. Real proxy semantics must be verified against the pinned Nexus version and exact cached asset state.

12. Analyze logical versus physical storage

Stage Expected logical state Expected physical state
Before retention 4 tags; 3 manifests; shared base + 3 version layers. All live blob bytes present.
After retiring 1.0 1.0 absent; remaining tags intact. Old blobs may still consume space.
After Docker GC Unreferenced 1.0 manifest/layer marked/removed per current task semantics. Soft-deleted bytes may still occupy storage.
After compaction Logical repository unchanged from post-GC. Eligible soft-deleted file blobs physically reclaimed.

If the measured bytes do not match a simple arithmetic expectation, first consider layer sharing, filesystem allocation, task age thresholds, and unrelated blob-store content. Do not force the number by deleting files.

13. Inject one intentional client failure

Choose exactly one safe failure:

  • Attempt a pull using a wrong disposable tag to get a controlled 404.
  • Use a disposable reader without pull privilege and capture the authorization failure.
  • Stop the local upstream fixture for the proxy drill.

Diagnose it with the Chapter 24 sequence and restore the correct state. Do not disable TLS, grant admin, or modify database/blob files.

14. Required evidence and handoff

Your 15-runbook-handoff.md must state:

  • Version/edition/runtime/database/blob assumptions.
  • Repository/routing/authentication topology.
  • Retention SLA and why 1.0 was selected.
  • Before/after tag→digest mapping.
  • Which layers remained shared/live.
  • Task names, scope, schedule, and logs.
  • Logical versus physical storage measurements.
  • Proxy-failure diagnosis and recovery.
  • Current 3.95.2 Last Downloaded warning.
  • What would require backup/restore approval in production.

15. Verification checklist

  • Exactly the intended 1.0 reference was retired.
  • 1.1, 1.2, and stable remain pullable/readable in the live path.
  • 1.2 and stable resolve to the same manifest digest as predicted.
  • The shared base layer remains referenced.
  • Docker GC and compaction were recorded as separate stages.
  • No Last Downloaded–based Docker policy was used on 3.95.2.
  • Proxy outage targeted only a local disposable upstream.
  • No secrets were printed or archived.
  • No blob/database internals were modified.
  • The evidence packet explains every discrepancy between logical state and disk bytes.

16. Cleanup / rollback

Restore the local upstream fixture, remove only the explicitly named disposable Docker repositories/content through supported Nexus UI/API, remove temporary client tags/images if desired, and delete the local checkpoint directory after preserving any learning evidence. If the lab blob store was dedicated solely to Chapter 24, remove it through supported Nexus administration only after all repository references are gone and scope is verified.

Do not delete shared blob-store directories from the filesystem. In production, rollback and recovery would be backed by the consistency/backup discipline taught next in Chapter 25.

17. Knowledge check

Why must the base layer remain after 1.0 is retired?

What proves that stable and 1.2 are byte-identical at the manifest level?

Why separate storage-before, post-delete, post-GC, and post-compaction measurements?

What is the safe proxy-outage target?

What does Chapter 25 add to this operating model?

18. Chapter summary and bridge to Chapter 25

Chapter 24 turns Docker from a package-client topic into an operated registry system. You learned to anchor mutable tags to immutable digests, reason about manifests and shared layers, separate retention from garbage collection and compaction, diagnose proxy/cache/auth/network failures with evidence, and preserve storage/recovery boundaries. Chapter 25 now adds the recovery side of the contract: backup, restore, configuration state, blob consistency, disaster scenarios, and proof that a restored repository actually serves the expected artifacts.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.