Chapter 18Lesson 03190–250 min

Cleanup Policies, Retention, Component Age, Last-Downloaded Rules, Preview, and Storage Reclamation: Configuration, Design Choices, and Tradeoffs

Turn retention into an operating policy by balancing rollback value, artifact velocity, usage quality, format semantics, repository topology, task load, recovery guarantees, and licensing boundaries.

Retention SLATradeoffsRollbackSchedulingGovernance

Learning objectives

  • Design different retention windows for development, release, and proxy-cache content.
  • Evaluate whether last-downloaded data is trustworthy enough for a given format/version.
  • Choose between shared policies and format/repository-specific policies.
  • Schedule cleanup and compaction according to IO/recovery risk rather than convenience.
  • Document a retention SLA that can be reviewed, measured, and reversed through backup/recovery controls.

1. Retention is a service-level decision

A cleanup policy converts business and delivery assumptions into deletion eligibility. That means the first design question is not “30 or 90 days?” It is “what evidence must remain available for rollback, incident reconstruction, compliance, offline development, rebuild reproducibility, and vulnerability response?” Different artifact classes have different answers.

A repository platform with no cleanup eventually loses capacity headroom. A repository platform with aggressive uniform cleanup can lose the exact release needed during an incident. Production retention therefore needs a declared retention SLA: class of content, minimum survival window, usage exception, release protection, reclaim delay, owner, review cadence, and recovery source.

Version baseline (26 August 2026). Sonatype's current 3.95.x release notes list Nexus Repository 3.95.2 (released 21 August 2026) as the newest patch in that line. The 3.95 release expanded cleanup-policy administration and retain-last-N coverage. Current Nexus Repository 3.95.x system requirements use Java 21; official bundles include the recommended runtime. Record the exact version, edition, database, and blob-store type of your lab before applying any cleanup rule because cleanup behavior and available criteria are version- and format-sensitive.

2. Five dimensions shape a safe policy

Dimension Question Observable evidence
Lifecycle value How long could this artifact be needed for rollback/debug/audit? Release support window, incident history, deployment records
Artifact velocity How quickly does the repository grow? components/day, GiB/day, CI build frequency
Usage quality Does recorded download activity represent real consumption? lastDownloaded behavior, client request patterns, known issues
Reclamation cost How expensive is cleanup/compaction for this storage backend? task duration, IO latency, free-space curve
Recoverability What exact restore point exists if retention is wrong? validated backup, RPO/RTO, immutable release mirror if applicable

3. Development artifacts and releases should rarely share the same retention rule

CI can create hundreds of ephemeral builds per day. Most are valuable only until the next successful build or for a short debugging window. Releases can remain operational dependencies for months or years because rollback, long-lived branches, customer installations, or downstream systems still reference them.

Different lifecycle curves
flowchart TD
CI[CI build] --> D[Development repository]
D -->|short retention| X[Cleanup candidate]
CI -->|approved immutable build| R[Release repository]
R -->|support / rollback window| K[Longer retention]
K -->|policy + validated recovery| Y[Later cleanup]
X --> C[Compaction after verification]
Y --> C

Separating repositories can make policy easier to reason about because repository boundary becomes lifecycle boundary. A single repository can still work when format metadata and criteria distinguish prereleases/releases safely, but operational evidence becomes harder to audit.

4. Last-downloaded is a signal, not a universal truth

Usage-based cleanup is attractive because it retains what people actually consume. It also inherits every assumption in request tracking. Package managers may use metadata requests, HEAD requests, local caches, group repositories, proxy caches, or offline build caches in ways that make “downloaded” differ from “needed.” A release can be critical yet rarely fetched because clients already cache it.

Before relying on last-downloaded, test representative clients against the exact Nexus version and format. The 3.94.0–3.95.2 Docker known issue is a concrete example: some manifest HEAD requests do not update lastDownloaded, so active images can appear stale. In such a condition, usage-based Docker cleanup should be disabled rather than compensated with guesswork.

5. One broad policy versus repository-specific policies

Approach Advantages Risks Use when
One age/usage policy across many repositories Low configuration count; consistent default Formats/lifecycles may differ; broad blast radius Repositories truly share lifecycle and supported criteria
Format-specific policy Matches format metadata semantics Still may blur dev/release teams Format behavior dominates retention logic
Lifecycle-specific policy Clear dev/release intent More policy objects to own Repository topology already separates lifecycle
Repository-specific policy Smallest blast radius; easiest preview Configuration sprawl High-value or exceptional repositories

In 3.95, policy administration can associate a policy with multiple repositories from the policy screen. That improves usability but does not reduce the blast radius of a broad rule. Treat bulk association as configuration convenience, not as proof that those repositories belong under one lifecycle contract.

6. Choose criteria to encode intent, not to chase a storage number

Suppose a team keeps 14 days of debug builds but wants unused artifacts gone after 7 days. If builds are generated daily, an age threshold around 14 days may approximate the minimum debug window, while usage can protect older builds still being consumed. For release content, a fixed age may be inappropriate unless the release support policy has an explicit end-of-life mechanism.

Artifact class Illustrative policy Reasoning
Ephemeral CI Age ≥ 14d AND usage ≥ 7d Keep recent builds; protect older builds still used
Snapshots/prereleases Age ≥ 30d AND usage ≥ 14d; prerelease only More rollback history than raw CI while protecting releases
Immutable releases No automated cleanup until support/EOL rule exists Availability and rollback outrank storage savings
Proxy cache Usage ≥ working-set window, optionally age Remove stale cache while retaining commonly used dependencies

These numbers are examples, not defaults. Production values must come from build cadence, release policy, recovery objective, and measured storage growth.

7. Optional Pro/PostgreSQL design: retain N recent versions

Current 3.95 documentation expands Retain Select Versions to versioned formats that have version semantics. It remains a Pro feature with PostgreSQL requirements in current cleanup documentation. Use it when a version floor materially reduces rollback risk—for example, “eligible if inactive for 90 days, but always keep the newest three release versions.”

Do not treat N as an integrity guarantee. A package may have three semantically newer versions that are all broken, revoked, or incompatible with a downstream system. Release governance should still record approved artifact identity, provenance, and support status.

8. Cleanup frequency and compaction frequency need not be equal

Policy evaluation and soft deletion can run relatively often, while blob compaction may be more IO-intensive and destructive. Current Nexus creates system cleanup tasks with regular schedules, but operators can reschedule them. Production design should consider repository traffic, database load, blob-storage latency, backup windows, and recovery acceptance.

Activity Typical concern Safer scheduling thought
Preview Query cost on large repositories Run before change window; capture results
Cleanup evaluation DB work and component mutation Separate from migrations/large imports when possible
Unused-asset cleanup Metadata/blob housekeeping Observe duration before increasing frequency
Blob compaction IO and permanent deletion Off peak, after backup/preview acceptance, one store at a time where practical
Backup Consistency and restore point Coordinate so recovery point corresponds to known repository state

9. Retention cost is larger than blob bytes

Long retention consumes blob capacity, database rows, indexes/search state, backup bandwidth, object-store requests, restore time, and operator attention. Short retention reduces those costs but increases rebuild risk, upstream dependence, and incident friction. A proxy cache with 30-day inactivity cleanup can reduce stale vulnerable components and storage, but it also means an upstream outage could matter more if needed dependencies were evicted.

Therefore tie proxy retention to resilience: if builds must continue during a public-registry outage, the working set must remain cached long enough and clients must use Nexus rather than bypassing it.

10. Security and retention have competing goals

Removing unused vulnerable packages reduces dormant attack surface and accidental reuse. Preserving releases and forensic evidence supports incident response. Neither goal implies “delete vulnerabilities immediately everywhere.” An artifact under investigation may need preservation even if blocked from normal consumption. Repository cleanup, quarantine/policy systems, and legal/forensic retention are separate controls.

Least privilege matters too. Current cleanup docs distinguish permissions for policy creation, repository attachment, and task modification. A CI publisher should not automatically be able to redefine cleanup rules or compact blob stores.

11. Worked design: internal Maven, npm, PyPI, and proxy caches

Assume 80 developers, daily CI, monthly releases, and a 60-day rollback requirement. Storage grows 8 GiB/day, mostly CI prereleases. The team uses Community Edition today but may move to Pro/PostgreSQL later.

Repository Decision Why
maven-dev-hosted Prerelease/snapshot cleanup after measured 21–30 day window High artifact velocity; releases elsewhere
npm-dev-hosted Separate dev lifecycle policy; validate semver prerelease behavior Package semantics differ from Maven
pypi-release-hosted No automated release deletion until EOL workflow exists Rollback/support requirement dominates storage
maven-central-proxy Usage-based cache cleanup only after confirming client request behavior Balance storage and outage resilience
docker-proxy No Last Downloaded policy while 3.95.2 known issue applies Usage metric can under-report active images

If Pro/PostgreSQL is introduced, retain-last-N can become an additional guardrail for versioned repositories, but it does not replace release/EOL governance.

12. Decision record template

repository: maven-dev-hosted
owner: build-platform
purpose: ephemeral CI and SNAPSHOT artifacts
format: maven2
repository_type: hosted
retention_intent:
  minimum_debug_window_days: 21
  inactivity_days: 14
  release_protection: "releases stored elsewhere"
execution:
  preview_required: true
  cleanup_window: "off-peak"
  compact_after_validation: true
  soft_delete_grace_days: 7
recovery:
  source: "validated Nexus backup"
  rpo: "24h"
  restore_test: "quarterly"
known_issues:
  - "re-check version-specific cleanup release notes before changes"

The YAML is documentation, not Nexus configuration. Its value is that a future operator can compare business intent with the actual policy before changing anything.

13. Design mistakes to avoid

  • Applying one retention age to development, releases, and proxies because it is administratively convenient.
  • Using last-downloaded without testing whether representative clients update it reliably.
  • Running compaction immediately after cleanup, eliminating a useful soft-delete grace window.
  • Assuming backups are valid because a job says “success” without restore testing.
  • Using Pro-only retain-last-N as a hidden dependency in a Community operating runbook.
  • Optimizing for reclaimed GiB while ignoring build-outage and rollback costs.

14. Knowledge check

Why might a proxy repository need longer retention than its raw storage cost suggests?

What is the advantage of separate development and release repositories?

Should cleanup and compaction always run at the same frequency?

What must happen before choosing a last-downloaded threshold?

Why does retain-last-N not guarantee a usable rollback?

15. Summary and next step

A production retention policy is an SLA encoded as criteria, scope, schedule, recovery window, and evidence—not a universal age number. Lesson 4 takes these design choices into failure analysis: broad regexes, misunderstood timestamps, ignored preview, delayed reclamation, task overlap, and direct filesystem deletion.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.