Cleanup Policies, Retention, Component Age, Last-Downloaded Rules, Preview, and Storage Reclamation: Configuration, Design Choices, and Tradeoffs
Turn retention into an operating policy by balancing rollback value, artifact velocity, usage quality, format semantics, repository topology, task load, recovery guarantees, and licensing boundaries.
Learning objectives
- Design different retention windows for development, release, and proxy-cache content.
- Evaluate whether last-downloaded data is trustworthy enough for a given format/version.
- Choose between shared policies and format/repository-specific policies.
- Schedule cleanup and compaction according to IO/recovery risk rather than convenience.
- Document a retention SLA that can be reviewed, measured, and reversed through backup/recovery controls.
1. Retention is a service-level decision
A cleanup policy converts business and delivery assumptions into deletion eligibility. That means the first design question is not “30 or 90 days?” It is “what evidence must remain available for rollback, incident reconstruction, compliance, offline development, rebuild reproducibility, and vulnerability response?” Different artifact classes have different answers.
A repository platform with no cleanup eventually loses capacity headroom. A repository platform with aggressive uniform cleanup can lose the exact release needed during an incident. Production retention therefore needs a declared retention SLA: class of content, minimum survival window, usage exception, release protection, reclaim delay, owner, review cadence, and recovery source.
Version baseline (26 August 2026). Sonatype's current 3.95.x release notes list Nexus Repository 3.95.2 (released 21 August 2026) as the newest patch in that line. The 3.95 release expanded cleanup-policy administration and retain-last-N coverage. Current Nexus Repository 3.95.x system requirements use Java 21; official bundles include the recommended runtime. Record the exact version, edition, database, and blob-store type of your lab before applying any cleanup rule because cleanup behavior and available criteria are version- and format-sensitive.
2. Five dimensions shape a safe policy
| Dimension | Question | Observable evidence |
|---|---|---|
| Lifecycle value | How long could this artifact be needed for rollback/debug/audit? | Release support window, incident history, deployment records |
| Artifact velocity | How quickly does the repository grow? | components/day, GiB/day, CI build frequency |
| Usage quality | Does recorded download activity represent real consumption? | lastDownloaded behavior, client request patterns, known issues |
| Reclamation cost | How expensive is cleanup/compaction for this storage backend? | task duration, IO latency, free-space curve |
| Recoverability | What exact restore point exists if retention is wrong? | validated backup, RPO/RTO, immutable release mirror if applicable |
3. Development artifacts and releases should rarely share the same retention rule
CI can create hundreds of ephemeral builds per day. Most are valuable only until the next successful build or for a short debugging window. Releases can remain operational dependencies for months or years because rollback, long-lived branches, customer installations, or downstream systems still reference them.
flowchart TD CI[CI build] --> D[Development repository] D -->|short retention| X[Cleanup candidate] CI -->|approved immutable build| R[Release repository] R -->|support / rollback window| K[Longer retention] K -->|policy + validated recovery| Y[Later cleanup] X --> C[Compaction after verification] Y --> C
Separating repositories can make policy easier to reason about because repository boundary becomes lifecycle boundary. A single repository can still work when format metadata and criteria distinguish prereleases/releases safely, but operational evidence becomes harder to audit.
4. Last-downloaded is a signal, not a universal truth
Usage-based cleanup is attractive because it retains what people actually consume. It also inherits every assumption in request tracking. Package managers may use metadata requests, HEAD requests, local caches, group repositories, proxy caches, or offline build caches in ways that make “downloaded” differ from “needed.” A release can be critical yet rarely fetched because clients already cache it.
Before relying on last-downloaded, test representative clients
against the exact Nexus version and format. The 3.94.0–3.95.2 Docker
known issue is a concrete example: some manifest HEAD requests do
not update lastDownloaded, so active images can appear
stale. In such a condition, usage-based Docker cleanup should be
disabled rather than compensated with guesswork.
5. One broad policy versus repository-specific policies
| Approach | Advantages | Risks | Use when |
|---|---|---|---|
| One age/usage policy across many repositories | Low configuration count; consistent default | Formats/lifecycles may differ; broad blast radius | Repositories truly share lifecycle and supported criteria |
| Format-specific policy | Matches format metadata semantics | Still may blur dev/release teams | Format behavior dominates retention logic |
| Lifecycle-specific policy | Clear dev/release intent | More policy objects to own | Repository topology already separates lifecycle |
| Repository-specific policy | Smallest blast radius; easiest preview | Configuration sprawl | High-value or exceptional repositories |
In 3.95, policy administration can associate a policy with multiple repositories from the policy screen. That improves usability but does not reduce the blast radius of a broad rule. Treat bulk association as configuration convenience, not as proof that those repositories belong under one lifecycle contract.
6. Choose criteria to encode intent, not to chase a storage number
Suppose a team keeps 14 days of debug builds but wants unused artifacts gone after 7 days. If builds are generated daily, an age threshold around 14 days may approximate the minimum debug window, while usage can protect older builds still being consumed. For release content, a fixed age may be inappropriate unless the release support policy has an explicit end-of-life mechanism.
| Artifact class | Illustrative policy | Reasoning |
|---|---|---|
| Ephemeral CI | Age ≥ 14d AND usage ≥ 7d | Keep recent builds; protect older builds still used |
| Snapshots/prereleases | Age ≥ 30d AND usage ≥ 14d; prerelease only | More rollback history than raw CI while protecting releases |
| Immutable releases | No automated cleanup until support/EOL rule exists | Availability and rollback outrank storage savings |
| Proxy cache | Usage ≥ working-set window, optionally age | Remove stale cache while retaining commonly used dependencies |
These numbers are examples, not defaults. Production values must come from build cadence, release policy, recovery objective, and measured storage growth.
7. Optional Pro/PostgreSQL design: retain N recent versions
Current 3.95 documentation expands Retain Select Versions to versioned formats that have version semantics. It remains a Pro feature with PostgreSQL requirements in current cleanup documentation. Use it when a version floor materially reduces rollback risk—for example, “eligible if inactive for 90 days, but always keep the newest three release versions.”
Do not treat N as an integrity guarantee. A package may have three semantically newer versions that are all broken, revoked, or incompatible with a downstream system. Release governance should still record approved artifact identity, provenance, and support status.
8. Cleanup frequency and compaction frequency need not be equal
Policy evaluation and soft deletion can run relatively often, while blob compaction may be more IO-intensive and destructive. Current Nexus creates system cleanup tasks with regular schedules, but operators can reschedule them. Production design should consider repository traffic, database load, blob-storage latency, backup windows, and recovery acceptance.
| Activity | Typical concern | Safer scheduling thought |
|---|---|---|
| Preview | Query cost on large repositories | Run before change window; capture results |
| Cleanup evaluation | DB work and component mutation | Separate from migrations/large imports when possible |
| Unused-asset cleanup | Metadata/blob housekeeping | Observe duration before increasing frequency |
| Blob compaction | IO and permanent deletion | Off peak, after backup/preview acceptance, one store at a time where practical |
| Backup | Consistency and restore point | Coordinate so recovery point corresponds to known repository state |
9. Retention cost is larger than blob bytes
Long retention consumes blob capacity, database rows, indexes/search state, backup bandwidth, object-store requests, restore time, and operator attention. Short retention reduces those costs but increases rebuild risk, upstream dependence, and incident friction. A proxy cache with 30-day inactivity cleanup can reduce stale vulnerable components and storage, but it also means an upstream outage could matter more if needed dependencies were evicted.
Therefore tie proxy retention to resilience: if builds must continue during a public-registry outage, the working set must remain cached long enough and clients must use Nexus rather than bypassing it.
10. Security and retention have competing goals
Removing unused vulnerable packages reduces dormant attack surface and accidental reuse. Preserving releases and forensic evidence supports incident response. Neither goal implies “delete vulnerabilities immediately everywhere.” An artifact under investigation may need preservation even if blocked from normal consumption. Repository cleanup, quarantine/policy systems, and legal/forensic retention are separate controls.
Least privilege matters too. Current cleanup docs distinguish permissions for policy creation, repository attachment, and task modification. A CI publisher should not automatically be able to redefine cleanup rules or compact blob stores.
11. Worked design: internal Maven, npm, PyPI, and proxy caches
Assume 80 developers, daily CI, monthly releases, and a 60-day rollback requirement. Storage grows 8 GiB/day, mostly CI prereleases. The team uses Community Edition today but may move to Pro/PostgreSQL later.
| Repository | Decision | Why |
|---|---|---|
| maven-dev-hosted | Prerelease/snapshot cleanup after measured 21–30 day window | High artifact velocity; releases elsewhere |
| npm-dev-hosted | Separate dev lifecycle policy; validate semver prerelease behavior | Package semantics differ from Maven |
| pypi-release-hosted | No automated release deletion until EOL workflow exists | Rollback/support requirement dominates storage |
| maven-central-proxy | Usage-based cache cleanup only after confirming client request behavior | Balance storage and outage resilience |
| docker-proxy | No Last Downloaded policy while 3.95.2 known issue applies | Usage metric can under-report active images |
If Pro/PostgreSQL is introduced, retain-last-N can become an additional guardrail for versioned repositories, but it does not replace release/EOL governance.
12. Decision record template
repository: maven-dev-hosted
owner: build-platform
purpose: ephemeral CI and SNAPSHOT artifacts
format: maven2
repository_type: hosted
retention_intent:
minimum_debug_window_days: 21
inactivity_days: 14
release_protection: "releases stored elsewhere"
execution:
preview_required: true
cleanup_window: "off-peak"
compact_after_validation: true
soft_delete_grace_days: 7
recovery:
source: "validated Nexus backup"
rpo: "24h"
restore_test: "quarterly"
known_issues:
- "re-check version-specific cleanup release notes before changes"
The YAML is documentation, not Nexus configuration. Its value is that a future operator can compare business intent with the actual policy before changing anything.
13. Design mistakes to avoid
- Applying one retention age to development, releases, and proxies because it is administratively convenient.
- Using last-downloaded without testing whether representative clients update it reliably.
- Running compaction immediately after cleanup, eliminating a useful soft-delete grace window.
- Assuming backups are valid because a job says “success” without restore testing.
- Using Pro-only retain-last-N as a hidden dependency in a Community operating runbook.
- Optimizing for reclaimed GiB while ignoring build-outage and rollback costs.
14. Knowledge check
Why might a proxy repository need longer retention than its raw storage cost suggests?
Cached dependencies can provide resilience during upstream outages. Aggressive eviction increases dependency on remote availability.
What is the advantage of separate development and release repositories?
The repository boundary itself expresses lifecycle intent, reducing the chance that an aggressive dev cleanup rule reaches long-lived releases.
Should cleanup and compaction always run at the same frequency?
No. Logical cleanup and physical reclamation have different performance and recovery consequences and can be scheduled independently.
What must happen before choosing a last-downloaded threshold?
Validate how the exact Nexus version and representative package clients update usage data, and review known issues.
Why does retain-last-N not guarantee a usable rollback?
Version ordering does not prove that the retained versions are approved, compatible, secure, or operationally deployable.
15. Summary and next step
A production retention policy is an SLA encoded as criteria, scope, schedule, recovery window, and evidence—not a universal age number. Lesson 4 takes these design choices into failure analysis: broad regexes, misunderstood timestamps, ignored preview, delayed reclamation, task overlap, and direct filesystem deletion.
Official references and version notes
- Sonatype: Cleanup Policies — current criteria, preview, system cleanup tasks, soft deletion, compact-blob-store reclamation, and format matrix.
- Sonatype: Nexus Repository 3.95.x Release Notes — 3.95 cleanup enhancements and current known issues.
-
Sonatype: Tasks
— current task types including
Admin - Compact blob store. - Sonatype: Keeping Disk Usage Low — supported cleanup/reclamation guidance and the separation between deletion and freed disk space.
- Sonatype: Administration Best Practices — cleanup cadence and storage-management recommendations.
- Sonatype: Staging Concepts — aligning repository lifecycle and cleanup.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.