Chapter 01Lesson 03~100 minutes

Artifact Repository Foundations, Package Supply Chains, Components, Assets, and Repository Managers: Configuration, Design Choices, and Tradeoffs

Design a Nexus repository operating model by making explicit tradeoffs around centralization, retention, immutability, endpoints, availability, edition boundaries, storage, and recovery.

Architecture decisionsRetentionImmutabilityCommunity vs ProRecovery

Learning objectives

  • Translate repository architecture choices into concrete Nexus/client/storage state.
  • Choose boundaries for centralization, retention, immutability, endpoint exposure, and availability.
  • Separate Community capabilities from Pro-only features without making the free path second-class.
  • Explain why one logical artifact hub still needs format-specific client endpoints.
  • Use capacity, recovery, security, developer experience, and upgrade cost as first-class design criteria.
  • Produce a defensible small-team repository decision record.

Design rule. Repository topology is policy encoded as URLs, repository types, member order, write rules, credentials, storage assignment, routing, retention, and recovery dependencies. If the organization cannot explain those choices, the topology will eventually become accidental infrastructure.

1. Centralization versus developer autonomy

Centralization reduces configuration drift: teams can share approved proxy endpoints, internal hosted namespaces, access policy, retention, audit evidence, and backup/recovery procedures. The tradeoff is that the artifact platform becomes critical shared infrastructure. Poorly designed centralization can also turn every harmless package request into a platform-team ticket.

Autonomy can speed experimentation, but dozens of unmanaged direct-upstream configurations create duplicated caches, inconsistent provenance, forgotten credentials, and weak incident response. The useful design target is usually central policy with delegated, format-appropriate self-service, not either extreme.

Choice Nexus/client state created Benefit Failure cost / control
Central shared proxy/group endpoints Common repository URLs, shared cache, shared routing/access policy Consistent dependency path and fewer duplicate upstream downloads Outage blast radius grows; requires capacity, observability, backups, and ownership.
Team-specific groups over shared proxies More group configs but reusable proxy members Different teams can expose only required repositories Group sprawl and member-order drift require naming/onboarding standards.
Direct upstream per developer/CI agent Minimal Nexus dependency Low initial platform effort Scattered trust decisions, no central cache/policy/evidence, harder outage containment.

2. Binary retention versus storage cost

Retaining every byte forever feels safe until blob stores become operationally expensive. Deleting aggressively feels efficient until a rollback, audit, incident investigation, or isolated rebuild requires a component that no longer exists. Retention must follow artifact class and business/recovery requirements.

Classify content before writing cleanup policy: immutable production releases, development snapshots/prereleases, proxy cache entries, large OCI/container layers, CI-generated temporary packages, third-party vendor components, and legal/audit retention sets have different value. Cleanup policy and blob compaction are separate mechanisms; later chapters teach them in detail.

A useful decision record states the retention objective, evidence that a component is safe to delete, recovery path if deletion is wrong, and how blob-space reclamation is verified. “Disk is full, delete old stuff” is not a policy.

3. Immutable releases versus mutable development versions

Stable release coordinates should identify stable bytes. If a pipeline can overwrite 1.0.0 after consumers have resolved it, the coordinate no longer tells you which binary was tested. Immutable release policy therefore protects traceability and rollback.

Development flows sometimes need mutable semantics—Maven SNAPSHOTs, pre-release channels, rapidly replaced test packages, or rolling integration tags. Those should live in explicitly development-oriented repositories or namespaces, with retention and client expectations that reflect mutability. Do not weaken a release repository because a development workflow is inconvenient.

Recommended identity split

release repository:
  com.example:ledger-api:1.4.0  -> write once / stable bytes
  evidence: SHA-256 + source commit + build ID

development repository:
  com.example:ledger-api:1.5.0-SNAPSHOT (format-specific semantics)
  evidence: timestamp/build identity + retention window

consumer rule:
  production promotion references an accepted immutable candidate,
  not "whatever bytes version 1.4.0 resolves to today".

4. One logical hub versus format-specific endpoints

“One repository URL” is often used loosely. A Maven group URL cannot serve as an npm registry or OCI registry simply because they live on the same Nexus server. Each format has protocol-specific requests and metadata. Centralization should mean one managed artifact service and predictable endpoint conventions—not protocol erasure.

One managed service, several protocol endpoints
flowchart TB
N[Nexus Repository service]
N --> M[Maven group URL]
N --> J[npm group URL]
N --> P[PyPI group URL]
N --> C[OCI or Docker endpoint]
N --> U[NuGet group URL]
M --> MC[Maven clients]
J --> NC[npm clients]
P --> PC[pip clients]
C --> OC[OCI clients]
U --> UC[NuGet clients]

Within one format, a group can reduce client configuration. Sonatype's current docs emphasize that group member order matters. Keep groups intentionally small enough to understand; every proxy member can add remote checks, cache state, credentials, and policy surface.

5. Availability versus strict governance

A repository platform must remain available enough that builds and deployments can function, but “availability” is not an excuse for bypass endpoints. If developers are told to switch directly to public registries whenever Nexus has a problem, the strongest security/routing controls disappear exactly during incidents.

Design failure modes explicitly: what can continue from proxy cache during an upstream outage; what happens when Nexus becomes read-only; which CI jobs may publish; which consumers are read-only; how credentials are rotated; what the restore objective is; and whether a Pro HA architecture is justified by business requirements. Community labs should learn the operating principles without pretending HA is free.

6. Community versus Pro: gate features precisely

As of this lesson's verification date, Sonatype's self-hosted feature matrix gives Community Edition substantial production-relevant basics: repository formats, REST/API coverage, component search, custom access controls, content selectors, routing rules, LDAP, external PostgreSQL, AWS S3 blob stores, and more. The same matrix marks features such as High Availability deployment options, SAML, user tokens, staging/build promotion, content replication, and certain blob-store capabilities as Pro.

Need Community/free path in this course Optional Pro/enterprise extension
Repository access control Users/roles/privileges/content selectors; LDAP where appropriate SAML/enterprise identity features where licensed.
Automation REST APIs and CI credentials Pro-specific staging/promotion capabilities when licensed.
Database H2 for bounded supported use; external PostgreSQL supported/recommended Same database fundamentals; HA topology adds licensing/architecture constraints.
Resilience Backups, tested recovery, capacity, runbooks, restore drills HA/resilient deployment features and zero-downtime patterns where licensed.
Supply-chain governance Repository routing, least privilege, checksums/digests, external scanners as optional integrations Repository Firewall/IQ policy capabilities where separately licensed.

Licensing can change. The correct habit is to verify the current matrix before promising an architecture. A course should label a Pro boundary, not quietly substitute a paid feature into mandatory instructions.

7. Runtime/database/blob choices affect operations even in a concepts chapter

Current Sonatype system requirements matter to design. Nexus Repository requires Java 21. New installations use H2 by default, but Sonatype recommends external PostgreSQL and documents H2 limits of 200,000 requests/day or 100,000 components. Container-based deployments are not supported with H2. Disk pressure also has service behavior: current docs require at least 4 GB free, below which the database can switch to read-only mode.

These facts do not mean every small lab needs PostgreSQL. They mean an architecture document must not treat the embedded database as infinitely scalable or container-neutral. Likewise, blob storage is not “just a folder”: throughput, capacity, backup, compaction, object-store support, and edition constraints affect recovery and performance.

8. Separate adjacent systems rather than configuring Nexus to compensate for them

Concern Primary owner Nexus boundary
Package-manager local cache Client/CI agent Nexus cannot make a stale local cache disappear; isolate or invalidate client state deliberately.
TLS/public ingress Reverse proxy/load balancer + Nexus HTTP config Nexus authorization is not a substitute for correct TLS and network exposure.
Enterprise identity lifecycle IdP/LDAP + Nexus realms/roles Map identities/roles; do not copy long-lived shared passwords into package configs.
CI orchestration CI platform CI invokes clients/API with scoped credentials; Nexus does not decide build/test success.
Database operations PostgreSQL/H2 operational layer Use supported backup/migration procedures; do not edit Nexus rows directly.
Blob storage Filesystem/object store operational layer Nexus tracks blob-backed content; storage replication/deletion outside supported procedures can break consistency.

9. Worked decision: a 40-developer platform team

Constraints: Maven and npm workloads; internal releases must be reproducible; public dependency access should go through Nexus; only free/self-hosted capabilities are mandatory; a four-hour recovery target is acceptable; the team cannot justify Pro HA yet.

Decision Choice Reason / evidence to retain
Topology Separate Maven and npm hosted/proxy/group sets Protocol-specific endpoints with one managed platform; predictable producer/consumer flow.
Release mutability Dedicated write-once-style release repos; development repos separate Prevents stable coordinates from silently changing.
Routing Private namespaces blocked from public proxies Reduces dependency-confusion exposure; routing policy is explicit rather than relying only on group order.
Database External PostgreSQL for production target Aligns with current Sonatype recommendation and leaves H2 for bounded lab/small use.
Resilience Tested backup/restore + documented RTO/RPO Community-compatible; no false claim of HA.
Authentication Least-privilege publisher/reader roles; no shared admin client credentials Limits blast radius and makes CI publication auditable.
Retention Long release retention; shorter dev/proxy cache lifecycle Balances rollback/audit needs with storage cost.

10. Design lab: write an architecture decision record

Create evidence/ch01-l3/repository-adr.md. Do not change Nexus. Your ADR must name two package ecosystems, define hosted/proxy/group endpoints for each, choose release mutability and retention principles, describe routing for private namespaces, identify Community/Pro boundaries, record database/blob assumptions, and define recovery ownership.

Then perform a read-only check against your training instance: list existing repositories and mark which parts of the ADR already exist, conflict, or are unknown. Unknowns are valid findings; do not “fix” a shared instance to make your paper design look correct.

Verification: every decision maps to observable state (URL, type, member order, role, storage assignment, policy, evidence); no cross-format group is invented; no Pro-only feature is required; backup/recovery includes both metadata/database and blob content conceptually; production secrets are absent.

Knowledge check

Why is “one Nexus URL for all package managers” an unsafe design statement?

A team needs SAML and active-active/high-availability behavior but has Community Edition. What should the design document do?

Why separate release and development repositories even when both use the same package format?

What is wrong with using group member order as the only defense against dependency confusion?

When is H2 a design red flag?

11. Summary

A repository platform is a collection of explicit tradeoffs: central control versus autonomy, retention versus cost, immutable release identity versus mutable development flows, one managed service versus protocol-specific endpoints, and availability versus governance. Edition, Java/database/blob support, client caches, identity, CI, and reverse-proxy concerns must remain correctly bounded.

Next lesson

Diagnose failures without destroying evidence

Lesson 4 turns the architecture into a troubleshooting sequence and exercises repository-type, provenance, cache, authorization, storage, and performance failure modes.

Official references and version notes

Version-sensitive statements were rechecked against Sonatype primary documentation on 2026-08-26. The mandatory path remains self-hosted, Community/free-compatible, and disposable; production credentials, production repositories, and paid-only capabilities are outside the lab boundary.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.