Chapter 05Lesson 03135–175 min

Blob Stores, Storage Layout, Database Choices, Capacity Planning, and Data Separation: Configuration, Design Choices, and Tradeoffs

Choose shared or separated blob stores, file or object storage, H2 or PostgreSQL, and capacity headroom with explicit reliability, security, performance, and recovery tradeoffs.

Storage architectureFile vs S3Capacity planningRecovery boundariesTradeoffs

Learning objectives

  • Choose between a single shared blob store and purposeful storage separation without creating unnecessary repository sprawl.
  • Compare local file blob stores with S3-style object storage using latency, security, cost, and recovery evidence.
  • Select H2 or PostgreSQL based on the supported workload and operating model rather than installation convenience.
  • Calculate capacity headroom from growth, retention, temporary work, cleanup behavior, and recovery needs.
  • Document neighboring-system boundaries so backups and client behavior are not mistakenly treated as Nexus storage.

Current lab baseline (reviewed 2026-08-26): Nexus Repository Community Edition 3.94.1-06, Java 21, loopback-only HTTP, a dedicated non-root/non-administrator process identity, and embedded H2 only for the disposable local instance. Sonatype currently recommends external PostgreSQL for production deployments. The lab never treats a blob store as a database backup or manipulates $data-dir/blobs or database files directly.

1. Start with failure domains, not folder names

A storage layout should make an operational promise explicit. Separate stores may isolate growth, retention, access policy, cost class, or recovery scope. A single store may reduce administrative overhead and make capacity easier to reason about. Neither is inherently “enterprise.” The design is good when the reason for each boundary is observable and supportable.

2. Shared versus purpose-specific blob stores

Pattern Strengths Costs / risks Good fit
One file blob store Simplest capacity, backup, permissions, and operations. One capacity/failure domain; all repositories compete for the same storage. Small instance with similar retention/performance needs.
Few purpose-specific stores Can isolate hot vs archive, internal vs proxy cache, or distinct storage classes. More monitoring, quota, backup, and capacity decisions. Different retention, performance, security, or recovery requirements.
Store per repository Maximum isolation in theory. Operational sprawl; inefficient cleanup/search; difficult capacity visibility; little value when requirements are identical. Only when a real boundary justifies it.
Group blob store Can combine member stores and spread writes under supported policies. Current Pro feature with extra topology/operations complexity. Licensed architectures that need multi-store growth/fill behavior.

More stores are not automatically safer. Sonatype’s planning guidance notes that excessive blob-store count can make cleanup/search and disk-capacity reasoning harder, especially when several file stores share one physical filesystem.

3. File storage versus S3-style object storage

Question File blob store S3-style object blob store
Latency Usually lowest when on fast local/nearby storage. Adds object API and network latency; same cloud/region placement matters.
Capacity growth Bound by filesystem/device planning. Cloud object capacity scales differently, but request/storage cost and provider limits matter.
Permissions OS path ownership/ACLs for Nexus process. IAM/bucket policy/credentials plus network reachability.
Failure evidence Filesystem errors, latency, free space, mount/permission state. HTTP/provider errors, IAM denial, endpoint/region/network latency.
Backup/recovery Protect blob content consistently with DB/config state. Provider durability/versioning is useful but still not a relational DB backup.
Edition boundary Core file blob stores are available. Current feature matrix supports AWS S3 in Community; other cloud/group blob options have edition-specific boundaries that must be rechecked.

The deciding question is not “cloud or not?” It is where Nexus runs, the latency budget, required durability, provider support, and the team’s ability to operate the storage backend. Cross-region or cross-cloud blob access can convert every package request into a network performance problem.

4. H2 versus PostgreSQL: convenience is not a production criterion

H2 keeps the disposable lab simple and is the current default for new installations. Current Sonatype requirements cap supported H2 workload and exclude H2 container deployments. PostgreSQL is recommended for production and separates database lifecycle, backup, capacity, monitoring, and availability from the Nexus process.

Decision signal Prefer disposable H2 Prefer external PostgreSQL
Learning / throwaway instance Yes, within documented limits. Usually unnecessary unless learning DB integration.
Production growth expected Poor choice as workload approaches documented bounds. Recommended production path.
Container deployment Current H2 container deployment is unsupported. Required direction for supported container/cloud-native patterns.
Recovery operations Simple topology but still must coordinate DB/blob consistency. Dedicated DB backup/restore, monitoring and low-latency network are explicit.
Scale / resiliency Single-node simplicity. Foundation for larger and HA-capable architectures, subject to edition/topology requirements.

5. Capacity headroom is a policy decision

Suppose a team retains 120 GiB of current blobs, adds 18 GiB/month, removes about 8 GiB/month through supported retention/cleanup, and wants twelve months before a capacity expansion. Net annual growth is:

\[(18-8)\times 12=120\text{ GiB}\]

Projected retained bytes become 240 GiB before overhead. If the organization requires storage to remain below 70% utilization to leave room for temporary work, cleanup lag, unexpected releases, and recovery operations, the capacity target is approximately:

\[240 / 0.70 \approx 343\text{ GiB}\]

Then add local data-directory/database/log/temporary requirements separately. The 4 GB read-only threshold is an emergency floor, not the design reserve.

6. Separation changes recovery—not the need for consistency

Separating a file blob store from the Nexus data disk can reduce the chance that artifact growth consumes application/database headroom. Separating PostgreSQL gives independent database operations. But every extra boundary creates a recovery dependency that must be documented: which backup point belongs with which blob snapshot? Which configuration recreates the repository-to-store mapping? Which credentials and IAM policies are required to read the object store?

A storage design is incomplete until its restore sequence is explainable without editing Nexus internals.

7. Keep neighboring state stores separate

Neighbor Not the same as Nexus storage Why confusion hurts
Maven/npm/pip/Docker client cache Local consumer cache. Can hide repository outages and is not an authoritative backup.
Reverse proxy / load balancer Network/TLS/routing layer. Changing it does not relocate repository blobs or DB state.
CI artifact retention CI platform-owned build outputs. May duplicate some bytes but does not recreate Nexus configuration/history.
Object-store replication Blob-backend durability feature. Does not protect PostgreSQL/H2 relational state by itself.
VM snapshot Infrastructure capture mechanism. Crash-consistency across DB and blob state must be validated; a snapshot label alone is not proof.

8. Worked architecture decision

A 35-developer team serves Maven, npm, and Raw artifacts from one office/private cloud. It expects 300 GiB in year one, has a fast local SSD array, and can operate PostgreSQL. Its proxy cache may be rebuilt from public upstreams, but internal release content must meet a stricter recovery objective.

A defensible design is:

  • External PostgreSQL for production relational state.
  • One file blob store for internal authoritative hosted content and a separate file blob store for reconstructable proxy cache only if the different recovery/retention policy justifies the boundary.
  • Capacity alerts well before 70–75% utilization plus the hard 4 GB emergency threshold.
  • Independent but coordinated database and blob protection with a tested restore runbook.
  • Repository mappings and storage assumptions recorded as configuration evidence.

Creating a separate blob store for every Maven/npm repository would add complexity without a stated failure or policy boundary.

9. Mini-lab: write a storage ADR

Write a one-page architecture decision record with these fields: workload assumptions, database choice, blob backend, repository-to-store mapping, capacity model, alert thresholds, process identity/permissions, backup dependencies, RPO/RTO assumptions, edition prerequisites, and a trigger for redesign. The acceptance criterion is that another operator can predict what fails when the database, one blob store, or the data filesystem is unavailable.

Knowledge check

Why can one blob store for many repositories be a good design?

When does splitting proxy-cache blobs from internal hosted blobs make sense?

Does S3 durability remove the need for database backup?

Why is the 4 GB disk rule not an adequate capacity threshold?

What should justify a blob-store boundary?

10. Summary

Storage architecture is a set of operating promises: where bytes live, where metadata lives, how fast each layer must be, how much headroom exists, and how recovery reconstructs a consistent service. The next lesson turns those promises into a diagnostic sequence for realistic storage failures.

Next lesson

Diagnose storage without destroying evidence

Interpret quota warnings, low disk, permissions, filesystem latency, and incomplete backups with the least destructive correction.

Official references and version notes

Version-sensitive statements were rechecked against Sonatype primary documentation on 2026-08-26. The current Download, versions-status, and 2026 release-notes pages list 3.94.1 as the newest GA/downloadable self-hosted line. Mandatory labs therefore pin Nexus Repository Community Edition 3.94.1-06 on loopback with Java 21 and embedded H2 only as a disposable learning database. Current system requirements recommend external PostgreSQL for supported production-scale deployments and require at least 4 GB of free disk at all times. Learners should re-check the live pages before executing the lab because Nexus support matrices evolve.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.