Chapter 05Lesson 04150–190 min

Blob Stores, Storage Layout, Database Choices, Capacity Planning, and Data Separation: Diagnostics, Failure Modes, Security, and Performance

Diagnose low disk, quota warnings, permissions, filesystem latency, incomplete backups, and accidental blob manipulation using an evidence-first, least-destructive sequence.

Storage diagnosticsSoft quotasRead-only modePermissionsBackup consistency

Learning objectives

  • Use the course-wide evidence-first sequence to separate storage symptoms from routing, authorization, cache, and client problems.
  • Interpret soft-quota warnings correctly and distinguish them from the 4 GB database read-only threshold.
  • Diagnose permissions and filesystem support without running Nexus as root or recursively changing internal ownership blindly.
  • Reject incomplete backup claims that protect only blob content or only database state.
  • Apply least-destructive corrections and verify recovery with a controlled repository request.

Current lab baseline (reviewed 2026-08-26): Nexus Repository Community Edition 3.94.1-06, Java 21, loopback-only HTTP, a dedicated non-root/non-administrator process identity, and embedded H2 only for the disposable local instance. Sonatype currently recommends external PostgreSQL for production deployments. The lab never treats a blob store as a database backup or manipulates $data-dir/blobs or database files directly.

Do not induce a real disk-full event. The low-disk case in this lesson uses a synthetic evidence fixture. Deliberately exhausting the filesystem can corrupt in-flight blobs and destabilize the lab host. The only live failure we induce is a harmless soft-quota warning on a disposable blob store.

1. Diagnostic sequence: prove the layer before touching storage

A 500 error or failed package install is not automatically a blob-store failure. Use the same disciplined sequence from earlier chapters:

  1. Preserve concise client/HTTP evidence and timestamp.
  2. Confirm Nexus version, edition, Java runtime, and process identity.
  3. Inspect client URL/auth and repository type/routing/group membership.
  4. Inspect authorization and component/asset/metadata state.
  5. For proxy requests, inspect cache/upstream state.
  6. Then inspect database, blob-store state, free disk, filesystem latency, and permissions.
  7. Inspect logs, tasks, and metrics around the same timestamp.
  8. Apply the least destructive correction and repeat one controlled request.

This order prevents an operator from deleting storage to fix what was actually a wrong repository endpoint or authorization denial.

2. Broken example: a soft quota warns but does not block writes

If Lesson 2 resources still exist, use academy-ch05-capacity. Otherwise recreate a disposable file blob store and Raw hosted repository mapped to it. In the UI, edit only the blob store’s soft quota and set Space Used to a value below its current used size, such as 1 MB after the 4 MiB test asset exists.

Prediction: quota status becomes violated and Nexus records a warning/status signal, but the soft quota itself does not block reads or writes.

# POSIX/Bash. Use only against the disposable loopback instance.
LAB="$HOME/nexus-ch05-lab"
NX_URL="http://127.0.0.1:8081"
mkdir -p "$LAB/evidence" "$LAB/payload" "$LAB/move"
umask 077
read -rsp "Disposable Nexus admin password: " NX_PASS; printf "\n"
printf 'machine 127.0.0.1 login admin password %s\n' "$NX_PASS" > "$LAB/nexus.netrc"
unset NX_PASS
NETRC="$LAB/nexus.netrc"

# Never print, commit, or reuse this temporary credential file.
curl --fail-with-body --silent --show-error --netrc-file "$NETRC"   "$NX_URL/service/rest/v1/blobstores/academy-ch05-capacity/quota-status"   | tee "$LAB/evidence/01-quota-violated.json"

printf 'quota-warning-does-not-block-write\n' > "$LAB/payload/tiny.txt"

curl --fail-with-body --silent --show-error --netrc-file "$NETRC"   --upload-file "$LAB/payload/tiny.txt"   "$NX_URL/repository/academy-ch05-archive/diagnostic/tiny.txt"

curl --fail-with-body --silent --show-error   "$NX_URL/repository/academy-ch05-archive/diagnostic/tiny.txt"   | tee "$LAB/evidence/02-tiny-read.txt"

Windows PowerShell: use the same loopback URLs with curl.exe. Keep the disposable credential in a user-only temporary credential mechanism rather than embedding a password in command history. Use Get-FileHash -Algorithm SHA256 for hashes and Get-Volume/Get-PSDrive for free-space evidence. Do not run Nexus as Administrator merely to bypass a permissions problem.

Interpretation: the configured soft quota is a monitoring control. If the upload succeeds, that is expected behavior. Restore/remove the disposable quota after the exercise instead of leaving a misleading warning.

3. Low disk is a different condition

Now interpret this synthetic evidence rather than filling your real disk:

# SYNTHETIC INCIDENT FIXTURE — do not try to reproduce by filling a disk.
Filesystem      Size  Used Avail Use% Mounted on
/dev/nvme0n1p3  500G  496G  3.6G 100% /srv/nexus-data

GET /service/rest/v1/status/writable
HTTP/1.1 503 Service Unavailable

Current requirements state that Nexus must maintain at least 4 GB free; below that the database switches to read-only. The correct response is not to delete arbitrary blob files. Preserve evidence, stop new workload where practical, identify supported reclaim/expansion options, restore safe headroom, verify database writable state, and then test one controlled repository operation.

Capacity monitoring should trigger well before 4 GB. A 500 GiB volume with only 3.6 GiB free has already violated a sensible operating reserve even before the product’s emergency behavior becomes visible.

4. Permission failures: fix the path, not the process identity

A Nexus process that cannot write a file blob path may log access-denied/permission errors. First prove the account and each path segment:

# POSIX evidence only; adapt paths to the disposable installation.
ps -eo user,pid,cmd | grep '[n]exus'
id nexus 2>/dev/null || true
namei -l /srv/nexus-blobs/academy-ch05-fast 2>/dev/null || true
findmnt -T /srv/nexus-blobs/academy-ch05-fast 2>/dev/null || true
ulimit -n

The wrong correction is sudo nexus, chmod -R 777, or a blind recursive ownership change over an unknown data tree. Fix the smallest intended path/ACL/mount configuration so the dedicated Nexus account has precisely the required access, then restart only if the supported configuration requires it and verify one write/read.

5. Reachable does not mean supported or fast enough

Symptom Evidence to collect Likely layer
Metadata/search/UI stalls while blob downloads look normal Database latency, embedded-data filesystem IO, JVM/log timing. Database / embedded-data storage.
Large downloads are slow but metadata queries are fast Blob backend throughput/latency, network, object-store metrics. Blob IO / network.
Intermittent “too many open files” errors Process file-descriptor limits and OS logs. OS process limits; integrity risk.
Access denied when creating blobs Process user, path owner/ACL, mount options. Filesystem permissions.
Cloud object requests return access denied IAM policy, bucket/container scope, credentials, provider logs. Object-store authorization.
Client seems fast despite Nexus outage Fresh/isolated client cache test. Client cache masking server behavior.

Do not “fix” slow storage by increasing JVM heap without measurements. Heap pressure, direct memory, database latency, blob IO, upstream latency, task load, and client cache are distinct causal layers.

6. Broken recovery claim: “the object store is replicated, so Nexus is backed up”

Object replication can protect blob objects from some storage failures. It does not capture PostgreSQL/H2 records, repository mappings, security/configuration state, or the relationship between a database point and blob content. Similarly, a database backup with missing blobs is incomplete.

Protected state Example evidence What remains missing
Blob store only S3 versioning/replication or file snapshot. Database/configuration consistency.
Database only PostgreSQL backup or supported H2 backup point. Artifact/package bytes in blob stores.
Application directory only Copied Nexus distribution. Nearly all repository persistent state.
Coordinated DB + blob + config Documented backup point plus tested restore validation. Still must prove the restored service and package reads.

Chapter 25 will teach full backup/restore mechanics. Here the required skill is recognizing the dependency so storage architecture does not create a false recovery claim.

7. Never use internal edits as a troubleshooting shortcut

Blob filenames and database structures are Nexus-managed implementation details. Removing a suspicious .bytes file, editing a database row, or copying one blob directory while Nexus is active can make the service internally inconsistent. If you need isolation, create a disposable repository/blob store or fresh client configuration first. If you need relocation, use the edition-appropriate supported operation.

8. Least-destructive correction matrix

Evidence First correction Verification
Soft quota violated, disk healthy Increase/adjust monitoring threshold or capacity plan; no emergency deletion. Quota status + controlled read/write.
Free disk below operating threshold Expand/reclaim through supported Nexus/storage procedures. Free space + writable status + controlled write.
Wrong blob path permissions Correct minimal ACL/ownership/mount for dedicated process user. Process user + path access + one upload.
Unsupported/slow filesystem Plan supported storage migration; avoid live manual copy. Supported topology + latency baseline + integrity check.
Only blobs were backed up Do not claim recoverability; create coordinated DB/config/blob recovery plan. Test restore later, not a file-count assertion.

9. Return the diagnostic lab to a neutral state

Remove the artificial soft quota through the Blob Store UI. Delete the tiny diagnostic asset/repository only through supported Nexus repository/component operations if you no longer need it. Keep or delete the Lesson 2 repositories according to your next exercise; if deleting them, delete repositories first and blob stores second.

rm -f "$NETRC"
unset NETRC

Knowledge check

A soft quota is exceeded and an upload still succeeds. Is Nexus broken?

What should you do when free disk falls below 4 GB?

Why is running Nexus as root a bad permissions fix?

A blob bucket is replicated across regions. What recovery state is still required?

Why can a warm client cache mislead a storage incident diagnosis?

10. Summary

Storage diagnosis is a process of separating layers and protecting evidence. Soft quotas, low disk, path permissions, file descriptors, database latency, blob IO, and backup consistency have different symptoms and different safe corrections. The checkpoint now combines capacity, placement, identity, and cleanup into one independent exercise.

Next lesson

Checkpoint: prove the storage operating model

Model growth, map repositories to stores, measure bytes, and document the recovery dependency.

Official references and version notes

Version-sensitive statements were rechecked against Sonatype primary documentation on 2026-08-26. The current Download, versions-status, and 2026 release-notes pages list 3.94.1 as the newest GA/downloadable self-hosted line. Mandatory labs therefore pin Nexus Repository Community Edition 3.94.1-06 on loopback with Java 21 and embedded H2 only as a disposable learning database. Current system requirements recommend external PostgreSQL for supported production-scale deployments and require at least 4 GB of free disk at all times. Learners should re-check the live pages before executing the lab because Nexus support matrices evolve.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.