Chapter 19Lesson 04225–310 min

Scheduled Tasks, Repair Jobs, Metadata Rebuilds, Blob Maintenance, and Operational Housekeeping: Diagnostics, Failure Modes, Security, and Performance

Diagnose maintenance incidents without escalating damage: preserve task evidence, identify the state layer, distinguish scheduler/privilege/disk failures from repository corruption, understand partial cancellation, and choose the least destructive supported correction.

DiagnosticsFailure analysisCancellationDisk pressureRecovery

Learning objectives

  • Apply a repeatable evidence-first diagnostic sequence to task failures.
  • Recognize stale/removed task guidance and version mismatch.
  • Diagnose privilege, destination, disk, concurrency, and log-volume failures.
  • Explain why cancellation may leave partial task effects.
  • Separate metadata/index repair from database/blob recovery.

Version baseline (26 August 2026). These lessons use Sonatype Nexus Repository 3.95.2 as a dated reference point and Java 21 as the current runtime requirement. New installations default to H2, while Sonatype recommends external PostgreSQL for production. The mandatory lab assumes a small, disposable, loopback-only self-hosted Community Edition instance installed from the distribution archive with H2 and a file blob store. Task names, task availability, edition entitlements, and repair guidance are version-sensitive, so inspect the task list on your exact instance before following any example.

Repair-task boundary. Current Sonatype task documentation says tasks prefixed Repair are intended for specific problems rather than routine schedules, and Sonatype recommends running them only with an identified need and, for production incidents, appropriate support guidance. In this chapter, routine labs use harmless Admin tasks. Repair tasks are inspected, reasoned about, or demonstrated only in disposable/simulated recovery contexts.

1. Preserve task evidence before changing the task

A failed maintenance job creates pressure to “try again.” Resist that reflex. Re-running an unknown task can overwrite the best evidence of the original failure or repeat a destructive partial action.

Capture: task name/type/ID, configuration, schedule, current state, last run/result, run-specific log, relevant nexus.log window, Nexus version/edition/runtime, database/blob store, free disk, repository scope, client symptom, and any concurrent jobs.

2. Diagnostic sequence

Task incident diagnosis
flowchart TB
E[Preserve task + log evidence] --> V[Confirm version / edition / runtime]
V --> T[Confirm task type still exists and is appropriate]
T --> P[Check privileges / task definition / scope]
P --> C[Check concurrent jobs + scheduler timing]
C --> D[Check DB / blob / disk / file handles]
D --> M[Identify state layer: search / browse / format metadata / DB / blob]
M --> L[Read task log + nexus.log]
L --> X[Least-destructive correction]
X --> Q[Independent repository/client verification]

3. Symptom-to-layer matrix

Symptom Likely layer First evidence Wrong shortcut
Task type missing from UI Version/edition/deprecation Release docs + live task menu Edit DB to recreate old task
Task creation/run denied Privilege Role/privilege matrix, HTTP/UI error Grant nx-admin
Task fails writing output Filesystem/destination/disk Task log, path ownership, free space Run Nexus as root
Builds slow during maintenance Resource contention Task overlap, DB/blob IO, request latency Blind heap increase
Search incomplete after cancellation Partial index rebuild Task log + exact download test Assume cancel rolled back
Blob bytes missing after manual deletion Unsupported storage mutation Blob/database consistency evidence Run every rebuild task

4. Failure mode: old documentation names an obsolete or unsafe task

An old runbook says: “Every Sunday run Reconcile Component Database From Blob Store.” On a current 3.95.x instance, this is a red flag. Current Sonatype guidance says the older reconcile workflow has been replaced by the Data Repair Plan/Execute tasks for recovery scenarios, and these should not be used during normal operation.

Correction:

  1. Disable the inherited schedule if it exists.
  2. Preserve why it was originally created.
  3. Check whether there is an actual restore inconsistency.
  4. If not, remove the recurring repair concept from the runbook.
  5. If yes, follow the current data-repair procedure for the exact version and recovery point.

5. Failure mode: task denied by privilege

A task operator receives a UI denial or an API 403. The correct response is to identify the required task privilege, not to switch to the administrator account. Current privileges distinguish read, run/start-stop, create, update, and delete.

Observer: nx-tasks-read
Runner:   nx-tasks-read + nx-tasks-run
Designer: add only the create/update/delete privileges actually required

Then add repository/blob-specific administration only where the task requires it. Validate with the least-privilege account before broadening access.

6. Intentionally broken lab example: H2 backup destination is not writable

Use only the disposable lab instance. Create a second H2 backup task whose destination is a lab directory that exists but is intentionally not writable by the Nexus process identity. Do not change permissions on the Nexus data directory, application directory, system directories, or shared production paths.

Run the task and capture the failure. Expected evidence is a failed result plus a task/nexus log message indicating inability to create/write backup output. Then restore the lab directory permissions, point the task to the validated writable backup destination, rerun, and verify success.

Portable alternative: if you cannot create a safely unwritable directory without administrative changes, use the provided synthetic failure fixture instead of altering OS permissions. The learning objective is diagnosis, not forcing a filesystem error at any cost.

2026-08-26 03:10:14 ERROR [h2.backup.task] destination=/lab/denied
java.nio.file.AccessDeniedException: /lab/denied
result=FAILED

7. Failure mode: overlapping cleanup, compaction, backup, or repair

A task can be technically successful and still degrade service because another heavy job is running. Look for overlapping timestamps in task history and logs, then correlate with database latency, blob IO, request latency, file handles, and free disk.

Correction is usually schedule arbitration, not heap tuning. Serialize jobs that contend for the same state, add buffer for duration variance, and keep a maintenance-window overrun rule.

8. Failure mode: metadata rebuild mistaken for blob recovery

Scenario: a component record exists, but the asset download fails because the blob file was manually deleted outside Nexus. Running Repair - Rebuild repository search cannot recreate missing bytes. Search is derived metadata; the binary is authoritative content.

Preserve the incident, stop unsupported manual deletion, identify the backup/restore state, and use the current recovery/data-repair procedure only if its prerequisites match. Chapter 25 goes deeper into backup/restore, but the key diagnostic rule is already clear: rebuild the layer that is broken.

9. Failure mode: long task generates excessive logs or fills disk

Task logs normally rotate on a time basis, but a long-running or repeatedly failing task can still generate significant output. If available disk approaches Nexus minimum requirements, the repository itself can enter protective/read-only behavior.

  1. Preserve the current task log and disk evidence.
  2. Stop launching duplicate retries.
  3. Determine whether the task is cancel-responsive and whether cancellation has partial-state consequences.
  4. Free space only through safe OS/log-retention actions and supported Nexus storage procedures.
  5. Do not delete database/blob content to make room.

10. Failure mode: operator cancels without understanding partial effects

The Tasks API explicitly warns that not all tasks respond to cancellation. Task-specific docs matter. For Repair - Rebuild repository search, Sonatype states that cancelling results in a partially rebuilt search index.

Assumption Reality
Cancel = transaction rollback False; completed work can remain.
Every task stops immediately False; some tasks may not respond to cancellation.
After cancel, old search index is restored False for search rebuild; index can be partial.
Stop endpoint success proves repository consistency False; verify affected state separately.

11. Data Repair Plan is a recovery workflow, not a troubleshooting hammer

Current data-repair guidance uses two phases: plan, review, then execute. It can be scoped by time and blob stores/repositories, and is intended to repair inconsistencies between database and blob evidence. It can take significant time and affect recovery timing.

A critical historical detail: Sonatype documented a 3.83.0–3.89.1 bug where Verify/Repair or Data Repair Plan tasks could incorrectly delete valid assets; the issue was fixed in 3.90.0. This is why recovery procedures must be version-gated rather than copied from memory.

12. Failure mode: manual blob/database deletion created corruption

Never “repair” this by making more manual edits. Preserve exact paths/rows already changed if known, isolate client traffic if required, validate backups, and follow supported recovery instructions. Direct database statements shown in specialized official procedures are not a license to invent your own cleanup SQL.

13. Performance diagnosis: measure the shared bottleneck

Metric What it can indicate Do not conclude automatically
CPU Indexing/compression/thread work High CPU means more heap is needed
Heap/direct memory Allocation pressure All slow tasks are memory-bound
DB latency Metadata scan/write contention Blob storage is healthy
Blob IO latency Scan/compaction/rebuild pressure Database is the cause
Free disk Backup/log/temp/reclamation headroom Deleting blobs manually is acceptable
Client latency User-visible contention The task itself failed

14. Evidence packet for support or postmortem

  • Nexus version/edition and Java/runtime.
  • Database and blob-store types.
  • Task name/type/ID/configuration/schedule.
  • Task log and relevant nexus.log slice.
  • Start/end/cancel timestamps and result.
  • Concurrent task list.
  • Repository/blob scope and representative component paths.
  • Disk/DB/blob/request metrics around the event.
  • Changes made after failure.
  • Redaction review before sharing support ZIPs or logs.

15. Knowledge check

An old runbook says to schedule a reconciliation Repair task weekly. What is your first action?

A backup task fails with AccessDenied. Should Nexus be run as root?

After cancelling search rebuild, why might Search be incomplete?

What task fixes blob bytes manually deleted from disk?

Why preserve task logs before retries?

16. Summary and next step

Task incidents are diagnosed by state layer and evidence, not by running more maintenance. You can now recognize version drift, privilege failures, disk/destination errors, task overlap, partial cancellation, and the boundary between metadata rebuild and true recovery. Lesson 5 integrates these skills into a maintenance-calendar checkpoint with two safe tasks and one controlled failure.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.