Production Capstone: Design, Secure, Automate, Migrate, and Recover an Enterprise Artifact Platform: Security, Governance, and Reliability Validation
A platform can be perfectly automated and still be unsafe. This lesson tries to disprove the capstone design: can an unauthorized identity publish? Can an internal namespace escape to a public registry? Can cleanup remove a required release? Can operators restore the platform within the stated RPO/RTO? Can a future upgrade or licensed policy extension be introduced without breaking the trust model?
Learning objectives
- Validate least privilege with positive and negative authorization tests.
- Test namespace ownership and alternate-route bypass rather than assuming Nexus sees all traffic.
- Validate retention, cleanup preview, backup/restore, and RPO/RTO evidence.
- Distinguish checksums, signatures, SBOMs, vulnerability intelligence, and policy decisions.
- Produce a production-readiness scorecard with explicit Community/Pro/licensing dependencies.
capstone-lab-*, and every change has
preflight, evidence, verification, and rollback.
1. Validation mindset: try to falsify every control
“Configured” is not evidence. A role is validated when allowed operations succeed and forbidden operations fail. A retention policy is validated when preview selects exactly the intended disposable content. A backup is validated when an isolated restore serves the same controlled artifacts. A bypass control is validated when the alternate path is actually blocked or detected.
2. Authorization matrix
| Identity | Read group | Publish hosted | Change repositories | Security admin |
|---|---|---|---|---|
| capstone-reader | Allow | Deny | Deny | Deny |
| capstone-publisher | Allow as needed | Allow only internal hosted namespace | Deny | Deny |
| capstone-reconciler | Inspect as needed | Deny unless justified | Allow only capstone configuration scope | Deny broad admin |
| platform-admin | Allow | Emergency/administrative | Allow | Allow |
Do not accept a single broad administrator role because it makes the lab simpler. The capstone is specifically about separating duties.
3. Namespace and dependency-confusion validation
For @learner-example/* and com.example.*,
verify that internal publication has deterministic precedence and
that clients do not silently fall through to uncontrolled public
endpoints. Group member order, proxy routing rules, package-manager
configuration, DNS/egress policy, and developer fallbacks can all
influence the result.
NEGATIVE TEST
1. Request a synthetic internal name that exists only in the public-fixture route.
2. Expected: approved internal client path blocks or fails closed for the owned namespace.
3. If direct public access succeeds, repository policy is bypassable outside Nexus.
4. Fix client/network governance, then repeat the exact request.
4. Artifact evidence is layered
| Evidence | What it answers | What it does not prove |
|---|---|---|
| SHA-256/digest | Are these bytes identical? | Trusted origin or vulnerability-free status |
| Signature | Did a key sign the referenced data? | Key trust, secure build, vulnerability safety |
| SBOM | What components were declared/inventoried? | How the artifact was built |
| Provenance | Where/how/when was an artifact produced? | Absence of vulnerabilities |
| Vulnerability intelligence | What is currently known about risk? | Immutable artifact identity |
| Policy disposition | What does governance permit now? | Proof of malware or universal safety |
5. Licensed policy boundary
Repository Firewall is separately licensed and IQ-powered; it can integrate with supported Nexus Repository editions when properly licensed/configured. Native policy/quarantine is therefore an optional production control, not a hidden prerequisite for this free capstone. The mandatory path uses a deterministic fixture:
components=[
{"name":"alpha","critical":0,"license":"approved"},
{"name":"beta","critical":1,"license":"approved"},
{"name":"gamma","critical":0,"license":"review"},
]
def disposition(c):
if c["critical"] > 0: return "quarantine"
if c["license"] == "review": return "hold"
return "allow"
print([(c["name"], disposition(c)) for c in components])
assert [disposition(c) for c in components] == ["allow","quarantine","hold"]
6. Retention versus rollback and forensics
Development artifacts may have short retention; releases, incident evidence, regulated builds, or rollback candidates may need longer retention. Last-downloaded data, age, version matching, and format-specific criteria have different semantics. Preview before destructive cleanup and distinguish logical deletion from later blob reclamation. Never reclaim disk by deleting blob files directly.
7. RPO/RTO validation
Assume the capstone business objectives are RPO ≤ 30 minutes and RTO ≤ 60 minutes for the release repository. Define them precisely:
- RPO: maximum acceptable data loss measured from the restored consistency point.
- RTO: elapsed time from declared recovery start until validated service meets the defined acceptance tests.
Do not claim these targets from architecture diagrams. Measure them in the Lesson 4 recovery drill.
8. Backup completeness gate
A complete recovery set includes coordinated database/configuration state and blob content plus the node-ID material/current custom configuration needed by the chosen topology. Repository export/import is not a full-instance backup. Record database type, blob locations, backup consistency point, Nexus version, and restore target before calling a backup “complete.”
9. Upgrade and migration governance
Every change ticket that alters Nexus version, Java/runtime, database, blob backend, plugin/extension, or HA topology must include:
- source version/database/blob/edition;
- target version/database/blob/edition;
- crossed-version upgrade procedures;
- release-note/known-issue review;
- space and runtime prerequisites;
- backup/rollback checkpoint;
- rehearsal result;
- post-change smoke/client/content checks.
10. Current runtime/upgrade gate
Nexus Repository requires Java 21 on the current line. Official distributions from 3.87+ use a bundled Java 21 runtime; old custom JVM truststore additions are not automatically inherited. A production validation checklist must therefore test proxy/LDAP/SAML/TLS trust after an upgrade rehearsal, not merely confirm that the process starts.
11. Availability versus recovery
If the production requirement justifies Pro HA, validate node/LB/database/blob failure behavior separately. A successful failover proves availability behavior, not data recoverability. A corrupted blob can remain highly available and still be wrong. Backup/restore remains mandatory for corruption/deletion scenarios.
12. Observability acceptance
Before go-live, prove that operators can correlate one controlled
request across the evidence layers appropriate to the topology:
client time/status, request.log, outbound request for a
proxy, Nexus application logs, Prometheus/service metrics,
PostgreSQL evidence, blob/storage timing, scheduled tasks, and
capacity headroom. Status API success alone does not prove DB/blob
health.
13. Governance validator fixture
controls={
"least_privilege_negative_tests": True,
"owned_namespace_bypass_blocked": True,
"cleanup_preview_reviewed": True,
"backup_restore_tested": True,
"artifact_hashes_recorded": True,
"upgrade_rehearsed": True,
"support_evidence_redacted": True,
"ha_is_not_backup": True,
"licensed_features_labeled": True,
}
failed=[k for k,v in controls.items() if not v]
print("GO" if not failed else "NO-GO", failed)
assert not failed
14. Decision table
| Constraint | Design choice | Evidence required |
|---|---|---|
| Small team, modest load, no HA requirement | Simple single-node topology; consider PostgreSQL for production discipline | capacity + backup/restore proof |
| Strict uptime requirement | Evaluate Pro HA + supported shared state | licensed prerequisites + failure tests + backup still present |
| Strong procurement policy | Optional Firewall/IQ + network bypass controls | policy disposition + waiver governance + bypass test |
| Frequent releases | immutable versioning + exact-byte promotion + CI service identity | build ID/hash/promotion record |
| High storage growth | format-aware retention + capacity forecasting | cleanup preview + growth trend + restore headroom |
Knowledge check
What is the strongest evidence that least privilege works?
Positive tests for required operations plus negative tests proving forbidden operations are denied.
Why can a correct Nexus proxy still fail to prevent dependency confusion?
A client may bypass Nexus through a direct public registry route unless client and network governance close that path.
Does a quarantined component prove malware?
No. It is a policy/governance disposition based on current intelligence and configured rules.
What proves an RTO target?
A timed recovery drill that reaches the defined validation criteria within the target, not an architecture estimate.
Why must HA and backup both exist when both availability and corruption recovery matter?
HA addresses selected service failures; backup/restore provides an earlier coherent state after corruption/deletion. They solve different failure classes.
Summary and next step
The capstone implementation has now been challenged with authorization, namespace, evidence, policy, retention, recovery, upgrade, availability, and observability validation gates.
Lesson 4 injects controlled failures across those exact boundaries and requires evidence-driven diagnosis plus recovery rather than optimistic design claims.
Official references and version notes
- Sonatype: Nexus Repository documentation — current self-hosted product entry point and supported-format documentation.
- Sonatype: 2026 self-hosted release notes — release-specific changes and known-issue/upgrade guidance.
- Sonatype: Self-Hosted Feature Matrix — Community versus Pro capability boundary.
- Sonatype: System Requirements — Java 21, H2 workload limits, PostgreSQL guidance, sizing, file handles, and deployment constraints.
- Sonatype: Repository Types — hosted, proxy, and group responsibilities.
- Sonatype: Content Selectors — fine-grained content authorization concepts.
- Sonatype: Cleanup Policies — retention criteria, preview/evaluation, deletion semantics, and blob reclamation boundary.
- Sonatype: REST and Integration API — documented REST/OpenAPI automation boundary.
- Sonatype: Webhooks — repository/global events and delivery semantics.
- Sonatype: Staging — current Pro-only staged component movement/promotion capability.
- Sonatype: Repository Firewall — separately licensed IQ-powered policy integration and supported Nexus editions.
- Sonatype: Prepare a Backup — coordinated database/blob backups and node-ID preservation.
- Sonatype: Database migration — current migration tooling and compatibility gates.
- Sonatype: Instance Migrator — current source/target migration workflow and constraints.
- Sonatype: Upgrade Paths — version thresholds and mandatory crossed-version procedures.
- Sonatype: Rolling Upgrades in HA — Pro/HA rolling-upgrade behavior and mixed-mode constraints.
- Sonatype: HA system requirements — Pro-only active-node topology, shared PostgreSQL/blob state, and failure-domain requirements.
- Sonatype: Logging — application/request/outbound/audit/JVM evidence.
- Sonatype: Prometheus — current metrics endpoint and privilege boundary.
- Sonatype: Support Features — support ZIP generation and evidence handling.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.