Chapter 30Lesson 04320–440 min

Production Capstone: Design, Secure, Automate, Migrate, and Recover an Enterprise Artifact Platform: Failure Injection, Troubleshooting, and Recovery Drill

Production confidence comes from controlled failure, not from a green dashboard. This lesson runs a disposable incident drill through the same diagnostic sequence used throughout the course: preserve concise evidence, confirm version/topology, isolate routing and authorization, inspect content/cache/upstream, inspect database/blob/disk, inspect logs/tasks/metrics, apply the least destructive correction, and verify with a controlled client request.

Failure injectionTroubleshootingRecoveryEvidenceRollback

Learning objectives

  • Diagnose credential, namespace, upstream, storage, cleanup, migration/upgrade, database/blob, network, policy, and observability failures.
  • Preserve evidence before mutation and reject fixes that do not match the observed layer.
  • Restore coherent database/blob state in an isolated recovery target.
  • Distinguish service failover from data recovery and prove artifact identity after correction.
  • Produce an incident timeline and recovery evidence bundle suitable for operational handoff.
Dated capstone baseline (27 August 2026). The lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. The live release notes, feature matrix, format support, migration/upgrade gates, and edition/licensing requirements must be re-checked before any production implementation.
Two targets, one operating model. The production architecture may include external PostgreSQL, Pro HA, licensed Firewall/IQ, enterprise identity, and shared/object storage. The mandatory lab remains Community/free-compatible and disposable, using synthetic namespaces and fixtures where a paid or infrastructure-heavy capability would otherwise be required.
Safety boundary. Never run this capstone against an employer/production Nexus instance, production database/blob store, real SSO, CI secrets, public DNS, valuable package namespace, or recovery backup. Destructive operations are limited to resources explicitly named capstone-lab-*, and every change has preflight, evidence, verification, and rollback.

1. Failure-injection rules

  • Use only synthetic capstone-lab-* resources and fixture files.
  • Never expose a real credential. “Leak” only token-FAKE_DO_NOT_USE.
  • Never attack or overload public registries; simulate upstream failure locally.
  • Never corrupt a real Nexus database/blob store; use copied fixtures.
  • Record a baseline fingerprint before every destructive simulation.
  • Apply one correction per hypothesis and repeat the same verification request.

2. Incident matrix

Incident Primary symptom First evidence Wrong shortcut
Credential leak secret appears in synthetic log scope/identity + access history hide line but keep credential valid
Namespace collision unexpected public artifact selected client endpoint + group/order/routing delete random public cache files
Upstream outage proxy miss fails; warm cache may work outbound request/cache state disable TLS/routing controls
Storage exhaustion writes/tasks fail or quota alerts blob/disk usage + task activity delete blob files manually
Cleanup misconfiguration needed disposable release candidate deleted policy criteria/preview/task history run compaction again
Upgrade failure startup/search/client regression version/runtime/release-note gate blind downgrade of upgraded data
DB/blob mismatch metadata exists but asset bytes missing component/asset + blob/hash evidence edit DB rows/files
Node/network failure service path unavailable LB/node/DB/blob dependency health call independent servers “HA”
Policy bypass direct upstream succeeds outside Nexus client/network route tighten Nexus policy only
Observability gap cannot correlate request logging/metrics configuration change multiple settings before preserving evidence

3. Diagnostic sequence as an executable decision discipline

PRESERVE EVIDENCE
→ VERSION / EDITION / JAVA / DB / BLOB / NODE TOPOLOGY
→ CLIENT URL + AUTH
→ REPOSITORY TYPE / GROUP / ROUTING
→ AUTHORIZATION
→ COMPONENT / ASSET / METADATA / DIGEST
→ PROXY CACHE / UPSTREAM
→ DATABASE / BLOB / DISK
→ LOGS / TASKS / METRICS
→ LEAST-DESTRUCTIVE CORRECTION
→ SAME CONTROLLED REQUEST + HASH/STATE VERIFICATION

4. Synthetic incident fixture

baseline={
 "version":"3.95.2-01","java":21,"client_url":"nexus-group",
 "auth":"reader-ok","route":"internal-first","db":"ok","blob":"ok",
 "upstream":"ok","policy_path":"nexus-only","artifact_sha":"abc123"
}
incidents={
 "credential_leak": {**baseline,"auth":"token-exposed"},
 "namespace_bypass": {**baseline,"client_url":"public-direct","policy_path":"bypassed"},
 "upstream_outage": {**baseline,"upstream":"down"},
 "blob_loss": {**baseline,"blob":"missing","artifact_sha":None},
 "upgrade_gate": {**baseline,"java":17},
}

def classify(s):
    if s["auth"] == "token-exposed": return "credential"
    if s["client_url"] == "public-direct": return "routing-bypass"
    if s["java"] != 21: return "runtime-upgrade-gate"
    if s["blob"] != "ok": return "storage-consistency"
    if s["upstream"] != "ok": return "proxy-upstream"
    return "unknown"

for name,state in incidents.items():
    print(name, classify(state))
assert classify(incidents["blob_loss"]) == "storage-consistency"

5. Incident A: synthetic credential leak

Evidence contains token-FAKE_DO_NOT_USE. Correct response: identify the owning service identity, revoke/rotate the disposable credential, search the synthetic evidence/history for use, replace it through the secret injection path, and rerun the least-privilege test. Deleting the log line without invalidating the credential does not remediate exposure.

6. Incident B: internal namespace bypass

The client config was edited to point directly to a public registry. Nexus logs show no request. That absence is itself evidence: changing Nexus roles or routing rules cannot control a request that bypasses Nexus. Correct the client/egress path, then repeat the request and confirm it traverses the approved group/proxy boundary.

7. Incident C: upstream outage with cache distinction

A known cached artifact remains available while a never-requested version fails during a simulated upstream outage. Diagnose proxy cache state, negative cache, auto-block/remote health, upstream/network evidence, and package-client cache separately. Do not globally disable TLS validation to “fix” remote access.

8. Incident D: storage exhaustion

When a disposable blob fixture reaches its operational threshold, stop new lab writes, preserve task/storage evidence, identify reclaimable versus live content, evaluate cleanup policy on synthetic data, then run supported reclamation only after logical deletion where appropriate. Direct filesystem deletion is never the recovery mechanism.

9. Incident E: cleanup deleted a needed candidate

This is a recovery/governance failure, not a reason to run more cleanup. Preserve the policy and task history, determine whether the artifact exists in a coherent backup/promotion target, restore or republish only from the authoritative build evidence, narrow the policy, preview again, and add a retention test to change review.

10. Incident F: failed migration/upgrade

If a rehearsal fails a Java/version/database/plugin/truststore gate, stop. Do not “see if production works anyway.” Restore the rehearsal clone or untouched source, correct the prerequisite, and rerun. If newer binaries already changed persistent state, rollback means restoring the verified pre-change checkpoint according to the supported procedure—not starting old binaries blindly.

11. Incident G: database/blob inconsistency

Fixture database metadata lists capstone-lib-1.0.0.jar but the corresponding blob is missing. A database-only restore cannot fix the content. Recover a coherent database/blob pair from the same recovery point and validate through Nexus/client requests plus SHA-256.

from hashlib import sha256
backup={"db":{"asset":"capstone-lib-1.0.0.jar","sha":""},"blob":b"capstone-release-bytes\n"}
backup["db"]["sha"] = sha256(backup["blob"]).hexdigest()
restored_blob = backup["blob"]
assert sha256(restored_blob).hexdigest() == backup["db"]["sha"]
print("coherent-restore: PASS")

12. Incident H: node/network failure versus data failure

In the optional HA simulation, a node failure can be masked by healthy peers. A database writer failure requires database-layer failover. A shared blob outage can affect every node. Logical corruption can be served successfully from every node. Diagnose failure domain before declaring “HA worked” or “HA failed.”

13. Incident I: insufficient observability

If you cannot correlate a controlled request, add the smallest bounded evidence source required for the next hypothesis—such as request logging, a narrow logger, or DB/blob timing—then reproduce once. Do not enable all DEBUG logging indefinitely. Review any support ZIP before sharing because it can contain sensitive operational context even after password sanitization.

14. Timed recovery drill

events=[
 (0,"declare incident"),
 (4,"freeze writes and preserve evidence"),
 (12,"select verified recovery point"),
 (24,"restore database and blob fixture"),
 (38,"start isolated target"),
 (46,"repository/client/hash validation complete"),
]
rto_minutes=events[-1][0]
backup_age_minutes=22
print("RTO",rto_minutes,"minutes")
print("observed data loss window",backup_age_minutes,"minutes")
assert rto_minutes <= 60
assert backup_age_minutes <= 30

This proves the synthetic drill meets the capstone's RTO ≤ 60 min and RPO ≤ 30 min assumptions. Real production targets require measured production-grade recovery rehearsals.

15. Incident evidence bundle

capstone-incident/
  timeline.md
  version-topology.txt
  client-request-redacted.txt
  routing-authz-state.json
  artifact-hashes.json
  cache-upstream-evidence.txt
  db-blob-capacity-evidence.txt
  logs-tasks-metrics-redacted.txt
  hypothesis-and-rejected-fixes.md
  correction.md
  verification.md
  recovery-rpo-rto.md

Knowledge check

Why is deleting an exposed secret from a log insufficient remediation?

What does “no request in Nexus logs” suggest when a dependency still downloads?

Why can a database-only restore leave repository content broken?

What is the safe rollback when an upgrade changed persistent state?

How did the fixture prove the capstone RPO/RTO assumptions?

Summary and next step

The failure drill proved that the capstone can classify incidents by layer, preserve evidence, reject folklore fixes, recover coherent state, verify immutable artifact identity, and measure synthetic RPO/RTO.

Lesson 5 turns that technical proof into a final production-readiness review and operational handoff: ownership, runbooks, evidence retention, change gates, capacity, security, recovery, and course completion.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.