Chapter 39Lesson 04~210 minutes

LTS Upgrades, Java Runtime Transitions, Plugin Upgrade Strategy, Compatibility Testing, and Rollback Planning: Diagnostics, Failure Modes, Security, and Performance

Preserve the first failed startup, plugin dependency error, agent rejection or integration regression, then locate the failing compatibility boundary before retrying.

diagnosticsstartup logsplugin dependencyagent Javarollbacksecurity

Learning objectives

  • Diagnose unsupported Java before touching plugins.
  • Separate core startup failure from plugin load/configuration failure.
  • Interpret agent incompatibility independently from controller health.
  • Recognize why no tested backup and blind plugin updates are incident multipliers.
  • Preserve security and performance evidence during upgrade troubleshooting.

1. Evidence-first diagnostic sequence

  1. Freeze the exact failed target: core, Java, plugin manifest, backup ID and change timestamp.
  2. Preserve first startup/controller logs before restart.
  3. Confirm the controller JVM is supported for the target core.
  4. Inspect failed/degraded plugin loading and dependency messages.
  5. Confirm JCasC/system properties/config migrations.
  6. Run minimal controller/API/auth checks.
  7. Inspect agent Java/Remoting/launcher separately.
  8. Run representative Pipeline/library/SCM/artifact tests.
  9. Compare performance/queue metrics to baseline.
  10. Choose the smallest repair, rollback from snapshot, or forward-fix.

2. Failure 1: target Jenkins on unsupported Java

Crossing 2.555.1 while the Jenkins process still uses Java 17 is a prerequisite failure, not a plugin failure. The correct repair is to run the Jenkins system on Java 21 or 25, then retry from the preserved state.

Observed layer: controller runtime
Expected target: Jenkins 2.568.3 on Java 21/25
Observed: Java 17
Action: stop; correct runtime; do not update random plugins
Evidence: java -version + controller startup log + exact image/package

3. Failure 2: core starts but a plugin fails to load

Do not infer that the whole core upgrade is invalid. Record plugin short name/version, dependency error, minimum-core requirement and whether the plugin has persisted configuration migrations. Resolve the dependency graph in staging. If the affected plugin controls authentication, credentials, Script Security, SCM or agents, treat it as a high-trust incident.

Never “solve” plugin failures by deleting .jpi files until you have proven what depends on them and have a recoverable snapshot.

4. Failure 3: yesterday’s plugin baseline is no longer secure

The 16 September 2026 advisory demonstrates why upgrade review includes same-day advisories. Script Security 1415, Pipeline: Groovy Libraries 805 and Pipeline: Multibranch 841 or earlier are below the fixes named in that advisory. A compatibility test that passes on a vulnerable plugin version is not an acceptable production gate.

Security fixes still require compatibility testing; “security update” does not mean “skip staging.”

5. Failure 4: controller healthy, agents offline after cutover

If the controller target requires Java 21/25 and an agent image still runs Java 17, classify the failure at the agent runtime layer. Preserve node name, launcher/Remoting output, Java version, image digest and label pool. Fix the agent image/runtime; do not rebuild the controller.

Evidence What it proves What it does not prove
Controller UI/API healthy Controller process is serving requests Agents are compatible
Agent log rejects Java/runtime Agent JVM boundary failed Pipeline itself is bad
Node comes online after Java 21 image Runtime repair worked Every tool/build JDK is compatible
Smoke job passes on one pool That pool/job path works All OS/privileged/signing pools work

6. Failure 5: naive downgrade after data migration

Suppose the target core/plugin starts, writes updated configuration, then an external integration fails. Starting the old core against that mutated target volume is not a controlled rollback. Older code may not understand newer serialized/plugin state. Stop the target and restore the pre-upgrade snapshot into a clean recovery location.

7. Failure 6: “update all plugins” destroys causality

If 35 plugins change at the same time as core and Java, a broken Pipeline gives you dozens of candidate causes. Preserve manifests and reduce the mutation set. On a clone, restore the known baseline and replay the planned batches until the failing batch is isolated.

8. Performance regression is a compatibility failure too

An upgrade can be functionally correct yet operationally unacceptable: queue wait doubles, controller GC pauses rise, or Pipeline persistence becomes slower. Re-run the Chapter 38 bounded workload and compare the same SLIs. Do not respond to a new performance regression by immediately adding executors/heap; first correlate the version/configuration change with measured behavior.

9. Security guardrails during troubleshooting

  • Do not disable CSRF, TLS, Script Security or authorization to make post-upgrade tests pass.
  • Do not expose old vulnerable controllers beyond loopback/disposable lab networks.
  • Do not paste plugin manifests, support bundles or logs publicly before redaction.
  • Do not use real production credentials in clone controllers unless the clone environment is explicitly authorized and isolated.
  • Do not print inbound-agent secrets while debugging connection problems.
  • Preserve an admin recovery path before changing authentication plugins.

10. Triage matrix

Symptom Likely first layer First evidence
Controller will not start Java/core/startup configuration java -version, startup log, target core
Plugin disabled/failed Plugin dependency/min-core/config plugin manifest, load error, plugin page
Login broken Auth plugin/config/external IdP controller log + provider response + break-glass test
Agents offline Agent Java/Remoting/launcher/network agent log, Java version, node log
Pipeline compile/sandbox failure Pipeline/Script Security/library plugin build log, plugin versions, Jenkinsfile/library SHA
SCM/artifact integration fails Plugin/API/TLS/credential scope request error, provider logs, exact plugin
Everything works but slower Performance/controller/agent/storage matched pre/post Chapter 38 metrics
Next

Execute the full upgrade checkpoint

Lesson 5 requires a complete evidence packet, a deliberate compatibility failure, a tested recovery path and an explicit accept/rollback decision.

Knowledge check

Answer before revealing the explanation.

1. A target controller will not start on Java 17. Which layer do you diagnose first?

2. Why is a controller rollback not enough after an external deployment side effect?

3. Why can a passing test on an advisory-affected plugin still fail the gate?

4. What should happen when only one agent pool is incompatible?

5. Why preserve first startup logs?

Official references and version notes

Upgrade behavior is version-specific. Always read every skipped LTS guide and the current security advisories for the exact day you plan an upgrade.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.