Chapter 18Lesson 04~140 minutes

SSH, Inbound, WebSocket, Windows, and Ephemeral Agents: Connectivity, Remoting, and Agent Lifecycle: Diagnostics, Failure Modes, Security, and Performance

Diagnose agent failures from first evidence: distinguish authentication material, host-key trust, Java compatibility, proxy/WebSocket transport, Remoting channel state, workspace residue, and queue effects before changing anything.

DiagnosticsJava mismatchProxy failureStale secretsWorkspace residueRemoting

Learning objectives

  • Preserve first-failure connection, queue, agent and external evidence before retrying.
  • Diagnose stale inbound identity, SSH host-key failures, Java mismatch and WebSocket proxy defects.
  • Separate connectivity recovery from workspace/process/credential cleanup.
  • Avoid insecure “fixes” such as no host-key verification, broad admin accounts or controller-local execution.
  • Use bounded reconnect tests and resource evidence instead of blind restart loops.

1. Evidence-first diagnostic sequence

  1. Preserve job/build/queue IDs, source SHA and the first disconnect/error timestamp.
  2. Confirm Jenkins core/Java/plugin/image baseline.
  3. Confirm node name, launcher, labels and configured remote root.
  4. Confirm agent Java/runtime and process identity.
  5. Confirm bootstrap identity reference: inbound node identity or SSH credential/host key.
  6. Confirm network path: DNS/TLS/firewall/proxy/WebSocket or SSH reachability.
  7. Read agent + controller/launcher logs and identify where Remoting stopped.
  8. Inspect workspace/process residue separately from connection state.
  9. Apply the smallest reversible repair and test one bounded reconnect.

2. Failure layers

Failure layers
flowchart TD
  A[Queued build or offline node] --> B{Correct node/launcher?}
  B -->|no| C[Fix configuration]
  B -->|yes| D{Supported Java and agent binary?}
  D -->|no| E[Fix runtime / agent.jar]
  D -->|yes| F{Bootstrap identity valid?}
  F -->|no| G[Rotate/replace identity]
  F -->|yes| H{Transport succeeds?}
  H -->|no| I[DNS/TLS/firewall/proxy/SSH]
  H -->|yes| J{Remoting channel stable?}
  J -->|no| K[Preserve channel logs; bounded reconnect]
  J -->|yes| L[Inspect workspace/tool/process state]

3. Failure: stale or compromised inbound secret

Symptom: the JVM starts but the configured node never authenticates, or a replaced node is being launched with material copied from an older identity.

Preserve: node name, creation/change timestamp, redacted agent authentication error, controller node state and which secret source/file was used—never the value.

Repair: stop the old process. If compromise is suspected, retire the affected node identity and create a replacement node name/material. Do not paste the secret into logs for comparison.

4. Failure: SSH host key does not match

Symptom: launcher refuses the host because its verified key/fingerprint changed.

This is a security signal, not friction to bypass. Determine whether the machine was legitimately rebuilt, compare the new host key through an independent trusted channel, then update the Jenkins trust configuration if authorized.

Invalid shortcut: do not disable SSH host-key checking. The plugin historically had a man-in-the-middle vulnerability when verification was absent. Preserve and repair trust instead.

5. Failure: Java mismatch

Broken example: a worker still uses Java 17 while the Jenkins 2.568.3 system baseline requires Java 21 or 25. The agent may fail before a stable channel exists.

# Read-only evidence on the worker
java -version
sha256sum agent.jar 2>/dev/null || true
ls -ld "$HOME/ch18-agent" "$HOME/ch18-agent/work" 2>/dev/null || true

Repair: install/select a supported Java runtime for the agent process, then use the controller-provided agent.jar or reviewed pinned image. Do not change the application build JDK to “fix” the Remoting JVM unless they are intentionally the same runtime.

6. Intentionally broken example: WebSocket upgrade fails

Assume the agent log shows an HTTP handshake failure or repeated close immediately after trying -webSocket. The JVM and secret are valid; direct HTTPS to Jenkins UI works.

Preserve the agent log, controller timestamp and reverse-proxy access/error log. Verify the agent URL and whether the proxy/load balancer permits WebSocket Upgrade and a sufficiently long-lived connection.

Repair: correct the proxy WebSocket path or use the documented inbound TCP path in the disposable lab. Do not disable TLS or authentication, and do not schedule the job on the controller to avoid the issue.

7. Failure: reconnect succeeds, stale workspace secret remains

A connection can be healthy while execution hygiene is broken. Suppose a synthetic build intentionally creates tmp/fake-token.txt, then the agent disconnects before cleanup. After reconnect, another build finds the file.

# Read-only inspection first
find "$WORKSPACE" -maxdepth 2 -type f -name 'fake-token.txt' -print

Preserve the path and owning build as synthetic evidence. Then fix the job so temporary secrets are created in controlled locations and removed in post { always { ... } } or by ephemeral teardown. For a disposable worker, recreation is a valid cleanup mechanism after root cause is preserved; it is not a substitute for correcting the job.

8. Queue effects are downstream evidence

When the only matching node disconnects, new work may queue. That does not mean the label is wrong or more executors are needed. Compare:

  • eligible node labels;
  • online/offline state;
  • connection log and offline cause;
  • busy executor count;
  • queue why text.

Repair transport/lifecycle first when eligibility is correct but the worker is offline.

9. Performance: do not tune Remoting before measuring

High latency, overloaded proxies, packet loss, constrained agent CPU/heap and heavy workspace I/O can all look like “slow Remoting.” Capture network timing, controller/agent CPU/memory and queue/build timing. Avoid undocumented or speculative JVM/Remoting tuning as the first response.

If an organization adjusts ping/timeouts, it should do so from current Jenkins/Remoting guidance and with a rollback plan. Longer timeouts do not fix broken routing, TLS or proxy upgrades.

10. Security-sensitive actions

Action Risk Safe lab rule
Regenerate/replace inbound identity Disconnects old worker; secret exposure Disposable node, value never logged
Change SSH known host Trusts a different server identity Verify fingerprint independently first
Change agent OS account Filesystem/network privilege change Least privilege; record before/after permissions
Change proxy/WebSocket config Affects multiple clients Disposable/local proxy or approved maintenance path
Delete workspace/worker Destroys first-failure evidence Capture reviewed evidence before cleanup

11. Smallest-safe-scope repair examples

  • Wrong Java → change agent runtime, not controller authorization.
  • Invalid inbound identity → replace that node identity, not all Jenkins credentials.
  • Proxy lacks WebSocket support → repair WebSocket route or use bounded inbound TCP, not TLS-off mode.
  • SSH host key changed → verify/update that host trust, not global “no verification.”
  • Stale workspace file → fix cleanup/isolation for that worker/job, not delete unrelated build history.
Next lesson

Checkpoint Lab

Connect a disposable remote-style agent, capture Remoting evidence, force one disconnect, recover it, then destroy/recreate the worker and prove retained build attribution.

Knowledge check

Answer before revealing the explanation.

1. An inbound agent suddenly reports an authentication failure after a node was recreated with the same display purpose. What should you compare first?

2. What evidence distinguishes a Java mismatch from a WebSocket proxy failure?

3. Why is disabling SSH host-key verification an invalid repair?

4. An agent reconnects successfully but a build reads an old fake token file. Which layer failed?

5. Why are blind reconnect loops dangerous troubleshooting?

Official references and version notes

  • Jenkins LTS changelog and Java support policy — lab baseline Jenkins 2.568.3 LTS; Jenkins system components, including agents, require Java 21 or 25 on this baseline.
  • Using Jenkins agents and Managing Nodes — distributed execution, node configuration, Windows agent examples, and lifecycle concepts.
  • Jenkins Remoting and Launching inbound agents — obtain agent.jar from the controller; inbound command, -webSocket, work directory, and direct TCP guidance.
  • SSH Build Agents plugin — baseline 3.1097.v868116049892; SSH launch requires verified server identity and least-privilege agent credentials.
  • Official Jenkins inbound-agent images — reviewed lab image line jenkins/inbound-agent:3391.va_37fa_a_305d6d-2-jdk21. Pin a reviewed tag/digest for repeatable labs rather than using latest.
  • Controller Isolation — keep the built-in node at zero executors and do not use controller-local builds as a connectivity workaround.
  • Installing Jenkins — standalone Jenkins uses the supported Jetty/Winstone stack; WebSocket-agent support depends on a compatible servlet/proxy path.
Assumption timestamp: 2026-09-17. Recheck Jenkins LTS/Java support, Remoting/inbound-agent versions, SSH Build Agents advisories, reverse-proxy WebSocket behavior, and Windows service guidance before repeating the lab later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.