SSH, Inbound, WebSocket, Windows, and Ephemeral Agents: Connectivity, Remoting, and Agent Lifecycle: Diagnostics, Failure Modes, Security, and Performance
Diagnose agent failures from first evidence: distinguish authentication material, host-key trust, Java compatibility, proxy/WebSocket transport, Remoting channel state, workspace residue, and queue effects before changing anything.
Learning objectives
- Preserve first-failure connection, queue, agent and external evidence before retrying.
- Diagnose stale inbound identity, SSH host-key failures, Java mismatch and WebSocket proxy defects.
- Separate connectivity recovery from workspace/process/credential cleanup.
- Avoid insecure “fixes” such as no host-key verification, broad admin accounts or controller-local execution.
- Use bounded reconnect tests and resource evidence instead of blind restart loops.
1. Evidence-first diagnostic sequence
- Preserve job/build/queue IDs, source SHA and the first disconnect/error timestamp.
- Confirm Jenkins core/Java/plugin/image baseline.
- Confirm node name, launcher, labels and configured remote root.
- Confirm agent Java/runtime and process identity.
- Confirm bootstrap identity reference: inbound node identity or SSH credential/host key.
- Confirm network path: DNS/TLS/firewall/proxy/WebSocket or SSH reachability.
- Read agent + controller/launcher logs and identify where Remoting stopped.
- Inspect workspace/process residue separately from connection state.
- Apply the smallest reversible repair and test one bounded reconnect.
2. Failure layers
flowchart TD
A[Queued build or offline node] --> B{Correct node/launcher?}
B -->|no| C[Fix configuration]
B -->|yes| D{Supported Java and agent binary?}
D -->|no| E[Fix runtime / agent.jar]
D -->|yes| F{Bootstrap identity valid?}
F -->|no| G[Rotate/replace identity]
F -->|yes| H{Transport succeeds?}
H -->|no| I[DNS/TLS/firewall/proxy/SSH]
H -->|yes| J{Remoting channel stable?}
J -->|no| K[Preserve channel logs; bounded reconnect]
J -->|yes| L[Inspect workspace/tool/process state]
3. Failure: stale or compromised inbound secret
Symptom: the JVM starts but the configured node never authenticates, or a replaced node is being launched with material copied from an older identity.
Preserve: node name, creation/change timestamp, redacted agent authentication error, controller node state and which secret source/file was used—never the value.
Repair: stop the old process. If compromise is suspected, retire the affected node identity and create a replacement node name/material. Do not paste the secret into logs for comparison.
4. Failure: SSH host key does not match
Symptom: launcher refuses the host because its verified key/fingerprint changed.
This is a security signal, not friction to bypass. Determine whether the machine was legitimately rebuilt, compare the new host key through an independent trusted channel, then update the Jenkins trust configuration if authorized.
5. Failure: Java mismatch
Broken example: a worker still uses Java 17 while the Jenkins 2.568.3 system baseline requires Java 21 or 25. The agent may fail before a stable channel exists.
# Read-only evidence on the worker
java -version
sha256sum agent.jar 2>/dev/null || true
ls -ld "$HOME/ch18-agent" "$HOME/ch18-agent/work" 2>/dev/null || true
Repair: install/select a supported Java runtime for
the agent process, then use the controller-provided
agent.jar or reviewed pinned image. Do not change the
application build JDK to “fix” the Remoting JVM unless they are
intentionally the same runtime.
6. Intentionally broken example: WebSocket upgrade fails
Assume the agent log shows an HTTP handshake failure or repeated
close immediately after trying -webSocket. The JVM and
secret are valid; direct HTTPS to Jenkins UI works.
Preserve the agent log, controller timestamp and reverse-proxy access/error log. Verify the agent URL and whether the proxy/load balancer permits WebSocket Upgrade and a sufficiently long-lived connection.
Repair: correct the proxy WebSocket path or use the documented inbound TCP path in the disposable lab. Do not disable TLS or authentication, and do not schedule the job on the controller to avoid the issue.
7. Failure: reconnect succeeds, stale workspace secret remains
A connection can be healthy while execution hygiene is broken.
Suppose a synthetic build intentionally creates
tmp/fake-token.txt, then the agent disconnects before
cleanup. After reconnect, another build finds the file.
# Read-only inspection first
find "$WORKSPACE" -maxdepth 2 -type f -name 'fake-token.txt' -print
Preserve the path and owning build as synthetic evidence. Then fix
the job so temporary secrets are created in controlled locations and
removed in post { always { ... } } or by ephemeral
teardown. For a disposable worker, recreation is a valid cleanup
mechanism after root cause is preserved; it is not a
substitute for correcting the job.
8. Queue effects are downstream evidence
When the only matching node disconnects, new work may queue. That does not mean the label is wrong or more executors are needed. Compare:
- eligible node labels;
- online/offline state;
- connection log and offline cause;
- busy executor count;
- queue
whytext.
Repair transport/lifecycle first when eligibility is correct but the worker is offline.
9. Performance: do not tune Remoting before measuring
High latency, overloaded proxies, packet loss, constrained agent CPU/heap and heavy workspace I/O can all look like “slow Remoting.” Capture network timing, controller/agent CPU/memory and queue/build timing. Avoid undocumented or speculative JVM/Remoting tuning as the first response.
If an organization adjusts ping/timeouts, it should do so from current Jenkins/Remoting guidance and with a rollback plan. Longer timeouts do not fix broken routing, TLS or proxy upgrades.
10. Security-sensitive actions
| Action | Risk | Safe lab rule |
|---|---|---|
| Regenerate/replace inbound identity | Disconnects old worker; secret exposure | Disposable node, value never logged |
| Change SSH known host | Trusts a different server identity | Verify fingerprint independently first |
| Change agent OS account | Filesystem/network privilege change | Least privilege; record before/after permissions |
| Change proxy/WebSocket config | Affects multiple clients | Disposable/local proxy or approved maintenance path |
| Delete workspace/worker | Destroys first-failure evidence | Capture reviewed evidence before cleanup |
11. Smallest-safe-scope repair examples
- Wrong Java → change agent runtime, not controller authorization.
- Invalid inbound identity → replace that node identity, not all Jenkins credentials.
- Proxy lacks WebSocket support → repair WebSocket route or use bounded inbound TCP, not TLS-off mode.
- SSH host key changed → verify/update that host trust, not global “no verification.”
- Stale workspace file → fix cleanup/isolation for that worker/job, not delete unrelated build history.
Knowledge check
Answer before revealing the explanation.
1. An inbound agent suddenly reports an authentication failure after a node was recreated with the same display purpose. What should you compare first?
Compare the node identity and the exact current agent secret/reference. Do not assume the old bootstrap material is valid for the new configured node.
2. What evidence distinguishes a Java mismatch from a WebSocket proxy failure?
A Java mismatch appears before or during agent startup with Java/class-version/runtime evidence; a proxy failure usually shows HTTP/WebSocket handshake, upgrade, close or network errors after the JVM starts. Preserve both agent and controller/proxy logs.
3. Why is disabling SSH host-key verification an invalid repair?
It converts a connectivity problem into an identity-bypass vulnerability. Repair the known-host/fingerprint trust data or correct the hostname/key transition with reviewed evidence.
4. An agent reconnects successfully but a build reads an old fake token file. Which layer failed?
Workspace/process cleanup and lifecycle policy failed. Connectivity is healthy; the stale file is agent filesystem state. Preserve it as synthetic evidence, then clean or recreate the worker and fix the job so secrets are short-lived.
5. Why are blind reconnect loops dangerous troubleshooting?
They overwrite timing evidence, can amplify proxy/controller load, and may repeat side effects or workspace corruption. Capture the first disconnect and channel/log evidence before a bounded retry.
Official references and version notes
-
Jenkins LTS changelog
and
Java support policy
— lab baseline
Jenkins 2.568.3 LTS; Jenkins system components, including agents, require Java 21 or 25 on this baseline. - Using Jenkins agents and Managing Nodes — distributed execution, node configuration, Windows agent examples, and lifecycle concepts.
-
Jenkins Remoting
and
Launching inbound agents
— obtain
agent.jarfrom the controller; inbound command,-webSocket, work directory, and direct TCP guidance. -
SSH Build Agents plugin
— baseline
3.1097.v868116049892; SSH launch requires verified server identity and least-privilege agent credentials. -
Official Jenkins inbound-agent images
— reviewed lab image line
jenkins/inbound-agent:3391.va_37fa_a_305d6d-2-jdk21. Pin a reviewed tag/digest for repeatable labs rather than usinglatest. - Controller Isolation — keep the built-in node at zero executors and do not use controller-local builds as a connectivity workaround.
- Installing Jenkins — standalone Jenkins uses the supported Jetty/Winstone stack; WebSocket-agent support depends on a compatible servlet/proxy path.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.