Chapter 33Lesson 04~245 minutes

Troubleshooting Out-of-Memory, Socket, SSL, DNS, and Distributed Failures: Diagnostics, Failure Modes, and Production Practices

Troubleshooting shortcuts make symptoms disappear by changing too many layers. Preserve the original failure, keep the safety boundary, and correct only the state that evidence identifies.

No blind heapNo blanket retryNo TLS bypassNo global hosts hacksFleet consistency

Learning objectives

  • Apply a preserve-first diagnostic sequence to failure modes involving Troubleshooting Out-of-Memory, Socket, SSL, DNS, and Distributed Failures.
  • Separate plan/configuration, generator, protocol/network, target, CI/container, and distributed-engine causes before changing settings.
  • Reproduce a failure with the smallest authorized workload and retain the original JTL, jmeter.log, and supporting evidence.
  • Reject shortcuts such as blanket retries, disabled verification, unbounded load increases, silent global overrides, or deleting first-failure artifacts.
  • Verify the least-invasive correction under the original bounded workload before declaring the problem resolved.
Runnable boundary: all executable network reproductions in this lesson stay on 127.0.0.1:8033 (or the Chapter 33 local TLS fixture on localhost:8443), use at most one negative sample or the normal 2×5 baseline, and retain raw JTL plus matching jmeter.log before correction.

1. Required diagnostic sequence

Production troubleshooting sequence

The failure message is a symptom. Preserve the first state, locate the owning layer, then make one falsifiable correction before rerunning.

flowchart TD
P[Preserve first-failure JTL / jmeter.log / engine / target evidence] --> V[Confirm JMeter / Java / plugin / tool versions]
V --> I[Confirm JMX / data / properties / CLI / authorized target]
I --> S[Validate tree scope + resolved variables/properties]
S --> D[Inspect protocol/session/data]
D --> G[Inspect generator JVM/OS/network]
G --> T[Inspect SUT telemetry]
T --> R[Inspect remote/CI/container state]
R --> F[Least destructive correction]
F --> M[Smallest controlled rerun]
M --> O[Original workload verification]
Do not turn knobs before this sequence. Runnable reproductions remain one sample or the normal10-sample loopback workload. Dumps/certs/truststores are synthetic/local and removed after evidence review.

2. Blind heap increases

An OOM occurs with a verbose listener, retained response data or a pathological script. Jumping from1GiB to8GiB may only delay the same retention problem and change GC behavior. Inspect listener/result settings, live-set trend and generator headroom first. Increase heap only when sizing evidence proves legitimate demand and host memory supports it.

3. Blanket retries

Connection refused, UnknownHost, PKIX and missing files do not become correct after five attempts. Retries add load and time. Keep httpclient4.retrycount=0 for diagnostics; model legitimate application retry behavior separately and explicitly.

4. Intentionally broken TLS example

# anti-patterns — do not use
server.rmi.ssl.disable=true
Trust All Certificates = true
# or weakening global Java trust/security settings for one lab certificate

If jmeter.log shows PKIX or hostname failure, inspect certificate/trust/hostname. Lesson2 repairs only the scoped test truststore. Current remote transport uses SSL by default; configure RMI keystore/network rather than disabling it.

5. Global hosts/DNS edit without rollback

A forgotten hosts entry can redirect browsers, CI and future JMeter runs. Prefer JMeter DNS Cache Manager for test-scoped HttpClient4 mappings. If an OS/DNS change is genuinely required, record exact before/after state, authorization and rollback.

6. Deleting logs before analysis

Deleting or overwriting results/ removes the exception/timing evidence that distinguishes the layer. Use a new directory for each rerun, hash the first failure, and clean up only after retention policy.

7. Blaming the SUT for generator socket exhaustion

Ephemeral-port/file-descriptor pressure, connection churn or TLS CPU can limit the injector while target telemetry stays quiet. Inspect generator socket counts/CPU/GC and achieved RPS before scaling the target.

8. Fixing one engine but not the fleet

Manually copying a CSV/JAR to Engine B while A/C/D remain different creates a non-reproducible fleet. Provision exact JMeter/Java/data/plugin/property hashes and make parity a preflight.

9. Controller/result-transfer pressure

Remote result transfer can overload the controller/network. That is generator/control-plane overhead, not automatically SUT latency. Preserve controller/engine logs and resource state. Change sender/result strategy only after evidence identifies the bottleneck.

10. CI/container startup versus SUT failure

A service container may not be ready, DNS/service names may differ, mounts can miss CSV/JARs, or the runner may have less heap/socket capacity than a laptop. Capture image/runner identity, service logs, mounted-file inventory and preflight. Pipeline timeout alone does not prove server slowness.

11. Causal symptom table

Observation Likely layer Discriminating evidence
OOM before target traffic Generator/JVM/listener/script jmeter.log, heap/process state, target count zero.
Connection refused to wrong local port Socket endpoint listener table/Test-NetConnection.
UnknownHost; zero target events DNS/resolver resolved property + resolver/cache/DNS Manager.
Handshake fails after connect TLS trust/identity cert SAN/issuer/validity + truststore.
One engine class/file failure Distributed parity per-engine JMeter/Java/JAR/data inventory.
p95 rises only under root DEBUG Generator diagnostic overhead generator CPU/I/O + normal-logging baseline rerun.

12. Shortcuts explicitly rejected

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not assign giant heaps without evidence.
  • Do not mass-disable listeners/evidence without measured cost.
  • Do not use global property hacks.
  • Do not disable TLS/RMI verification.
  • Do not edit production/public DNS/hosts for an experiment.
  • Do not increase workload while the layer is unresolved.
  • Do not delete result files before preserving first-failure evidence.

13. Security-sensitive evidence

Heap dumps/JFR/debug logs can retain credentials or bodies; cert private keys/truststores are identity material; environment/property files may contain secrets; RMI/packet capture broadens network exposure. This chapter uses synthetic data, one-day local certificates and no real RMI traffic.

Knowledge check

Why can a larger heap hide the cause?

What is the safe fix for self-signed local TLS?

Why new result directories per rerun?

Why can a CI timeout be runner/container state?

Why is fixing one engine manually insufficient?

Next lesson

Checkpoint: diagnose three cases from evidence

Lesson5 requires a written decision tree, preserved first failures, least-invasive fixes, and final original-workload verification.

Official references and version notes

Version and compatibility note

Checked against current primary documentation on 2026-09-05. Mandatory runtime: Apache JMeter 5.6.3, Java 17, no third-party plugin. Meaningful runs use CLI with raw CSV JTL and a matching jmeter.log. Temporary logging uses category-specific -L...=DEBUG only for minimal reproductions. The bounded OOM case overrides JVM heap only for one disposable process. DNS Cache Manager is the scoped fallback for the p33.invalid exercise; Java 17 negative DNS cache defaults to 10 seconds. Remote JMeter runs the complete plan on each engine; data files are not automatically copied. RMI uses SSL by default and is not disabled in this chapter.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.