Chapter 33Lesson 03~225 minutes

Troubleshooting Out-of-Memory, Socket, SSL, DNS, and Distributed Failures: Configuration, Design Patterns, and Trade-Offs

Diagnostics can change the system they observe. Root DEBUG adds I/O, heap dumps pause/write large artifacts, packet captures can expose payloads, and retries change traffic. Choose the smallest diagnostic capable of separating the current hypotheses.

Narrow DEBUGjcmdHeap-dump trade-offLocal vs remoteClient/server evidence

Learning objectives

  • Compare the principal configuration and design choices for Troubleshooting Out-of-Memory, Socket, SSL, DNS, and Distributed Failures without changing the workload question unintentionally.
  • Identify which settings belong to the JMeter plan, JVM, OS/network, target, extensions, CI/container, or distributed-engine layers.
  • Explain the trade-offs among performance cost, reliability, reproducibility, security, portability, and operational complexity.
  • Choose an appropriate pattern from measured evidence and explicit constraints rather than from convenience or folklore.
  • Preserve measurement validity and a stable evidence baseline before moving into failure diagnosis.
Runnable boundary: all executable network reproductions in this lesson stay on 127.0.0.1:8033 (or the Chapter 33 local TLS fixture on localhost:8443), use at most one negative sample or the normal 2×5 baseline, and retain raw JTL plus matching jmeter.log before correction.

1. Debug category and duration

Choice Good use Cost / risk
INFO + first-failure log Always keep as baseline. Lowest overhead.
One category DEBUG Short protocol/engine reproduction. Extra I/O/privacy/generator CPU.
Root DEBUG Only very short isolated cases when category is unknown. Large logs and measurement distortion.
Permanent verbose mode Rarely justified. Diagnostic cost becomes hidden test-plan cost.

Current JMeter CLI lets you override a logger with -Lorg.apache...=DEBUG. Return to normal logging for the final load verification.

2. Resource observation versus heap dump

Observation first Heap dump
Process CPU/RSS, JTL counts, jcmd VM.command_line, GC.heap_info, Thread.print. Full object graph in HPROF.
Low/medium impact and quick hypothesis screening. Can pause the JVM and consume significant disk.
Good for “is the generator near its limit?” Good for “what is retaining heap?” when evidence points there.
Repeatable during a run. May contain credentials/bodies; protect as sensitive evidence.

JDK17 documents GC.heap_dump as higher-impact than simple state commands. A dump is not a reflexive first step.

3. Retries versus root-cause analysis

Retry is a workload/protocol rule, not a general troubleshooting switch. Connection refused, UnknownHost, PKIX, missing CSV/JAR, or OOM are deterministic layer failures. Retrying them increases samples and can change configured/achieved load without fixing the root cause.

If the real application legitimately retries an idempotent operation, model that explicitly with a bounded count and separate retry metrics. Keep the diagnostic lab at httpclient4.retrycount=0.

4. Local reproduction versus distributed capture

Local minimal reproduction Distributed capture
Best when one node reproduces JMX/protocol/DNS/TLS failure. Needed when failure exists only on a specific engine/network/container.
Fewer moving parts; safer and cheaper. Preserves per-engine file/JAR/socket/clock/resource state.
Validates client/target behavior before RMI. Requires controller + all engine inventories/logs.
Cannot prove fleet parity. Start with a tiny authorized remote smoke, not full load.

5. Client-side versus server-side evidence

For connection refusal, client exception plus listener-table evidence is stronger than either alone. For TLS, client handshake error plus certificate/SAN/truststore state narrows the failure before HTTP. For target500s, pair JTL request/result state with SUT/dependency logs. Packet capture is optional and privacy-sensitive; it is not the default local step.

6. DNS diagnostic choices

Inspect the resolved JMeter host property, system/Java resolver behavior, DNS Cache Manager placement and HTTP implementation. Java caches negative results; JMeter DNS Cache Manager has an independent HttpClient4 cache. A fresh process plus a scoped static mapping is a cleaner experiment than editing the global hosts file.

7. TLS diagnostic choices

Separate TCP connectivity from TLS handshake and HTTP response. Inspect certificate subject, SAN, issuer, validity, truststore and JVM system properties. A self-signed training certificate is repaired by scoped trust. A hostname mismatch requires a certificate valid for that hostname; it is not repaired by trust-all.

8. Socket exhaustion versus one refused port

A known unused local port returning Connection refused is an endpoint/listener issue. Generator socket exhaustion appears under concurrency: ephemeral-port/file-descriptor pressure, many active/TIME_WAIT sockets, high generator CPU, or OS-level socket errors while the target is listening. Inspect generator state before blaming server capacity.

9. Remote-mode diagnosis

Remote JMeter executes the full plan on each engine, so a 100-thread plan on four engines can mean400 threads. Exact JMeter version, recommended Java parity, data/plugin files, RMI SSL, ports, controller result-transfer cost and engine-local properties are all part of the failure surface.

10. Configuration layers

Layer Examples
JMeter plan/core Sampler host/port/protocol, DNS manager, listeners, retries, result fields.
Java/JVM Heap/GC, truststore/system properties, DNS cache, jcmd/JFR/heap dump.
OS/network Socket limits, firewall, routes, resolver, ephemeral ports.
SUT Listener/service, certificate, rate limits, DB/cache/downstreams.
Plugin/driver HTTP/client libraries, extension/classpath compatibility.
CI provider Runner resources, service startup, timeout, artifact retention.
Container/orchestrator DNS/service/network policy, resource limits, mounted data/JARs.
Distributed RMI Controller/engine versions, SSL keystore, ports, data/JAR parity.

11. Worked scenario

At high concurrency, target CPU is30%, but JMeter CPU is95%, heap approaches its limit, sockets/TIME_WAIT rise, achieved RPS plateaus and connection errors appear. The leading hypothesis is generator exhaustion, not target saturation. Reduce generator overhead or scale injectors only after validating workload semantics and engine parity.

12. Decision table

Symptom First evidence Least-invasive next step
OOM jmeter.log + VM command/heap state + listener/result settings Minimal synthetic reproduction; dump only if retained-object question remains.
Connection refused resolved host/port + listener state Correct endpoint/start intended listener.
UnknownHost resolved property + resolver/cache scope Scoped DNS mapping or DNS service correction.
PKIX/hostname cert/SAN/issuer + truststore/system props Correct scoped trust/certificate; keep verification.
One engine fails per-engine log/version/data/JAR inventory Fix mismatched engine/image; parity-check before remote load.

13. Diagnostic validity

Heap, logging, retries, resolver state, truststore, engine count and listeners all can change generator cost or workload behavior. Record the diagnostic configuration separately. Before target capacity/regression conclusions, rerun the original plan with normal diagnostics and report configured and achieved load plus generator state.

Evidence baseline: Apache JMeter 5.6.3/Java17, CLI mode, raw CSV JTL + matching jmeter.log, loopback targets only, no third-party plugin, no real credential.

Knowledge check

Why prefer narrow DEBUG?

When is a heap dump justified?

Why can blanket retries invalidate the experiment?

What points to generator socket exhaustion?

What must happen after diagnostic changes?

Next lesson

Reject destructive shortcuts

Lesson4 diagnoses blind heap increases, blanket retries, TLS bypass, global hosts edits, deleted logs, generator socket exhaustion and partial fleet fixes.

Official references and version notes

Version and compatibility note

Checked against current primary documentation on 2026-09-05. Mandatory runtime: Apache JMeter 5.6.3, Java 17, no third-party plugin. Meaningful runs use CLI with raw CSV JTL and a matching jmeter.log. Temporary logging uses category-specific -L...=DEBUG only for minimal reproductions. The bounded OOM case overrides JVM heap only for one disposable process. DNS Cache Manager is the scoped fallback for the p33.invalid exercise; Java 17 negative DNS cache defaults to 10 seconds. Remote JMeter runs the complete plan on each engine; data files are not automatically copied. RMI uses SSL by default and is not disabled in this chapter.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.