Chapter 26Lesson 04~230 minutes

Load Generator Sizing, JVM Tuning, OS Limits, and Network Capacity: Diagnostics, Failure Modes, and Production Practices

Generator bottlenecks are easy to misdiagnose because they alter the same metrics you use to judge the target: latency, throughput and errors. The diagnostic rule is therefore causal: preserve evidence, locate where time/resource pressure is accumulating, and change the smallest layer that can explain it.

CPU/GCEphemeral portsDNSOS tuningFalse bottlenecks

Learning objectives

  • Diagnose blindly increased heap and memory-allocation problems.
  • Diagnose ephemeral-port/socket exhaustion without immediately editing OS ranges.
  • Identify DNS delay and resolver/cache mistakes separately from server latency.
  • Recognize generator CPU saturation while the SUT remains healthy.
  • Reject unmanaged security/OS changes and universal legacy sizing rules.
  • Repair a deliberate keep-alive failure without hiding the original evidence.

1. Preserve first-failure evidence

All runnable failure diagnosis remains on 127.0.0.1:8026. Preserve first-failure JTL, jmeter.log, JMX/properties/CLI, JVM flags/heap/GC, process CPU/handles/FDs, socket states/dynamic-port range, DNS state, NIC/disk evidence and target JSONL before correction. Do not delete evidence or increase load.

2. Diagnostic sequence

Generator-bottleneck diagnostic sequence

A JMeter sample is generated by an injector that has finite CPU, heap, sockets, resolver, disk and network capacity. If one of those resources saturates, the configured workload and the achieved workload diverge even when the target itself is healthy.

flowchart TD
E[Preserve JTL + jmeter.log + generator + target evidence] --> V[Confirm JMeter / Java / plugin / tool versions]
V --> C[Confirm exact JMX / data / properties / CLI + authorized target]
C --> S[Validate tree scope / listeners / timers / scripts / resolved properties]
S --> P[Inspect HTTP keepalive / retries / DNS / TLS / payload / session state]
P --> G[Inspect JVM heap/GC + CPU + handles/FDs + sockets + disk/NIC]
G --> T[Inspect SUT service/resource telemetry]
T --> D[Inspect distributed/CI/container quotas/network if relevant]
D --> F[Least-destructive correction]
F --> R[Small identical rerun]

3. Failure mode: blindly increasing heap

Symptom: CPU is pegged by JSON parsing/Groovy but the operator changes -Xmx1g to -Xmx8g. Throughput does not improve because memory was not the constraint.

Evidence: jcmd GC.heap_info shows comfortable occupancy and jstat shows little GC time while process CPU is saturated.

Repair: restore the original heap, profile/remove CPU-heavy plan logic or add CPU/engines. Heap sizing follows live-set/GC evidence.

4. Failure mode: OS tuning without measurement/rollback

An engineer copies a blog command that raises FD/port/socket limits on a shared host. The original values were not recorded, the bottleneck was actually connection churn, and rollback is unknown.

Repair: revert from configuration management/backup, reproduce on an isolated generator, measure the real limit and document old/new/rollback. The mandatory lab never changes kernel/registry limits.

5. Failure mode: disabling antivirus/firewall on an unmanaged system

Security tools can affect performance, but disabling them ad hoc changes the host's protection state and can invalidate organizational trust assumptions. It is also usually not reproducible across CI/generator fleets.

Use an approved dedicated generator image/network policy or IT-managed exclusion if evidence proves scanning/firewall overhead matters. Do not instruct learners to turn protection off globally.

6. Failure mode: ephemeral-port exhaustion

Symptoms include connect failures, rapidly rising TIME_WAIT, high unique local-port use and target service time that remains healthy. Common causes include:

  • keep-alive disabled;
  • short connection TTL;
  • retries creating more connects;
  • many target endpoints/connection pools;
  • very high new-connection rate.

Inspect the actual dynamic port range and TIME_WAIT count. First repair unrealistic connection lifecycle; only then consider an authorized OS port-range/TCP policy change.

7. Failure mode: ignoring DNS latency

JTL connect/elapsed time rises intermittently, target service time is stable, and hostname lookup failures appear in jmeter.log. The team blames the API.

Compare an authorized hostname test with resolver metrics/cache behavior and, if appropriate, a stable-IP control. If DNS is part of production behavior, fix/model the resolver path rather than bypassing it permanently.

8. Failure mode: generator CPU at 100%, server blamed

Configured threads keep rising; achieved RPS plateaus; JMeter process CPU stays near machine capacity; target CPU/service time remains low. This is an injector-capacity ceiling.

Repair: reduce avoidable JMeter CPU work, use CLI/lean listeners/compiled Groovy, add CPU or another measured engine, and rerun below/around the knee. Do not report the plateau as target capacity.

9. Failure mode: universal hardware rules

Statements such as “one JMeter thread needs X MB” or “one core supports Y users” are not portable limits. Thread cost varies with protocol, payload, response speed, scripting, timers, TLS, listeners, connection reuse and Java/OS versions.

Use a repeatable profile of the actual plan on the actual generator image and step load while recording headroom.

10. Intentionally broken example: connection churn

The baseline tree has Use KeepAlive unchecked:

Test Plan
├── HTTP Request Defaults
│   host=${__P(target.host,127.0.0.1)}
│   port=${__P(target.port,8026)}
│   implementation=HttpClient4
│   connect/response timeouts from properties
└── Thread Group
    threads=${__P(threads,1)}
    loops=${__P(loops,1)}
    ├── Counter -> SEQ (per user)
    └── HTTP Request — Work — CONNECTION CHURN BASELINE
        GET /work
          run_id=${__P(run.id,p26-local)}
          thread=T${__threadNum}
          seq=${SEQ}
          payload=${__P(payload.bytes,8192)}
        Use KeepAlive = unchecked
        ├── Constant Timer ${__P(pacing.ms,100)} ms
        └── Response Assertion: response code = 200

Expected evidence: target still succeeds in ~10 ms, but target JSONL shows hundreds of unique client ports, OS shows many short-lived/TIME_WAIT connections, and connect/CPU overhead can rise. This is a generator/protocol configuration symptom, not server saturation.

Repair: preserve the run, check Use KeepAlive, reset only the run-specific target state, use a new run ID, and rerun identical 10×100. Do not increase ephemeral-port range, heap, target timeout or retries first.

11. Failure mode: verbose result retention consumes disk/CPU

If response bodies/headers/samplerData are enabled for the 8 KiB payload at scale, generator disk/serialization/privacy cost grows without changing the SUT. The repair is the Chapter 22 lean result contract and a tiny separate forensic reproduction.

12. Causal symptom table

Observation Generator-side cause Target cause to distinguish Evidence
RPS plateaus + generator CPU pegged + target service stable CPU-bound injector SUT CPU/queue saturation process CPU + target CPU/service time + achieved count.
Connect failures + TIME_WAIT/local ports spike socket churn/port pressure server connection refusal local socket states/range + target accept/service evidence.
Latency spikes + GC time rises + heap pressure JVM allocation/GC SUT latency GC telemetry + target service time.
Lookup errors/delay DNS/resolver/cache path server slow resolver/JMeter log + stable-IP control + target service time.
JTL/disk growth, target stable result retention/storage large server response itself save config + JTL bytes + response bytes + disk I/O.
NIC near line rate generator network limit SUT application bottleneck NIC counters + payload bytes + target server timing.

13. Distributed/CI/container implications

Chapter 25 engines need the same per-engine sizing gate. Containers/CI can introduce CPU quotas, memory limits, virtual NIC/DNS/overlay-disk bottlenecks even when the physical host is idle. Record cgroup/job/container limits separately from Java -Xmx and OS host capacity.

14. Troubleshooting shortcuts to reject

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not assign a giant heap without heap/GC evidence.
  • Do not mass-disable evidence/listeners without identifying their cost and required diagnostics.
  • Do not use global properties to fake workload success.
  • Do not disable TLS/RMI verification.
  • Do not run tuning/failure experiments against production/public targets.
  • Do not increase workload while the injector is already saturated.
  • Do not delete first-failure JTL/log/generator/target evidence.

Knowledge check

High generator CPU + low target CPU/service time indicates what first?

Why not enlarge the ephemeral-port range first?

What evidence distinguishes DNS delay from target processing?

Why can a giant heap make diagnosis worse?

What is the least-destructive repair for the chapter's broken example?

Next lesson

Checkpoint: prove whether generator headroom improved

Lesson 5 formalizes the 1,000-sample connection-churn versus keep-alive comparison, records exact JVM/OS/network/JTL evidence, and states what can—and cannot—be concluded about usable capacity.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK. JMeter 5.6.3 requires Java 8+; the 5.6.x changes page recommends Java 17 or later. Current JMeter launcher scripts default to a 1 GiB heap (-Xms1g -Xmx1g plus a 256 MiB metaspace cap) and G1GC with -XX:MaxGCPauseMillis=250/-XX:G1ReservePercent=20. Those defaults are a starting point, not a universal sizing rule. JMeter's HTTP sampler default is HttpClient4. Its retry count defaults to 0; its documented connection TTL defaults to 60 seconds. The HTTP Request Use KeepAlive option is effective with the Apache HttpComponents implementation and is the only plan change used in the checkpoint. JMeter's own best-practice guidance says effective thread capacity depends on hardware, plan design and how fast the target responds; CLI mode, minimal listeners, CSV and only required fields reduce generator cost. The mandatory lab never changes OS ephemeral-port ranges, file-descriptor limits, firewall/antivirus state or global DNS/JDK security settings.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.