Load Generator Sizing, JVM Tuning, OS Limits, and Network Capacity: Diagnostics, Failure Modes, and Production Practices
Generator bottlenecks are easy to misdiagnose because they alter the same metrics you use to judge the target: latency, throughput and errors. The diagnostic rule is therefore causal: preserve evidence, locate where time/resource pressure is accumulating, and change the smallest layer that can explain it.
Learning objectives
- Diagnose blindly increased heap and memory-allocation problems.
- Diagnose ephemeral-port/socket exhaustion without immediately editing OS ranges.
- Identify DNS delay and resolver/cache mistakes separately from server latency.
- Recognize generator CPU saturation while the SUT remains healthy.
- Reject unmanaged security/OS changes and universal legacy sizing rules.
- Repair a deliberate keep-alive failure without hiding the original evidence.
1. Preserve first-failure evidence
127.0.0.1:8026.
Preserve first-failure JTL, jmeter.log,
JMX/properties/CLI, JVM flags/heap/GC, process CPU/handles/FDs,
socket states/dynamic-port range, DNS state, NIC/disk evidence and
target JSONL before correction. Do not delete evidence or increase
load.
2. Diagnostic sequence
A JMeter sample is generated by an injector that has finite CPU, heap, sockets, resolver, disk and network capacity. If one of those resources saturates, the configured workload and the achieved workload diverge even when the target itself is healthy.
flowchart TD E[Preserve JTL + jmeter.log + generator + target evidence] --> V[Confirm JMeter / Java / plugin / tool versions] V --> C[Confirm exact JMX / data / properties / CLI + authorized target] C --> S[Validate tree scope / listeners / timers / scripts / resolved properties] S --> P[Inspect HTTP keepalive / retries / DNS / TLS / payload / session state] P --> G[Inspect JVM heap/GC + CPU + handles/FDs + sockets + disk/NIC] G --> T[Inspect SUT service/resource telemetry] T --> D[Inspect distributed/CI/container quotas/network if relevant] D --> F[Least-destructive correction] F --> R[Small identical rerun]
3. Failure mode: blindly increasing heap
Symptom: CPU is pegged by JSON parsing/Groovy but the operator
changes -Xmx1g to -Xmx8g. Throughput does
not improve because memory was not the constraint.
Evidence: jcmd GC.heap_info shows comfortable occupancy
and jstat shows little GC time while process CPU is
saturated.
Repair: restore the original heap, profile/remove CPU-heavy plan logic or add CPU/engines. Heap sizing follows live-set/GC evidence.
4. Failure mode: OS tuning without measurement/rollback
An engineer copies a blog command that raises FD/port/socket limits on a shared host. The original values were not recorded, the bottleneck was actually connection churn, and rollback is unknown.
Repair: revert from configuration management/backup, reproduce on an isolated generator, measure the real limit and document old/new/rollback. The mandatory lab never changes kernel/registry limits.
5. Failure mode: disabling antivirus/firewall on an unmanaged system
Security tools can affect performance, but disabling them ad hoc changes the host's protection state and can invalidate organizational trust assumptions. It is also usually not reproducible across CI/generator fleets.
Use an approved dedicated generator image/network policy or IT-managed exclusion if evidence proves scanning/firewall overhead matters. Do not instruct learners to turn protection off globally.
6. Failure mode: ephemeral-port exhaustion
Symptoms include connect failures, rapidly rising TIME_WAIT, high unique local-port use and target service time that remains healthy. Common causes include:
- keep-alive disabled;
- short connection TTL;
- retries creating more connects;
- many target endpoints/connection pools;
- very high new-connection rate.
Inspect the actual dynamic port range and TIME_WAIT count. First repair unrealistic connection lifecycle; only then consider an authorized OS port-range/TCP policy change.
7. Failure mode: ignoring DNS latency
JTL connect/elapsed time rises intermittently, target service time
is stable, and hostname lookup failures appear in
jmeter.log. The team blames the API.
Compare an authorized hostname test with resolver metrics/cache behavior and, if appropriate, a stable-IP control. If DNS is part of production behavior, fix/model the resolver path rather than bypassing it permanently.
8. Failure mode: generator CPU at 100%, server blamed
Configured threads keep rising; achieved RPS plateaus; JMeter process CPU stays near machine capacity; target CPU/service time remains low. This is an injector-capacity ceiling.
Repair: reduce avoidable JMeter CPU work, use CLI/lean listeners/compiled Groovy, add CPU or another measured engine, and rerun below/around the knee. Do not report the plateau as target capacity.
9. Failure mode: universal hardware rules
Statements such as “one JMeter thread needs X MB” or “one core supports Y users” are not portable limits. Thread cost varies with protocol, payload, response speed, scripting, timers, TLS, listeners, connection reuse and Java/OS versions.
Use a repeatable profile of the actual plan on the actual generator image and step load while recording headroom.
10. Intentionally broken example: connection churn
The baseline tree has Use KeepAlive unchecked:
Test Plan
├── HTTP Request Defaults
│ host=${__P(target.host,127.0.0.1)}
│ port=${__P(target.port,8026)}
│ implementation=HttpClient4
│ connect/response timeouts from properties
└── Thread Group
threads=${__P(threads,1)}
loops=${__P(loops,1)}
├── Counter -> SEQ (per user)
└── HTTP Request — Work — CONNECTION CHURN BASELINE
GET /work
run_id=${__P(run.id,p26-local)}
thread=T${__threadNum}
seq=${SEQ}
payload=${__P(payload.bytes,8192)}
Use KeepAlive = unchecked
├── Constant Timer ${__P(pacing.ms,100)} ms
└── Response Assertion: response code = 200
Expected evidence: target still succeeds in ~10 ms, but target JSONL shows hundreds of unique client ports, OS shows many short-lived/TIME_WAIT connections, and connect/CPU overhead can rise. This is a generator/protocol configuration symptom, not server saturation.
Repair: preserve the run, check Use KeepAlive, reset only the run-specific target state, use a new run ID, and rerun identical 10×100. Do not increase ephemeral-port range, heap, target timeout or retries first.
11. Failure mode: verbose result retention consumes disk/CPU
If response bodies/headers/samplerData are enabled for the 8 KiB payload at scale, generator disk/serialization/privacy cost grows without changing the SUT. The repair is the Chapter 22 lean result contract and a tiny separate forensic reproduction.
12. Causal symptom table
| Observation | Generator-side cause | Target cause to distinguish | Evidence |
|---|---|---|---|
| RPS plateaus + generator CPU pegged + target service stable | CPU-bound injector | SUT CPU/queue saturation | process CPU + target CPU/service time + achieved count. |
| Connect failures + TIME_WAIT/local ports spike | socket churn/port pressure | server connection refusal | local socket states/range + target accept/service evidence. |
| Latency spikes + GC time rises + heap pressure | JVM allocation/GC | SUT latency | GC telemetry + target service time. |
| Lookup errors/delay | DNS/resolver/cache path | server slow | resolver/JMeter log + stable-IP control + target service time. |
| JTL/disk growth, target stable | result retention/storage | large server response itself | save config + JTL bytes + response bytes + disk I/O. |
| NIC near line rate | generator network limit | SUT application bottleneck | NIC counters + payload bytes + target server timing. |
13. Distributed/CI/container implications
Chapter 25 engines need the same per-engine sizing gate.
Containers/CI can introduce CPU quotas, memory limits, virtual
NIC/DNS/overlay-disk bottlenecks even when the physical host is
idle. Record cgroup/job/container limits separately from Java
-Xmx and OS host capacity.
14. Troubleshooting shortcuts to reject
- Do not add blanket retries or arbitrary long sleeps.
- Do not assign a giant heap without heap/GC evidence.
- Do not mass-disable evidence/listeners without identifying their cost and required diagnostics.
- Do not use global properties to fake workload success.
- Do not disable TLS/RMI verification.
- Do not run tuning/failure experiments against production/public targets.
- Do not increase workload while the injector is already saturated.
- Do not delete first-failure JTL/log/generator/target evidence.
Knowledge check
High generator CPU + low target CPU/service time indicates what first?
An injector-side CPU ceiling, not proof of target saturation.
Why not enlarge the ephemeral-port range first?
If the plan is creating unrealistic short-lived connections, the root cause is connection lifecycle; OS tuning would hide it.
What evidence distinguishes DNS delay from target processing?
Resolver/JMeter errors or lookup timing plus stable target service timing and a controlled IP/hostname comparison.
Why can a giant heap make diagnosis worse?
It can hide retention problems, increase memory commitment and make rare full-GC/dump events larger without fixing CPU/socket/disk limits.
What is the least-destructive repair for the chapter's broken example?
Restore realistic HTTP keep-alive and rerun the identical bounded workload while preserving the original churn evidence.
Official references and version notes
- Apache JMeter downloads — current stable release and Java requirement.
- Apache JMeter current changes — Java guidance for the 5.6.x line.
- JMeter Getting Started — load-generator sizing, CLI guidance, default heap/GC launcher settings and JVM environment variables.
- JMeter Best Practices — effective thread count, CLI execution, listener/result minimization, CSV output and generator resource reduction.
- Component Reference — HTTP Request — HttpClient4 default, keep-alive behavior and response-body MD5 option.
- Component Reference — DNS Cache Manager — JVM DNS cache behavior and HttpClient4-only per-thread DNS control.
- JMeter Properties Reference — HttpClient4 retry, connection TTL/validation and result-save properties.
-
JDK 17
jcmd— JVM process/heap inspection. -
JDK 17
jstat— GC/heap utilization statistics; documented as experimental/unsupported.
Version-sensitive statements were rechecked against current
primary documentation on 2026-09-05. The course baseline remains
Apache JMeter 5.6.3 with a Java 17 JDK. JMeter
5.6.3 requires Java 8+; the 5.6.x changes page recommends Java 17
or later. Current JMeter launcher scripts default to a 1 GiB heap
(-Xms1g -Xmx1g plus a 256 MiB metaspace cap) and G1GC
with
-XX:MaxGCPauseMillis=250/-XX:G1ReservePercent=20. Those defaults are a
starting point, not a universal sizing rule. JMeter's HTTP sampler
default is HttpClient4. Its retry count defaults to 0; its
documented connection TTL defaults to 60 seconds. The HTTP Request
Use KeepAlive option is effective with the Apache
HttpComponents implementation and is the only plan change used in
the checkpoint. JMeter's own best-practice guidance says effective
thread capacity depends on hardware, plan design and how fast the
target responds; CLI mode, minimal listeners, CSV and only
required fields reduce generator cost. The mandatory lab never
changes OS ephemeral-port ranges, file-descriptor limits,
firewall/antivirus state or global DNS/JDK security settings.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.