Chapter 26Lesson 05~310 minutes

Checkpoint Lab — Load Generator Sizing, JVM Tuning, OS Limits, and Network Capacity

The checkpoint profiles one generator like an infrastructure component. Both variants run the same 1,000-sample localhost workload. The baseline intentionally disables HTTP connection reuse; the corrected run enables it. You must predict the resource changes, verify target work independently, quantify JVM/CPU/socket/network/JTL evidence, and decide whether usable generator headroom improved without claiming an unmeasured maximum capacity.

Checkpoint1,000 samplesSocket pressureKeep-alive correctionCapacity validity

Learning objectives

  • Inventory hardware/JVM/OS/network/result settings before execution.
  • Predict the connection/JVM/RPS effects of disabled versus enabled keep-alive.
  • Execute two identical bounded CLI workloads.
  • Cross-check JTL with target events and unique client-port evidence.
  • Measure heap/GC, CPU, handles/FDs, sockets, NIC and JTL size.
  • State whether usable generator headroom/capacity evidence improved and record rollback.

1. Exact assumptions and ceilings

Item Checkpoint baseline
JMeter Apache JMeter 5.6.3.
Java Java 17 JDK; JMeter 5.6.3 requires Java 8+.
Plugins None.
HTTP implementation HttpClient4 explicit; retrycount=0.
Target http://127.0.0.1:8026 only.
Target service ~10 ms synthetic work; 8 KiB bounded JSON payload.
Workload 10 threads ×100 loops ×1 Work sampler = 1,000 samples/run.
Pacing 100 ms Constant Timer.
Baseline plan Use KeepAlive unchecked.
Corrected plan Use KeepAlive checked; no other behavioral change.
JTL lean CSV, no response/header/sampler body retention.
Runtime ceiling ≤20 seconds/run; two normal runs.
OS tuning None; read-only dynamic-port/FD/handle/DNS/NIC/disk inspection.
Evidence hardware/JVM inventory, heap/GC, CPU, handles/FDs, sockets/ports, DNS snapshot, NIC, JTL size/RPS, target events/ports, rollback.
Abort: non-loopback traffic, >1,000 target work events/run, target errors/service instability, OOM/full-GC loop, sustained generator CPU exhaustion, unexpected file/port growth, low disk space, resolver/network anomaly contaminating the IP-based benchmark, or missing first-run evidence.

2. Preflight inventory

Record before either run:

  • CPU model/physical/logical cores, RAM, disk free space, NIC/link state;
  • JMeter 5.6.3 and exact Java 17 build;
  • effective VM.flags/heap/GC after JMeter starts;
  • Windows dynamic TCP port ranges or Linux ip_local_port_range;
  • Windows handle/process state or Linux ulimit -n/fs.file-max;
  • DNS cache/resolver snapshot even though benchmark uses IP;
  • hash/diff proving the two JMXs differ only in KeepAlive.

3. Predictions before execution

P1 — workload: both variants produce exactly 1,000 successful JTL rows and 1,000 target events.

P2 — sockets: connection-close baseline uses a high fraction of unique client ports; keep-alive run uses dramatically fewer unique ports/connections and less TIME_WAIT growth.

P3 — target: target service_wall_ms remains approximately unchanged because target code/payload/workload are identical.

P4 — generator: keep-alive may reduce connect/CPU overhead and improve or stabilize achieved RPS, but the exact percentage is machine-dependent and must be measured.

P5 — heap/JTL: neither variant should require materially different heap live-set or JTL bytes/sample; if they do, investigate another confounder.

4. Run the connection-churn baseline

curl.exe --fail --silent -X POST `
  "http://127.0.0.1:8026/reset?run_id=p26-check-close"

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -n -t .\plans\generator-close.jmx `
  -q .\config\generator-local.properties `
  -Jrun.id=p26-check-close `
  -l .\results\p26-check-close\results.jtl `
  -j .\results\p26-check-close\jmeter.log

Capture the platform/JVM inspections from Lesson 2 during the run and immediately after. Save the outputs under the run directory.

5. Baseline verification gate

python tools/profile_analyzer.py results/p26-check-close/results.jtl
python tools/analyze_target.py results/server-events.jsonl p26-check-close

Require 1,000 rows/events, zero failures, stable target service time. Record unique-client-port count, socket states, CPU/heap/GC, NIC counters and JTL file bytes.

6. Identify the dominant observed generator pressure

Use evidence, not the expected answer. The intended baseline should show connection/local-port churn while target service time remains stable. But if CPU or GC is clearly the first limiting signal on your machine, record that instead and explain why the keep-alive experiment may not address the dominant constraint.

Do not manufacture a “port exhaustion” conclusion merely because the lesson demonstrates connection churn.

7. Reversible change and rollback record

Change only:

HTTP Request — Work
Use KeepAlive:
  BEFORE = unchecked
  AFTER  = checked

Rollback:
  reopen generator-close.jmx, or uncheck Use KeepAlive in the copied plan.

No HEAP, GC, OS, DNS, firewall, antivirus, timeout, retry, thread, loop, pacing, payload, target or JTL setting changes.

8. Run the corrected profile identically

curl.exe --fail --silent -X POST `
  "http://127.0.0.1:8026/reset?run_id=p26-check-keepalive"

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -n -t .\plans\generator-keepalive.jmx `
  -q .\config\generator-local.properties `
  -Jrun.id=p26-check-keepalive `
  -l .\results\p26-check-keepalive\results.jtl `
  -j .\results\p26-check-keepalive\jmeter.log

Capture exactly the same process/JVM/socket/NIC/JTL measurements at comparable points.

9. Before/after comparison

python tools/compare_profiles.py   results/p26-check-close/results.jtl   results/p26-check-keepalive/results.jtl

python tools/analyze_target.py results/server-events.jsonl p26-check-close
python tools/analyze_target.py results/server-events.jsonl p26-check-keepalive
Evidence Connection-close baseline Keep-alive corrected Interpretation
JTL / target samples expect 1000 / 1000 expect 1000 / 1000 Workload equivalence gate.
Failures 0 expected 0 expected Protocol/target correctness gate.
Unique client ports high much lower expected Connection reuse / ephemeral-port pressure.
TIME_WAIT/socket count higher expected lower expected Kernel/socket-state cost.
Target service time ~10 ms ~10 ms Target control should remain stable.
Achieved RPS measure measure Generator completion capacity at this workload.
JMeter CPU / heap / GC measure measure Injector cost/headroom.
NIC bytes similar payload bytes similar payload bytes Application payload unchanged; handshake overhead differs.
JTL bytes/sample similar similar Result contract unchanged.

10. Decide whether usable capacity improved

Use this wording discipline:

  • If keep-alive maintains the same 1,000 samples with dramatically lower connection/TIME_WAIT pressure and equal/better CPU/RPS, state that generator headroom improved for this workload.
  • If achieved RPS materially improves while target service time and all workload inputs stay stable, state the measured improvement.
  • If RPS is unchanged, do not claim a higher maximum. State that the plan reduced socket pressure but the checkpoint did not reach either variant's sustainable-load ceiling.
  • If another resource (CPU/GC/disk) dominates, state that keep-alive was not the limiting fix and identify the next evidence-backed experiment.

11. Optional safe capacity step

Only if both 1,000-sample runs are valid and you need evidence of additional usable capacity, perform one small step such as 12 threads ×100 loops (1,200 samples) with keep-alive enabled. Keep the target/monitoring identical and abort on the same resource gates. This optional step is not required to complete the chapter and should not be run if the baseline already showed generator pressure.

12. Required evidence packet

Artifact Required content
Hardware/JVM inventory CPU/RAM/NIC/disk, JMeter/Java, effective heap/GC flags.
JVM evidence heap snapshots, GC counts/time, process CPU/working set/thread count.
OS sockets/limits handles/FDs, socket states, dynamic-port range; read-only.
Network/DNS NIC counters/drops and resolver/cache snapshot.
JTL 1,000 rows/run, achieved RPS/latency, total result bytes/bytes per sample.
Target JSONL 1,000 events/run, service-wall timing, unique source-port counts.
Before/after table close versus keepalive under identical settings.
Rollback exact JMX/KeepAlive reversion; no OS changes to undo.
Validity note configured vs achieved work and generator/target headroom/limitations.

13. Example validity statement

Example: “Apache JMeter 5.6.3 on Java 17 executed two identical 10-thread ×100-loop localhost HttpClient4 workloads with an 8 KiB response and lean CSV JTL. The baseline differed only by Use KeepAlive=off; the corrected plan enabled it. Both produced 1,000 successful JTL rows and 1,000 target events with stable target service time. The baseline consumed substantially more unique client ports/TIME_WAIT state; the corrected run reused far fewer connections. JVM heap/GC, process CPU, NIC counters, socket/handle state, JTL size and achieved RPS were recorded for both runs. Therefore the correction improved connection/socket headroom; any maximum-capacity or percentage-throughput claim is limited to the measured values and requires an additional safely stepped experiment.”

14. Verification checklist

  • Only 127.0.0.1:8026.
  • JMeter 5.6.3 / Java 17 / no plugins.
  • Same threads/loops/pacing/payload/target/result schema.
  • Only KeepAlive differs.
  • 1,000 JTL +1,000 target events each; zero failures.
  • Heap/GC + CPU + handles/FDs + socket states recorded.
  • Dynamic-port range inspected, not changed.
  • DNS/NIC/JTL size/RPS captured.
  • Target service time stable enough for comparison.
  • Conclusion distinguishes improved headroom from unmeasured maximum capacity.

15. Cleanup / rollback

  1. Stop JMeter and the localhost fixture after evidence capture.
  2. Keep close/keepalive JTL, matching jmeter.log, profile outputs and target JSONL until review.
  3. Rollback is simply the original generator-close.jmx; no OS/JVM global tuning was changed.
  4. Delete disposable artifacts only after the retention decision.
  5. No production/public target, real credential, firewall/antivirus state, RMI service, managed cloud, paid CI, OS port/FD/DNS setting or system-wide JVM security setting was modified.

16. What Chapter 26 adds to the operating model

The production performance-testing operating model now has a load-generator capacity contract: hardware/JVM inventory, effective heap/GC flags, measured heap/GC/CPU, socket/handle/FD/local-port state, DNS strategy, NIC/disk/JTL cost, HTTP connection/retry policy, configured-versus-achieved load, target service control, per-engine headroom criteria, before/after tuning evidence and rollback are required before an injector's workload is considered trustworthy.

Chapter 27 moves to Containers, Kubernetes, Ephemeral Injectors, and Infrastructure Automation. It carries this capacity contract into container CPU/memory quotas, cgroups, ephemeral filesystems, pod/service DNS, virtual networking, image/version pinning and reproducible disposable generator fleets.

Knowledge check

What must remain identical for the checkpoint comparison?

What independently verifies socket churn?

If keepalive lowers ports but CPU remains 100% and RPS unchanged, what should you conclude?

Why is changing the OS port range outside this checkpoint?

What is Chapter 27's bridge?

Next chapter

Containers, Kubernetes, Ephemeral Injectors, and Infrastructure Automation

Chapter 27 turns measured injectors into reproducible disposable infrastructure without losing JVM/OS/network capacity evidence.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK. JMeter 5.6.3 requires Java 8+; the 5.6.x changes page recommends Java 17 or later. Current JMeter launcher scripts default to a 1 GiB heap (-Xms1g -Xmx1g plus a 256 MiB metaspace cap) and G1GC with -XX:MaxGCPauseMillis=250/-XX:G1ReservePercent=20. Those defaults are a starting point, not a universal sizing rule. JMeter's HTTP sampler default is HttpClient4. Its retry count defaults to 0; its documented connection TTL defaults to 60 seconds. The HTTP Request Use KeepAlive option is effective with the Apache HttpComponents implementation and is the only plan change used in the checkpoint. JMeter's own best-practice guidance says effective thread capacity depends on hardware, plan design and how fast the target responds; CLI mode, minimal listeners, CSV and only required fields reduce generator cost. The mandatory lab never changes OS ephemeral-port ranges, file-descriptor limits, firewall/antivirus state or global DNS/JDK security settings.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.