CLI Mode, Headless Execution, Result Files, and Reproducible Runs: Configuration, Design Patterns, and Trade-Offs
A mature CLI workflow optimizes for reproducibility and diagnostic value rather than for the shortest command. The right choices depend on sample volume, report timing, artifact retention, CI behavior and whether a failure belongs to the JMeter engine or to the measured system.
Learning objectives
- Choose CSV or XML JTL based on result volume and diagnostic requirements.
- Choose report-at-end or later report generation based on pipeline design.
- Choose unique run IDs over fixed overwrite-prone output paths.
- Choose console summary and artifact analysis for their different purposes.
- Choose a wrapper script versus raw CLI based on repeatability needs.
- Separate engine/process failures from sample/performance threshold gates.
1. Mandatory examples remain local/free
127.0.0.1:8021, ≤2×4
samples.
Paid CI, load clouds, centralized observability, containers and
remote RMI are optional future execution contexts only.
2. CSV versus XML JTL
| CSV JTL | XML JTL |
|---|---|
| Smaller; JMeter recommends CSV for lots of samples. | Larger; supports richer structured fields including some data CSV cannot store. |
| Dashboard path is designed around compatible CSV sample logs. | Useful for a tiny diagnostic/functional case requiring XML-only content. |
| No response body storage in CSV. | Can retain response data, which can greatly increase disk/memory cost. |
| Better production/load default. | Use deliberately and bounded; never as a reflex for all load evidence. |
3. Report at end versus later -g
| <code>-e -o</code> after load | Later <code>-g JTL -o dir</code> |
|---|---|
| One command produces raw JTL plus dashboard. | Load process ends after raw results; reporting is a separate job. |
| Convenient developer/small CI path. | Useful when reporting CPU/time should not extend injector job or when reports are regenerated. |
| Report failure is close to test execution. | Requires preserved compatible CSV/properties and a second command. |
| Run directory contains complete immediate artifact set. | Lets one raw JTL feed multiple report configurations later. |
Both require a new/empty report directory.
4. Fixed output path versus run ID
Fixed results/results.jtl encourages overwrite, race
conditions and ambiguous provenance. A unique
results/<run-id>/ makes each run an immutable
evidence unit.
Use human-readable run IDs (timestamp/build SHA/job ID) plus a collision check. Do not automatically delete the previous directory to “make the command work.”
5. Console summary versus artifacts
JMeter CLI provides a running summary by default. It is useful for operator awareness: progress, throughput, errors while the test is active. It is not a substitute for raw JTL/log/target evidence because terminal output can be truncated, interleaved or lost.
Use console summary for observation; use artifacts for gates/audits/comparisons.
6. Wrapper script versus raw CLI
| Raw CLI | Wrapper |
|---|---|
| Best for learning/troubleshooting a single explicit command. | Best for repeated scheduled/CI runs with consistent safeguards. |
| Easy to see every flag directly. | Can validate paths, refuse overwrite, capture versions/hashes/exits/manifests. |
| Operator must remember report/gate/path rules. | Encodes the execution contract once. |
| Good diagnostic fallback when wrapper itself is suspect. | Another code asset that must be reviewed/tested cross-platform. |
Keep the wrapper thin: it should orchestrate JMeter and evidence, not reimplement workload logic hidden from the JMX.
7. Engine failure versus performance threshold failure
Use distinct states:
| Failure class | Example | Primary evidence | Automation action |
|---|---|---|---|
| Preflight/wrapper | missing JMX/property file; existing run dir | wrapper stderr/manifest-pre | no load; non-zero immediately. |
| Engine/process | invalid JMX, initialization/fatal engine problem | engine exit + jmeter.log | non-zero; preserve partial artifacts. |
| Sample correctness | HTTP/assertion failures | JTL success/error fields + target evidence | gate non-zero even if engine exit is 0. |
| Performance SLO | p95/error rate/sample count outside threshold | JTL + configured expectation + generator/target state | gate non-zero; do not rewrite engine status. |
8. -j versus -L
-j chooses the run log file; it should be present on
every run. -L changes logging level/category and is
diagnostic. Broad DEBUG can create generator CPU/disk overhead and
huge logs. Use narrow categories only for a small reproduction,
record the override in the manifest, then restore normal logging.
9. Force delete versus fail-if-exists
-f exists specifically to delete old result/report
paths. In a tightly controlled disposable CI workspace it can be
appropriate if the path is statically validated. The mandatory
launcher does not use it because immutable run directories are safer
and preserve historical evidence by default.
10. CLI properties versus secrets
Current JMeter documentation warns that command-line proxy credentials can be visible to other users on the system. The same operational principle applies broadly: command lines and manifests are not secret stores.
Use -J for non-secret workload/environment knobs. Real
credentials should come from an approved secret injection mechanism
that avoids command echo/log/manifest exposure. This local chapter
uses no credentials.
11. Pinned versus floating tools
Record JMeter/Java versions every run and standardize an approved pair. A “same JMX” comparison across different JMeter/Java/plugin versions is not automatically a regression comparison. If a team later uses containers, pin image digest/version and still capture runtime versions.
12. Configuration-layer boundaries
| Layer | Examples | Do not confuse with |
|---|---|---|
| JMeter core/test plan | CLI flags, JMX, properties, save service, report generator | Java heap/GC or SUT behavior. |
| Java/JVM | Java version, heap, GC, system properties | JTL sample failure semantics. |
| OS/network | cwd, file permissions, disk, DNS/sockets | Target server processing. |
| SUT | service latency/errors/saturation | JMeter process exit or report generation cost. |
| Plugin/driver | none mandatory; future versions must be recorded | Core CLI behavior. |
| CI/container | workspace/run IDs/artifact retention/image startup | Measured target response time unless independently proven. |
13. Worked scenario
A scheduled run must execute 200,000 samples and preserve seven days of evidence.
- CSV JTL, not XML full-body retention.
- unique build/run ID directory; no overwrite;
- raw JTL + jmeter.log + manifest retained as primary evidence;
-
HTML report can be generated in a separate reporting job with
-gif injector job time/resources matter; - engine exit and performance gate stored separately;
- generator disk/CPU/network monitored before claiming target capacity.
14. Decision table
| Requirement | Preferred choice | Evidence |
|---|---|---|
| High-volume load results | CSV JTL | smaller artifacts + dashboard-compatible required fields. |
| Tiny diagnostic needs response body | bounded XML JTL | explicit response-data scope and disk budget. |
| Immediate local analysis | -e -o |
dashboard beside raw JTL. |
| Separate CI reporting stage | -g from retained JTL |
raw run finished before report task. |
| Repeatable scheduled execution | thin wrapper + unique run dir | command/version/hash/exit/gate manifest. |
| Business SLO failure | post-run JTL gate | engine/process status preserved separately. |
15. Configured versus achieved load
Configured load is resolved threads/loops/timers/duration in the input contract. Achieved load is valid sample count/rate observed in JTL and target events. A run with the expected process exit but fewer samples, generator saturation or target errors is not equivalent. Preserve generator CPU/heap/network/disk evidence and experiment limitations before capacity/regression claims.
Knowledge check
Why prefer CSV for high-volume JTL?
It is much smaller than XML and avoids response-body storage while retaining normal timing/status fields.
When is -g preferable to -e?
When reporting should occur later/separately from injector execution or the dashboard must be regenerated from retained raw JTL.
What is the main value of a unique run directory?
It prevents evidence overwrite and gives every JTL/log/report/manifest one unambiguous provenance boundary.
Why can engine exit 0 and CI still fail?
The independent result gate may find sample errors, wrong sample count, or SLO threshold violations.
What evidence is needed before blaming target latency?
Stable configured/achieved workload plus generator CPU/GC/disk/network, JTL, jmeter.log and target telemetry.
Official references and version notes
-
JMeter Getting Started — CLI mode
—
-n,-t,-l,-j,-g,-e,-o,-J,-L, remote flags, and CLI/load guidance. -
JMeter Listeners / Result files
— CSV versus XML, save-service defaults, CLI
-llistener, result fields, and memory guidance. -
JMeter Generating Dashboard Report
— dashboard-required CSV fields,
-g,-e -o, output-folder rules, graphs, and report properties. - JMeter Best Practices — GUI authoring/debugging and CLI execution for load.
- Apache JMeter downloads — current stable release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a Java 17 JDK;
JMeter 5.6.3 requires Java 8+. JMeter's own manual says GUI mode
is for building/debugging while CLI mode must be used for load
testing. Current CLI flags include -n (CLI),
-t (JMX), -l (JTL), -j (run
log), -g <CSV> (report only),
-e (report after test), and -o (report
output). The report output folder must not exist or must be empty.
-f force-deletes an existing result file/report
folder before a test, so the safe launcher in this chapter
intentionally does not use it; instead it refuses to reuse an
existing run directory. CSV result files are smaller than XML and
are the normal large-run choice. Current CSV defaults include
timing/status/thread and byte fields such as
timeStamp, elapsed, label,
responseCode, success,
bytes, sentBytes,
grpThreads, and allThreads when their
save-service fields are enabled (the current defaults required by
the dashboard are documented as correct unless changed). Response
data is not supported in CSV. JMeter process/engine completion is
therefore treated separately from a post-run SLO/result gate: a
launcher first records the engine exit code, then validates JTL
sample counts/failures/latency. A zero engine exit is not used as
evidence that every sample met an SLO.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.