HTML Dashboard Reports and Performance Result Interpretation: Diagnostics, Failure Modes, and Production Practices
Dashboard failure modes are often interpretation failures rather than software crashes. A technically correct graph can still support a wrong conclusion if averages hide tails, APDEX thresholds are arbitrary, warm-up is mixed into steady state, controller and child samples are counted as one unit, or correlation is presented as a bottleneck diagnosis. Preserve the raw experiment first; then fix the interpretation/report configuration.
Learning objectives
- Reject average-only performance reporting.
- Diagnose APDEX scores whose thresholds do not match service objectives.
- Separate warm-up/ramp and steady-state measurement windows.
- Prevent Transaction Controller parent/child double counting.
- Repair dashboard generation from incompatible/missing CSV fields.
- Use generator/SUT evidence before calling a correlated graph a bottleneck proof.
1. Preserve first-failure evidence
127.0.0.1:8023, ≤80 samples/run.
Preserve original JTL, jmeter.log, JMX/properties/CLI, report
directory/properties, generator state and target events before
changing filters or rerunning. Never delete the source JTL because a
report is embarrassing or incomplete.
2. Diagnostic sequence
A dashboard is downstream of the experiment. It transforms retained sample rows into tables and graphs; it cannot repair an invalid workload, missing sample fields, generator saturation, ambiguous labels, or missing target telemetry.
flowchart TD E[Preserve JTL + jmeter.log + report props + target evidence] --> V[Confirm JMeter / Java / plugin / tool versions] V --> C[Confirm JMX / data / properties / CLI + authorized target] C --> S[Validate tree scope / labels / transaction-parent policy / resolved properties] S --> P[Inspect protocol / session / data / retries / response codes] P --> G[Inspect generator CPU / GC / disk / network / result cost] G --> T[Inspect SUT service/resource telemetry] T --> X[Inspect distributed / CI / container timing/artifact state if relevant] X --> F[Least destructive report/config/experiment correction] F --> R[Small controlled rerun or report regeneration]
3. Failure mode: reporting averages only
Broken conclusion: “Checkout average is 90 ms, therefore performance is fine.” Averages can hide a slow tail or error population.
Repair: report sample/error count plus distribution metrics that match the objective (commonly p90/p95/p99 where volume supports them), response codes and time-series context. Compare raw JTL if a dashboard statistic surprises you.
4. Failure mode: APDEX treated as universal truth
Broken: “APDEX=0.91 means the API passes.” The score is meaningless without satisfied/tolerated thresholds and transaction scope.
Repair: publish thresholds, justify them from service/product objectives, use per-transaction thresholds where appropriate, and keep explicit latency/error/throughput SLOs.
5. Failure mode: warm-up mixed with steady state without disclosure
A JVM/service/cache warms during the first minutes; full-run p95 is then compared with a steady-state-only baseline. The comparison is invalid even if both dashboards look professional.
Repair: retain full-run evidence, predefine phases/windows, and generate an additional start/end filtered report or raw-JTL gate for the exact steady window. Never select the “best” window after seeing the candidate result.
6. Failure mode: transaction parent/children double-counted
Broken JTL contains Checkout Journey plus
Catalog/Checkout children and a report
author adds all rows together to claim “120 transactions.” Units
have been mixed: journey transactions and HTTP requests.
Repair: define business-transaction and request metrics separately; use stable filters/labels and the recommended/default transaction configuration appropriate to the report. Remember dashboard graphs differ in whether controller samples are included.
7. Intentionally broken example: dashboard from incompatible CSV
Create a copy of a valid JTL and remove required fields
such as latency, thread counts or label:
Copy-Item .\results\p23-baseline\results.jtl .\results\broken-missing-fields.jtl
# Edit ONLY the copied diagnostic file and remove a required column/header.
# Keep the original baseline JTL untouched.
& "$env:JMETER_HOME\bin\jmeter.bat" `
-q .\config\local.properties `
-g .\results\broken-missing-fields.jtl `
-o .\results\broken-report
Expected: report generation fails or cannot create trustworthy graphs/statistics because the CSV schema no longer satisfies report requirements. Preserve the failed report/log/console evidence.
Repair: regenerate the report from the untouched valid JTL or rerun the smallest test with dashboard-required save fields restored. Do not fabricate missing timing/thread columns or copy values from another run.
8. Failure mode: correlation presented as bottleneck proof
A Response Time vs Request graph rises with request rate, so the report says “the database is saturated.” JMeter has not measured the database.
Possible causes include target CPU/DB/lock saturation, generator CPU/GC, connection-pool behavior, DNS/network, changing payload mix, retries, queueing, or coordinated workload behavior. Correlate with target/generator telemetry and controlled experiments.
9. Failure mode: throughput comparison ignores closed-loop behavior
Degraded Checkout increases loop cycle time, so Catalog hits/sec also drops. A naive report says “Catalog throughput regressed.” But Catalog target processing time is stable; fewer loop cycles reach it because threads are blocked longer on Checkout.
Repair: interpret throughput with workload model and sequence, not as an independent server capacity measure.
10. Failure mode: silent filter changes
Baseline uses no sample filter; candidate excludes error-prone
Checkout through sample_filter. The resulting reports
are incomparable.
Repair: version/report filter properties beside each run, compare filters before metrics, and preserve full raw JTL.
11. Failure mode: report output directory/property conflict
The output directory already contains files, or a configured HTML
output_dir property overrides the expected
-o destination. Report generation then fails or writes
somewhere surprising.
Repair: use run-specific new/empty output directories and inspect
report properties before rerunning -g. This is a report
artifact problem; no target load is required to fix it.
12. Causal symptom table
| Observation | Possible cause | Evidence needed before conclusion |
|---|---|---|
| p95 rises, target service time rises for same label | SUT/controlled target change likely | target telemetry + stable generator/workload/filter. |
| p95 rises, target service time stable, generator CPU/GC high | generator-side delay/result/connection/network issue | generator JVM/OS/network + JTL timing components. |
| throughput drops while service time unchanged on control label | closed-loop blocking elsewhere | sequence/workload model + per-label target service time. |
| 503 errors appear | target/service/proxy or assertion semantics | response codes/body + target/proxy logs + protocol config. |
| candidate APDEX drops only after threshold change | report policy changed | threshold properties; explicit SLO unchanged. |
| time-series graph empties | series/sample filter mismatch | filter regex + raw labels + report log. |
13. Security-sensitive/disruptive boundaries
Do not use real credentials/PII in labels or error messages. Sustained traffic, recorder certificates, environment variables, database/message/API targets, RMI remote engines, containers, CI secrets and OS/JVM tuning remain separate privileged concerns. This chapter uses fake local errors only and changes no global system setting.
14. Troubleshooting shortcuts to reject
- Do not add blanket retries or arbitrary long sleeps.
- Do not assign giant heap to make invalid report generation “finish.”
- Do not mass-disable listeners/evidence without understanding the source data.
- Do not use global properties to rewrite sample outcomes.
- Do not disable TLS/RMI verification.
- Do not reproduce report issues against a production target.
- Do not increase workload to get “smoother graphs” before validity is resolved.
- Do not delete raw/first-failure result files.
Knowledge check
Why is an average-only report weak?
The mean can hide tail latency, distribution shape and errors; percentiles/error/time-series/raw evidence are needed.
What makes an APDEX comparison invalid even with identical JTL?
Changing satisfied/tolerated thresholds or per-transaction threshold policy changes the score.
Why is a missing-field report a data-contract failure rather than a server failure?
The report generator cannot derive required statistics from an incompatible CSV; no new target behavior is involved.
Why can transaction parent+children totals be misleading?
They represent different units—business transaction versus component requests—and some graphs include/exclude controllers differently.
What additional evidence is required before calling a rising response-time-vs-load graph a database bottleneck?
Database/SUT telemetry plus generator/network/workload evidence and preferably a controlled experiment; the graph alone shows correlation.
Official references and version notes
-
JMeter User Manual — Generating Dashboard Report
— required CSV fields, APDEX, statistics/percentiles, errors,
graphs, filters, date windows, granularity,
-g,-e,-o, and output-folder rules. - JMeter Getting Started — CLI/load analysis — GUI authoring/debugging, CLI load execution, CSV/XML results and HTML analysis.
- JMeter Listeners / Result files — sample-log fields and raw-result interpretation.
- Component Reference — Transaction Controller — parent/additional transaction samples and reporting semantics.
- JMeter Properties Reference — report/save-service configuration and defaults.
- Apache JMeter downloads — current stable release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a Java 17 JDK;
JMeter 5.6.3 requires Java 8+. The dashboard generator reads
compatible CSV sample logs and can run at the end of a load test
with -e -o or later with
-g <CSV> -o <dir>. Required CSV data
include bytes, label, latency, response code/message, success,
thread counts/name, elapsed time, connect time, assertion failure
message, and a timestamp format containing time; current defaults
are suitable unless changed. The HTML output directory must be
empty/new. Dashboard statistics expose three configurable
percentile levels via aggregate_rpt_pct1/2/3,
defaulting to 90/95/99. The report-generator general APDEX
defaults are 500 ms satisfied and 1500 ms tolerated; this
chapter's local lab deliberately overrides them to 100/300 ms.
sample_filter removes sample data before report
calculations. HTML series_filter filters displayed
series/rows after calculations. start_date/end_date
constrain the report measurement window. The default over-time
granularity is 60000 ms and must remain above 1000 ms; the tiny
local lab uses 2000 ms. Dashboard percentile estimates can differ
from GUI Aggregate Report, especially with few/widely distributed
samples, because estimator formulas differ. Several dashboard
graphs include Transaction Controller sample results while others
explicitly ignore/exclude them, so labels and controller-sample
policy must be stated before totals are interpreted.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.