Chapter 23Lesson 04~215 minutes

HTML Dashboard Reports and Performance Result Interpretation: Diagnostics, Failure Modes, and Production Practices

Dashboard failure modes are often interpretation failures rather than software crashes. A technically correct graph can still support a wrong conclusion if averages hide tails, APDEX thresholds are arbitrary, warm-up is mixed into steady state, controller and child samples are counted as one unit, or correlation is presented as a bottleneck diagnosis. Preserve the raw experiment first; then fix the interpretation/report configuration.

Average-onlyAPDEX misuseWarm-up mixingDouble countingCorrelation ≠ causation

Learning objectives

  • Reject average-only performance reporting.
  • Diagnose APDEX scores whose thresholds do not match service objectives.
  • Separate warm-up/ramp and steady-state measurement windows.
  • Prevent Transaction Controller parent/child double counting.
  • Repair dashboard generation from incompatible/missing CSV fields.
  • Use generator/SUT evidence before calling a correlated graph a bottleneck proof.

1. Preserve first-failure evidence

All runnable diagnosis stays local at 127.0.0.1:8023, ≤80 samples/run. Preserve original JTL, jmeter.log, JMX/properties/CLI, report directory/properties, generator state and target events before changing filters or rerunning. Never delete the source JTL because a report is embarrassing or incomplete.

2. Diagnostic sequence

Dashboard diagnostic sequence

A dashboard is downstream of the experiment. It transforms retained sample rows into tables and graphs; it cannot repair an invalid workload, missing sample fields, generator saturation, ambiguous labels, or missing target telemetry.

flowchart TD
E[Preserve JTL + jmeter.log + report props + target evidence] --> V[Confirm JMeter / Java / plugin / tool versions]
V --> C[Confirm JMX / data / properties / CLI + authorized target]
C --> S[Validate tree scope / labels / transaction-parent policy / resolved properties]
S --> P[Inspect protocol / session / data / retries / response codes]
P --> G[Inspect generator CPU / GC / disk / network / result cost]
G --> T[Inspect SUT service/resource telemetry]
T --> X[Inspect distributed / CI / container timing/artifact state if relevant]
X --> F[Least destructive report/config/experiment correction]
F --> R[Small controlled rerun or report regeneration]

3. Failure mode: reporting averages only

Broken conclusion: “Checkout average is 90 ms, therefore performance is fine.” Averages can hide a slow tail or error population.

Repair: report sample/error count plus distribution metrics that match the objective (commonly p90/p95/p99 where volume supports them), response codes and time-series context. Compare raw JTL if a dashboard statistic surprises you.

4. Failure mode: APDEX treated as universal truth

Broken: “APDEX=0.91 means the API passes.” The score is meaningless without satisfied/tolerated thresholds and transaction scope.

Repair: publish thresholds, justify them from service/product objectives, use per-transaction thresholds where appropriate, and keep explicit latency/error/throughput SLOs.

5. Failure mode: warm-up mixed with steady state without disclosure

A JVM/service/cache warms during the first minutes; full-run p95 is then compared with a steady-state-only baseline. The comparison is invalid even if both dashboards look professional.

Repair: retain full-run evidence, predefine phases/windows, and generate an additional start/end filtered report or raw-JTL gate for the exact steady window. Never select the “best” window after seeing the candidate result.

6. Failure mode: transaction parent/children double-counted

Broken JTL contains Checkout Journey plus Catalog/Checkout children and a report author adds all rows together to claim “120 transactions.” Units have been mixed: journey transactions and HTTP requests.

Repair: define business-transaction and request metrics separately; use stable filters/labels and the recommended/default transaction configuration appropriate to the report. Remember dashboard graphs differ in whether controller samples are included.

7. Intentionally broken example: dashboard from incompatible CSV

Create a copy of a valid JTL and remove required fields such as latency, thread counts or label:

Copy-Item .\results\p23-baseline\results.jtl .\results\broken-missing-fields.jtl
# Edit ONLY the copied diagnostic file and remove a required column/header.
# Keep the original baseline JTL untouched.

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -q .\config\local.properties `
  -g .\results\broken-missing-fields.jtl `
  -o .\results\broken-report

Expected: report generation fails or cannot create trustworthy graphs/statistics because the CSV schema no longer satisfies report requirements. Preserve the failed report/log/console evidence.

Repair: regenerate the report from the untouched valid JTL or rerun the smallest test with dashboard-required save fields restored. Do not fabricate missing timing/thread columns or copy values from another run.

8. Failure mode: correlation presented as bottleneck proof

A Response Time vs Request graph rises with request rate, so the report says “the database is saturated.” JMeter has not measured the database.

Possible causes include target CPU/DB/lock saturation, generator CPU/GC, connection-pool behavior, DNS/network, changing payload mix, retries, queueing, or coordinated workload behavior. Correlate with target/generator telemetry and controlled experiments.

9. Failure mode: throughput comparison ignores closed-loop behavior

Degraded Checkout increases loop cycle time, so Catalog hits/sec also drops. A naive report says “Catalog throughput regressed.” But Catalog target processing time is stable; fewer loop cycles reach it because threads are blocked longer on Checkout.

Repair: interpret throughput with workload model and sequence, not as an independent server capacity measure.

10. Failure mode: silent filter changes

Baseline uses no sample filter; candidate excludes error-prone Checkout through sample_filter. The resulting reports are incomparable.

Repair: version/report filter properties beside each run, compare filters before metrics, and preserve full raw JTL.

11. Failure mode: report output directory/property conflict

The output directory already contains files, or a configured HTML output_dir property overrides the expected -o destination. Report generation then fails or writes somewhere surprising.

Repair: use run-specific new/empty output directories and inspect report properties before rerunning -g. This is a report artifact problem; no target load is required to fix it.

12. Causal symptom table

Observation Possible cause Evidence needed before conclusion
p95 rises, target service time rises for same label SUT/controlled target change likely target telemetry + stable generator/workload/filter.
p95 rises, target service time stable, generator CPU/GC high generator-side delay/result/connection/network issue generator JVM/OS/network + JTL timing components.
throughput drops while service time unchanged on control label closed-loop blocking elsewhere sequence/workload model + per-label target service time.
503 errors appear target/service/proxy or assertion semantics response codes/body + target/proxy logs + protocol config.
candidate APDEX drops only after threshold change report policy changed threshold properties; explicit SLO unchanged.
time-series graph empties series/sample filter mismatch filter regex + raw labels + report log.

13. Security-sensitive/disruptive boundaries

Do not use real credentials/PII in labels or error messages. Sustained traffic, recorder certificates, environment variables, database/message/API targets, RMI remote engines, containers, CI secrets and OS/JVM tuning remain separate privileged concerns. This chapter uses fake local errors only and changes no global system setting.

14. Troubleshooting shortcuts to reject

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not assign giant heap to make invalid report generation “finish.”
  • Do not mass-disable listeners/evidence without understanding the source data.
  • Do not use global properties to rewrite sample outcomes.
  • Do not disable TLS/RMI verification.
  • Do not reproduce report issues against a production target.
  • Do not increase workload to get “smoother graphs” before validity is resolved.
  • Do not delete raw/first-failure result files.

Knowledge check

Why is an average-only report weak?

What makes an APDEX comparison invalid even with identical JTL?

Why is a missing-field report a data-contract failure rather than a server failure?

Why can transaction parent+children totals be misleading?

What additional evidence is required before calling a rising response-time-vs-load graph a database bottleneck?

Next lesson

Checkpoint: identify a known regression through multiple independent views

Lesson 5 repeats baseline/degraded runs as a formal checkpoint, verifies predicted errors/latency/throughput/APDEX changes against raw JTL and target telemetry, and produces a concise interpretation with uncertainty.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. The dashboard generator reads compatible CSV sample logs and can run at the end of a load test with -e -o or later with -g <CSV> -o <dir>. Required CSV data include bytes, label, latency, response code/message, success, thread counts/name, elapsed time, connect time, assertion failure message, and a timestamp format containing time; current defaults are suitable unless changed. The HTML output directory must be empty/new. Dashboard statistics expose three configurable percentile levels via aggregate_rpt_pct1/2/3, defaulting to 90/95/99. The report-generator general APDEX defaults are 500 ms satisfied and 1500 ms tolerated; this chapter's local lab deliberately overrides them to 100/300 ms. sample_filter removes sample data before report calculations. HTML series_filter filters displayed series/rows after calculations. start_date/end_date constrain the report measurement window. The default over-time granularity is 60000 ms and must remain above 1000 ms; the tiny local lab uses 2000 ms. Dashboard percentile estimates can differ from GUI Aggregate Report, especially with few/widely distributed samples, because estimator formulas differ. Several dashboard graphs include Transaction Controller sample results while others explicitly ignore/exclude them, so labels and controller-sample policy must be stated before totals are interpreted.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.