HTML Dashboard Reports and Performance Result Interpretation: Core Concepts and Mental Model
Chapter 22 established a trustworthy result-collection contract:
lean CSV JTL, explicit fields, preserved jmeter.log,
understood generator cost, and a privacy-safe retention policy. A
dashboard is the next layer—not a replacement for that evidence.
JMeter's HTML report transforms raw sample rows into statistics,
percentiles, error summaries, APDEX and time-series graphs. Those
summaries are useful only when you can explain which samples they
contain, which workload phase they represent, and whether the
generator and target were valid.
Learning objectives
- Model the dashboard as a transformation of retained CSV/JTL rows.
- Define average, percentile, throughput, error rate, APDEX, label series, and measurement window.
- Understand that different dashboard graphs include/exclude controller samples differently.
- Inspect raw source rows, report properties, labels and time boundaries before interpretation.
- Separate observed correlation from a proven bottleneck/cause.
- Connect every graph to configured load, achieved load, generator state and target telemetry.
1. The practical problem: polished graphs can make weak experiments look authoritative
A dashboard can show a clean p95 line and an APDEX score even if warm-up is mixed into steady state, the load generator saturated, required CSV fields were changed, Transaction Controller parent/child samples are being interpreted inconsistently, or the target received fewer requests than configured. The report has no way to repair those experiment-design problems.
http://127.0.0.1:8023, maximum 2 threads ×20
iterations ×2 samplers = 80 samples/run, 50 ms pacing, synthetic
responses/errors only, no credentials/plugins/remote
engines/containers/public targets.
Never substitute a public/shared/production endpoint
for any mandatory lab or reporting experiment. Dashboards are
interpreted only after raw JTL, generator state and target event
counts pass validity checks.
2. Mental model: clean JTL → report generator → summaries → hypothesis evaluation
A dashboard is downstream of the experiment. It transforms retained sample rows into tables and graphs; it cannot repair an invalid workload, missing sample fields, generator saturation, ambiguous labels, or missing target telemetry.
flowchart TD W[Workload manifest: threads / loops / pacing / labels] --> J[JMeter engine] J --> P[Protocol requests to local target] J --> C[Compatible CSV JTL rows] P --> T[Target service/event telemetry] C --> G[Report generator + properties] G --> S[Statistics + percentiles + APDEX + errors] G --> O[Time-series / distribution / throughput graphs] F[Filters + date window + label/controller policy] --> G S --> H[Interpretation / hypothesis evaluation] O --> H T --> H R[Generator CPU / GC / network / disk + jmeter.log] --> H
The workload manifest defines what JMeter was supposed to do. The engine issues protocol requests and writes compatible CSV sample rows. The target separately records service-side events/timing. The report generator reads the CSV plus report properties and filters to produce tables/graphs. Filters and time windows therefore change the dataset represented by the report. The final interpretation combines dashboard summaries with target telemetry and generator health. The dashboard itself is evidence visualization; it is not causal proof.
3. Core metrics before you use them
| Metric | Meaning | Common interpretation trap |
|---|---|---|
| Average / mean | Arithmetic mean elapsed time for included samples. | Can hide a slow tail or bimodal distribution. |
| Percentile pN | A threshold below which approximately N% of included elapsed times fall. | Estimator/sample-count details matter; p95 is not '95% faster than' anything. |
| Error rate | Failed samples divided by included samples. | Depends on label/controller/filter scope and assertion/HTTP failure semantics. |
| Throughput | Rate of completed requests/transactions in a measurement interval. | Closed workloads can reduce throughput when latency rises even with unchanged configured threads. |
| APDEX | Satisfied/tolerated/frustrated experience score relative to configured thresholds. | Threshold choice is policy; score is not a universal SLO. |
| Active threads | Number of active virtual-user threads over time. | Does not prove achieved request arrival rate. |
| Latency/connect time | JMeter timing subcomponents versus total elapsed. | Must not be conflated with server-only processing time. |
4. Why p90/p95/p99 are usually more informative than average alone
The mean answers an aggregate question; percentiles describe distribution/tail. If 95 of 100 responses are 50 ms and five are 900 ms, the average moves, but p95/p99 make the tail structure much easier to see. Report at least the percentiles that match your engineering/SLO questions—not every available number.
Current JMeter dashboard defaults are p90/p95/p99. The dashboard estimator can differ from GUI Aggregate Report, especially with small or widely spread samples, so small raw-JTL checks should tolerate minor estimator differences rather than assuming one implementation is wrong.
5. APDEX is threshold-dependent policy
APDEX compresses response-time experience relative to a satisfied threshold T and tolerated threshold. JMeter lets you configure global thresholds and per-transaction overrides. Its general defaults are 500 ms satisfied / 1500 ms tolerated.
The mandatory lab overrides those thresholds to 100/300 ms so a deliberate checkout shift from ~45 ms to ~180 ms becomes visible. That does not mean 100/300 is a production SLO. Production thresholds must come from product/service objectives.
6. Sample labels are your reporting schema
The dashboard groups many statistics/series by sample label. Stable
labels such as Catalog and Checkout let
you compare runs. Dynamic labels containing IDs/timestamps create
high-cardinality reports and make regression comparisons difficult.
Transaction Controller samples add another reporting layer. Some dashboard graphs include controller samples; others explicitly ignore/exclude them. Never sum “Journey” + child sampler counts unless that is the metric you intentionally defined.
7. Measurement window is part of every claim
A full-run report may contain startup/warm-up, ramp, steady state,
spike/recovery and shutdown. If you claim “steady-state p95,” define
the time range or run phase used. JMeter report generator supports
start_date/end_date; its
date_format determines parsing.
Filtering the measurement window is legitimate only when disclosed and motivated before interpreting the result—not after seeing an unfavorable graph.
8. sample_filter versus HTML series_filter
jmeter.reportgenerator.sample_filter filters source
samples before calculations. It can change
statistics/APDEX/errors. HTML series_filter is applied
after calculations to simplify displayed series/summary rows. These
mechanisms answer different questions.
An empty/misconfigured series filter can also create empty-looking graphs. Always retain the raw JTL and filter configuration.
9. Graphs answer different questions
- Response Times Over Time: how elapsed behavior changes through the run.
- Response Time Percentiles Over Time: how tail estimates evolve by time bucket.
- Active Threads Over Time: workload thread activity, not request arrival by itself.
- Hits/Transactions per Second: completed request/transaction rates; controller inclusion differs.
- Response Codes per Second: when HTTP/status failures appear.
- Response Time vs Request/Threads: correlation with concurrent activity, not proof that concurrency caused the latency.
- Response Time Distribution/Percentiles: whole-window distribution rather than a chronology.
10. Read-only inspection before report generation
PowerShell:
Get-Content .\results\run\results.jtl -TotalCount 6
(Get-Content .\results\run\results.jtl | Measure-Object -Line).Lines
Get-Content .\config\local.properties |
Select-String -Pattern "save.saveservice|reportgenerator|aggregate_rpt_pct"
Get-Item .\results\run\results.jtl | Select-Object FullName,Length,LastWriteTime
Also record unique labels/response codes and JTL start/end
timestamps with the raw summarizer in Lesson 2. You should know the
source dataset before opening index.html.
11. State checklist before interpretation
| State | Question |
|---|---|
| Generator | JMeter/Java versions; CPU/heap/GC/network/disk healthy; report generation after load? |
| Thread/arrival | Resolved threads/loops/timers/ramp/duration and expected sample count? |
| Component scope | Which sampler/controller labels exist; parent transaction samples generated? |
| Variables/properties/data | Which run/mode properties and datasets produced the JTL? |
| Protocol/session | Any retries, redirects, connection/session behavior altering sample meaning? |
| Target | Exact authorized target and service-side latency/error/event evidence? |
| Artifacts | Raw CSV, jmeter.log, dashboard dir, report properties, manifest? |
| Trust/privacy | Do labels/error tables/URLs expose sensitive data? |
| Validity | Configured versus achieved samples/rate, generator headroom, window/filter disclosed? |
12. DevOps connection
Dashboards are decision aids. A trustworthy pull-request/release note should point from “p95 rose” to the exact run IDs, JTL, measurement window, workload manifest, generator state and service telemetry. That traceability turns a chart into an engineering argument rather than a screenshot.
Knowledge check
Why can two dashboards from the same JTL show different statistics?
Different report properties such as sample_filter or start/end measurement windows can change the source rows included in calculations.
Why isn't APDEX a universal pass/fail metric?
Its meaning depends on satisfied/tolerated thresholds that must reflect the service's actual objectives.
Why may throughput fall when Checkout latency rises in a closed thread model?
Threads spend longer waiting for Checkout, so they complete fewer loop iterations/requests per unit time even with unchanged thread count.
Why must transaction parent/child policy be stated?
Different report tables/graphs include or exclude controller samples differently, so careless summing can double-count logical work.
What does a response-time-vs-request correlation prove?
Association in the recorded dataset, not causality; generator/SUT/network evidence is needed to identify a bottleneck.
Official references and version notes
-
JMeter User Manual — Generating Dashboard Report
— required CSV fields, APDEX, statistics/percentiles, errors,
graphs, filters, date windows, granularity,
-g,-e,-o, and output-folder rules. - JMeter Getting Started — CLI/load analysis — GUI authoring/debugging, CLI load execution, CSV/XML results and HTML analysis.
- JMeter Listeners / Result files — sample-log fields and raw-result interpretation.
- Component Reference — Transaction Controller — parent/additional transaction samples and reporting semantics.
- JMeter Properties Reference — report/save-service configuration and defaults.
- Apache JMeter downloads — current stable release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a Java 17 JDK;
JMeter 5.6.3 requires Java 8+. The dashboard generator reads
compatible CSV sample logs and can run at the end of a load test
with -e -o or later with
-g <CSV> -o <dir>. Required CSV data
include bytes, label, latency, response code/message, success,
thread counts/name, elapsed time, connect time, assertion failure
message, and a timestamp format containing time; current defaults
are suitable unless changed. The HTML output directory must be
empty/new. Dashboard statistics expose three configurable
percentile levels via aggregate_rpt_pct1/2/3,
defaulting to 90/95/99. The report-generator general APDEX
defaults are 500 ms satisfied and 1500 ms tolerated; this
chapter's local lab deliberately overrides them to 100/300 ms.
sample_filter removes sample data before report
calculations. HTML series_filter filters displayed
series/rows after calculations. start_date/end_date
constrain the report measurement window. The default over-time
granularity is 60000 ms and must remain above 1000 ms; the tiny
local lab uses 2000 ms. Dashboard percentile estimates can differ
from GUI Aggregate Report, especially with few/widely distributed
samples, because estimator formulas differ. Several dashboard
graphs include Transaction Controller sample results while others
explicitly ignore/exclude them, so labels and controller-sample
policy must be stated before totals are interpreted.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.