Chapter 23Lesson 03~195 minutes

HTML Dashboard Reports and Performance Result Interpretation: Configuration, Design Patterns, and Trade-Offs

Dashboard configuration is part of the measurement contract. Changing percentile levels, APDEX thresholds, sample filters, time windows, transaction-parent policy or series visibility can change the report—or its meaning—without changing a single target request. Production reporting therefore needs explicit choices that are stable across comparable runs.

Percentile policyAPDEX vs SLOTransaction labelsTime windowsDashboard vs telemetry

Learning objectives

  • Select percentile levels based on sample volume and decision/SLO needs.
  • Separate APDEX threshold policy from explicit latency/error/throughput SLOs.
  • Choose transaction-level or low-level sampler labels deliberately.
  • Use full-run/time-window/sample/series filtering without hiding data.
  • Choose HTML dashboards or time-series backends based on run duration and analysis needs.
  • Control configuration layers so report comparisons remain portable and valid.

1. Mandatory path remains local/free

Runnable examples stay 127.0.0.1:8023. Backend telemetry is optional discussion; Chapter 24 provides a dedicated local InfluxDB/Graphite/Grafana path. No paid platform is required.

2. Which percentiles should you report?

JMeter's dashboard exposes three configurable percentile slots, default p90/p95/p99. Choose percentiles that match:

  • the service objective/regression policy;
  • sample volume—p99 from 20 samples is unstable and low-resolution;
  • traffic criticality—high-tail objectives may require far more samples;
  • comparison consistency—the same percentile policy across baseline/candidate runs.

A useful dashboard may report p50 (via raw analysis) plus p90/p95/p99, but do not turn every percentile into a release gate.

3. APDEX versus explicit SLOs

APDEX Explicit SLO/gate
One normalized satisfaction score driven by configured satisfied/tolerated thresholds. Specific requirements such as p95≤250 ms, error rate≤0.5%, throughput≥X under workload Y.
Convenient for high-level experience comparison. More direct for contractual/release engineering decisions.
Can differ per transaction via apdex_per_transaction. Usually defined per user journey/API tier/operation.
Can hide which dimension changed; inspect source statistics/errors. Should still be tied to valid workload/generator/target evidence.

Use APDEX as a summary, not a substitute for latency/error/throughput objectives.

4. Per-transaction APDEX thresholds

If Search and Checkout have different objectives, current JMeter supports:

jmeter.reportgenerator.apdex_per_transaction=Catalog:75|200;\
Checkout:100|300

Names/regex must remain stable. Document the thresholds beside the report because a score cannot be interpreted without them.

5. Transaction-controller labels versus low-level sampler labels

A business journey such as “Checkout Journey” can be easier to report than five internal HTTP requests. Transaction Controller can add a transaction sample, but then label/count semantics must be explicit.

Current dashboard behavior varies by graph: request summary/hits/bytes/codes and some overview graphs ignore/exclude Transaction Controller samples, while response-time, latency, transactions/sec, distribution and percentile graphs may include them. The JMeter dashboard documentation recommends the default Generate parent sample unchecked for accurate reporting in normal transaction-controller usage.

6. Avoid double-counting parent and children

If a transaction sample and its child samplers all exist in JTL, “total samples” is not automatically “business transactions.” Define:

  • request rate: low-level protocol samplers only;
  • business transaction rate: transaction/controller label only;
  • latency: whichever layer your SLO is defined on.

Never add parent + children to produce a business throughput count unless you intentionally want mixed units.

7. Full run versus time-window filtering

Full-run report Filtered measurement window
Shows all startup/ramp/steady/recovery/shutdown behavior. Useful for a predeclared steady-state or incident interval.
Best for operational chronology and completeness. Best for a specific hypothesis with a defined window.
Can mix warm-up and steady state in summary percentiles. Can hide transitions if chosen after seeing the result.
Always retain it or raw JTL even when producing a filtered view. Publish exact start/end/date_format and rationale.

8. Sample filter versus series filter

Example pre-calculation source filter:

jmeter.reportgenerator.sample_filter=^(Catalog|Checkout)$

Example HTML display filter:

jmeter.reportgenerator.exporter.html.series_filter=^(Catalog|Checkout)(-success|-failure)?$
jmeter.reportgenerator.exporter.html.filters_only_sample_series=true

The first changes the data used for calculations. The second primarily changes what series/rows are displayed after calculations. Record which was used.

9. Time-series granularity

The current default over-time granularity is 60 seconds. That is reasonable for longer tests but would collapse the short local Chapter 23 run. The lab uses 2000 ms; current JMeter warns the granularity must remain above 1000 ms or throughput graphs become incorrect.

Do not use tiny buckets to create noisy “precision.” Choose a bucket size that produces enough samples per bucket for your workload duration/rate.

10. HTML dashboard versus time-series backend

HTML dashboard Time-series backend
Static post-run artifact from retained CSV. Near-real-time long-run/fleet telemetry.
Excellent for portable evidence/review. Excellent for live correlation with SUT/generator metrics.
No external service required. Requires backend/network/schema/retention operations.
Report generation cost occurs after/around load depending workflow. Metric serialization/network cost occurs during load.
Best with raw JTL preserved. Still retain enough raw/forensic evidence for errors/regression.

Chapter 24 will implement the backend path; Chapter 23 remains fully local with HTML.

11. Labels/error tables can leak data

Dynamic labels, response messages and assertion failure messages can contain identifiers or sensitive values. Keep labels stable and non-sensitive. Treat dashboard directories as derived performance evidence with the same access/retention review as their source JTL.

12. Layer boundaries

Layer Examples Do not confuse with
JMeter core/report save fields, report properties, filters, labels, APDEX/percentiles Java heap/GC or SUT processing.
Java/JVM heap, GC, report-generator memory target error rate.
OS/network disk I/O, cwd/output permissions, DNS/network dashboard sample semantics.
SUT service timings/errors/resource saturation report filter/configuration effects.
Plugin/backend future Backend Listener client/schema core HTML generator.
CI/container workspace/report artifact path/retention measured production capacity.

13. Worked scenario

A 30-minute test has 5-minute ramp, 20-minute steady state and 5-minute recovery. Release SLO is Checkout p95≤250 ms and errors≤1% during steady state.

  • Retain full raw JTL and full dashboard for chronology.
  • Create a second explicitly filtered 20-minute report using predetermined start_date/end_date.
  • Gate Checkout p95/error from raw rows in that window.
  • Report APDEX only as an additional summary with published thresholds.
  • Compare generator/SUT resource telemetry for the same window.

14. Decision table

Decision Preferred choice Why/evidence
Regression tail metric p95 plus p99 only if sample volume supports it Stable, policy-linked percentiles.
Business summary APDEX + explicit thresholds, not alone Readable but policy-dependent.
Request debugging low-level sampler labels Maps directly to protocol operations/errors.
Business-journey SLO Transaction label with documented parent/child policy Matches user-level objective.
Steady-state gate predeclared start/end window + raw-JTL gate Avoids post-hoc warm-up mixing.
Long soak/live triage time-series backend + lean JTL HTML alone is post-run/static; Chapter24 covers backend cost.

15. Evidence contract for comparable reports

Every dashboard comparison must retain the source CSV JTL, the matching jmeter.log, the exact report properties/filter/window configuration, workload manifest, JMX/property hashes, generator resource evidence, and independent target telemetry. The HTML directory is derived evidence and must never be the only retained artifact.

16. Configured versus achieved load

Dashboard throughput/errors/latency describe completed recorded samples, not configured intent. Before a regression claim, verify expected sample count/rate, generator headroom, target event count and the exact filter/window. A candidate run that completed fewer requests because the generator saturated is not comparable merely because its p95 looks lower.

Knowledge check

Why can p99 be misleading with a tiny sample count?

When should APDEX thresholds differ per transaction?

What is the difference between sample_filter and series_filter?

Why retain a full-run dashboard if the gate uses a steady-state window?

Why isn't a static HTML dashboard sufficient for a long soak incident?

Next lesson

Diagnose report interpretation failures

Lesson 4 engineers average-only, APDEX, warm-up, transaction double-counting, incompatible-CSV and false-correlation mistakes and repairs them without hiding the original JTL.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. The dashboard generator reads compatible CSV sample logs and can run at the end of a load test with -e -o or later with -g <CSV> -o <dir>. Required CSV data include bytes, label, latency, response code/message, success, thread counts/name, elapsed time, connect time, assertion failure message, and a timestamp format containing time; current defaults are suitable unless changed. The HTML output directory must be empty/new. Dashboard statistics expose three configurable percentile levels via aggregate_rpt_pct1/2/3, defaulting to 90/95/99. The report-generator general APDEX defaults are 500 ms satisfied and 1500 ms tolerated; this chapter's local lab deliberately overrides them to 100/300 ms. sample_filter removes sample data before report calculations. HTML series_filter filters displayed series/rows after calculations. start_date/end_date constrain the report measurement window. The default over-time granularity is 60000 ms and must remain above 1000 ms; the tiny local lab uses 2000 ms. Dashboard percentile estimates can differ from GUI Aggregate Report, especially with few/widely distributed samples, because estimator formulas differ. Several dashboard graphs include Transaction Controller sample results while others explicitly ignore/exclude them, so labels and controller-sample policy must be stated before totals are interpreted.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.