Chapter 23Lesson 02~270 minutes

HTML Dashboard Reports and Performance Result Interpretation: Guided Hands-On Workflow

The workflow uses two identical JMeter workloads. Only the target's Checkout behavior changes: baseline checkout sleeps about 45 ms and succeeds; degraded checkout sleeps about 180 ms and returns a deterministic HTTP 503 every fifth sequence. Catalog stays at ~20 ms. That design gives the dashboard a known signal and gives target telemetry an independent causal reference.

Baseline/degraded-e -o-gRaw JTL verificationAPDEX/errors/graphs

Learning objectives

  • Run a bounded baseline and degraded local experiment from one JMX.
  • Generate a dashboard at end of a run and later from retained JTL.
  • Inspect dashboard statistics, percentiles, errors, APDEX, throughput and time-series graphs.
  • Map dashboard labels back to JMX sampler labels.
  • Compare dashboard claims with raw JTL calculations and target events.
  • Write an interpretation note that separates observations from causal evidence.

1. Safety envelope

Only http://127.0.0.1:8023. 2 threads ×20 loops ×2 HTTP samplers = 80 samples/run, 50 ms pacing, ≤15 seconds/run, two runs only for the mandatory workflow, synthetic 503 failures only in degraded Checkout, no retries/credentials/plugins/remote engines. Abort on non-loopback target, >80 samples/run, unexpected labels, generator saturation, more than the expected 8 degraded Checkout failures, or target/JTL count disagreement.

2. Create the deterministic local target

Save fixtures/dashboard_fixture.py:

from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.parse import urlparse, parse_qs
import argparse
import json
import re
import threading
import time

FIXTURE_VERSION = "prompt23-dashboard-fixture-v1"
SAFE_TOKEN = re.compile(r"^[A-Za-z0-9_.-]{1,64}$")

lock = threading.Lock()
event_log = None
metrics = {
    "requests": 0,
    "errors": 0,
    "catalog": 0,
    "checkout": 0,
    "by_run": {},
    "by_mode": {},
}

def now_ms():
    return int(time.time() * 1000)

def log_event(event):
    if event_log is None:
        return
    with lock:
        with event_log.open("a", encoding="utf-8") as handle:
            handle.write(json.dumps(event, sort_keys=True) + "\n")

class Handler(BaseHTTPRequestHandler):
    protocol_version = "HTTP/1.1"

    def send_json(self, status, payload):
        raw = json.dumps(payload, sort_keys=True).encode("utf-8")
        self.send_response(status)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(raw)))
        self.send_header("X-Fixture-Version", FIXTURE_VERSION)
        self.end_headers()
        self.wfile.write(raw)

    def record(self, started, operation, status, **extra):
        with lock:
            metrics["requests"] += 1
            if status >= 400:
                metrics["errors"] += 1
            if operation == "catalog" and status == 200:
                metrics["catalog"] += 1
            if operation == "checkout":
                metrics["checkout"] += 1
            run_id = extra.get("run_id", "")
            mode = extra.get("mode", "")
            if run_id:
                metrics["by_run"][run_id] = metrics["by_run"].get(run_id, 0) + 1
            if mode:
                metrics["by_mode"][mode] = metrics["by_mode"].get(mode, 0) + 1

        event = {
            "ts_ms": now_ms(),
            "operation": operation,
            "status": status,
            "service_wall_ms": now_ms() - started,
        }
        event.update(extra)
        log_event(event)

    def parse_common(self, parsed):
        q = parse_qs(parsed.query)
        run_id = q.get("run_id", [""])[0]
        mode = q.get("mode", [""])[0]
        thread_id = q.get("thread", [""])[0]
        seq_raw = q.get("seq", [""])[0]

        if not SAFE_TOKEN.fullmatch(run_id) or not SAFE_TOKEN.fullmatch(thread_id):
            return None, (400, {"status": "invalid_metadata"})
        if mode not in {"baseline", "degraded"}:
            return None, (400, {"status": "invalid_mode"})
        try:
            seq = int(seq_raw)
        except ValueError:
            seq = -1
        if not 1 <= seq <= 1000:
            return None, (400, {"status": "invalid_seq"})

        return {
            "run_id": run_id,
            "mode": mode,
            "thread": thread_id,
            "seq": seq,
        }, None

    def do_GET(self):
        started = now_ms()
        parsed = urlparse(self.path)

        if parsed.path == "/health":
            self.send_json(200, {"status": "ok", "fixture_version": FIXTURE_VERSION})
            self.record(started, "health", 200)
            return

        if parsed.path == "/stats":
            with lock:
                snapshot = json.loads(json.dumps(metrics))
            self.send_json(200, {"fixture_version": FIXTURE_VERSION, "metrics": snapshot})
            self.record(started, "stats", 200)
            return

        if parsed.path not in {"/catalog", "/checkout"}:
            self.send_json(404, {"status": "not_found"})
            self.record(started, "unknown", 404)
            return

        common, error = self.parse_common(parsed)
        if error:
            status, payload = error
            self.send_json(status, payload)
            self.record(started, parsed.path.lstrip("/"), status)
            return

        if parsed.path == "/catalog":
            # Intentionally unchanged between modes.
            time.sleep(0.020)
            self.send_json(200, {
                "status": "ok",
                "operation": "catalog",
                **common,
            })
            self.record(started, "catalog", 200, **common)
            return

        # Checkout is the only deliberately changed target behavior.
        if common["mode"] == "baseline":
            time.sleep(0.045)
            status = 200
            payload = {"status": "ok", "operation": "checkout", **common}
        else:
            time.sleep(0.180)
            if common["seq"] % 5 == 0:
                status = 503
                payload = {
                    "status": "synthetic_failure",
                    "error": "synthetic_checkout_failure",
                    "operation": "checkout",
                    **common,
                }
            else:
                status = 200
                payload = {"status": "ok", "operation": "checkout", **common}

        self.send_json(status, payload)
        self.record(started, "checkout", status, **common)

    def log_message(self, format, *args):
        return

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--host", default="127.0.0.1")
    parser.add_argument("--port", type=int, default=8023)
    parser.add_argument("--log", default="results/server-events.jsonl")
    args = parser.parse_args()

    global event_log
    event_log = Path(args.log).resolve()
    event_log.parent.mkdir(parents=True, exist_ok=True)
    event_log.write_text("", encoding="utf-8")

    print(f"fixture_version={FIXTURE_VERSION}")
    print(f"listen=http://{args.host}:{args.port}")
    print(f"event_log={event_log}")
    ThreadingHTTPServer((args.host, args.port), Handler).serve_forever()

if __name__ == "__main__":
    main()

Start:

python .\fixtures\dashboard_fixture.py `
  --host 127.0.0.1 `
  --port 8023 `
  --log .\results\server-events.jsonl

The mode is supplied by each request. This lets one fixture process record both runs with explicit run_id/mode metadata.

3. Target/generator preflight

curl --fail --silent http://127.0.0.1:8023/health
curl --fail --silent http://127.0.0.1:8023/stats

Record JMeter 5.6.3, Java 17, free disk, generator CPU/heap baseline and target event count. No load runs in GUI.

4. Create workload/report properties

config/local.properties:

# Prompt 23 local workload + dashboard input contract
target.host=127.0.0.1
target.port=8023
threads=2
loops=20
pacing.ms=50
connect.timeout.ms=500
response.timeout.ms=2000

# Dashboard-compatible lean CSV fields.
jmeter.save.saveservice.output_format=csv
jmeter.save.saveservice.print_field_names=true
jmeter.save.saveservice.bytes=true
jmeter.save.saveservice.sent_bytes=true
jmeter.save.saveservice.label=true
jmeter.save.saveservice.latency=true
jmeter.save.saveservice.response_code=true
jmeter.save.saveservice.response_message=true
jmeter.save.saveservice.successful=true
jmeter.save.saveservice.thread_counts=true
jmeter.save.saveservice.thread_name=true
jmeter.save.saveservice.time=true
jmeter.save.saveservice.connect_time=true
jmeter.save.saveservice.assertion_results_failure_message=true
jmeter.save.saveservice.timestamp_format=ms

# Keep privacy-heavy content out of the JTL.
jmeter.save.saveservice.response_data=false
jmeter.save.saveservice.response_data.on_error=false
jmeter.save.saveservice.samplerData=false
jmeter.save.saveservice.responseHeaders=false
jmeter.save.saveservice.requestHeaders=false

# Lab report settings. These intentionally differ from JMeter defaults
# so the known 180ms checkout degradation is visible in APDEX.
jmeter.reportgenerator.report_title=Prompt 23 Local Dashboard Lab
jmeter.reportgenerator.overall_granularity=2000
jmeter.reportgenerator.apdex_satisfied_threshold=100
jmeter.reportgenerator.apdex_tolerated_threshold=300
aggregate_rpt_pct1=90
aggregate_rpt_pct2=95
aggregate_rpt_pct3=99

Two details are intentional: overall_granularity=2000 because the runs last only a few seconds, and APDEX 100/300 ms because the known 180 ms degraded Checkout should move from “satisfied” toward “tolerated/frustrated.” These are lab thresholds, not production objectives.

5. Build one JMX in GUI

Test Plan
├── HTTP Request Defaults
│   host=${__P(target.host,127.0.0.1)}
│   port=${__P(target.port,8023)}
│   timeouts from properties
└── Thread Group
    threads=${__P(threads,1)}
    loops=${__P(loops,1)}
    ├── Counter -> SEQ (start=1, per-user=true)
    ├── HTTP Request — Catalog
    │   GET /catalog
    │   run_id=${__P(run.id,p23-local)}
    │   mode=${__P(mode,baseline)}
    │   thread=T${__threadNum}
    │   seq=${SEQ}
    │   └── Constant Timer ${__P(pacing.ms,50)} ms
    └── HTTP Request — Checkout
        GET /checkout
        run_id=${__P(run.id,p23-local)}
        mode=${__P(mode,baseline)}
        thread=T${__threadNum}
        seq=${SEQ}
        └── Constant Timer ${__P(pacing.ms,50)} ms

The Counter value is shared by the two samplers within an iteration, so degraded Checkout fails at sequence 5/10/15/20 in each thread: 8 checkout failures total. Catalog remains successful.

6. Predict before execution

  • Both runs: 40 Catalog + 40 Checkout = 80 JTL/target operations.
  • Baseline: zero expected HTTP failures.
  • Degraded: 8 Checkout HTTP 503 failures = 20% Checkout error rate and 10% overall error rate.
  • Catalog elapsed/service time should remain near baseline.
  • Checkout p90/p95/p99 and average should increase materially.
  • Checkout APDEX should decline under 100/300 ms lab thresholds.
  • Closed-loop throughput may decline in degraded mode because threads wait longer in Checkout.

7. Baseline — generate dashboard at end

PowerShell:

New-Item -ItemType Directory -Force .\results\p23-baseline | Out-Null

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -n `
  -t .\plans\dashboard-local.jmx `
  -q .\config\local.properties `
  -Jrun.id=p23-baseline `
  -Jmode=baseline `
  -l .\results\p23-baseline\results.jtl `
  -j .\results\p23-baseline\jmeter.log `
  -e `
  -o .\results\p23-baseline\html-report

Do not create html-report beforehand. Preserve raw JTL even though the dashboard is generated immediately.

8. Degraded — retain JTL first

New-Item -ItemType Directory -Force .\results\p23-degraded | Out-Null

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -n `
  -t .\plans\dashboard-local.jmx `
  -q .\config\local.properties `
  -Jrun.id=p23-degraded `
  -Jmode=degraded `
  -l .\results\p23-degraded\results.jtl `
  -j .\results\p23-degraded\jmeter.log

Expected engine completion does not make the run “green”: the JTL intentionally contains eight HTTP failures.

9. Generate degraded dashboard later from JTL

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -q .\config\local.properties `
  -g .\results\p23-degraded\results.jtl `
  -o .\results\p23-degraded\html-report

No target traffic occurs during -g. The command reads the existing CSV plus report properties and writes a new/empty report directory.

10. Inspect raw JTL before opening dashboards

Save tools/summarize_jtl.py:

import csv
import json
import math
import sys
from collections import defaultdict
from pathlib import Path

def nearest_rank(values, pct):
    data = sorted(values)
    if not data:
        return 0
    idx = max(0, min(len(data) - 1, math.ceil((pct / 100.0) * len(data)) - 1))
    return data[idx]

def summarize(path):
    rows = list(csv.DictReader(Path(path).open(newline="", encoding="utf-8")))
    required = {"timeStamp", "elapsed", "label", "responseCode", "success", "threadName"}
    missing = required - set(rows[0].keys() if rows else [])
    if missing:
        raise SystemExit(f"{path}: missing required columns {sorted(missing)}")

    groups = defaultdict(list)
    for row in rows:
        groups[row["label"]].append(row)
        groups["TOTAL"].append(row)

    out = {}
    for label, items in groups.items():
        elapsed = [int(float(r["elapsed"])) for r in items]
        failures = sum(r["success"].lower() != "true" for r in items)
        start = min(int(r["timeStamp"]) for r in items)
        end = max(int(r["timeStamp"]) + int(float(r["elapsed"])) for r in items)
        span_s = max((end - start) / 1000.0, 0.001)
        out[label] = {
            "samples": len(items),
            "failures": failures,
            "error_pct": round(100.0 * failures / len(items), 3),
            "avg_ms": round(sum(elapsed) / len(elapsed), 3),
            "min_ms": min(elapsed),
            "max_ms": max(elapsed),
            "p50_nearest_rank_ms": nearest_rank(elapsed, 50),
            "p90_nearest_rank_ms": nearest_rank(elapsed, 90),
            "p95_nearest_rank_ms": nearest_rank(elapsed, 95),
            "p99_nearest_rank_ms": nearest_rank(elapsed, 99),
            "completed_samples_per_second_over_label_span": round(len(items) / span_s, 3),
            "response_codes": sorted({r["responseCode"] for r in items}),
        }
    return out

if len(sys.argv) < 2:
    raise SystemExit("usage: summarize_jtl.py <results1.jtl> [results2.jtl ...]")

for raw in sys.argv[1:]:
    print(json.dumps({"path": raw, "summary": summarize(raw)}, indent=2))
print("NOTE: nearest-rank percentiles are an independent raw-JTL check; JMeter dashboard percentile estimates may differ slightly.")
python tools/summarize_jtl.py   results/p23-baseline/results.jtl   results/p23-degraded/results.jtl

Require exactly two labels, 40 rows each, 80 total. The script uses nearest-rank percentiles as an independent teaching check; small differences from JMeter dashboard estimates are acceptable/documented.

11. Inspect target telemetry

Save tools/analyze_events.py:

import json
import math
import sys
from collections import Counter
from pathlib import Path

path = Path(sys.argv[1])
run_id = sys.argv[2]
events = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
work = [
    e for e in events
    if e.get("operation") in {"catalog", "checkout"} and e.get("run_id") == run_id
]

print(f"run_id={run_id}")
print(f"events={len(work)}")
print(f"operations={dict(Counter(e.get('operation') for e in work))}")
print(f"statuses={dict(Counter(e.get('status') for e in work))}")
print(f"modes={dict(Counter(e.get('mode') for e in work))}")

for op in ("catalog", "checkout"):
    vals = sorted(int(e.get("service_wall_ms", 0)) for e in work if e.get("operation") == op)
    if vals:
        p95 = vals[max(0, min(len(vals)-1, math.ceil(0.95 * len(vals)) - 1))]
        print(
            f"{op}: n={len(vals)} avg_service_wall_ms={sum(vals)/len(vals):.2f} "
            f"p95_nearest_rank_ms={p95} max_ms={max(vals)}"
        )
python tools/analyze_events.py results/server-events.jsonl p23-baseline
python tools/analyze_events.py results/server-events.jsonl p23-degraded

Baseline target should report 40 Catalog/40 Checkout, all 200. Degraded target should still report 40/40 but Checkout includes eight 503s and a much higher service-wall time; Catalog stays near 20 ms.

12. Read the Statistics table

For Catalog and Checkout, record sample count, error %, average, min/max, p90/p95/p99 and throughput. Compare counts/error rate/average to raw JTL.

Interpretation:

  • Catalog stable latency supports “unchanged catalog service behavior.”
  • Checkout higher p95/average localizes the direct latency change to Checkout.
  • Checkout 20% error rate should match eight 503s out of 40 Checkout samples.
  • Overall/label throughput may fall because the closed-loop threads spend longer in Checkout.

13. Read APDEX with the threshold contract visible

Write “APDEX thresholds = 100/300 ms (lab)” beside the score. Baseline labels should score better; degraded Checkout should fall. Do not say “APDEX below X means bad” without the business objective that chose X/T thresholds.

14. Inspect Request Summary / Errors

Baseline should show 100% successful requests. Degraded should show the expected failed proportion. The error table/top errors should map failures to Checkout and HTTP 503/Service Unavailable-type evidence.

If the dashboard shows a different failure count than JTL, stop interpretation and resolve filter/controller/result-field differences first.

15. Inspect selected graphs in a fixed order

  1. Active Threads Over Time: prove similar virtual-user shape.
  2. Response Times Over Time: locate the Checkout shift.
  3. Response Time Percentiles Over Time: examine tail movement by bucket.
  4. Response Codes per Second: locate degraded 503 events.
  5. Hits/Transactions per Second: observe completion-rate change.
  6. Response Time Distribution/Percentiles: compare full-window distributions.

Never conclude “thread count caused latency” solely from Response Time vs Request/Times vs Threads. Use the known fixture change and service-side timings as causal evidence.

16. Sample-label mapping note

JMX label Target operation Expected rows/run Interpretation role
Catalog GET /catalog 40 Control: target behavior intentionally unchanged.
Checkout GET /checkout 40 Treatment: only degraded mode adds latency/errors.

Stable label mapping is part of the workload manifest and must be preserved across runs.

17. Automated raw comparison

Save tools/compare_runs.py:

import csv
import json
import math
import sys
from collections import defaultdict
from pathlib import Path

def percentile(values, pct):
    data = sorted(values)
    if not data:
        return 0
    return data[max(0, min(len(data)-1, math.ceil(len(data) * pct / 100.0) - 1))]

def load(path):
    groups = defaultdict(list)
    rows = list(csv.DictReader(Path(path).open(newline="", encoding="utf-8")))
    for r in rows:
        groups[r["label"]].append(r)
    out = {}
    for label, items in groups.items():
        elapsed = [int(float(r["elapsed"])) for r in items]
        failures = sum(r["success"].lower() != "true" for r in items)
        out[label] = {
            "samples": len(items),
            "failures": failures,
            "error_pct": 100.0 * failures / len(items),
            "avg_ms": sum(elapsed) / len(elapsed),
            "p95_nearest_rank_ms": percentile(elapsed, 95),
        }
    return out

if len(sys.argv) != 3:
    raise SystemExit("usage: compare_runs.py baseline.jtl degraded.jtl")

base = load(sys.argv[1])
deg = load(sys.argv[2])

labels = sorted(set(base) | set(deg))
for label in labels:
    b = base.get(label, {})
    d = deg.get(label, {})
    print(label)
    print("  baseline:", json.dumps(b, sort_keys=True))
    print("  degraded:", json.dumps(d, sort_keys=True))
    if b and d:
        print(
            f"  delta_avg_ms={d['avg_ms']-b['avg_ms']:.2f} "
            f"delta_p95_ms={d['p95_nearest_rank_ms']-b['p95_nearest_rank_ms']:.2f} "
            f"delta_error_pct={d['error_pct']-b['error_pct']:.2f}"
        )
python tools/compare_runs.py   results/p23-baseline/results.jtl   results/p23-degraded/results.jtl

18. Write the interpretation note

A strong note separates observations, causal evidence and uncertainty:

“Under the same configured 2-thread ×20-loop closed workload, both runs produced 80 target/JTL samples. Degraded Checkout shows a material increase in average/tail latency and eight HTTP 503 failures while Catalog service timing remains approximately unchanged. Target telemetry independently records the injected Checkout sleep/error behavior, so the Checkout regression is causally attributable to the controlled fixture change. Overall completion throughput also decreases, which is expected in this closed workload because threads wait longer on Checkout. APDEX decline is specific to the lab's 100/300 ms thresholds. The run is small/local and does not estimate production capacity.”

19. Challenge

The degraded report shows Catalog hits/sec lower even though Catalog p95 and server processing time are unchanged. Is Catalog now the bottleneck?

No. In a closed sequential workload, slower Checkout occupies threads longer, so fewer loop cycles reach Catalog per second. Stable Catalog elapsed/service time plus the known Checkout change argue against Catalog as the bottleneck.

Knowledge check

Why is degraded overall error rate 10% but Checkout error rate 20%?

Why generate one dashboard with -e -o and another with -g?

Why can nearest-rank p95 differ slightly from dashboard p95?

What evidence establishes the known Checkout change causally?

Why must html-report be new/empty?

Next lesson

Make reporting choices explicit

Lesson 3 compares percentile sets, APDEX versus SLOs, transaction versus sampler labels, full-run versus time-window filtering, and HTML dashboards versus long-run time-series telemetry.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. The dashboard generator reads compatible CSV sample logs and can run at the end of a load test with -e -o or later with -g <CSV> -o <dir>. Required CSV data include bytes, label, latency, response code/message, success, thread counts/name, elapsed time, connect time, assertion failure message, and a timestamp format containing time; current defaults are suitable unless changed. The HTML output directory must be empty/new. Dashboard statistics expose three configurable percentile levels via aggregate_rpt_pct1/2/3, defaulting to 90/95/99. The report-generator general APDEX defaults are 500 ms satisfied and 1500 ms tolerated; this chapter's local lab deliberately overrides them to 100/300 ms. sample_filter removes sample data before report calculations. HTML series_filter filters displayed series/rows after calculations. start_date/end_date constrain the report measurement window. The default over-time granularity is 60000 ms and must remain above 1000 ms; the tiny local lab uses 2000 ms. Dashboard percentile estimates can differ from GUI Aggregate Report, especially with few/widely distributed samples, because estimator formulas differ. Several dashboard graphs include Transaction Controller sample results while others explicitly ignore/exclude them, so labels and controller-sample policy must be stated before totals are interpreted.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.