HTML Dashboard Reports and Performance Result Interpretation: Guided Hands-On Workflow
The workflow uses two identical JMeter workloads. Only the target's
Checkout behavior changes: baseline checkout sleeps
about 45 ms and succeeds; degraded checkout sleeps about 180 ms and
returns a deterministic HTTP 503 every fifth sequence.
Catalog stays at ~20 ms. That design gives the
dashboard a known signal and gives target telemetry an independent
causal reference.
Learning objectives
- Run a bounded baseline and degraded local experiment from one JMX.
- Generate a dashboard at end of a run and later from retained JTL.
- Inspect dashboard statistics, percentiles, errors, APDEX, throughput and time-series graphs.
- Map dashboard labels back to JMX sampler labels.
- Compare dashboard claims with raw JTL calculations and target events.
- Write an interpretation note that separates observations from causal evidence.
1. Safety envelope
http://127.0.0.1:8023. 2 threads
×20 loops ×2 HTTP samplers = 80 samples/run, 50 ms pacing, ≤15
seconds/run, two runs only for the mandatory workflow, synthetic 503
failures only in degraded Checkout, no
retries/credentials/plugins/remote engines. Abort on non-loopback
target, >80 samples/run, unexpected labels, generator saturation,
more than the expected 8 degraded Checkout failures, or target/JTL
count disagreement.
2. Create the deterministic local target
Save fixtures/dashboard_fixture.py:
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.parse import urlparse, parse_qs
import argparse
import json
import re
import threading
import time
FIXTURE_VERSION = "prompt23-dashboard-fixture-v1"
SAFE_TOKEN = re.compile(r"^[A-Za-z0-9_.-]{1,64}$")
lock = threading.Lock()
event_log = None
metrics = {
"requests": 0,
"errors": 0,
"catalog": 0,
"checkout": 0,
"by_run": {},
"by_mode": {},
}
def now_ms():
return int(time.time() * 1000)
def log_event(event):
if event_log is None:
return
with lock:
with event_log.open("a", encoding="utf-8") as handle:
handle.write(json.dumps(event, sort_keys=True) + "\n")
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def send_json(self, status, payload):
raw = json.dumps(payload, sort_keys=True).encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.send_header("X-Fixture-Version", FIXTURE_VERSION)
self.end_headers()
self.wfile.write(raw)
def record(self, started, operation, status, **extra):
with lock:
metrics["requests"] += 1
if status >= 400:
metrics["errors"] += 1
if operation == "catalog" and status == 200:
metrics["catalog"] += 1
if operation == "checkout":
metrics["checkout"] += 1
run_id = extra.get("run_id", "")
mode = extra.get("mode", "")
if run_id:
metrics["by_run"][run_id] = metrics["by_run"].get(run_id, 0) + 1
if mode:
metrics["by_mode"][mode] = metrics["by_mode"].get(mode, 0) + 1
event = {
"ts_ms": now_ms(),
"operation": operation,
"status": status,
"service_wall_ms": now_ms() - started,
}
event.update(extra)
log_event(event)
def parse_common(self, parsed):
q = parse_qs(parsed.query)
run_id = q.get("run_id", [""])[0]
mode = q.get("mode", [""])[0]
thread_id = q.get("thread", [""])[0]
seq_raw = q.get("seq", [""])[0]
if not SAFE_TOKEN.fullmatch(run_id) or not SAFE_TOKEN.fullmatch(thread_id):
return None, (400, {"status": "invalid_metadata"})
if mode not in {"baseline", "degraded"}:
return None, (400, {"status": "invalid_mode"})
try:
seq = int(seq_raw)
except ValueError:
seq = -1
if not 1 <= seq <= 1000:
return None, (400, {"status": "invalid_seq"})
return {
"run_id": run_id,
"mode": mode,
"thread": thread_id,
"seq": seq,
}, None
def do_GET(self):
started = now_ms()
parsed = urlparse(self.path)
if parsed.path == "/health":
self.send_json(200, {"status": "ok", "fixture_version": FIXTURE_VERSION})
self.record(started, "health", 200)
return
if parsed.path == "/stats":
with lock:
snapshot = json.loads(json.dumps(metrics))
self.send_json(200, {"fixture_version": FIXTURE_VERSION, "metrics": snapshot})
self.record(started, "stats", 200)
return
if parsed.path not in {"/catalog", "/checkout"}:
self.send_json(404, {"status": "not_found"})
self.record(started, "unknown", 404)
return
common, error = self.parse_common(parsed)
if error:
status, payload = error
self.send_json(status, payload)
self.record(started, parsed.path.lstrip("/"), status)
return
if parsed.path == "/catalog":
# Intentionally unchanged between modes.
time.sleep(0.020)
self.send_json(200, {
"status": "ok",
"operation": "catalog",
**common,
})
self.record(started, "catalog", 200, **common)
return
# Checkout is the only deliberately changed target behavior.
if common["mode"] == "baseline":
time.sleep(0.045)
status = 200
payload = {"status": "ok", "operation": "checkout", **common}
else:
time.sleep(0.180)
if common["seq"] % 5 == 0:
status = 503
payload = {
"status": "synthetic_failure",
"error": "synthetic_checkout_failure",
"operation": "checkout",
**common,
}
else:
status = 200
payload = {"status": "ok", "operation": "checkout", **common}
self.send_json(status, payload)
self.record(started, "checkout", status, **common)
def log_message(self, format, *args):
return
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--host", default="127.0.0.1")
parser.add_argument("--port", type=int, default=8023)
parser.add_argument("--log", default="results/server-events.jsonl")
args = parser.parse_args()
global event_log
event_log = Path(args.log).resolve()
event_log.parent.mkdir(parents=True, exist_ok=True)
event_log.write_text("", encoding="utf-8")
print(f"fixture_version={FIXTURE_VERSION}")
print(f"listen=http://{args.host}:{args.port}")
print(f"event_log={event_log}")
ThreadingHTTPServer((args.host, args.port), Handler).serve_forever()
if __name__ == "__main__":
main()
Start:
python .\fixtures\dashboard_fixture.py `
--host 127.0.0.1 `
--port 8023 `
--log .\results\server-events.jsonl
The mode is supplied by each request. This lets one fixture process
record both runs with explicit run_id/mode
metadata.
3. Target/generator preflight
curl --fail --silent http://127.0.0.1:8023/health
curl --fail --silent http://127.0.0.1:8023/stats
Record JMeter 5.6.3, Java 17, free disk, generator CPU/heap baseline and target event count. No load runs in GUI.
4. Create workload/report properties
config/local.properties:
# Prompt 23 local workload + dashboard input contract
target.host=127.0.0.1
target.port=8023
threads=2
loops=20
pacing.ms=50
connect.timeout.ms=500
response.timeout.ms=2000
# Dashboard-compatible lean CSV fields.
jmeter.save.saveservice.output_format=csv
jmeter.save.saveservice.print_field_names=true
jmeter.save.saveservice.bytes=true
jmeter.save.saveservice.sent_bytes=true
jmeter.save.saveservice.label=true
jmeter.save.saveservice.latency=true
jmeter.save.saveservice.response_code=true
jmeter.save.saveservice.response_message=true
jmeter.save.saveservice.successful=true
jmeter.save.saveservice.thread_counts=true
jmeter.save.saveservice.thread_name=true
jmeter.save.saveservice.time=true
jmeter.save.saveservice.connect_time=true
jmeter.save.saveservice.assertion_results_failure_message=true
jmeter.save.saveservice.timestamp_format=ms
# Keep privacy-heavy content out of the JTL.
jmeter.save.saveservice.response_data=false
jmeter.save.saveservice.response_data.on_error=false
jmeter.save.saveservice.samplerData=false
jmeter.save.saveservice.responseHeaders=false
jmeter.save.saveservice.requestHeaders=false
# Lab report settings. These intentionally differ from JMeter defaults
# so the known 180ms checkout degradation is visible in APDEX.
jmeter.reportgenerator.report_title=Prompt 23 Local Dashboard Lab
jmeter.reportgenerator.overall_granularity=2000
jmeter.reportgenerator.apdex_satisfied_threshold=100
jmeter.reportgenerator.apdex_tolerated_threshold=300
aggregate_rpt_pct1=90
aggregate_rpt_pct2=95
aggregate_rpt_pct3=99
Two details are intentional:
overall_granularity=2000 because the runs last only a
few seconds, and APDEX 100/300 ms because the known 180 ms degraded
Checkout should move from “satisfied” toward “tolerated/frustrated.”
These are lab thresholds, not production objectives.
5. Build one JMX in GUI
Test Plan
├── HTTP Request Defaults
│ host=${__P(target.host,127.0.0.1)}
│ port=${__P(target.port,8023)}
│ timeouts from properties
└── Thread Group
threads=${__P(threads,1)}
loops=${__P(loops,1)}
├── Counter -> SEQ (start=1, per-user=true)
├── HTTP Request — Catalog
│ GET /catalog
│ run_id=${__P(run.id,p23-local)}
│ mode=${__P(mode,baseline)}
│ thread=T${__threadNum}
│ seq=${SEQ}
│ └── Constant Timer ${__P(pacing.ms,50)} ms
└── HTTP Request — Checkout
GET /checkout
run_id=${__P(run.id,p23-local)}
mode=${__P(mode,baseline)}
thread=T${__threadNum}
seq=${SEQ}
└── Constant Timer ${__P(pacing.ms,50)} ms
The Counter value is shared by the two samplers within an iteration, so degraded Checkout fails at sequence 5/10/15/20 in each thread: 8 checkout failures total. Catalog remains successful.
6. Predict before execution
- Both runs: 40 Catalog + 40 Checkout = 80 JTL/target operations.
- Baseline: zero expected HTTP failures.
- Degraded: 8 Checkout HTTP 503 failures = 20% Checkout error rate and 10% overall error rate.
- Catalog elapsed/service time should remain near baseline.
- Checkout p90/p95/p99 and average should increase materially.
- Checkout APDEX should decline under 100/300 ms lab thresholds.
- Closed-loop throughput may decline in degraded mode because threads wait longer in Checkout.
7. Baseline — generate dashboard at end
PowerShell:
New-Item -ItemType Directory -Force .\results\p23-baseline | Out-Null
& "$env:JMETER_HOME\bin\jmeter.bat" `
-n `
-t .\plans\dashboard-local.jmx `
-q .\config\local.properties `
-Jrun.id=p23-baseline `
-Jmode=baseline `
-l .\results\p23-baseline\results.jtl `
-j .\results\p23-baseline\jmeter.log `
-e `
-o .\results\p23-baseline\html-report
Do not create html-report beforehand. Preserve raw JTL
even though the dashboard is generated immediately.
8. Degraded — retain JTL first
New-Item -ItemType Directory -Force .\results\p23-degraded | Out-Null
& "$env:JMETER_HOME\bin\jmeter.bat" `
-n `
-t .\plans\dashboard-local.jmx `
-q .\config\local.properties `
-Jrun.id=p23-degraded `
-Jmode=degraded `
-l .\results\p23-degraded\results.jtl `
-j .\results\p23-degraded\jmeter.log
Expected engine completion does not make the run “green”: the JTL intentionally contains eight HTTP failures.
9. Generate degraded dashboard later from JTL
& "$env:JMETER_HOME\bin\jmeter.bat" `
-q .\config\local.properties `
-g .\results\p23-degraded\results.jtl `
-o .\results\p23-degraded\html-report
No target traffic occurs during -g. The command reads
the existing CSV plus report properties and writes a new/empty
report directory.
10. Inspect raw JTL before opening dashboards
Save tools/summarize_jtl.py:
import csv
import json
import math
import sys
from collections import defaultdict
from pathlib import Path
def nearest_rank(values, pct):
data = sorted(values)
if not data:
return 0
idx = max(0, min(len(data) - 1, math.ceil((pct / 100.0) * len(data)) - 1))
return data[idx]
def summarize(path):
rows = list(csv.DictReader(Path(path).open(newline="", encoding="utf-8")))
required = {"timeStamp", "elapsed", "label", "responseCode", "success", "threadName"}
missing = required - set(rows[0].keys() if rows else [])
if missing:
raise SystemExit(f"{path}: missing required columns {sorted(missing)}")
groups = defaultdict(list)
for row in rows:
groups[row["label"]].append(row)
groups["TOTAL"].append(row)
out = {}
for label, items in groups.items():
elapsed = [int(float(r["elapsed"])) for r in items]
failures = sum(r["success"].lower() != "true" for r in items)
start = min(int(r["timeStamp"]) for r in items)
end = max(int(r["timeStamp"]) + int(float(r["elapsed"])) for r in items)
span_s = max((end - start) / 1000.0, 0.001)
out[label] = {
"samples": len(items),
"failures": failures,
"error_pct": round(100.0 * failures / len(items), 3),
"avg_ms": round(sum(elapsed) / len(elapsed), 3),
"min_ms": min(elapsed),
"max_ms": max(elapsed),
"p50_nearest_rank_ms": nearest_rank(elapsed, 50),
"p90_nearest_rank_ms": nearest_rank(elapsed, 90),
"p95_nearest_rank_ms": nearest_rank(elapsed, 95),
"p99_nearest_rank_ms": nearest_rank(elapsed, 99),
"completed_samples_per_second_over_label_span": round(len(items) / span_s, 3),
"response_codes": sorted({r["responseCode"] for r in items}),
}
return out
if len(sys.argv) < 2:
raise SystemExit("usage: summarize_jtl.py <results1.jtl> [results2.jtl ...]")
for raw in sys.argv[1:]:
print(json.dumps({"path": raw, "summary": summarize(raw)}, indent=2))
print("NOTE: nearest-rank percentiles are an independent raw-JTL check; JMeter dashboard percentile estimates may differ slightly.")
python tools/summarize_jtl.py results/p23-baseline/results.jtl results/p23-degraded/results.jtl
Require exactly two labels, 40 rows each, 80 total. The script uses nearest-rank percentiles as an independent teaching check; small differences from JMeter dashboard estimates are acceptable/documented.
11. Inspect target telemetry
Save tools/analyze_events.py:
import json
import math
import sys
from collections import Counter
from pathlib import Path
path = Path(sys.argv[1])
run_id = sys.argv[2]
events = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
work = [
e for e in events
if e.get("operation") in {"catalog", "checkout"} and e.get("run_id") == run_id
]
print(f"run_id={run_id}")
print(f"events={len(work)}")
print(f"operations={dict(Counter(e.get('operation') for e in work))}")
print(f"statuses={dict(Counter(e.get('status') for e in work))}")
print(f"modes={dict(Counter(e.get('mode') for e in work))}")
for op in ("catalog", "checkout"):
vals = sorted(int(e.get("service_wall_ms", 0)) for e in work if e.get("operation") == op)
if vals:
p95 = vals[max(0, min(len(vals)-1, math.ceil(0.95 * len(vals)) - 1))]
print(
f"{op}: n={len(vals)} avg_service_wall_ms={sum(vals)/len(vals):.2f} "
f"p95_nearest_rank_ms={p95} max_ms={max(vals)}"
)
python tools/analyze_events.py results/server-events.jsonl p23-baseline
python tools/analyze_events.py results/server-events.jsonl p23-degraded
Baseline target should report 40 Catalog/40 Checkout, all 200. Degraded target should still report 40/40 but Checkout includes eight 503s and a much higher service-wall time; Catalog stays near 20 ms.
12. Read the Statistics table
For Catalog and Checkout, record sample
count, error %, average, min/max, p90/p95/p99 and throughput.
Compare counts/error rate/average to raw JTL.
Interpretation:
- Catalog stable latency supports “unchanged catalog service behavior.”
- Checkout higher p95/average localizes the direct latency change to Checkout.
- Checkout 20% error rate should match eight 503s out of 40 Checkout samples.
- Overall/label throughput may fall because the closed-loop threads spend longer in Checkout.
13. Read APDEX with the threshold contract visible
Write “APDEX thresholds = 100/300 ms (lab)” beside the score. Baseline labels should score better; degraded Checkout should fall. Do not say “APDEX below X means bad” without the business objective that chose X/T thresholds.
14. Inspect Request Summary / Errors
Baseline should show 100% successful requests. Degraded should show
the expected failed proportion. The error table/top errors should
map failures to Checkout and HTTP 503/Service
Unavailable-type evidence.
If the dashboard shows a different failure count than JTL, stop interpretation and resolve filter/controller/result-field differences first.
15. Inspect selected graphs in a fixed order
- Active Threads Over Time: prove similar virtual-user shape.
- Response Times Over Time: locate the Checkout shift.
- Response Time Percentiles Over Time: examine tail movement by bucket.
- Response Codes per Second: locate degraded 503 events.
- Hits/Transactions per Second: observe completion-rate change.
- Response Time Distribution/Percentiles: compare full-window distributions.
Never conclude “thread count caused latency” solely from Response Time vs Request/Times vs Threads. Use the known fixture change and service-side timings as causal evidence.
16. Sample-label mapping note
| JMX label | Target operation | Expected rows/run | Interpretation role |
|---|---|---|---|
| Catalog | GET /catalog | 40 | Control: target behavior intentionally unchanged. |
| Checkout | GET /checkout | 40 | Treatment: only degraded mode adds latency/errors. |
Stable label mapping is part of the workload manifest and must be preserved across runs.
17. Automated raw comparison
Save tools/compare_runs.py:
import csv
import json
import math
import sys
from collections import defaultdict
from pathlib import Path
def percentile(values, pct):
data = sorted(values)
if not data:
return 0
return data[max(0, min(len(data)-1, math.ceil(len(data) * pct / 100.0) - 1))]
def load(path):
groups = defaultdict(list)
rows = list(csv.DictReader(Path(path).open(newline="", encoding="utf-8")))
for r in rows:
groups[r["label"]].append(r)
out = {}
for label, items in groups.items():
elapsed = [int(float(r["elapsed"])) for r in items]
failures = sum(r["success"].lower() != "true" for r in items)
out[label] = {
"samples": len(items),
"failures": failures,
"error_pct": 100.0 * failures / len(items),
"avg_ms": sum(elapsed) / len(elapsed),
"p95_nearest_rank_ms": percentile(elapsed, 95),
}
return out
if len(sys.argv) != 3:
raise SystemExit("usage: compare_runs.py baseline.jtl degraded.jtl")
base = load(sys.argv[1])
deg = load(sys.argv[2])
labels = sorted(set(base) | set(deg))
for label in labels:
b = base.get(label, {})
d = deg.get(label, {})
print(label)
print(" baseline:", json.dumps(b, sort_keys=True))
print(" degraded:", json.dumps(d, sort_keys=True))
if b and d:
print(
f" delta_avg_ms={d['avg_ms']-b['avg_ms']:.2f} "
f"delta_p95_ms={d['p95_nearest_rank_ms']-b['p95_nearest_rank_ms']:.2f} "
f"delta_error_pct={d['error_pct']-b['error_pct']:.2f}"
)
python tools/compare_runs.py results/p23-baseline/results.jtl results/p23-degraded/results.jtl
18. Write the interpretation note
A strong note separates observations, causal evidence and uncertainty:
“Under the same configured 2-thread ×20-loop closed workload, both runs produced 80 target/JTL samples. Degraded Checkout shows a material increase in average/tail latency and eight HTTP 503 failures while Catalog service timing remains approximately unchanged. Target telemetry independently records the injected Checkout sleep/error behavior, so the Checkout regression is causally attributable to the controlled fixture change. Overall completion throughput also decreases, which is expected in this closed workload because threads wait longer on Checkout. APDEX decline is specific to the lab's 100/300 ms thresholds. The run is small/local and does not estimate production capacity.”
19. Challenge
The degraded report shows Catalog hits/sec lower even though Catalog p95 and server processing time are unchanged. Is Catalog now the bottleneck?
No. In a closed sequential workload, slower Checkout occupies threads longer, so fewer loop cycles reach Catalog per second. Stable Catalog elapsed/service time plus the known Checkout change argue against Catalog as the bottleneck.
Knowledge check
Why is degraded overall error rate 10% but Checkout error rate 20%?
There are 8 failed Checkout samples out of 40 Checkout samples, but 80 total samples including 40 successful Catalog samples.
Why generate one dashboard with -e -o and another with -g?
To prove both supported workflows and show that later report generation reads retained CSV without creating target traffic.
Why can nearest-rank p95 differ slightly from dashboard p95?
JMeter dashboard uses its configured percentile estimator/window; estimator formulas differ, especially for small/spread samples.
What evidence establishes the known Checkout change causally?
The fixture configuration and target event service_wall_ms/errors independently record the deliberate Checkout sleep/503 behavior while Catalog is unchanged.
Why must html-report be new/empty?
The JMeter report generator requires an empty/new output directory.
Official references and version notes
-
JMeter User Manual — Generating Dashboard Report
— required CSV fields, APDEX, statistics/percentiles, errors,
graphs, filters, date windows, granularity,
-g,-e,-o, and output-folder rules. - JMeter Getting Started — CLI/load analysis — GUI authoring/debugging, CLI load execution, CSV/XML results and HTML analysis.
- JMeter Listeners / Result files — sample-log fields and raw-result interpretation.
- Component Reference — Transaction Controller — parent/additional transaction samples and reporting semantics.
- JMeter Properties Reference — report/save-service configuration and defaults.
- Apache JMeter downloads — current stable release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a Java 17 JDK;
JMeter 5.6.3 requires Java 8+. The dashboard generator reads
compatible CSV sample logs and can run at the end of a load test
with -e -o or later with
-g <CSV> -o <dir>. Required CSV data
include bytes, label, latency, response code/message, success,
thread counts/name, elapsed time, connect time, assertion failure
message, and a timestamp format containing time; current defaults
are suitable unless changed. The HTML output directory must be
empty/new. Dashboard statistics expose three configurable
percentile levels via aggregate_rpt_pct1/2/3,
defaulting to 90/95/99. The report-generator general APDEX
defaults are 500 ms satisfied and 1500 ms tolerated; this
chapter's local lab deliberately overrides them to 100/300 ms.
sample_filter removes sample data before report
calculations. HTML series_filter filters displayed
series/rows after calculations. start_date/end_date
constrain the report measurement window. The default over-time
granularity is 60000 ms and must remain above 1000 ms; the tiny
local lab uses 2000 ms. Dashboard percentile estimates can differ
from GUI Aggregate Report, especially with few/widely distributed
samples, because estimator formulas differ. Several dashboard
graphs include Transaction Controller sample results while others
explicitly ignore/exclude them, so labels and controller-sample
policy must be stated before totals are interpreted.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.