Checkpoint Lab — HTML Dashboard Reports and Performance Result Interpretation
The checkpoint is a controlled interpretation exercise. Two runs use the same JMX, threads, loops, pacing, labels, generator and localhost target process. The only treatment is fixture mode: degraded Checkout adds a deterministic sleep and HTTP 503 pattern. You must predict the dashboard effects before execution, prove workload equivalence, verify every major claim against raw JTL and target events, and state what remains uncertain.
Learning objectives
- Create baseline/degraded dashboards from identical bounded workloads.
- Identify the known Checkout change through statistics, percentiles, errors, APDEX and selected time-series graphs.
- Verify sample/error/latency claims against raw JTL.
- Verify causality against fixture event telemetry and unchanged Catalog control behavior.
- Preserve workload/report configuration and generator-validity evidence.
- Write a short interpretation that separates findings, mechanism and uncertainty.
1. Exact assumptions and hard ceilings
| Item | Checkpoint baseline |
|---|---|
| JMeter | Apache JMeter 5.6.3. |
| Java | Java 17 JDK; JMeter 5.6.3 requires Java 8+. |
| Plugins | None. |
| Target | http://127.0.0.1:8023 only. |
| Fixture |
Python stdlib prompt23-dashboard-fixture-v1.
|
| Workload | 2 threads ×20 loops ×2 samplers = 80 samples/run. |
| Pacing | 50 ms child timer on each sampler. |
| Labels | Catalog and Checkout only. |
| Baseline treatment | Catalog ~20 ms; Checkout ~45 ms; no intentional failures. |
| Degraded treatment | Catalog unchanged; Checkout ~180 ms; seq multiple-of-5 returns HTTP 503. |
| Expected degraded failures | 8 Checkout failures / 40 Checkout = 20%; 8/80 overall = 10%. |
| JTL | Dashboard-compatible lean CSV. |
| APDEX | Lab thresholds 100 ms satisfied / 300 ms tolerated. |
| Percentiles | Dashboard p90/p95/p99; raw checker uses nearest-rank approximation. |
| Time granularity | 2000 ms. |
| Max duration | ≤15 seconds/run. |
2. Setup and target authorization/preflight
- Start one fresh fixture process and event log.
- GET health/stats; record baseline request counters.
- Record JMeter/Java versions and generator CPU/heap/disk/network headroom.
-
Hash/version the JMX and
local.properties; confirm both runs will use the same inputs except-Jrun.id/-Jmode. - Confirm new/empty result/report directories.
3. Predictions before execution
P1 — workload: both runs will produce 80 JTL rows and 80 target work events (40 Catalog + 40 Checkout).
P2 — error: baseline error rate ≈0; degraded Checkout error rate=20% and total error rate=10%.
P3 — latency: Catalog p95/target service time remains roughly unchanged; degraded Checkout average/p90/p95/p99 increases materially.
P4 — APDEX: Checkout APDEX decreases under the fixed 100/300 ms lab thresholds.
P5 — throughput: degraded overall/hit throughput likely decreases because this is a closed sequential workload whose threads wait longer on Checkout.
4. Execute baseline and degraded runs
Baseline at-end report:
& "$env:JMETER_HOME\bin\jmeter.bat" `
-n -t .\plans\dashboard-local.jmx `
-q .\config\local.properties `
-Jrun.id=p23-check-baseline -Jmode=baseline `
-l .\results\p23-check-baseline\results.jtl `
-j .\results\p23-check-baseline\jmeter.log `
-e -o .\results\p23-check-baseline\html-report
Degraded JTL then report:
& "$env:JMETER_HOME\bin\jmeter.bat" `
-n -t .\plans\dashboard-local.jmx `
-q .\config\local.properties `
-Jrun.id=p23-check-degraded -Jmode=degraded `
-l .\results\p23-check-degraded\results.jtl `
-j .\results\p23-check-degraded\jmeter.log
& "$env:JMETER_HOME\bin\jmeter.bat" `
-q .\config\local.properties `
-g .\results\p23-check-degraded\results.jtl `
-o .\results\p23-check-degraded\html-report
5. Raw-JTL verification gate
python tools/summarize_jtl.py results/p23-check-baseline/results.jtl results/p23-check-degraded/results.jtl
python tools/compare_runs.py results/p23-check-baseline/results.jtl results/p23-check-degraded/results.jtl
Do not open with a conclusion until row counts, labels and degraded error counts match P1/P2.
6. Independent target verification gate
python tools/analyze_events.py results/server-events.jsonl p23-check-baseline
python tools/analyze_events.py results/server-events.jsonl p23-check-degraded
Require:
- 80 events each;
- 40 Catalog +40 Checkout each;
- baseline all 200;
- degraded exactly eight 503 Checkout events;
- Catalog service-wall distribution similar across modes;
- Checkout service-wall distribution near the deliberately injected mode shift.
7. Dashboard verification worksheet
| Dashboard view | Baseline expectation | Degraded expectation | Raw/target cross-check |
|---|---|---|---|
| Statistics: sample count | Catalog40 / Checkout40 | same | JTL count + target events. |
| Statistics: Error % | 0/0 | Checkout20%, total10% | JTL success/code + target statuses. |
| p90/p95/p99 | Checkout low | Checkout much higher; Catalog stable | raw nearest-rank + target service time. |
| APDEX | both high relative to lab thresholds | Checkout lower | threshold properties + elapsed/success rows. |
| Request Summary / Errors | all success | eight checkout failures | JTL responseCode + target 503. |
| Response Times Over Time | stable low labels | Checkout series shifted upward | target service-wall timeline. |
| Response Codes/sec | 200 only | 503 points appear | target statuses/timestamps. |
| Hits/Transactions/sec | baseline completion rate | often lower | closed-loop sequence + run spans. |
8. Sample-label mapping
Attach this to the evidence packet:
Catalog -> GET /catalog -> control operation, unchanged target behavior
Checkout -> GET /checkout -> treatment operation, mode-controlled latency/errors
No Transaction Controller parent sample in checkpoint
This removes ambiguity when reviewing table/graph labels later.
9. Workload/report manifest
Record:
- JMX/property SHA-256 and JMeter/Java versions;
- threads=2, loops=20, two samplers/loop, 50 ms pacing;
- run IDs/modes;
- labels/control/treatment mapping;
- save-service required fields and CSV format;
- APDEX=100/300 ms, p90/p95/p99, granularity=2000 ms;
- no sample_filter/start_date/end_date for the main checkpoint reports;
- generator CPU/heap/network/disk headroom;
- result/report directories and target event log.
10. Write the interpretation with uncertainty
Observation: Checkout latency distribution and error rate worsen in degraded mode; Catalog latency remains stable. Overall completion throughput falls.
Controlled mechanism: The fixture intentionally changes only Checkout from ~45 ms/success to ~180 ms with deterministic HTTP 503s every fifth sequence. Target telemetry confirms that exact operation-level change.
Interpretation: The Checkout regression is attributable to the controlled target change. The overall throughput reduction is consistent with a closed workload spending longer blocked on Checkout; it does not prove Catalog capacity changed.
Uncertainty/limits: This is a tiny loopback experiment with 40 samples/label; extreme-tail percentiles are low-resolution, network variability is minimal, and the result says nothing about production capacity. APDEX conclusions apply only to the lab's 100/300 ms thresholds.
11. Check one reporting failure without rerunning load
Copy the baseline JTL and remove one required field as in Lesson 4.
Attempt -g into a new directory, preserve the failure,
then regenerate from the untouched JTL. This proves report failures
can be diagnosed at the artifact/configuration layer without sending
more traffic.
12. Required evidence packet
| Artifact | Required content |
|---|---|
| Raw JTL | baseline/degraded CSVs with expected labels/counts. |
| Dashboards | two separate new/empty report directories. |
| Statistics/percentile table | Catalog/Checkout counts/errors/avg/p90/p95/p99. |
| APDEX/error table | thresholds published; degraded errors mapped to Checkout. |
| Selected graphs | threads, response time, response codes, throughput/distribution. |
| Sample-label map | Catalog control / Checkout treatment; no parent transaction sample. |
| Workload manifest | versions/hashes/load/report properties/filter/window state. |
| Target telemetry | 80 events/run, expected statuses and service times. |
| Generator validity | CPU/heap/network/disk headroom and jmeter.log. |
| Interpretation note | observations, causal mechanism, uncertainty/limitations. |
13. Final validity statement
prompt23-dashboard-fixture-v1. Both produced 80 JTL
rows and 80 target events. The only treatment was
mode=degraded, which target telemetry confirmed added
Checkout service delay and eight deterministic HTTP 503s while
Catalog remained unchanged. Dashboards used compatible lean CSV
fields, p90/p95/p99, APDEX 100/300 ms and 2-second over-time
granularity with no source/window filter. Dashboard
counts/errors/latency trends were cross-checked against raw JTL and
target events; generator headroom and jmeter.log were preserved.
This establishes the local controlled-regression interpretation
method, not production capacity or a universal APDEX/SLO threshold.”
14. Verification checklist
- Only 127.0.0.1:8023.
- JMeter 5.6.3 / Java 17 / no plugins.
- Same JMX/hash/workload/report properties except run ID/mode.
- 80 JTL +80 target events each.
- Degraded exactly eight Checkout 503s; Catalog no intentional failures.
- Dashboard statistics/error table match raw JTL counts.
- Checkout latency/APDEX move in expected direction; Catalog control remains stable.
- Throughput interpretation states the closed-loop dependency.
- Generator headroom/filter/window/controller policy documented.
- Interpretation states uncertainty/sample-size/local-only limitations.
15. Cleanup / rollback
- Stop JMeter processes and localhost fixture after target evidence is saved.
- Keep raw JTL, jmeter.log, dashboard dirs, report/workload properties, manifests and event evidence until review.
- Delete only disposable local artifacts after the defined retention period; do not delete first-failure evidence to rerun.
- No public/production target, real credential/PII, recorder CA, remote RMI, container, paid telemetry/CI, database/message service or OS/JVM global tuning was changed.
16. What Chapter 23 adds to the operating model
The production performance-testing operating model now has a dashboard interpretation contract: report source JTL/schema, labels/controller policy, percentile set, APDEX thresholds, sample/series filters, measurement window, time granularity, workload manifest, configured-versus-achieved count/rate, generator validity, target telemetry, raw-to-dashboard cross-check and uncertainty statement must accompany any performance conclusion.
Chapter 24 moves to Backend Listener, InfluxDB/Graphite, Grafana, and Real-Time Telemetry. It extends this interpretation discipline into live time-series pipelines and explicitly measures the network/backend cost that a dashboard generated after the run does not impose during sampling.
Knowledge check
What is the strongest evidence that Checkout caused the controlled regression?
Only Checkout target behavior was intentionally changed, and target events independently confirm the increased service time/503s while Catalog remains the control.
Why can degraded Catalog hits/sec fall without a Catalog latency regression?
Closed threads spend longer in Checkout, so fewer iterations reach Catalog per second even though Catalog processing remains stable.
Why must APDEX thresholds be included in the evidence packet?
The score depends on those thresholds; changing them can change the conclusion without changing JTL.
What invalidates the comparison even if the dashboards look similar?
Different sample counts/filters/windows/labels, generator saturation, target-event mismatch, or changed JMX/report inputs.
What is Chapter 24's bridge?
Move from post-run static dashboard evidence to real-time Backend Listener/time-series telemetry while preserving workload/label/validity contracts and measuring telemetry overhead.
Official references and version notes
-
JMeter User Manual — Generating Dashboard Report
— required CSV fields, APDEX, statistics/percentiles, errors,
graphs, filters, date windows, granularity,
-g,-e,-o, and output-folder rules. - JMeter Getting Started — CLI/load analysis — GUI authoring/debugging, CLI load execution, CSV/XML results and HTML analysis.
- JMeter Listeners / Result files — sample-log fields and raw-result interpretation.
- Component Reference — Transaction Controller — parent/additional transaction samples and reporting semantics.
- JMeter Properties Reference — report/save-service configuration and defaults.
- Apache JMeter downloads — current stable release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a Java 17 JDK;
JMeter 5.6.3 requires Java 8+. The dashboard generator reads
compatible CSV sample logs and can run at the end of a load test
with -e -o or later with
-g <CSV> -o <dir>. Required CSV data
include bytes, label, latency, response code/message, success,
thread counts/name, elapsed time, connect time, assertion failure
message, and a timestamp format containing time; current defaults
are suitable unless changed. The HTML output directory must be
empty/new. Dashboard statistics expose three configurable
percentile levels via aggregate_rpt_pct1/2/3,
defaulting to 90/95/99. The report-generator general APDEX
defaults are 500 ms satisfied and 1500 ms tolerated; this
chapter's local lab deliberately overrides them to 100/300 ms.
sample_filter removes sample data before report
calculations. HTML series_filter filters displayed
series/rows after calculations. start_date/end_date
constrain the report measurement window. The default over-time
granularity is 60000 ms and must remain above 1000 ms; the tiny
local lab uses 2000 ms. Dashboard percentile estimates can differ
from GUI Aggregate Report, especially with few/widely distributed
samples, because estimator formulas differ. Several dashboard
graphs include Transaction Controller sample results while others
explicitly ignore/exclude them, so labels and controller-sample
policy must be stated before totals are interpreted.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.