Chapter 23Lesson 05~270 minutes

Checkpoint Lab — HTML Dashboard Reports and Performance Result Interpretation

The checkpoint is a controlled interpretation exercise. Two runs use the same JMX, threads, loops, pacing, labels, generator and localhost target process. The only treatment is fixture mode: degraded Checkout adds a deterministic sleep and HTTP 503 pattern. You must predict the dashboard effects before execution, prove workload equivalence, verify every major claim against raw JTL and target events, and state what remains uncertain.

CheckpointControlled changeDashboard/JTL cross-checkCausal evidenceUncertainty

Learning objectives

  • Create baseline/degraded dashboards from identical bounded workloads.
  • Identify the known Checkout change through statistics, percentiles, errors, APDEX and selected time-series graphs.
  • Verify sample/error/latency claims against raw JTL.
  • Verify causality against fixture event telemetry and unchanged Catalog control behavior.
  • Preserve workload/report configuration and generator-validity evidence.
  • Write a short interpretation that separates findings, mechanism and uncertainty.

1. Exact assumptions and hard ceilings

Item Checkpoint baseline
JMeter Apache JMeter 5.6.3.
Java Java 17 JDK; JMeter 5.6.3 requires Java 8+.
Plugins None.
Target http://127.0.0.1:8023 only.
Fixture Python stdlib prompt23-dashboard-fixture-v1.
Workload 2 threads ×20 loops ×2 samplers = 80 samples/run.
Pacing 50 ms child timer on each sampler.
Labels Catalog and Checkout only.
Baseline treatment Catalog ~20 ms; Checkout ~45 ms; no intentional failures.
Degraded treatment Catalog unchanged; Checkout ~180 ms; seq multiple-of-5 returns HTTP 503.
Expected degraded failures 8 Checkout failures / 40 Checkout = 20%; 8/80 overall = 10%.
JTL Dashboard-compatible lean CSV.
APDEX Lab thresholds 100 ms satisfied / 300 ms tolerated.
Percentiles Dashboard p90/p95/p99; raw checker uses nearest-rank approximation.
Time granularity 2000 ms.
Max duration ≤15 seconds/run.
Abort: non-loopback target, >80 samples/run, unexpected label/filter, >8 intentional Checkout failures, Catalog failures, generator saturation, JTL/target count mismatch, missing dashboard-required CSV fields, report path collision, or any real credential/PII.

2. Setup and target authorization/preflight

  1. Start one fresh fixture process and event log.
  2. GET health/stats; record baseline request counters.
  3. Record JMeter/Java versions and generator CPU/heap/disk/network headroom.
  4. Hash/version the JMX and local.properties; confirm both runs will use the same inputs except -Jrun.id/-Jmode.
  5. Confirm new/empty result/report directories.

3. Predictions before execution

P1 — workload: both runs will produce 80 JTL rows and 80 target work events (40 Catalog + 40 Checkout).

P2 — error: baseline error rate ≈0; degraded Checkout error rate=20% and total error rate=10%.

P3 — latency: Catalog p95/target service time remains roughly unchanged; degraded Checkout average/p90/p95/p99 increases materially.

P4 — APDEX: Checkout APDEX decreases under the fixed 100/300 ms lab thresholds.

P5 — throughput: degraded overall/hit throughput likely decreases because this is a closed sequential workload whose threads wait longer on Checkout.

4. Execute baseline and degraded runs

Baseline at-end report:

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -n -t .\plans\dashboard-local.jmx `
  -q .\config\local.properties `
  -Jrun.id=p23-check-baseline -Jmode=baseline `
  -l .\results\p23-check-baseline\results.jtl `
  -j .\results\p23-check-baseline\jmeter.log `
  -e -o .\results\p23-check-baseline\html-report

Degraded JTL then report:

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -n -t .\plans\dashboard-local.jmx `
  -q .\config\local.properties `
  -Jrun.id=p23-check-degraded -Jmode=degraded `
  -l .\results\p23-check-degraded\results.jtl `
  -j .\results\p23-check-degraded\jmeter.log

& "$env:JMETER_HOME\bin\jmeter.bat" `
  -q .\config\local.properties `
  -g .\results\p23-check-degraded\results.jtl `
  -o .\results\p23-check-degraded\html-report

5. Raw-JTL verification gate

python tools/summarize_jtl.py   results/p23-check-baseline/results.jtl   results/p23-check-degraded/results.jtl

python tools/compare_runs.py   results/p23-check-baseline/results.jtl   results/p23-check-degraded/results.jtl

Do not open with a conclusion until row counts, labels and degraded error counts match P1/P2.

6. Independent target verification gate

python tools/analyze_events.py results/server-events.jsonl p23-check-baseline
python tools/analyze_events.py results/server-events.jsonl p23-check-degraded

Require:

  • 80 events each;
  • 40 Catalog +40 Checkout each;
  • baseline all 200;
  • degraded exactly eight 503 Checkout events;
  • Catalog service-wall distribution similar across modes;
  • Checkout service-wall distribution near the deliberately injected mode shift.

7. Dashboard verification worksheet

Dashboard view Baseline expectation Degraded expectation Raw/target cross-check
Statistics: sample count Catalog40 / Checkout40 same JTL count + target events.
Statistics: Error % 0/0 Checkout20%, total10% JTL success/code + target statuses.
p90/p95/p99 Checkout low Checkout much higher; Catalog stable raw nearest-rank + target service time.
APDEX both high relative to lab thresholds Checkout lower threshold properties + elapsed/success rows.
Request Summary / Errors all success eight checkout failures JTL responseCode + target 503.
Response Times Over Time stable low labels Checkout series shifted upward target service-wall timeline.
Response Codes/sec 200 only 503 points appear target statuses/timestamps.
Hits/Transactions/sec baseline completion rate often lower closed-loop sequence + run spans.

8. Sample-label mapping

Attach this to the evidence packet:

Catalog  -> GET /catalog  -> control operation, unchanged target behavior
Checkout -> GET /checkout -> treatment operation, mode-controlled latency/errors
No Transaction Controller parent sample in checkpoint

This removes ambiguity when reviewing table/graph labels later.

9. Workload/report manifest

Record:

  • JMX/property SHA-256 and JMeter/Java versions;
  • threads=2, loops=20, two samplers/loop, 50 ms pacing;
  • run IDs/modes;
  • labels/control/treatment mapping;
  • save-service required fields and CSV format;
  • APDEX=100/300 ms, p90/p95/p99, granularity=2000 ms;
  • no sample_filter/start_date/end_date for the main checkpoint reports;
  • generator CPU/heap/network/disk headroom;
  • result/report directories and target event log.

10. Write the interpretation with uncertainty

Observation: Checkout latency distribution and error rate worsen in degraded mode; Catalog latency remains stable. Overall completion throughput falls.

Controlled mechanism: The fixture intentionally changes only Checkout from ~45 ms/success to ~180 ms with deterministic HTTP 503s every fifth sequence. Target telemetry confirms that exact operation-level change.

Interpretation: The Checkout regression is attributable to the controlled target change. The overall throughput reduction is consistent with a closed workload spending longer blocked on Checkout; it does not prove Catalog capacity changed.

Uncertainty/limits: This is a tiny loopback experiment with 40 samples/label; extreme-tail percentiles are low-resolution, network variability is minimal, and the result says nothing about production capacity. APDEX conclusions apply only to the lab's 100/300 ms thresholds.

11. Check one reporting failure without rerunning load

Copy the baseline JTL and remove one required field as in Lesson 4. Attempt -g into a new directory, preserve the failure, then regenerate from the untouched JTL. This proves report failures can be diagnosed at the artifact/configuration layer without sending more traffic.

12. Required evidence packet

Artifact Required content
Raw JTL baseline/degraded CSVs with expected labels/counts.
Dashboards two separate new/empty report directories.
Statistics/percentile table Catalog/Checkout counts/errors/avg/p90/p95/p99.
APDEX/error table thresholds published; degraded errors mapped to Checkout.
Selected graphs threads, response time, response codes, throughput/distribution.
Sample-label map Catalog control / Checkout treatment; no parent transaction sample.
Workload manifest versions/hashes/load/report properties/filter/window state.
Target telemetry 80 events/run, expected statuses and service times.
Generator validity CPU/heap/network/disk headroom and jmeter.log.
Interpretation note observations, causal mechanism, uncertainty/limitations.

13. Final validity statement

Example: “Apache JMeter 5.6.3/Java 17 executed two identical 2-thread ×20-loop ×2-sampler localhost workloads against prompt23-dashboard-fixture-v1. Both produced 80 JTL rows and 80 target events. The only treatment was mode=degraded, which target telemetry confirmed added Checkout service delay and eight deterministic HTTP 503s while Catalog remained unchanged. Dashboards used compatible lean CSV fields, p90/p95/p99, APDEX 100/300 ms and 2-second over-time granularity with no source/window filter. Dashboard counts/errors/latency trends were cross-checked against raw JTL and target events; generator headroom and jmeter.log were preserved. This establishes the local controlled-regression interpretation method, not production capacity or a universal APDEX/SLO threshold.”

14. Verification checklist

  • Only 127.0.0.1:8023.
  • JMeter 5.6.3 / Java 17 / no plugins.
  • Same JMX/hash/workload/report properties except run ID/mode.
  • 80 JTL +80 target events each.
  • Degraded exactly eight Checkout 503s; Catalog no intentional failures.
  • Dashboard statistics/error table match raw JTL counts.
  • Checkout latency/APDEX move in expected direction; Catalog control remains stable.
  • Throughput interpretation states the closed-loop dependency.
  • Generator headroom/filter/window/controller policy documented.
  • Interpretation states uncertainty/sample-size/local-only limitations.

15. Cleanup / rollback

  1. Stop JMeter processes and localhost fixture after target evidence is saved.
  2. Keep raw JTL, jmeter.log, dashboard dirs, report/workload properties, manifests and event evidence until review.
  3. Delete only disposable local artifacts after the defined retention period; do not delete first-failure evidence to rerun.
  4. No public/production target, real credential/PII, recorder CA, remote RMI, container, paid telemetry/CI, database/message service or OS/JVM global tuning was changed.

16. What Chapter 23 adds to the operating model

The production performance-testing operating model now has a dashboard interpretation contract: report source JTL/schema, labels/controller policy, percentile set, APDEX thresholds, sample/series filters, measurement window, time granularity, workload manifest, configured-versus-achieved count/rate, generator validity, target telemetry, raw-to-dashboard cross-check and uncertainty statement must accompany any performance conclusion.

Chapter 24 moves to Backend Listener, InfluxDB/Graphite, Grafana, and Real-Time Telemetry. It extends this interpretation discipline into live time-series pipelines and explicitly measures the network/backend cost that a dashboard generated after the run does not impose during sampling.

Knowledge check

What is the strongest evidence that Checkout caused the controlled regression?

Why can degraded Catalog hits/sec fall without a Catalog latency regression?

Why must APDEX thresholds be included in the evidence packet?

What invalidates the comparison even if the dashboards look similar?

What is Chapter 24's bridge?

Next chapter

Backend Listener, InfluxDB/Graphite, Grafana, and Real-Time Telemetry

Chapter 24 builds a local real-time telemetry pipeline and compares live metrics with raw JTL/dashboard evidence.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. The dashboard generator reads compatible CSV sample logs and can run at the end of a load test with -e -o or later with -g <CSV> -o <dir>. Required CSV data include bytes, label, latency, response code/message, success, thread counts/name, elapsed time, connect time, assertion failure message, and a timestamp format containing time; current defaults are suitable unless changed. The HTML output directory must be empty/new. Dashboard statistics expose three configurable percentile levels via aggregate_rpt_pct1/2/3, defaulting to 90/95/99. The report-generator general APDEX defaults are 500 ms satisfied and 1500 ms tolerated; this chapter's local lab deliberately overrides them to 100/300 ms. sample_filter removes sample data before report calculations. HTML series_filter filters displayed series/rows after calculations. start_date/end_date constrain the report measurement window. The default over-time granularity is 60000 ms and must remain above 1000 ms; the tiny local lab uses 2000 ms. Dashboard percentile estimates can differ from GUI Aggregate Report, especially with few/widely distributed samples, because estimator formulas differ. Several dashboard graphs include Transaction Controller sample results while others explicitly ignore/exclude them, so labels and controller-sample policy must be stated before totals are interpreted.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.