Checkpoint Lab — Capstone: Design and Operate a Production Performance-Testing Program
The final checkpoint delivers the complete local performance-testing platform and uses it to make a release decision. The target has two controlled builds: baseline checkout service time35 ms and regression checkout55 ms. Everything else—authorization, data, JMX, threads/loops/pacing, runtime, analysis and gate policy—stays fixed.
Learning objectives
- Complete the chapter checkpoint for Capstone: Design and Operate a Production Performance-Testing Program as one reviewable, bounded experiment.
- State the workload, predictions, acceptance criteria, authorization boundary, and abort conditions before execution.
- Reconcile configured versus achieved work with JTL, jmeter.log, target evidence, and generator validity before making a conclusion.
- Produce an evidence packet that records the exact inputs, results, diagnosis or gate outcome, and any material limitations.
- Perform cleanup or rollback and explain how the checkpoint evidence hands off to the next chapter or operating practice.
1. Required platform deliverables
| Area | Evidence |
|---|---|
| Architecture/trust boundary | Lifecycle diagram + charter showing authorized loopback target and synthetic session/data boundaries. |
| Project/version identity | Complete tree, versions.lock, JMX/config/data/charter hashes. |
| Workload | 2×5 local profile,10 journeys,40 HTTP samples,75ms pacing, unique10-row CSV. |
| Functional state | Cookie/session token correlation, catalog item correlation, cart/checkout assertions. |
| Execution | CLI command/runner, raw JTL, matching jmeter.log, HTML dashboard. |
| Telemetry | Target JSONL/stats + generator process/resource observation. |
| Scaling | Two-engine simulated preflight:1×5 each, unique5-row shards, fleet total40. |
| Gate/governance | Valid baseline/current summaries, p35-policy-v1 gate, immutable baseline reference. |
| Diagnostics | At least one preserved controlled failure drill and its repair. |
| Operations | RUNBOOK.md + archive/index metadata + cleanup/rollback. |
2. Exact assumptions, authorization and stop conditions
- Apache JMeter5.6.3; Java17; HttpClient4; no third-party plugin.
- Python3 standard library service/analysis only.
-
Target exactly
127.0.0.1:8035; allowed load endpoints session/catalog/cart/checkout only. - Local run=2 threads×5 loops=10 journeys=40 HTTP samples;75ms timer.
- Target guard=50 requests/process; normal measured run=40.
- Baseline build checkout=35ms; regression build checkout=55ms.
- Gate requires valid40-sample evidence, zero errors, checkout absolute p95≤80ms, relative checkout p95 regression≤25%.
3. Predict before running
P1: baseline and current both execute40 samples/10 complete journeys with zero errors because only checkout service delay changes.
P2: Session/Catalog/Cart JTL and target service metrics should remain similar; Checkout JTL and target service p95 should increase together.
P3: current checkout p95 should remain below the absolute80ms objective but regress by more than25% versus baseline, so technical gate should FAIL and release disposition should BLOCK.
P4: if generator state becomes invalid or any sample is missing, gate should be INVALID rather than interpreting the target regression.
4. Preflight
- Read charter/lock/governance/runbook.
- Verify JMeter5.6.3/Java17.
- Hash JMX/config/data/charter/lock.
- Verify exactly10 unique CSV rows.
- Confirm port8035 is free, then start the intended fixture build.
- Verify
/healthbuild ID/checkout delay/max50. - Confirm a fresh immutable run directory.
5. Establish baseline evidence
Start baseline build:
python .\fixtures\capstone_service.py `
--host 127.0.0.1 --port 8035 `
--build-id build-baseline `
--checkout-ms 35 `
--max-requests 50 `
--log .\runs\baseline-target-events.jsonl
.\tools\run_capstone.ps1 -RunId p35-baseline -Profile baseline
python .\tools\analyze_capstone.py `
--jtl .\runs\p35-baseline\results.jtl `
--target-events .\runs\baseline-target-events.jsonl `
--build-id build-baseline `
--run-id p35-baseline `
--generator-valid true `
--out .\baselines\baseline-v1.json
Verify40/40,10 per label/endpoint, zero errors/rejections, healthy generator observation and consistent dashboard/target telemetry. Freeze baseline-v1 and its input hashes; do not edit it after current results are known.
6. Required controlled failure drill before current run
Copy the plan to a disposable debug copy and change only the JSON
Extractor path for ITEM_ID from
$.item_id to $.missing_item. Run1 thread×1
loop against a freshly started baseline fixture into
runs/failure-first/.
Expected: Catalog itself returns200, extractor assigns
CORRELATION_MISSING; Cart then fails
validation/protocol (or an assertion catches the missing variable,
depending on your exact tree), and Stop Thread prevents Checkout.
Preserve JTL/jmeter.log/target events/changed JMX hash
before repairing. Restore $.item_id, rerun1×1 and
require four successful steps.
This drill demonstrates that correlation/data/scope failure must be diagnosed before interpreting latency or gate output.
7. Execute controlled regression build
Stop/reset the fixture, then:
python .\fixtures\capstone_service.py `
--host 127.0.0.1 --port 8035 `
--build-id build-regression `
--checkout-ms 55 `
--max-requests 50 `
--log .\runs\current-target-events.jsonl
.\tools\run_capstone.ps1 -RunId p35-current -Profile current
python .\tools\analyze_capstone.py `
--jtl .\runs\p35-current\results.jtl `
--target-events .\runs\current-target-events.jsonl `
--build-id build-regression `
--run-id p35-current `
--generator-valid true `
--out .\runs\p35-current\summary.json
Independently verify target checkout service time increased while other endpoint service times remain stable. Confirm the same plan/config/data/lock hashes and40/40 achieved work.
8. Run the CI-style gate
python .\tools\gate_capstone.py `
--baseline .\baselines\baseline-v1.json `
--current .\runs\p35-current\summary.json `
--out .\runs\p35-current\gate.json
$GateExit=$LASTEXITCODE
# Expected for designed regression: 2 = technical FAIL / release BLOCK.
Interpret causally: if current is valid and checkout target service/JTL p95 both rise while generator remains healthy, the controlled build regression is the leading cause. If evidence does not match those predictions, do not force the expected gate outcome—diagnose the actual run.
9. Verify the distributed profile without RMI
Create two engine manifests with identical
JMeter5.6.3/Java17/plan/config/analysis hashes, each
threads=1/loops=5. Split users001–005 and006–010 into separate
data/users.csv files. Run
fleet_preflight.py; require PASS/fleet40 samples/no
overlapping user IDs. This proves the scaling math before any real
RMI environment is authorized.
10. Archive the final evidence packet
Create archive/p35-final/index.json containing paths
and SHA-256 for:
- architecture/trust diagram source and charter;
- versions.lock, governance policy, runbook;
- JMX/config/data;
- baseline/current JTL, jmeter.log, dashboard directories and summaries;
- target JSONL/stats and generator observations;
- failure-first/fixed drill evidence;
- fleet preflight output;
- gate output and final decision statement.
The archive should contain no raw bearer token, production secret or uncontrolled target data.
11. Final decision statement
12. Experimental-validity limitations
- Only ten journeys/ten checkout samples: tail-percentile precision is coarse.
- Loopback removes real network/DNS/TLS variability.
- Closed two-thread workload couples response time and achieved throughput.
- Synthetic in-memory service is not a capacity model of a distributed production application.
- Single local injector does not validate actual RMI/controller/network behavior.
- The lab proves operating discipline and causality, not production capacity.
13. Cleanup / rollback
- Stop fixture and verify port8035 closes.
- Keep baseline/current/failure/fleet/gate/governance evidence per lab retention.
- Delete disposable session/cart/order memory with the fixture process; all data is synthetic.
- Do not overwrite baseline-v1 after the failed current run.
- No real remote engine, cloud service, container registry, CI secret, public target or system-wide setting was changed.
14. What the full course adds to a production operating model
A production performance-testing program now has: risk-driven test types; correct JMeter tree/scope and workload models; parameterization/correlation/session handling; protocol-specific patterns; modular/custom extension discipline; GUI-debug/CLI execution; lean results and dashboards; telemetry; distributed/container/CI scaling; generator/JVM/OS sizing; bottleneck analysis and experimental validity; authorization/security/privacy; preserve-first troubleshooting; versioned baselines/SLO/regression budgets/exceptions; and an evidence/runbook/ownership loop that makes decisions reproducible.
End of course: ongoing performance governance means the program keeps evolving with service risk. Review SLOs and workload representativeness, promote baselines only through change control, rehearse failure/rollback paths, update runtime/plugin locks deliberately, trend valid results, retire stale tests/data, and preserve enough evidence that every release decision can be explained later.
Knowledge check
What makes the final regression claim causal rather than just correlated?
Only checkout service delay changes; workload/runtime/data stay fixed, generator remains valid, and both target checkout service time and JMeter checkout elapsed rise while other endpoints stay stable.
Why must the correlation failure drill run before performance interpretation?
A broken ITEM_ID chain creates functional/session errors and incomplete journeys, making latency/gate conclusions invalid.
Why is the final release disposition BLOCK even if checkout stays below80ms?
The valid current run exceeds the separate25% relative regression budget; absolute and regression policies answer different risks.
Why not test the distributed profile with2×5 on both engines?
Each engine runs the full plan; that would double the authorized journeys and requests.
What should continue after the course ends?
Governance and improvement: review risks/SLOs/workloads, trend valid runs, control baselines/exceptions/runtime upgrades, rehearse failures, maintain evidence/runbooks and retire stale tests safely.
Official references and version notes
- Apache JMeter downloads — current stable JMeter 5.6.3, Java 8+.
- JMeter current changes — Java 17+ recommended for the 5.6.x line.
- JMeter Getting Started — GUI authoring/debug, CLI flags/properties and runtime configuration.
- JMeter Component Reference — HTTP Request, Cookie Manager, CSV Data Set Config, JSON Extractor/assertions and component semantics.
- JMeter Best Practices — non-GUI load execution, listener cost and cached JSR223 Groovy guidance.
- JMeter Dashboard — CSV-backed HTML report and report generation.
- JMeter Remote Testing — whole-plan fan-out, exact JMeter parity, Java/data requirements, controller overhead and RMI SSL.
Current behavior was rechecked against primary Apache JMeter documentation on 2026-09-06. Mandatory path: Apache JMeter 5.6.3, Java 17, built-in HttpClient4, no third-party plugin, Python 3 standard library fixture/tools, loopback target only. JMeter 5.6.3 requires Java 8+ and the 5.6.x line recommends Java 17+. Meaningful load is CLI; GUI is for authoring/debug. HTTP retry remains disabled; correlation uses built-in JSON Extractor; data uses CSV Data Set Config; per-thread cookies use HTTP Cookie Manager. The HTML dashboard is generated from raw CSV JTL and is corroborating evidence rather than the only source. Real remote mode is optional: every server runs the whole plan, so configured load multiplies unless per-engine properties are adjusted; engines require exact JMeter parity, should use the same Java, need their own data files, and RMI SSL remains enabled. No external container, CI provider, plugin, telemetry server, or paid service is mandatory.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.