Chapter 35Lesson 05~390 minutes

Checkpoint Lab — Capstone: Design and Operate a Production Performance-Testing Program

The final checkpoint delivers the complete local performance-testing platform and uses it to make a release decision. The target has two controlled builds: baseline checkout service time35 ms and regression checkout55 ms. Everything else—authorization, data, JMX, threads/loops/pacing, runtime, analysis and gate policy—stays fixed.

Final platformControlled regressionEvidence packetRunbookOperating model

Learning objectives

  • Complete the chapter checkpoint for Capstone: Design and Operate a Production Performance-Testing Program as one reviewable, bounded experiment.
  • State the workload, predictions, acceptance criteria, authorization boundary, and abort conditions before execution.
  • Reconcile configured versus achieved work with JTL, jmeter.log, target evidence, and generator validity before making a conclusion.
  • Produce an evidence packet that records the exact inputs, results, diagnosis or gate outcome, and any material limitations.
  • Perform cleanup or rollback and explain how the checkpoint evidence hands off to the next chapter or operating practice.

1. Required platform deliverables

Area Evidence
Architecture/trust boundary Lifecycle diagram + charter showing authorized loopback target and synthetic session/data boundaries.
Project/version identity Complete tree, versions.lock, JMX/config/data/charter hashes.
Workload 2×5 local profile,10 journeys,40 HTTP samples,75ms pacing, unique10-row CSV.
Functional state Cookie/session token correlation, catalog item correlation, cart/checkout assertions.
Execution CLI command/runner, raw JTL, matching jmeter.log, HTML dashboard.
Telemetry Target JSONL/stats + generator process/resource observation.
Scaling Two-engine simulated preflight:1×5 each, unique5-row shards, fleet total40.
Gate/governance Valid baseline/current summaries, p35-policy-v1 gate, immutable baseline reference.
Diagnostics At least one preserved controlled failure drill and its repair.
Operations RUNBOOK.md + archive/index metadata + cleanup/rollback.

2. Exact assumptions, authorization and stop conditions

  • Apache JMeter5.6.3; Java17; HttpClient4; no third-party plugin.
  • Python3 standard library service/analysis only.
  • Target exactly 127.0.0.1:8035; allowed load endpoints session/catalog/cart/checkout only.
  • Local run=2 threads×5 loops=10 journeys=40 HTTP samples;75ms timer.
  • Target guard=50 requests/process; normal measured run=40.
  • Baseline build checkout=35ms; regression build checkout=55ms.
  • Gate requires valid40-sample evidence, zero errors, checkout absolute p95≤80ms, relative checkout p95 regression≤25%.
Abort/stop: non-loopback target, request guard rejection, any real credential, correlation default, unexpected401/409/5xx, achieved samples≠40, target endpoint counts≠10 each, generator invalidity, missing JTL/log/dashboard/telemetry, changed plan/data/config/hash between baseline/current, or any unbounded workload increase.

3. Predict before running

P1: baseline and current both execute40 samples/10 complete journeys with zero errors because only checkout service delay changes.

P2: Session/Catalog/Cart JTL and target service metrics should remain similar; Checkout JTL and target service p95 should increase together.

P3: current checkout p95 should remain below the absolute80ms objective but regress by more than25% versus baseline, so technical gate should FAIL and release disposition should BLOCK.

P4: if generator state becomes invalid or any sample is missing, gate should be INVALID rather than interpreting the target regression.

4. Preflight

  1. Read charter/lock/governance/runbook.
  2. Verify JMeter5.6.3/Java17.
  3. Hash JMX/config/data/charter/lock.
  4. Verify exactly10 unique CSV rows.
  5. Confirm port8035 is free, then start the intended fixture build.
  6. Verify /health build ID/checkout delay/max50.
  7. Confirm a fresh immutable run directory.

5. Establish baseline evidence

Start baseline build:

python .\fixtures\capstone_service.py `
  --host 127.0.0.1 --port 8035 `
  --build-id build-baseline `
  --checkout-ms 35 `
  --max-requests 50 `
  --log .\runs\baseline-target-events.jsonl

.\tools\run_capstone.ps1 -RunId p35-baseline -Profile baseline

python .\tools\analyze_capstone.py `
  --jtl .\runs\p35-baseline\results.jtl `
  --target-events .\runs\baseline-target-events.jsonl `
  --build-id build-baseline `
  --run-id p35-baseline `
  --generator-valid true `
  --out .\baselines\baseline-v1.json

Verify40/40,10 per label/endpoint, zero errors/rejections, healthy generator observation and consistent dashboard/target telemetry. Freeze baseline-v1 and its input hashes; do not edit it after current results are known.

6. Required controlled failure drill before current run

Copy the plan to a disposable debug copy and change only the JSON Extractor path for ITEM_ID from $.item_id to $.missing_item. Run1 thread×1 loop against a freshly started baseline fixture into runs/failure-first/.

Expected: Catalog itself returns200, extractor assigns CORRELATION_MISSING; Cart then fails validation/protocol (or an assertion catches the missing variable, depending on your exact tree), and Stop Thread prevents Checkout. Preserve JTL/jmeter.log/target events/changed JMX hash before repairing. Restore $.item_id, rerun1×1 and require four successful steps.

This drill demonstrates that correlation/data/scope failure must be diagnosed before interpreting latency or gate output.

7. Execute controlled regression build

Stop/reset the fixture, then:

python .\fixtures\capstone_service.py `
  --host 127.0.0.1 --port 8035 `
  --build-id build-regression `
  --checkout-ms 55 `
  --max-requests 50 `
  --log .\runs\current-target-events.jsonl

.\tools\run_capstone.ps1 -RunId p35-current -Profile current

python .\tools\analyze_capstone.py `
  --jtl .\runs\p35-current\results.jtl `
  --target-events .\runs\current-target-events.jsonl `
  --build-id build-regression `
  --run-id p35-current `
  --generator-valid true `
  --out .\runs\p35-current\summary.json

Independently verify target checkout service time increased while other endpoint service times remain stable. Confirm the same plan/config/data/lock hashes and40/40 achieved work.

8. Run the CI-style gate

python .\tools\gate_capstone.py `
  --baseline .\baselines\baseline-v1.json `
  --current .\runs\p35-current\summary.json `
  --out .\runs\p35-current\gate.json

$GateExit=$LASTEXITCODE
# Expected for designed regression: 2 = technical FAIL / release BLOCK.

Interpret causally: if current is valid and checkout target service/JTL p95 both rise while generator remains healthy, the controlled build regression is the leading cause. If evidence does not match those predictions, do not force the expected gate outcome—diagnose the actual run.

9. Verify the distributed profile without RMI

Create two engine manifests with identical JMeter5.6.3/Java17/plan/config/analysis hashes, each threads=1/loops=5. Split users001–005 and006–010 into separate data/users.csv files. Run fleet_preflight.py; require PASS/fleet40 samples/no overlapping user IDs. This proves the scaling math before any real RMI environment is authorized.

10. Archive the final evidence packet

Create archive/p35-final/index.json containing paths and SHA-256 for:

  • architecture/trust diagram source and charter;
  • versions.lock, governance policy, runbook;
  • JMX/config/data;
  • baseline/current JTL, jmeter.log, dashboard directories and summaries;
  • target JSONL/stats and generator observations;
  • failure-first/fixed drill evidence;
  • fleet preflight output;
  • gate output and final decision statement.

The archive should contain no raw bearer token, production secret or uncontrolled target data.

11. Final decision statement

Expected decision shape: “The run is valid because the same JMeter5.6.3/Java17/JMX/config/data/workload executed40/40 samples with10 events per endpoint, zero errors, and acceptable generator state on the authorized loopback service. Checkout target service time and JMeter checkout p95 increased together after the only intentional target change, checkout35→55ms; Session/Catalog/Cart remained stable. The current build remains below the absolute80ms checkout p95 objective but exceeds the25% relative regression budget, so the technical gate is FAIL and release disposition is BLOCK. The baseline is not changed. Remediation or a separately governed Chapter34 exception would be required before release.”

12. Experimental-validity limitations

  • Only ten journeys/ten checkout samples: tail-percentile precision is coarse.
  • Loopback removes real network/DNS/TLS variability.
  • Closed two-thread workload couples response time and achieved throughput.
  • Synthetic in-memory service is not a capacity model of a distributed production application.
  • Single local injector does not validate actual RMI/controller/network behavior.
  • The lab proves operating discipline and causality, not production capacity.

13. Cleanup / rollback

  1. Stop fixture and verify port8035 closes.
  2. Keep baseline/current/failure/fleet/gate/governance evidence per lab retention.
  3. Delete disposable session/cart/order memory with the fixture process; all data is synthetic.
  4. Do not overwrite baseline-v1 after the failed current run.
  5. No real remote engine, cloud service, container registry, CI secret, public target or system-wide setting was changed.

14. What the full course adds to a production operating model

A production performance-testing program now has: risk-driven test types; correct JMeter tree/scope and workload models; parameterization/correlation/session handling; protocol-specific patterns; modular/custom extension discipline; GUI-debug/CLI execution; lean results and dashboards; telemetry; distributed/container/CI scaling; generator/JVM/OS sizing; bottleneck analysis and experimental validity; authorization/security/privacy; preserve-first troubleshooting; versioned baselines/SLO/regression budgets/exceptions; and an evidence/runbook/ownership loop that makes decisions reproducible.

End of course: ongoing performance governance means the program keeps evolving with service risk. Review SLOs and workload representativeness, promote baselines only through change control, rehearse failure/rollback paths, update runtime/plugin locks deliberately, trend valid results, retire stale tests/data, and preserve enough evidence that every release decision can be explained later.

Knowledge check

What makes the final regression claim causal rather than just correlated?

Why must the correlation failure drill run before performance interpretation?

Why is the final release disposition BLOCK even if checkout stays below80ms?

Why not test the distributed profile with2×5 on both engines?

What should continue after the course ends?

Course complete

Continue with governed performance engineering

Use the capstone architecture as a living template: scale only when evidence and authorization require it, keep experiments valid, and make every important performance decision reproducible.

Official references and version notes

Version and compatibility note

Current behavior was rechecked against primary Apache JMeter documentation on 2026-09-06. Mandatory path: Apache JMeter 5.6.3, Java 17, built-in HttpClient4, no third-party plugin, Python 3 standard library fixture/tools, loopback target only. JMeter 5.6.3 requires Java 8+ and the 5.6.x line recommends Java 17+. Meaningful load is CLI; GUI is for authoring/debug. HTTP retry remains disabled; correlation uses built-in JSON Extractor; data uses CSV Data Set Config; per-thread cookies use HTTP Cookie Manager. The HTML dashboard is generated from raw CSV JTL and is corroborating evidence rather than the only source. Real remote mode is optional: every server runs the whole plan, so configured load multiplies unless per-engine properties are adjusted; engines require exact JMeter parity, should use the same Java, need their own data files, and RMI SSL remains enabled. No external container, CI provider, plugin, telemetry server, or paid service is mandatory.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.