Chapter 19Lesson 05~250 minutes

Checkpoint Lab — Scripting with JSR223 and Groovy for Advanced Test Logic

The checkpoint is a refactoring experiment, not a race for the highest requests/second. The over-scripted plan and the refactored plan must produce the same target behavior first. Only then may generator CPU/heap/logging and achieved throughput be compared. The final plan should expose most logic in built-in JMeter elements and keep one small cache-safe Groovy file.

CheckpointRefactorCacheThread safetyGenerator overhead

Learning objectives

  • Prove exact local setup/tool versions/cache assumptions.
  • Unit-test the canonicalization helper outside JMeter.
  • Run over-scripted and refactored plans against identical local workload.
  • Prove semantic equivalence through target name/thread/sequence evidence.
  • Measure configured versus achieved samples plus generator CPU/heap/log evidence.
  • Demonstrate thread-local vars and immutable props with no global mutable user state.

1. Exact assumptions and hard ceilings

Item Checkpoint baseline
JMeter Apache JMeter 5.6.3.
Java Java 17 JDK; JMeter 5.6.3 requires Java 8+.
Groovy Bundled JMeter 5.6.3 line: Groovy 3.0.20.
Plugins None.
Target http://127.0.0.1:8019 only.
Fixture Python stdlib prompt19-groovy-fixture-v1.
Input 4 synthetic RAW_NAME rows.
Authoring check 1 thread ×4 loops.
Benchmark 2 threads ×300 loops max/variant + 10 ms Constant Timer; duration cap 15 s.
Cache Default compiled cache 100 unless explicitly overridden.
Refactored helper scripts/canonicalize.groovy as Script File; no runtime-value substitution in source.
Results Lean CSV JTL + matching jmeter.log; no per-sample response-body/log spam.
Exception test Separate 1 thread ×1 JSR223 Sampler only.
Abort: non-loopback target, >600 work samples/variant, script exception loop, failed helper self-test, name/thread/sequence mismatch, mutable per-user state in props, log growth from per-sample INFO, or sustained generator saturation. Never “fix” the checkpoint by hiding failures or raising load.

2. Setup and target preflight

  1. Create data/names.csv, scripts/canonicalize.groovy, and the fixture from Lesson 2.
  2. Start fixture at 127.0.0.1:8019 with a fresh event log.
  3. GET /health//stats; record fixture version and zero work count.
  4. Record JMeter 5.6.3, Java 17, Groovy 3.0.20 and effective/default cache note.
  5. Verify no extra plugin/ScriptEngine dependency is used.

3. Helper self-test is a gate

java -cp "$env:JMETER_HOME\lib\*" `
  groovy.ui.GroovyMain `
  .\scripts\canonicalize.groovy

Require PASS: prompt19 canonical helper self-test. If it fails, do not start JMeter load. Fix pure logic first.

4. State ownership before execution

State Owner Allowed behavior
RAW_NAME vars / one virtual user CSV changes it.
REQUEST_NAME vars / one virtual user Groovy helper overwrites from current RAW_NAME.
SEQ vars / one virtual user Counter increments independently per thread.
SERVER_* vars / one virtual user Built-in extractor writes after each sample.
RUN_ID props / one JVM CLI sets once; scripts/components read only.
VARIANT props / one JVM Separate process/run sets once; read only.
Shared mutable list/map/token none Forbidden in checkpoint.

5. Exact before/after trees

Before — over-scripted:

Thread Group
├── Counter -> SEQ
├── Constant Timer 10 ms
└── HTTP Work
    ├── JSR223 PreProcessor
    │   inline Groovy, cache OFF, ${RAW_NAME} substitution
    └── JSR223 PostProcessor
        inline Groovy, cache OFF, JsonSlurper + manual validation

After — refactored:

Thread Group
├── Counter -> SEQ
├── Constant Timer 10 ms
└── HTTP Work
    ├── JSR223 PreProcessor -> scripts/canonicalize.groovy
    ├── JSON JMESPath Assertion -> status == ok
    ├── JMESPath Extractors -> SERVER_NAME / SCORE / SIGNATURE
    └── Response Assertion -> SERVER_NAME == ${REQUEST_NAME}

Target URL/data/timer/thread count/result policy remain identical.

6. Predictions before running

Prediction A — behavior: both plans produce the same canonical-name distribution and zero unexpected target/JTL failures. Refactoring changes generator implementation, not business requests.

Prediction B — state isolation: sequence values are independent per thread because Counter and variables are thread-local; neither plan needs a mutable property for user state.

Prediction C — overhead: the refactored plan should eliminate repeated uncached Groovy compilation/interpreting and manual JSON parsing. It may improve achieved throughput/lower generator CPU, but the magnitude must be measured rather than asserted.

Prediction D — target: server service_wall_ms should remain broadly similar because target code/work is unchanged. If target timing changes materially, investigate before attributing the difference to Groovy.

7. Functional equivalence gate

Run each plan 1×4 with separate event/result directories. Require all four canonical names once (scheduler/order can differ when threads increase), correct status/name, zero unexpected failures and no Groovy exception.

Do not benchmark until this gate is green.

8. Bounded benchmark runs

Run over-scripted:

jmeter.bat -n `
  -t plans\over-scripted.jmx `
  -l results\over-scripted\results.jtl `
  -j results\over-scripted\jmeter.log `
  -JRUN_ID=p19-over `
  -JVARIANT=over-scripted

Restart/freshen fixture evidence as needed, then refactored:

jmeter.bat -n `
  -t plans\refactored.jmx `
  -l results\refactored\results.jtl `
  -j results\refactored\jmeter.log `
  -JRUN_ID=p19-ref `
  -JVARIANT=refactored

Each plan must enforce 2 threads ×300 loops max, 10 ms timer, ≤15-second duration safety cap.

9. Generator evidence during each variant

Resolve the current JMeter Java PID and record multiple CPU/memory snapshots plus jcmd PID GC.heap_info. Record jmeter.log size. Do not change heap, GC, log level, listener set, JVM flags, or OS priority between variants.

If the run is too short for reliable process sampling, repeat the same bounded benchmark a few times as separate runs rather than increasing threads/payload/target scope. Compare medians and keep each run under the stated cap.

10. Analyze configured versus achieved load

python tools/compare_jtl.py   results/over-scripted/results.jtl   results/refactored/results.jtl

python tools/analyze_server_events.py   results/server-events.jsonl

Configured per variant = at most 600 HTTP samples. Achieved = successful samples divided by measured result span, plus actual target event count. If a plan ends early, has failures, or sends a different request distribution, do not compare throughput as if the workloads were equal.

11. Exception evidence gate

Run a separate one-shot JSR223 Sampler:

throw new IllegalStateException("prompt19 synthetic exception - diagnostic only")

Require one failed sample and corresponding jmeter.log exception. Then disable/remove this diagnostic element. The purpose is to prove your evidence path catches script errors; it must not contaminate benchmark JTL.

12. Cache/configuration proof

Evidence packet must include:

  • Groovy 3.0.20 output;
  • documented/default jsr223.compiled_scripts_cache_size=100 or explicit override;
  • refactored JMX showing Script File, not dynamic inline source;
  • script source containing vars.get('RAW_NAME') and no direct JMeter variable replacement in script text;
  • over-scripted source/cache-off configuration preserved for comparison.

13. Prove target behavior independently

Target event analysis must show:

  • variant event counts consistent with successful JTL samples;
  • only canonical names alice, bob_2, carol-3, delta;
  • thread/sequence metadata in valid ranges;
  • HTTP status 200 for valid workload;
  • no hidden increase in target work for one variant.

14. Required evidence packet

Artifact Required content
Versions JMeter 5.6.3, Java 17, Groovy 3.0.20, no plugins.
Groovy source canonicalize.groovy + over-scripted inline source.
Cache note Script File/compiled cache behavior + default/override value.
vars/props table Thread-local mutable state versus immutable process-wide run settings.
Unit test PASS from GroovyMain running the same helper file.
Functional gate 1×4 equivalent target/request behavior.
Benchmark JTL Separate raw JTL + matching jmeter.log for both variants.
Generator evidence CPU/working set/heap/log size snapshots with unchanged JVM settings.
Target evidence Event counts/name/thread/seq distribution + service_wall_ms.
Exception evidence One separate failed JSR223 sample + stack trace.
Validity statement Configured/achieved load and limitations; no production-capacity claim.

15. Validity statement

Example: “Apache JMeter 5.6.3/Java 17 with bundled Groovy 3.0.20 tested the local prompt19-groovy-fixture-v1 at 127.0.0.1:8019. The over-scripted plan used cache-off inline Groovy for canonicalization plus JsonSlurper/manual validation; the refactored plan used Counter/JMESPath/Response Assertion and one Groovy Script File that reads RAW_NAME via thread-local vars. The same helper passed standalone Groovy assertions. Both variants first passed a 1×4 semantic-equivalence gate, then ran the same bounded 2×300/10 ms workload. Achieved throughput, generator CPU/heap/log state, JTL failures, and target event/service-time evidence were compared; no result was accepted if request distributions differed. A separate one-shot synthetic exception produced a failed sample and jmeter.log stack trace. This validates the refactoring/caching/thread-state methodology on a tiny local target; it does not prove production service capacity or guarantee that every script refactor improves performance by the same percentage.”

16. Verification checklist

  • Target exactly 127.0.0.1:8019.
  • Helper self-test passes outside JMeter.
  • Refactored source uses vars.get, not direct JMeter variable replacement in cached script text.
  • RUN_ID/VARIANT props are immutable; per-user values stay in vars.
  • Functional 1×4 behavior matches before benchmark.
  • Both benchmark variants use identical threads/loops/timer/target/data/results.
  • Achieved sample counts/throughput and target event counts agree.
  • Generator CPU/heap/log evidence recorded; no tuning changes between variants.
  • One exception sample is preserved separately and remains failed.
  • No script marks real HTTP failures successful.

17. Cleanup / rollback

  1. Stop JMeter runs and the localhost fixture.
  2. Preserve JTL/jmeter.log/script/unit-test/target-event evidence until review completes.
  3. Delete only disposable local results and fixture files after review.
  4. No external API, real credential, environment secret, plugin, remote engine, recorder certificate, container, paid service, database/message target, OS/JVM global tuning, or production system was changed.

18. What Chapter 19 adds to the operating model

The production performance-testing operating model now has a scripting contract: built-in-first justification, JSR223 element lifecycle/scope, pinned JMeter/Java/Groovy engine, compiled-cache policy, cache-safe runtime bindings, vars/props ownership, script-file deployment/user.dir, pure-helper tests, exception/logging policy, SampleResult integrity, generator CPU/heap/log budget, target-equivalence evidence, and configured-versus-achieved load comparison are reviewed before scripted results become release evidence.

Chapter 20 moves to Functions, Custom Properties, Reusable Fragments, and Modular Plans. It builds on this chapter by standardizing reusable configuration and composition while keeping state ownership and execution semantics explicit.

Knowledge check

What must be proven before comparing over-scripted and refactored throughput?

What thread-safety rule does the checkpoint demonstrate?

Why preserve a one-shot exception sample separately?

If refactored HTTP p95 is unchanged but throughput rises and generator CPU falls, what does that suggest?

What is the Chapter 20 bridge?

Next chapter

Functions, Custom Properties, Reusable Fragments, and Modular Plans

Chapter 20 focuses on reuse and composition: functions, property layering, fragments/controllers, shared configuration, portability, and module boundaries.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. The JMeter 5.6.3 binary line uses Groovy 3.0.20. Groovy's JSR223 engine implements Compilable. JMeter recommends script files (compiled/cached when supported) or inline script text with Cache compiled script if available checked. The compiled-script cache defaults to jsr223.compiled_scripts_cache_size=100. Do not put changing ${VAR} or JMeter function replacement directly inside cached script text: JMeter expands it before the script reaches the engine, so the first replacement can be captured by the cache. Read runtime state through vars, props, Parameters/args, or other supplied bindings instead. JSR223 Sampler exposes log, Label, FileName, Parameters, args, SampleResult, sampler, ctx, vars, props, and OUT; assertions/processors expose the appropriate current/previous-sample bindings. JMeter explicitly recommends migration from BeanShell to JSR223 + Groovy for performance/support, and Groovy is preferred over non-Compilable scripting engines for intensive load paths.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.