Scripting with JSR223 and Groovy for Advanced Test Logic: Guided Hands-On Workflow
The workflow deliberately starts with a plan that is correct but over-scripted: Groovy canonicalizes input, parses JSON, and validates the response with compilation caching disabled. We then move extraction/validation back to built-in JMeter elements and keep only the one transformation that benefits from a tiny Groovy helper. The target, data, threads, pacing, and result policy remain constant so the comparison is meaningful.
Learning objectives
- Run a bounded localhost JSON fixture and synthetic CSV input.
- Inspect JMeter/Java/Groovy/cache state before load.
- Build an over-scripted but behaviorally correct JSR223 variant.
- Move parsing/validation to JSON JMESPath and Response Assertions.
- Move only canonicalization to one Groovy script file and unit-test the same file outside JMeter.
- Compare JTL throughput plus generator CPU/heap/log evidence without assuming the refactor is faster.
1. Safety envelope
http://127.0.0.1:8019. Synthetic
names only, no credentials/environment-secret access, maximum 2
threads, benchmark maximum 300 loops/thread (600 HTTP
samples/variant), 10 ms Constant Timer, ≤15-second duration cap, no
plugins/containers/remote engines. Abort on unexpected target,
script exception storm, >600 work requests/variant, repeated
failures, or sustained generator saturation.
2. Create the disposable JSON fixture
Save as fixtures/groovy_fixture.py:
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.parse import urlparse, parse_qs
import argparse
import json
import re
import threading
import time
FIXTURE_VERSION = "prompt19-groovy-fixture-v1"
SAFE_NAME = re.compile(r"^[a-z0-9_-]{1,32}$")
SAFE_TOKEN = re.compile(r"^[A-Za-z0-9_.-]{1,64}$")
lock = threading.Lock()
event_log = None
metrics = {
"requests": 0,
"errors": 0,
"by_variant": {},
"by_name": {},
}
def now_ms():
return int(time.time() * 1000)
def signature(name, score):
return f"{name.upper()}:{score:03d}"
def log_event(event):
if event_log is None:
return
with lock:
with event_log.open("a", encoding="utf-8") as handle:
handle.write(json.dumps(event, sort_keys=True) + "\n")
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def send_json(self, status, payload):
raw = json.dumps(payload, sort_keys=True).encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.send_header("X-Fixture-Version", FIXTURE_VERSION)
self.end_headers()
self.wfile.write(raw)
def record(self, started, operation, status, **extra):
with lock:
metrics["requests"] += 1
if status >= 400:
metrics["errors"] += 1
variant = extra.get("variant", "")
name = extra.get("name", "")
if variant:
metrics["by_variant"][variant] = metrics["by_variant"].get(variant, 0) + 1
if name:
metrics["by_name"][name] = metrics["by_name"].get(name, 0) + 1
event = {
"ts_ms": now_ms(),
"operation": operation,
"status": status,
"service_wall_ms": now_ms() - started,
}
event.update(extra)
log_event(event)
def do_GET(self):
started = now_ms()
parsed = urlparse(self.path)
if parsed.path == "/health":
self.send_json(200, {
"status": "ok",
"fixture_version": FIXTURE_VERSION,
})
self.record(started, "health", 200)
return
if parsed.path == "/stats":
with lock:
snapshot = json.loads(json.dumps(metrics))
self.send_json(200, {
"fixture_version": FIXTURE_VERSION,
"metrics": snapshot,
})
self.record(started, "stats", 200)
return
if parsed.path != "/work":
self.send_json(404, {"status": "not_found"})
self.record(started, "unknown", 404)
return
q = parse_qs(parsed.query)
name = q.get("name", [""])[0]
thread_id = q.get("thread", [""])[0]
seq_raw = q.get("seq", [""])[0]
run_id = q.get("run_id", [""])[0]
variant = q.get("variant", [""])[0]
if not SAFE_NAME.fullmatch(name):
self.send_json(400, {"status": "invalid_name"})
self.record(started, "work", 400, name=name, variant=variant)
return
if not SAFE_TOKEN.fullmatch(thread_id) or not SAFE_TOKEN.fullmatch(run_id) or not SAFE_TOKEN.fullmatch(variant):
self.send_json(400, {"status": "invalid_metadata"})
self.record(started, "work", 400, name=name, variant=variant)
return
try:
seq = int(seq_raw)
except ValueError:
seq = -1
if not 0 <= seq <= 1_000_000:
self.send_json(400, {"status": "invalid_seq"})
self.record(started, "work", 400, name=name, variant=variant)
return
score = (sum(name.encode("ascii")) + seq) % 1000
response = {
"status": "ok",
"name": name,
"thread": thread_id,
"seq": seq,
"score": score,
"signature": signature(name, score),
"run_id": run_id,
"variant": variant,
}
self.send_json(200, response)
self.record(
started,
"work",
200,
name=name,
thread=thread_id,
seq=seq,
run_id=run_id,
variant=variant,
signature=response["signature"],
)
def log_message(self, format, *args):
return
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--host", default="127.0.0.1")
parser.add_argument("--port", type=int, default=8019)
parser.add_argument("--log", default="results/server-events.jsonl")
args = parser.parse_args()
global event_log
event_log = Path(args.log).resolve()
event_log.parent.mkdir(parents=True, exist_ok=True)
event_log.write_text("", encoding="utf-8")
print(f"fixture_version={FIXTURE_VERSION}")
print(f"listen=http://{args.host}:{args.port}")
print(f"event_log={event_log}")
ThreadingHTTPServer((args.host, args.port), Handler).serve_forever()
if __name__ == "__main__":
main()
Start:
python .\fixtures\groovy_fixture.py `
--host 127.0.0.1 `
--port 8019 `
--log .\results\server-events.jsonl
Bash uses the same arguments with backslash continuation. The fixture has no mutable business state; it records only synthetic name/thread/sequence/run/variant metadata and service wall time.
3. Read-only preflight
curl --fail --silent http://127.0.0.1:8019/health
curl --fail --silent http://127.0.0.1:8019/stats
Confirm the fixture version and zero work requests. Record JMeter/Java versions and Groovy version:
PowerShell:
& "$env:JMETER_HOME\bin\jmeter.bat" -v
java -version
java -cp "$env:JMETER_HOME\lib\*" groovy.ui.GroovyMain -e "println GroovySystem.version"
Select-String -Path "$env:JMETER_HOME\bin\jmeter.properties" `
-Pattern "jsr223.compiled_scripts_cache_size"
If the property is not explicitly overridden, use the documented default cache size 100. The 5.6.3 baseline should report Groovy 3.0.20.
4. Create synthetic input
data/names.csv:
RAW_NAME
Alice
BOB_2
carol-3
delta
CSV Data Set Config: header from first row, Recycle=true, Stop
Thread=false, All threads. RAW_NAME becomes a
thread-local JMeter variable.
5. Common plan state for both variants
-
HTTP Request Defaults →
127.0.0.1:8019, connect 500 ms, response 2000 ms. -
Counter → variable
SEQ, start 1, increment 1, track counter independently for each user = true. - Constant Timer → 10 ms.
-
Properties from CLI:
RUN_IDandVARIANT; treat as immutable process-wide configuration. -
HTTP GET
/work?name=${REQUEST_NAME}&thread=T${__threadNum}&seq=${SEQ}&run_id=${__P(RUN_ID)}&variant=${__P(VARIANT)}.
Because REQUEST_NAME varies by user/iteration, it
belongs in vars, not props.
6. Variant A — intentionally over-scripted but correct
JSR223 PreProcessor, Groovy, Cache compiled script unchecked:
// Intentionally over-scripted benchmark variant.
// Cache compiled script is OFF, so ${RAW_NAME} is re-expanded before each interpretation.
def raw = '${RAW_NAME}'
vars.put('REQUEST_NAME', raw.trim().toLowerCase(java.util.Locale.ROOT))
This works because the script is reinterpreted with each newly
expanded ${RAW_NAME}, but repeated parsing/compilation
is unnecessary generator work.
JSR223 PostProcessor, Groovy, cache unchecked:
// Intentionally over-scripted: JSON extraction + validation are both coded manually.
import groovy.json.JsonSlurper
def doc = new JsonSlurper().parseText(prev.getResponseDataAsString())
vars.put('SERVER_NAME', doc.name?.toString() ?: '__MISSING__')
vars.put('SERVER_SCORE', doc.score?.toString() ?: '-1')
vars.put('SERVER_SIGNATURE', doc.signature?.toString() ?: '__MISSING__')
if (doc.status != 'ok' || doc.name != vars.get('REQUEST_NAME')) {
prev.setSuccessful(false)
prev.setResponseMessage('prompt19 scripted validation failed')
}
This manually recreates work already provided by JSON JMESPath extractors/assertions and ordinary Response Assertions.
7. Over-scripted tree
Test Plan
├── HTTP Request Defaults
├── CSV Data Set Config
└── Thread Group — 1 thread × 4 loops authoring, later 2×300 benchmark
├── Counter -> SEQ (per user)
├── Constant Timer 10 ms
└── HTTP Work
├── JSR223 PreProcessor — inline Groovy, cache OFF
└── JSR223 PostProcessor — JsonSlurper + validation, cache OFF
Run 1×4 first. Preserve JTL/jmeter.log and confirm all
four canonical names succeed.
8. Refactor the irreducible transformation to one script file
Save scripts/canonicalize.groovy:
import java.util.Locale
String canonicalize(String raw) {
if (raw == null) {
throw new IllegalArgumentException("RAW_NAME is missing")
}
String value = raw.trim().toLowerCase(Locale.ROOT)
if (!(value ==~ /[a-z0-9_-]{1,32}/)) {
throw new IllegalArgumentException("RAW_NAME is outside the synthetic lab grammar")
}
return value
}
String makeSignature(String canonicalName, int score) {
if (!(canonicalName ==~ /[a-z0-9_-]{1,32}/)) {
throw new IllegalArgumentException("canonicalName is invalid")
}
if (score < 0 || score > 999) {
throw new IllegalArgumentException("score is outside 0..999")
}
return canonicalName.toUpperCase(Locale.ROOT) + ":" + String.format(Locale.ROOT, "%03d", score)
}
// The same file is executable outside JMeter for pure helper tests.
if (!binding.hasVariable("vars")) {
assert canonicalize(" Alice ") == "alice"
assert canonicalize("BOB_2") == "bob_2"
assert makeSignature("alice", 7) == "ALICE:007"
boolean rejected = false
try {
canonicalize("../unsafe")
} catch (IllegalArgumentException expected) {
rejected = true
}
assert rejected
println("PASS: prompt19 canonical helper self-test")
return
}
// JMeter execution path: only thread-local variable state is mutated.
String raw = vars.get("RAW_NAME")
String canonical = canonicalize(raw)
vars.put("REQUEST_NAME", canonical)
The pure functions live at the top. When the file runs outside
JMeter it executes assertions. When JMeter supplies
vars, the same file canonicalizes only the current
thread's RAW_NAME and writes thread-local
REQUEST_NAME. There is no per-sample logging and no
mutable global object.
9. Unit-test the same helper outside JMeter
PowerShell:
java -cp "$env:JMETER_HOME\lib\*" `
groovy.ui.GroovyMain `
.\scripts\canonicalize.groovy
Bash:
java -cp "$JMETER_HOME/lib/*" groovy.ui.GroovyMain scripts/canonicalize.groovy
Required output:
PASS: prompt19 canonical helper self-test. This tests
the pure transformation without threads, HTTP, JTL, or the fixture.
10. Variant B — built-ins plus one cached helper
Replace the inline PreProcessor with JSR223 PreProcessor:
- Language: Groovy.
- Script File:
scripts/canonicalize.groovy. -
No inline script text and no
${RAW_NAME}substitution in source. -
Run JMeter from the lab root so the relative Script File path
resolves from
user.dir.
Replace the scripted JSON PostProcessor with built-in elements:
| Element | Configuration |
|---|---|
| JSON JMESPath Assertion |
Expression status, expected ok.
|
| JSON JMESPath Extractor | SERVER_NAME ← name. |
| JSON JMESPath Extractor | SERVER_SCORE ← score. |
| JSON JMESPath Extractor |
SERVER_SIGNATURE ← signature.
|
| Response Assertion |
Field = JMeter Variable SERVER_NAME; equals
${{REQUEST_NAME}}.
|
Now the JMeter tree exposes extraction/validation, while Groovy performs only canonicalization.
11. vars/props evidence table
| Name | Store | Owner/lifetime | Mutation policy |
|---|---|---|---|
RAW_NAME |
vars | current JMeter thread / current CSV row | CSV updates per iteration. |
REQUEST_NAME |
vars | current thread | cached helper writes each invocation. |
SEQ |
vars | current thread | built-in Counter increments. |
SERVER_NAME |
vars | current thread | JMESPath extractor overwrites after sample. |
RUN_ID |
props | whole JMeter JVM | set once with -J; read-only during run. |
VARIANT |
props | whole JMeter JVM | set once per separate benchmark process; read-only. |
12. Refactored tree
Test Plan
├── HTTP Request Defaults
├── CSV Data Set Config
└── Thread Group
├── Counter -> SEQ (per user)
├── Constant Timer 10 ms
└── HTTP Work
├── JSR223 PreProcessor -> scripts/canonicalize.groovy
├── JSON JMESPath Assertion -> status == ok
├── JSON JMESPath Extractors -> SERVER_NAME / SCORE / SIGNATURE
└── Response Assertion -> SERVER_NAME == ${REQUEST_NAME}
13. Prove behavior unchanged before benchmarking
Run each variant at 1 thread ×4 loops with separate RUN_ID/VARIANT. Require:
- 4 successful work samples per variant;
- same canonical-name set: alice, bob_2, carol-3, delta;
- server status=200 for every request;
- no unexpected Groovy exception in
jmeter.log; - same response schema and assertions.
Only after semantic equivalence is proven should you compare overhead.
14. Bounded benchmark
Use identical 2 threads ×300 loops, 10 ms timer, duration cap 15 seconds, same fixture/data/result settings. Run the variants separately:
jmeter.bat -n `
-t plans\over-scripted.jmx `
-l results\over-scripted\results.jtl `
-j results\over-scripted\jmeter.log `
-JRUN_ID=p19-over `
-JVARIANT=over-scripted
jmeter.bat -n `
-t plans\refactored.jmx `
-l results\refactored\results.jtl `
-j results\refactored\jmeter.log `
-JRUN_ID=p19-ref `
-JVARIANT=refactored
Do not assume the percentage improvement. Measure it on your machine.
15. Record generator CPU/heap while each run is active
On Windows, resolve the JMeter Java PID by command line rather than choosing an arbitrary Java process:
Get-CimInstance Win32_Process |
Where-Object { $_.Name -eq "java.exe" -and $_.CommandLine -like "*ApacheJMeter.jar*" } |
Select-Object ProcessId, CommandLine
Then sample:
Get-Process -Id <PID> |
Select-Object Id, CPU, WorkingSet64, PrivateMemorySize64
jcmd <PID> GC.heap_info
Linux/macOS alternatives include
ps/pidstat/top plus
jcmd <PID> GC.heap_info. Record several snapshots
or use your normal local process monitor; do not change heap/tuning
during the comparison.
16. Compare achieved throughput and JTL timing
Save tools/compare_jtl.py:
import csv
import math
import sys
from pathlib import Path
def percentile(values, pct):
data = sorted(values)
if not data:
return 0
idx = max(0, min(len(data)-1, math.ceil((pct/100) * len(data)) - 1))
return data[idx]
def load(path):
path = Path(path)
rows = list(csv.DictReader(path.open(newline="", encoding="utf-8")))
if not rows:
raise SystemExit(f"No JTL rows in {path}")
elapsed = [int(float(r["elapsed"])) for r in rows]
start = min(int(r["timeStamp"]) for r in rows)
end = max(int(r["timeStamp"]) + int(float(r["elapsed"])) for r in rows)
duration_s = max((end - start) / 1000.0, 0.001)
failures = sum(r["success"].lower() != "true" for r in rows)
return {
"file": str(path),
"samples": len(rows),
"failures": failures,
"duration_s": duration_s,
"throughput_s": len(rows) / duration_s,
"p50_ms": percentile(elapsed, 50),
"p95_ms": percentile(elapsed, 95),
"jtl_bytes": path.stat().st_size,
}
if len(sys.argv) != 3:
raise SystemExit("usage: compare_jtl.py over-scripted.jtl refactored.jtl")
for result in (load(sys.argv[1]), load(sys.argv[2])):
print(
f"{result['file']}: samples={result['samples']} failures={result['failures']} "
f"duration_s={result['duration_s']:.3f} throughput_s={result['throughput_s']:.2f} "
f"p50_ms={result['p50_ms']} p95_ms={result['p95_ms']} jtl_bytes={result['jtl_bytes']}"
)
python tools/compare_jtl.py results/over-scripted/results.jtl results/refactored/results.jtl
Sampler p50/p95 mainly represent HTTP timing; JSR223 Pre/PostProcessor work can also appear as reduced overall achieved throughput/generator CPU rather than as target service time. This is why throughput and generator state are required.
17. Verify target workload equivalence
Save tools/analyze_server_events.py:
import json
import sys
from collections import Counter
from pathlib import Path
path = Path(sys.argv[1])
events = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
work = [e for e in events if e.get("operation") == "work"]
print(f"events={len(events)}")
print(f"work_events={len(work)}")
print(f"statuses={dict(Counter(e.get('status') for e in work))}")
print(f"variants={dict(Counter(e.get('variant') for e in work))}")
print(f"names={dict(Counter(e.get('name') for e in work))}")
print(f"threads={dict(Counter(e.get('thread') for e in work))}")
print(f"unique_seq={len(set((e.get('variant'), e.get('thread'), e.get('seq')) for e in work))}")
print(f"max_service_wall_ms={max([e.get('service_wall_ms',0) for e in work] or [0])}")
Use separate event logs or filter by variant. The target should observe the same number and distribution of successful work requests for completed benchmark runs. If variants send different work, the CPU/throughput comparison is invalid.
18. Preserve one controlled exception sample
In a separate 1-thread ×1 JSR223 Sampler (no HTTP target), use:
throw new IllegalStateException("prompt19 synthetic exception - diagnostic only")
Run once, preserve the failed JTL row and
jmeter.log stack trace, then disable/remove the
element. Do not swallow the exception or mark it successful. This
provides a known-good diagnostic pattern for script failures.
19. Challenge
A team has a Groovy PostProcessor that parses JSON, extracts five fields, compares status, and increments a JVM-global property counter. What should you keep in Groovy?
Use JMESPath extractors/assertions for JSON and a built-in Counter/normal metrics for counting. Keep Groovy only if one genuinely custom transformation remains. Per-user counters belong in thread-local vars; cross-thread metrics should use purpose-built result/telemetry mechanisms rather than ad-hoc mutable properties.
Knowledge check
Why is the over-scripted plan cache-off rather than cache-on with ${RAW_NAME}?
It keeps changing values correct for the deliberately inefficient baseline; cache-on with direct substitution can capture the first expanded value.
What proves the refactor preserved behavior?
Both variants send the same canonical-name/request distribution and pass the same target/status/name correctness checks.
Why unit-test canonicalize.groovy outside JMeter?
Pure transformation correctness can be tested without threads, target traffic, JTL, or JMeter execution complexity.
Why are RUN_ID and VARIANT properties but REQUEST_NAME a variable?
RUN_ID/VARIANT are immutable process-wide run settings; REQUEST_NAME changes per virtual user/iteration and must remain thread-local.
Why isn't lower HTTP p95 alone proof that the refactored script costs less CPU?
Sampler elapsed is a different boundary; compare achieved throughput plus generator CPU/heap and stable target service evidence.
Official references and version notes
- JMeter Component Reference — JSR223 Sampler — compilation caching, bindings, script-file paths, SampleResult behavior, and Groovy guidance.
- JMeter Component Reference — JSR223 PreProcessor — scoped pre-sample scripting and bindings.
- JMeter Component Reference — JSR223 PostProcessor — post-sample scripting and previous-sample access.
- JMeter Component Reference — JSR223 Assertion — scripted assertion scope and bindings.
- JMeter Functions — __groovy — Groovy function bindings and cache-safe variable access.
-
JMeter Properties Reference
—
jsr223.compiled_scripts_cache_sizeand advanced Groovy/JSR223 configuration. - JMeter Best Practices — JSR223/Groovy recommendations, compiled scripting, lean results, and non-GUI load execution.
- Apache JMeter downloads — current stable release and Java requirement.
- Apache JMeter issue #6402 — JMeter 5.6.3 stack traces identify bundled Groovy 3.0.20 and document newer-JDK compatibility context.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a
Java 17 JDK; JMeter 5.6.3 requires Java 8+. The
JMeter 5.6.3 binary line uses Groovy 3.0.20.
Groovy's JSR223 engine implements Compilable. JMeter
recommends script files (compiled/cached when supported) or inline
script text with
Cache compiled script if available checked. The
compiled-script cache defaults to
jsr223.compiled_scripts_cache_size=100. Do not put
changing ${VAR} or JMeter function replacement
directly inside cached script text: JMeter expands it before the
script reaches the engine, so the first replacement can be
captured by the cache. Read runtime state through
vars, props,
Parameters/args, or other supplied
bindings instead. JSR223 Sampler exposes log,
Label, FileName,
Parameters, args,
SampleResult, sampler, ctx,
vars, props, and OUT;
assertions/processors expose the appropriate
current/previous-sample bindings. JMeter explicitly recommends
migration from BeanShell to JSR223 + Groovy for
performance/support, and Groovy is preferred over non-Compilable
scripting engines for intensive load paths.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.