Pre-Processors, Post-Processors, Extractors, and Correlation: Guided Hands-On Workflow
The guided workflow uses only synthetic loopback state. Each Start Session response creates a unique JSON token. A JSON JMESPath Extractor stores it in the current thread, a Groovy PreProcessor derives an uppercase form immediately before Use Session, and the fixture validates both values.
Learning objectives
- Start a reproducible loopback service that emits unique synthetic session tokens.
-
Extract
session.tokenintoCORR_TOKENafter Start Session. -
Prepare
PREPARED_TOKENbefore Use Session using a small cached Groovy PreProcessor. - Verify synthetic variables with one-thread Debug Sampler authoring.
- Run two threads and prove each keeps an independent correlated value.
-
Preserve failure samples,
jmeter.log, and redacted/hashed target evidence.
1. Safety envelope
http://127.0.0.1:8000. Maximum 3 threads, maximum 3
loops, no more than 18 measured session requests in any mandatory
run. Tokens are synthetic; the fixture logs only a short SHA-256
fingerprint rather than the token itself. Abort on target mismatch,
unexpected 5xx, or unsafe generator pressure.
2. Start the disposable session fixture
Save as fixtures/correlation_fixture.py:
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import urlparse, parse_qs
from pathlib import Path
import argparse
import hashlib
import json
import threading
import time
lock = threading.Lock()
issued = {}
total = 0
active = 0
max_active = 0
event_log = None
def token_fingerprint(token):
return hashlib.sha256(token.encode('utf-8')).hexdigest()[:12]
def write_event(event):
if event_log is None:
return
with lock:
with event_log.open('a', encoding='utf-8') as handle:
handle.write(json.dumps(event, sort_keys=True) + '\n')
class Handler(BaseHTTPRequestHandler):
protocol_version = 'HTTP/1.1'
def _send_json(self, status, payload, extra_headers=None):
body = json.dumps(payload, sort_keys=True).encode('utf-8')
self.send_response(status)
self.send_header('Content-Type', 'application/json')
self.send_header('Content-Length', str(len(body)))
if extra_headers:
for name, value in extra_headers:
self.send_header(name, value)
self.end_headers()
self.wfile.write(body)
def do_GET(self):
global total, active, max_active
parsed = urlparse(self.path)
path = parsed.path
if path == '/health':
self._send_json(200, {'status': 'ok'})
return
if path == '/stats':
with lock:
snapshot = {
'issued_sessions': len(issued),
'total': total,
'active': active,
'max_active': max_active,
}
self._send_json(200, snapshot)
return
with lock:
total += 1
active += 1
max_active = max(max_active, active)
request_no = total
active_now = active
started_ms = int(time.time() * 1000)
status = 200
event = {'path': path, 'request_no': request_no, 'active_at_start': active_now}
try:
if path == '/session/start':
token = f'tok-{request_no:06d}-{int(time.time_ns()) % 1000000:06d}'
session_id = f'sess-{request_no:06d}'
with lock:
issued[token] = session_id
payload = {
'status': 'issued',
'session': {
'id': session_id,
'token': token,
},
'next': '/session/use',
}
event.update({
'status': status,
'session_id': session_id,
'token_fp': token_fingerprint(token),
})
self._send_json(200, payload, [('X-Synthetic-Session', session_id)])
elif path == '/session/use':
query = {k: v[-1] for k, v in parse_qs(parsed.query).items()}
token = query.get('token', '')
prepared = query.get('prepared', '')
client_thread = query.get('thread', '')
with lock:
session_id = issued.get(token)
expected_prepared = token.upper() if token else ''
if not session_id:
status = 401
payload = {'status': 'rejected', 'reason': 'unknown_token'}
elif prepared != expected_prepared:
status = 400
payload = {'status': 'rejected', 'reason': 'bad_prepared_value'}
else:
payload = {
'status': 'accepted',
'session_id': session_id,
'thread_echo': client_thread,
'token_fingerprint': token_fingerprint(token),
}
event.update({
'status': status,
'session_id': session_id,
'thread_echo': client_thread,
'token_fp': token_fingerprint(token) if token else 'missing',
'prepared_ok': bool(token and prepared == expected_prepared),
})
self._send_json(status, payload)
elif path == '/noise':
payload = {'status': 'noise', 'message': 'No session token is present here'}
event.update({'status': 200})
self._send_json(200, payload)
elif path == '/session/broken':
payload = {'status': 'issued-but-token-renamed', 'session': {'id': f'broken-{request_no}'}, 'sessionToken': 'different-field'}
event.update({'status': 200})
self._send_json(200, payload)
else:
status = 404
event.update({'status': status})
self._send_json(status, {'error': 'not_found', 'path': path})
finally:
event['started_ms'] = started_ms
event['finished_ms'] = int(time.time() * 1000)
write_event(event)
with lock:
active -= 1
def log_message(self, format, *args):
return
if __name__ == '__main__':
parser = argparse.ArgumentParser()
parser.add_argument('--log', default='results/server-events.jsonl')
args = parser.parse_args()
event_log = Path(args.log).resolve()
event_log.parent.mkdir(parents=True, exist_ok=True)
event_log.write_text('', encoding='utf-8')
print('fixture=http://127.0.0.1:8000')
print(f'event_log={event_log}')
ThreadingHTTPServer(('127.0.0.1', 8000), Handler).serve_forever()
Start and preflight:
python fixtures/correlation_fixture.py --log results/server-events.jsonl
curl --fail --silent http://127.0.0.1:8000/health
curl --fail --silent http://127.0.0.1:8000/stats
3. Build the authoring tree
Test Plan — Chapter 09 Correlation Lab
└── Thread Group — start with 1 user × 1 loop
├── HTTP Request Defaults — HttpClient4 / 127.0.0.1:8000
├── Start Session — GET /session/start
│ ├── JSON JMESPath Extractor — CORR_TOKEN
│ │ JMESPath: session.token
│ │ Match No.: 1
│ │ Default: CORR_MISSING
│ └── JSR223 Assertion — extraction guard
├── Debug Sampler — AUTHORING ONLY
└── Use Session — GET /session/use
├── JSR223 PreProcessor — Prepare synthetic token
└── Response / JSON assertion — accepted
Attach the extractor directly under Start Session so it cannot run against Use Session or Debug Sampler responses.
4. Configure the built-in extractor
- Apply to: Main sample only
- Name of created variable:
CORR_TOKEN - JMESPath:
session.token - Match No.:
1 - Default:
CORR_MISSING
Expected Start Session fragment:
{
"status": "issued",
"session": {
"id": "sess-000001",
"token": "tok-000001-123456"
},
"next": "/session/use"
}
5. Fail the producer sample when extraction is missing
Add a tiny JSR223 Assertion under Start Session (Chapter 08 correctness pattern):
def token = vars.get('CORR_TOKEN')
if (token == null || token == 'CORR_MISSING' || token.trim().isEmpty()) {
AssertionResult.setFailure(true)
AssertionResult.setFailureMessage('correlation token missing after Start Session')
}
This keeps the correlation failure attached to the response that was supposed to produce the value.
6. Inspect the variable once with Debug Sampler
Configure Debug Sampler to show JMeter variables. Run 1 thread × 1
loop in GUI and verify CORR_TOKEN=tok-.... Because the
token is synthetic, displaying it in this local authoring step is
acceptable. In real environments, do not dump all variables when
they may contain credentials/session values.
7. Prepare the next request immediately before sampling
Under Use Session add JSR223 PreProcessor, Language: Groovy, with compiled-script caching enabled:
def token = vars.get('CORR_TOKEN')
if (token == null || token == 'CORR_MISSING' || token.trim().isEmpty()) {
vars.put('PREPARED_TOKEN', 'PREP_MISSING')
vars.put('CORRELATION_STATE', 'missing')
} else {
// Synthetic transformation used only to prove the PreProcessor runs before /session/use.
vars.put('PREPARED_TOKEN', token.toUpperCase(Locale.ROOT))
vars.put('CORRELATION_STATE', 'ready')
}
The script reads vars, so it uses the current thread's
token. It does not use shared props.
8. Build the downstream request from thread-local state
Configure Use Session:
GET /session/use?token=${CORR_TOKEN}&prepared=${PREPARED_TOKEN}&thread=${__threadNum}
The local server verifies that the token was previously issued and
that prepared equals the uppercase form. It returns
only a token fingerprint, never the full token:
{
"status": "accepted",
"session_id": "sess-000001",
"thread_echo": "1",
"token_fingerprint": "9d2f4f8a1c20"
}
9. Assert the downstream result
Add JSON JMESPath Assertion under Use Session:
status must equal accepted. This makes a
correlation failure visible as a failed sample even when some
intermediate protocol response is otherwise parseable.
10. Prove per-thread independence
Disable Debug Sampler/View Results Tree. Set Thread Group to 2 threads × 2 loops. Every iteration issues a new token and uses it in the same thread. The server log should show four Start Session events and four Use Session events with four issued session IDs/fingerprints.
Expected invariants:
- each Use Session has
prepared_ok=true; - every token fingerprint belongs to a token issued earlier in the same workflow;
- thread echo values include both 1 and 2;
- no global property carries a token.
11. Run the load copy from CLI
mkdir -p results/two-users
jmeter -n \
-t plans/correlation-load.jmx \
-l results/two-users/results.jtl \
-j results/two-users/jmeter.log \
-Jjmeter.save.saveservice.print_field_names=true \
-Jjmeter.save.saveservice.response_code=true \
-Jjmeter.save.saveservice.response_message=true \
-Jjmeter.save.saveservice.assertion_results_failure_message=true
python tools/analyze_correlation.py results/two-users/results.jtl
curl --fail --silent http://127.0.0.1:8000/stats
PowerShell uses jmeter.bat and backtick continuation.
Keep token variables out of JTL/sample-variable settings; use the
synthetic server fingerprint log for correlation evidence.
12. Minimal JTL analyzer
import csv
import sys
from collections import Counter
from pathlib import Path
path = Path(sys.argv[1] if len(sys.argv) > 1 else 'results/run/results.jtl')
rows = list(csv.DictReader(path.open(encoding='utf-8')))
if not rows:
raise SystemExit('No sample rows found')
required = {'label', 'success', 'responseCode', 'elapsed', 'timeStamp'}
missing = required.difference(rows[0])
if missing:
raise SystemExit(f'Missing JTL fields: {sorted(missing)}')
print(f'samples={len(rows)}')
print(f'failures={sum(r["success"].lower() != "true" for r in rows)}')
print(f'labels={dict(Counter(r["label"] for r in rows))}')
print(f'response_codes={dict(Counter(r["responseCode"] for r in rows))}')
for row in rows:
if row['success'].lower() != 'true':
print(
'failure '
f'label={row["label"]!r} code={row["responseCode"]} '
f'message={row.get("responseMessage", "")!r} '
f'assertion={row.get("failureMessage", "")!r}'
)
13. Scope experiment: move the extractor broad, then inspect
Copy the debug plan only. Move JSON JMESPath Extractor from Start
Session to the Thread Group. It is now in scope for Start Session,
Debug Sampler, and Use Session. A later response without
session.token can overwrite
CORR_TOKEN with CORR_MISSING. This is why
response-owned extractors should normally be children of their
producer sampler.
Do not use this broad version for the main load run.
14. Challenge: cookie or extractor?
The real application sets a standard session cookie and all later HTTP requests should automatically send it. Should you write a regex extractor plus Header Manager?
Normally no. Use HTTP Cookie Manager for ordinary cookie semantics. Correlate manually only when the value must be read/transformed outside normal cookie behavior or when another protocol field requires it.
Knowledge check
Why is the extractor a child of Start Session?
It restricts Post-Processor scope to the response that actually owns session.token, preventing later samples from overwriting the variable.
What proves the PreProcessor runs before Use Session?
The fixture accepts the request only when prepared equals the uppercase form of the just-extracted token.
How do two threads avoid sharing CORR_TOKEN?
The extractor writes a JMeter variable in each thread’s variable map; the PreProcessor reads vars from the same thread.
Why is token fingerprint logging safer than logging the full token?
It allows correlation evidence without persisting the bearer/session value itself; real systems should minimize sensitive correlated data in artifacts.
When should Cookie Manager replace manual correlation?
When the application uses standard HTTP cookie semantics and later HTTP requests simply need the cookie carried automatically.
Official references and version notes
- Elements of a Test Plan — Pre-Processor/Post-Processor purpose, scope, and execution order: configuration → pre-processors → timers → sampler → post-processors → assertions → listeners.
- Component Reference — Regular Expression, JSON/JMESPath, Boundary, JSR223 Pre/Post Processor, User Parameters, and Result Status Action Handler semantics.
- Functions and Variables — thread-local JMeter variables versus process-wide JMeter properties.
- Best Practices — CLI load execution and scripting/performance guidance.
- Apache JMeter downloads — current stable release and Java requirement.
Version-sensitive behavior was rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a Java 17 JDK
for labs and no third-party plugins; JMeter 5.6.3 requires Java
8+. Pre-Processors execute before their in-scope sampler and
Post-Processors execute after the sampler but before Assertions.
Processor behavior is scope-driven rather than determined by
visual sibling order. Built-in Post-Processor extractors store
results in JMeter variables, which are normally thread-local. JSON
JMESPath Extractor accepts one JMESPath expression, can select a
match, and can set an explicit default when nothing matches.
Regular Expression and Boundary Extractors likewise support
explicit defaults, which are especially useful during debugging so
a missing extraction is distinguishable from a processor that
never ran. JSR223 Pre/Post Processors provide
vars (thread variables), props (shared
JMeter properties), and the relevant sampler/result context;
Groovy with compiled-script caching is preferred over BeanShell
when scripting is actually necessary. Mandatory labs use a
built-in extractor for correlation and only a tiny Groovy
PreProcessor for a deliberately simple transformation; Chapter 10
covers extractor families in greater depth.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.