Chapter 09Lesson 02~205 minutes

Pre-Processors, Post-Processors, Extractors, and Correlation: Guided Hands-On Workflow

The guided workflow uses only synthetic loopback state. Each Start Session response creates a unique JSON token. A JSON JMESPath Extractor stores it in the current thread, a Groovy PreProcessor derives an uppercase form immediately before Use Session, and the fixture validates both values.

JSON JMESPath ExtractorJSR223 PreProcessorDebug SamplerTwo threadsIndependent sessions

Learning objectives

  • Start a reproducible loopback service that emits unique synthetic session tokens.
  • Extract session.token into CORR_TOKEN after Start Session.
  • Prepare PREPARED_TOKEN before Use Session using a small cached Groovy PreProcessor.
  • Verify synthetic variables with one-thread Debug Sampler authoring.
  • Run two threads and prove each keeps an independent correlated value.
  • Preserve failure samples, jmeter.log, and redacted/hashed target evidence.

1. Safety envelope

Target allow-list: only http://127.0.0.1:8000. Maximum 3 threads, maximum 3 loops, no more than 18 measured session requests in any mandatory run. Tokens are synthetic; the fixture logs only a short SHA-256 fingerprint rather than the token itself. Abort on target mismatch, unexpected 5xx, or unsafe generator pressure.

2. Start the disposable session fixture

Save as fixtures/correlation_fixture.py:

from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import urlparse, parse_qs
from pathlib import Path
import argparse
import hashlib
import json
import threading
import time

lock = threading.Lock()
issued = {}
total = 0
active = 0
max_active = 0
event_log = None

def token_fingerprint(token):
    return hashlib.sha256(token.encode('utf-8')).hexdigest()[:12]

def write_event(event):
    if event_log is None:
        return
    with lock:
        with event_log.open('a', encoding='utf-8') as handle:
            handle.write(json.dumps(event, sort_keys=True) + '\n')

class Handler(BaseHTTPRequestHandler):
    protocol_version = 'HTTP/1.1'

    def _send_json(self, status, payload, extra_headers=None):
        body = json.dumps(payload, sort_keys=True).encode('utf-8')
        self.send_response(status)
        self.send_header('Content-Type', 'application/json')
        self.send_header('Content-Length', str(len(body)))
        if extra_headers:
            for name, value in extra_headers:
                self.send_header(name, value)
        self.end_headers()
        self.wfile.write(body)

    def do_GET(self):
        global total, active, max_active
        parsed = urlparse(self.path)
        path = parsed.path

        if path == '/health':
            self._send_json(200, {'status': 'ok'})
            return

        if path == '/stats':
            with lock:
                snapshot = {
                    'issued_sessions': len(issued),
                    'total': total,
                    'active': active,
                    'max_active': max_active,
                }
            self._send_json(200, snapshot)
            return

        with lock:
            total += 1
            active += 1
            max_active = max(max_active, active)
            request_no = total
            active_now = active

        started_ms = int(time.time() * 1000)
        status = 200
        event = {'path': path, 'request_no': request_no, 'active_at_start': active_now}
        try:
            if path == '/session/start':
                token = f'tok-{request_no:06d}-{int(time.time_ns()) % 1000000:06d}'
                session_id = f'sess-{request_no:06d}'
                with lock:
                    issued[token] = session_id
                payload = {
                    'status': 'issued',
                    'session': {
                        'id': session_id,
                        'token': token,
                    },
                    'next': '/session/use',
                }
                event.update({
                    'status': status,
                    'session_id': session_id,
                    'token_fp': token_fingerprint(token),
                })
                self._send_json(200, payload, [('X-Synthetic-Session', session_id)])

            elif path == '/session/use':
                query = {k: v[-1] for k, v in parse_qs(parsed.query).items()}
                token = query.get('token', '')
                prepared = query.get('prepared', '')
                client_thread = query.get('thread', '')
                with lock:
                    session_id = issued.get(token)
                expected_prepared = token.upper() if token else ''
                if not session_id:
                    status = 401
                    payload = {'status': 'rejected', 'reason': 'unknown_token'}
                elif prepared != expected_prepared:
                    status = 400
                    payload = {'status': 'rejected', 'reason': 'bad_prepared_value'}
                else:
                    payload = {
                        'status': 'accepted',
                        'session_id': session_id,
                        'thread_echo': client_thread,
                        'token_fingerprint': token_fingerprint(token),
                    }
                event.update({
                    'status': status,
                    'session_id': session_id,
                    'thread_echo': client_thread,
                    'token_fp': token_fingerprint(token) if token else 'missing',
                    'prepared_ok': bool(token and prepared == expected_prepared),
                })
                self._send_json(status, payload)

            elif path == '/noise':
                payload = {'status': 'noise', 'message': 'No session token is present here'}
                event.update({'status': 200})
                self._send_json(200, payload)

            elif path == '/session/broken':
                payload = {'status': 'issued-but-token-renamed', 'session': {'id': f'broken-{request_no}'}, 'sessionToken': 'different-field'}
                event.update({'status': 200})
                self._send_json(200, payload)

            else:
                status = 404
                event.update({'status': status})
                self._send_json(status, {'error': 'not_found', 'path': path})
        finally:
            event['started_ms'] = started_ms
            event['finished_ms'] = int(time.time() * 1000)
            write_event(event)
            with lock:
                active -= 1

    def log_message(self, format, *args):
        return

if __name__ == '__main__':
    parser = argparse.ArgumentParser()
    parser.add_argument('--log', default='results/server-events.jsonl')
    args = parser.parse_args()
    event_log = Path(args.log).resolve()
    event_log.parent.mkdir(parents=True, exist_ok=True)
    event_log.write_text('', encoding='utf-8')
    print('fixture=http://127.0.0.1:8000')
    print(f'event_log={event_log}')
    ThreadingHTTPServer(('127.0.0.1', 8000), Handler).serve_forever()

Start and preflight:

python fixtures/correlation_fixture.py --log results/server-events.jsonl
curl --fail --silent http://127.0.0.1:8000/health
curl --fail --silent http://127.0.0.1:8000/stats

3. Build the authoring tree

Test Plan — Chapter 09 Correlation Lab
└── Thread Group — start with 1 user × 1 loop
    ├── HTTP Request Defaults — HttpClient4 / 127.0.0.1:8000
    ├── Start Session — GET /session/start
    │   ├── JSON JMESPath Extractor — CORR_TOKEN
    │   │   JMESPath: session.token
    │   │   Match No.: 1
    │   │   Default: CORR_MISSING
    │   └── JSR223 Assertion — extraction guard
    ├── Debug Sampler — AUTHORING ONLY
    └── Use Session — GET /session/use
        ├── JSR223 PreProcessor — Prepare synthetic token
        └── Response / JSON assertion — accepted

Attach the extractor directly under Start Session so it cannot run against Use Session or Debug Sampler responses.

4. Configure the built-in extractor

  • Apply to: Main sample only
  • Name of created variable: CORR_TOKEN
  • JMESPath: session.token
  • Match No.: 1
  • Default: CORR_MISSING

Expected Start Session fragment:

{
  "status": "issued",
  "session": {
    "id": "sess-000001",
    "token": "tok-000001-123456"
  },
  "next": "/session/use"
}

5. Fail the producer sample when extraction is missing

Add a tiny JSR223 Assertion under Start Session (Chapter 08 correctness pattern):

def token = vars.get('CORR_TOKEN')
if (token == null || token == 'CORR_MISSING' || token.trim().isEmpty()) {
    AssertionResult.setFailure(true)
    AssertionResult.setFailureMessage('correlation token missing after Start Session')
}

This keeps the correlation failure attached to the response that was supposed to produce the value.

6. Inspect the variable once with Debug Sampler

Configure Debug Sampler to show JMeter variables. Run 1 thread × 1 loop in GUI and verify CORR_TOKEN=tok-.... Because the token is synthetic, displaying it in this local authoring step is acceptable. In real environments, do not dump all variables when they may contain credentials/session values.

7. Prepare the next request immediately before sampling

Under Use Session add JSR223 PreProcessor, Language: Groovy, with compiled-script caching enabled:

def token = vars.get('CORR_TOKEN')
if (token == null || token == 'CORR_MISSING' || token.trim().isEmpty()) {
    vars.put('PREPARED_TOKEN', 'PREP_MISSING')
    vars.put('CORRELATION_STATE', 'missing')
} else {
    // Synthetic transformation used only to prove the PreProcessor runs before /session/use.
    vars.put('PREPARED_TOKEN', token.toUpperCase(Locale.ROOT))
    vars.put('CORRELATION_STATE', 'ready')
}

The script reads vars, so it uses the current thread's token. It does not use shared props.

8. Build the downstream request from thread-local state

Configure Use Session:

GET /session/use?token=${CORR_TOKEN}&prepared=${PREPARED_TOKEN}&thread=${__threadNum}

The local server verifies that the token was previously issued and that prepared equals the uppercase form. It returns only a token fingerprint, never the full token:

{
  "status": "accepted",
  "session_id": "sess-000001",
  "thread_echo": "1",
  "token_fingerprint": "9d2f4f8a1c20"
}

9. Assert the downstream result

Add JSON JMESPath Assertion under Use Session: status must equal accepted. This makes a correlation failure visible as a failed sample even when some intermediate protocol response is otherwise parseable.

10. Prove per-thread independence

Disable Debug Sampler/View Results Tree. Set Thread Group to 2 threads × 2 loops. Every iteration issues a new token and uses it in the same thread. The server log should show four Start Session events and four Use Session events with four issued session IDs/fingerprints.

Expected invariants:

  • each Use Session has prepared_ok=true;
  • every token fingerprint belongs to a token issued earlier in the same workflow;
  • thread echo values include both 1 and 2;
  • no global property carries a token.

11. Run the load copy from CLI

mkdir -p results/two-users
jmeter -n \
  -t plans/correlation-load.jmx \
  -l results/two-users/results.jtl \
  -j results/two-users/jmeter.log \
  -Jjmeter.save.saveservice.print_field_names=true \
  -Jjmeter.save.saveservice.response_code=true \
  -Jjmeter.save.saveservice.response_message=true \
  -Jjmeter.save.saveservice.assertion_results_failure_message=true
python tools/analyze_correlation.py results/two-users/results.jtl
curl --fail --silent http://127.0.0.1:8000/stats

PowerShell uses jmeter.bat and backtick continuation. Keep token variables out of JTL/sample-variable settings; use the synthetic server fingerprint log for correlation evidence.

12. Minimal JTL analyzer

import csv
import sys
from collections import Counter
from pathlib import Path

path = Path(sys.argv[1] if len(sys.argv) > 1 else 'results/run/results.jtl')
rows = list(csv.DictReader(path.open(encoding='utf-8')))
if not rows:
    raise SystemExit('No sample rows found')
required = {'label', 'success', 'responseCode', 'elapsed', 'timeStamp'}
missing = required.difference(rows[0])
if missing:
    raise SystemExit(f'Missing JTL fields: {sorted(missing)}')

print(f'samples={len(rows)}')
print(f'failures={sum(r["success"].lower() != "true" for r in rows)}')
print(f'labels={dict(Counter(r["label"] for r in rows))}')
print(f'response_codes={dict(Counter(r["responseCode"] for r in rows))}')
for row in rows:
    if row['success'].lower() != 'true':
        print(
            'failure '
            f'label={row["label"]!r} code={row["responseCode"]} '
            f'message={row.get("responseMessage", "")!r} '
            f'assertion={row.get("failureMessage", "")!r}'
        )

13. Scope experiment: move the extractor broad, then inspect

Copy the debug plan only. Move JSON JMESPath Extractor from Start Session to the Thread Group. It is now in scope for Start Session, Debug Sampler, and Use Session. A later response without session.token can overwrite CORR_TOKEN with CORR_MISSING. This is why response-owned extractors should normally be children of their producer sampler.

Do not use this broad version for the main load run.

14. Challenge: cookie or extractor?

The real application sets a standard session cookie and all later HTTP requests should automatically send it. Should you write a regex extractor plus Header Manager?

Normally no. Use HTTP Cookie Manager for ordinary cookie semantics. Correlate manually only when the value must be read/transformed outside normal cookie behavior or when another protocol field requires it.

Knowledge check

Why is the extractor a child of Start Session?

What proves the PreProcessor runs before Use Session?

How do two threads avoid sharing CORR_TOKEN?

Why is token fingerprint logging safer than logging the full token?

When should Cookie Manager replace manual correlation?

Next lesson

Choose the correlation mechanism from the response contract

Lesson 3 compares body/header/cookie extraction, dedicated extractors versus JSR223, variables versus properties, sentinel/fail-fast handling, and narrow versus broad processor scope.

Official references and version notes

  • Elements of a Test Plan — Pre-Processor/Post-Processor purpose, scope, and execution order: configuration → pre-processors → timers → sampler → post-processors → assertions → listeners.
  • Component Reference — Regular Expression, JSON/JMESPath, Boundary, JSR223 Pre/Post Processor, User Parameters, and Result Status Action Handler semantics.
  • Functions and Variables — thread-local JMeter variables versus process-wide JMeter properties.
  • Best Practices — CLI load execution and scripting/performance guidance.
  • Apache JMeter downloads — current stable release and Java requirement.
Version and compatibility note

Version-sensitive behavior was rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK for labs and no third-party plugins; JMeter 5.6.3 requires Java 8+. Pre-Processors execute before their in-scope sampler and Post-Processors execute after the sampler but before Assertions. Processor behavior is scope-driven rather than determined by visual sibling order. Built-in Post-Processor extractors store results in JMeter variables, which are normally thread-local. JSON JMESPath Extractor accepts one JMESPath expression, can select a match, and can set an explicit default when nothing matches. Regular Expression and Boundary Extractors likewise support explicit defaults, which are especially useful during debugging so a missing extraction is distinguishable from a processor that never ran. JSR223 Pre/Post Processors provide vars (thread variables), props (shared JMeter properties), and the relevant sampler/result context; Groovy with compiled-script caching is preferred over BeanShell when scripting is actually necessary. Mandatory labs use a built-in extractor for correlation and only a tiny Groovy PreProcessor for a deliberately simple transformation; Chapter 10 covers extractor families in greater depth.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.