Pre-Processors, Post-Processors, Extractors, and Correlation: Configuration, Design Patterns, and Trade-Offs
Correlation design is mostly an ownership decision: where did the dynamic value come from, who owns it, how long should it live, which request consumes it, and what is the cheapest observable mechanism that preserves that causality?
Learning objectives
- Choose extraction source from the actual response contract.
- Prefer dedicated extractors over custom code where possible.
- Keep per-user dynamic state in variables rather than global properties.
- Select explicit missing-value handling instead of silent defaults.
- Choose narrow processor scope that matches causality.
- Relate correlation design to generator cost, privacy, CI stability, and measurement validity.
1. Correlation flexibility does not weaken target controls
http://127.0.0.1:8000.
Never use correlation practice to scrape or replay
authentication/session material from uncontrolled services.
2. Extract from body, header, or cookie?
| Source | Best mechanism | Trade-off |
|---|---|---|
| JSON body | JSON JMESPath/JSON Extractor | Structured, readable, avoids regex when the API contract is JSON. |
| Simple text/HTML body | Boundary or narrow Regular Expression Extractor | Simple but can be brittle if surrounding markup changes. |
| Response header | Header-scoped Regex/Boundary extractor | Useful for IDs/nonces in headers; avoid dumping all headers under load. |
| Standard Set-Cookie | HTTP Cookie Manager | Models browser-like cookie storage/sending; often no manual correlation needed. |
| Complex transformation/custom format | Dedicated extractor first, then JSR223 if required | More power adds code review and generator cost. |
3. Dedicated extractor versus JSR223
Use a built-in extractor when it can state the contract directly. JSR223/Groovy becomes justified for cross-field transformation, cryptographic signing in an authorized synthetic test, custom binary parsing, or logic not expressible with built-ins.
When scripting, use vars.get()/vars.put()
for per-thread data and cache compiled Groovy scripts. Avoid
BeanShell for new high-throughput logic when Groovy/JSR223 is
available.
4. Local variable versus global property
| Value | Variable | Property |
|---|---|---|
| Per-user token/session ID | Correct owner. Thread-local and causally tied to one user. | Wrong for mutable per-user state; shared/racy. |
| Run label/environment name | May be copied for convenience. | Natural process-wide configuration. |
| Response-derived order ID | Correct owner if downstream request belongs to same user. | Do not publish via props just for convenience. |
| Global immutable feature switch | Possible but redundant. | Appropriate run-wide property. |
5. Fail immediately or silently default?
An extractor default such as CORR_MISSING is useful for
diagnostics because it tells you the extractor ran and found no
match. It should not become a value that the test quietly sends for
thousands of requests.
A strong pattern is:
- extract with an explicit sentinel during development;
- assert the producer sample created a non-sentinel value;
- choose a thread error action that skips/stops dependent work if correlation is mandatory;
- preserve the first failure before retrying.
6. Beware stale variables when no default is configured
Some extractors can leave an existing variable unchanged when no match occurs if no default is supplied. That can be intentional when several elements conditionally populate the same variable, but it is dangerous for ordinary session correlation: iteration 2 may reuse iteration 1's token after the producer response changed.
For mandatory session correlation, clear/set a sentinel before extraction or use an extractor default and assert it.
7. Narrow versus broad processor scope
A processor under Start Session is causally obvious. The same processor under Thread Group may run for every sampler and overwrite the variable on responses that do not contain the field. Broad scope is appropriate only when all in-scope responses share the same extraction/preparation contract.
8. PreProcessor scope can also surprise you
A JSR223 PreProcessor under Thread Group runs before each descendant
sampler. If it mutates PREPARED_TOKEN, that mutation
occurs even before Start Session—before a token exists. Attach it to
Use Session if only that request needs the transformation.
9. Main samples and sub-samples are separate populations
Extractors have Apply-to controls because HTTP embedded resources and some controllers can produce sub-samples. Select Main sample only when the token lives in the primary JSON response. Searching main + sub-samples can pick the wrong match when the same marker appears in assets/redirects.
10. Correlation data is often sensitive
Real correlation variables may be bearer credentials, CSRF tokens,
account IDs, payment IDs, or signed URLs. Avoid Debug
Sampler/property dumps, JTL sample-variable capture, response-body
retention, and log.info(token) in load mode. Prefer
redacted fingerprints and synthetic data.
11. Correlation has generator cost
Parsing a small JSON field is usually cheap. Repeated full-document XPath, body-unescaping, large-document extraction, or custom scripts can become injector work. If adding correlation reduces achieved rate, compare generator CPU/GC, sample throughput, and target service time before blaming the SUT.
12. Keep correlation separate from surrounding configuration layers
| Layer | Examples | Do not confuse with |
|---|---|---|
| JMeter correlation | extractor, vars, Pre/Post-Processor scope | JVM heap or OS environment. |
| HTTP session | cookies, response tokens, headers | JMeter global properties. |
| JVM/generator | Groovy engine, CPU, heap, GC | Server token expiry. |
| SUT | session store, token validity, response contract | Extractor match failure caused by wrong scope. |
| CI/container | workspace, CPU quota, secrets injection | Thread-local variable semantics. |
13. Worked scenario: JSON token then signed request
Suppose Login returns JSON token auth.token, and Submit
requires a derived signature. Preferred design:
-
JSON JMESPath Extractor under Login creates
AUTH_TOKEN. -
Assertion under Login rejects
AUTH_TOKEN=AUTH_MISSING. -
Cached Groovy PreProcessor under Submit derives signature into
SIGNATUREif built-in functions cannot. -
Submit reads
${AUTH_TOKEN}/${SIGNATURE}. - Neither value is copied to
propsor logged.
14. Decision table
| Question | Preferred choice | Evidence |
|---|---|---|
| JSON scalar from one response? | JSON JMESPath Extractor child of producer sampler. | Debug one thread; downstream accepted; target fingerprint. |
| Normal session cookie? | Cookie Manager. | Cookie-backed request succeeds per thread. |
| Simple text between stable markers? | Boundary Extractor. | Exact extracted value in synthetic debug. |
| Custom transform required before request? | Cached JSR223/Groovy PreProcessor child of consumer. | Target validates transformed value. |
| Missing token? | Sentinel + assertion/fail policy. | Producer failureMessage + no misleading dependent success. |
| Per-user token sharing? | Never via process-global property. | Two-thread independent fingerprints/state. |
15. Preserve correlation evidence without preserving secrets
For meaningful load runs, preserve JTL plus matching
jmeter.log, JMX/extractor settings, synthetic or
redacted target event evidence, and generator observations. The JTL
tells you which sampler failed; jmeter.log shows
engine/script/extractor diagnostics; target logs confirm the request
actually carried a valid session mapping.
Knowledge check
Why can leaving an old value unchanged after a missing extraction be dangerous?
A later iteration can reuse stale dynamic state and make the test look correlated even though the current response did not produce a valid token.
When is a JSR223 PreProcessor justified?
When the consumer request needs a transformation/custom logic not cleanly expressible with built-in functions/components; basic JSON extraction should remain a dedicated extractor.
Why is a Thread-Group-level extractor usually wrong for one login response?
It runs after every in-scope sampler and can overwrite the correlation variable using unrelated responses.
Why should real correlation tokens not be stored as JTL sample variables?
JTL artifacts are routinely shared/archived and can leak credentials/session material; use synthetic values or redacted evidence.
What is the correct owner for a per-user order ID used in the next request?
A thread-local JMeter variable created from that user’s producer response.
Official references and version notes
- Elements of a Test Plan — Pre-Processor/Post-Processor purpose, scope, and execution order: configuration → pre-processors → timers → sampler → post-processors → assertions → listeners.
- Component Reference — Regular Expression, JSON/JMESPath, Boundary, JSR223 Pre/Post Processor, User Parameters, and Result Status Action Handler semantics.
- Functions and Variables — thread-local JMeter variables versus process-wide JMeter properties.
- Best Practices — CLI load execution and scripting/performance guidance.
- Apache JMeter downloads — current stable release and Java requirement.
Version-sensitive behavior was rechecked against current Apache
JMeter primary documentation on 2026-09-05. The course baseline
remains Apache JMeter 5.6.3 with a Java 17 JDK
for labs and no third-party plugins; JMeter 5.6.3 requires Java
8+. Pre-Processors execute before their in-scope sampler and
Post-Processors execute after the sampler but before Assertions.
Processor behavior is scope-driven rather than determined by
visual sibling order. Built-in Post-Processor extractors store
results in JMeter variables, which are normally thread-local. JSON
JMESPath Extractor accepts one JMESPath expression, can select a
match, and can set an explicit default when nothing matches.
Regular Expression and Boundary Extractors likewise support
explicit defaults, which are especially useful during debugging so
a missing extraction is distinguishable from a processor that
never ran. JSR223 Pre/Post Processors provide
vars (thread variables), props (shared
JMeter properties), and the relevant sampler/result context;
Groovy with compiled-script caching is preferred over BeanShell
when scripting is actually necessary. Mandatory labs use a
built-in extractor for correlation and only a tiny Groovy
PreProcessor for a deliberately simple transformation; Chapter 10
covers extractor families in greater depth.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.