Chapter 09Lesson 03~150 minutes

Pre-Processors, Post-Processors, Extractors, and Correlation: Configuration, Design Patterns, and Trade-Offs

Correlation design is mostly an ownership decision: where did the dynamic value come from, who owns it, how long should it live, which request consumes it, and what is the cheapest observable mechanism that preserves that causality?

Design patternsBody/header/cookieExtractor vs JSR223Fail-fastScope

Learning objectives

  • Choose extraction source from the actual response contract.
  • Prefer dedicated extractors over custom code where possible.
  • Keep per-user dynamic state in variables rather than global properties.
  • Select explicit missing-value handling instead of silent defaults.
  • Choose narrow processor scope that matches causality.
  • Relate correlation design to generator cost, privacy, CI stability, and measurement validity.

1. Correlation flexibility does not weaken target controls

All executable variants stay on http://127.0.0.1:8000. Never use correlation practice to scrape or replay authentication/session material from uncontrolled services.

2. Extract from body, header, or cookie?

Source Best mechanism Trade-off
JSON body JSON JMESPath/JSON Extractor Structured, readable, avoids regex when the API contract is JSON.
Simple text/HTML body Boundary or narrow Regular Expression Extractor Simple but can be brittle if surrounding markup changes.
Response header Header-scoped Regex/Boundary extractor Useful for IDs/nonces in headers; avoid dumping all headers under load.
Standard Set-Cookie HTTP Cookie Manager Models browser-like cookie storage/sending; often no manual correlation needed.
Complex transformation/custom format Dedicated extractor first, then JSR223 if required More power adds code review and generator cost.

3. Dedicated extractor versus JSR223

Use a built-in extractor when it can state the contract directly. JSR223/Groovy becomes justified for cross-field transformation, cryptographic signing in an authorized synthetic test, custom binary parsing, or logic not expressible with built-ins.

When scripting, use vars.get()/vars.put() for per-thread data and cache compiled Groovy scripts. Avoid BeanShell for new high-throughput logic when Groovy/JSR223 is available.

4. Local variable versus global property

Value Variable Property
Per-user token/session ID Correct owner. Thread-local and causally tied to one user. Wrong for mutable per-user state; shared/racy.
Run label/environment name May be copied for convenience. Natural process-wide configuration.
Response-derived order ID Correct owner if downstream request belongs to same user. Do not publish via props just for convenience.
Global immutable feature switch Possible but redundant. Appropriate run-wide property.

5. Fail immediately or silently default?

An extractor default such as CORR_MISSING is useful for diagnostics because it tells you the extractor ran and found no match. It should not become a value that the test quietly sends for thousands of requests.

A strong pattern is:

  1. extract with an explicit sentinel during development;
  2. assert the producer sample created a non-sentinel value;
  3. choose a thread error action that skips/stops dependent work if correlation is mandatory;
  4. preserve the first failure before retrying.

6. Beware stale variables when no default is configured

Some extractors can leave an existing variable unchanged when no match occurs if no default is supplied. That can be intentional when several elements conditionally populate the same variable, but it is dangerous for ordinary session correlation: iteration 2 may reuse iteration 1's token after the producer response changed.

For mandatory session correlation, clear/set a sentinel before extraction or use an extractor default and assert it.

7. Narrow versus broad processor scope

A processor under Start Session is causally obvious. The same processor under Thread Group may run for every sampler and overwrite the variable on responses that do not contain the field. Broad scope is appropriate only when all in-scope responses share the same extraction/preparation contract.

8. PreProcessor scope can also surprise you

A JSR223 PreProcessor under Thread Group runs before each descendant sampler. If it mutates PREPARED_TOKEN, that mutation occurs even before Start Session—before a token exists. Attach it to Use Session if only that request needs the transformation.

9. Main samples and sub-samples are separate populations

Extractors have Apply-to controls because HTTP embedded resources and some controllers can produce sub-samples. Select Main sample only when the token lives in the primary JSON response. Searching main + sub-samples can pick the wrong match when the same marker appears in assets/redirects.

10. Correlation data is often sensitive

Real correlation variables may be bearer credentials, CSRF tokens, account IDs, payment IDs, or signed URLs. Avoid Debug Sampler/property dumps, JTL sample-variable capture, response-body retention, and log.info(token) in load mode. Prefer redacted fingerprints and synthetic data.

11. Correlation has generator cost

Parsing a small JSON field is usually cheap. Repeated full-document XPath, body-unescaping, large-document extraction, or custom scripts can become injector work. If adding correlation reduces achieved rate, compare generator CPU/GC, sample throughput, and target service time before blaming the SUT.

12. Keep correlation separate from surrounding configuration layers

Layer Examples Do not confuse with
JMeter correlation extractor, vars, Pre/Post-Processor scope JVM heap or OS environment.
HTTP session cookies, response tokens, headers JMeter global properties.
JVM/generator Groovy engine, CPU, heap, GC Server token expiry.
SUT session store, token validity, response contract Extractor match failure caused by wrong scope.
CI/container workspace, CPU quota, secrets injection Thread-local variable semantics.

13. Worked scenario: JSON token then signed request

Suppose Login returns JSON token auth.token, and Submit requires a derived signature. Preferred design:

  1. JSON JMESPath Extractor under Login creates AUTH_TOKEN.
  2. Assertion under Login rejects AUTH_TOKEN=AUTH_MISSING.
  3. Cached Groovy PreProcessor under Submit derives signature into SIGNATURE if built-in functions cannot.
  4. Submit reads ${AUTH_TOKEN} / ${SIGNATURE}.
  5. Neither value is copied to props or logged.

14. Decision table

Question Preferred choice Evidence
JSON scalar from one response? JSON JMESPath Extractor child of producer sampler. Debug one thread; downstream accepted; target fingerprint.
Normal session cookie? Cookie Manager. Cookie-backed request succeeds per thread.
Simple text between stable markers? Boundary Extractor. Exact extracted value in synthetic debug.
Custom transform required before request? Cached JSR223/Groovy PreProcessor child of consumer. Target validates transformed value.
Missing token? Sentinel + assertion/fail policy. Producer failureMessage + no misleading dependent success.
Per-user token sharing? Never via process-global property. Two-thread independent fingerprints/state.

15. Preserve correlation evidence without preserving secrets

For meaningful load runs, preserve JTL plus matching jmeter.log, JMX/extractor settings, synthetic or redacted target event evidence, and generator observations. The JTL tells you which sampler failed; jmeter.log shows engine/script/extractor diagnostics; target logs confirm the request actually carried a valid session mapping.

Knowledge check

Why can leaving an old value unchanged after a missing extraction be dangerous?

When is a JSR223 PreProcessor justified?

Why is a Thread-Group-level extractor usually wrong for one login response?

Why should real correlation tokens not be stored as JTL sample variables?

What is the correct owner for a per-user order ID used in the next request?

Next lesson

Diagnose correlation by preserving causality

Lesson 4 engineers hard-coded replay, wrong-sample extraction, cross-thread properties, unresolved literals, broad scope, and sensitive logging, then repairs each without hiding the first failure.

Official references and version notes

  • Elements of a Test Plan — Pre-Processor/Post-Processor purpose, scope, and execution order: configuration → pre-processors → timers → sampler → post-processors → assertions → listeners.
  • Component Reference — Regular Expression, JSON/JMESPath, Boundary, JSR223 Pre/Post Processor, User Parameters, and Result Status Action Handler semantics.
  • Functions and Variables — thread-local JMeter variables versus process-wide JMeter properties.
  • Best Practices — CLI load execution and scripting/performance guidance.
  • Apache JMeter downloads — current stable release and Java requirement.
Version and compatibility note

Version-sensitive behavior was rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK for labs and no third-party plugins; JMeter 5.6.3 requires Java 8+. Pre-Processors execute before their in-scope sampler and Post-Processors execute after the sampler but before Assertions. Processor behavior is scope-driven rather than determined by visual sibling order. Built-in Post-Processor extractors store results in JMeter variables, which are normally thread-local. JSON JMESPath Extractor accepts one JMESPath expression, can select a match, and can set an explicit default when nothing matches. Regular Expression and Boundary Extractors likewise support explicit defaults, which are especially useful during debugging so a missing extraction is distinguishable from a processor that never ran. JSR223 Pre/Post Processors provide vars (thread variables), props (shared JMeter properties), and the relevant sampler/result context; Groovy with compiled-script caching is preferred over BeanShell when scripting is actually necessary. Mandatory labs use a built-in extractor for correlation and only a tiny Groovy PreProcessor for a deliberately simple transformation; Chapter 10 covers extractor families in greater depth.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.