Chapter 10Lesson 05~220 minutes

Checkpoint Lab — Regular Expressions, CSS/JQuery, XPath, JSONPath, and Boundary Extraction

The checkpoint tests selector judgment rather than memorization. Four response formats expose synthetic primary IDs; you must select a format-aware extractor, verify each downstream request, quantify small-scale generator overhead, and repair one intentionally brittle selector after XML structure evolves.

CheckpointFour formatsMatch evidenceOverheadSelector repair

Learning objectives

  • Extract required primary IDs from JSON, XML, HTML, and text using appropriate built-in components.
  • Use both Boundary and Regex Extractors on the same plain-text primary ID and compare maintainability.
  • Predict and verify match/default/downstream behavior before running.
  • Measure each extractor against its own source-only baseline with small local workloads.
  • Break an absolute XPath through a harmless wrapper change and restore the intended selector contract.
  • Produce a complete evidence packet without logging real credentials or using paid/remote infrastructure.

1. Assumptions and hard ceilings

Item Checkpoint baseline
JMeter Apache JMeter 5.6.3.
Java Java 17 JDK lab baseline; JMeter 5.6.3 requires Java 8+.
Plugins None; all extractors are JMeter core components.
Target http://127.0.0.1:8000 only.
Functional profile 1 thread × 2 loops; source + downstream verification requests.
Multi-user proof Optional 2 threads × 2 loops only.
Overhead profiles 1 thread × 20 loops per source-only/extractor pair.
Debug defaults __NOT_FOUND__ for required values.
Debug listeners View Results Tree/Debug Sampler authoring only.
Abort: external/wrong target, >2 threads, unbounded loops, unexpected response shape outside the deliberate XML-v2 step, unexpected errors, or unsafe generator pressure.

2. Start and preflight the fixture

mkdir -p results/checkpoint results/broken-xml results/repaired-xml results/overhead
python fixtures/extraction_fixture.py --log results/checkpoint/server-events.jsonl

curl --fail --silent http://127.0.0.1:8000/health
curl --fail --silent http://127.0.0.1:8000/json
curl --fail --silent http://127.0.0.1:8000/xml
curl --fail --silent http://127.0.0.1:8000/html
curl --fail --silent http://127.0.0.1:8000/text

3. Select the extractor for each response

Format Extractor Selector Expected primary ID
JSON JSON Extractor $.primary.id JSON-PRIMARY-100
XML XPath2 Extractor //primary/id/text() XML-PRIMARY-200
HTML CSS Selector Extractor #primary[data-id] + attribute data-id HTML-PRIMARY-300
Text Boundary Extractor left PRIMARY_ID=, right ; TEXT-PRIMARY-400
Text comparison Regex Extractor PRIMARY_ID=([A-Z0-9-]+); → $1$ TEXT-PRIMARY-400

Set Match Number 1 and default __NOT_FOUND__ for every required primary extraction.

4. Write predictions before execution

  1. Each source sampler returns HTTP 200 and creates exactly one expected primary variable.
  2. Every /verify request returns 200 only if the extractor selected the correct ID for its source.
  3. Changing XML from /xml to /xml-v2 will break the deliberately absolute XPath /response/primary/id/text() but not //primary/id/text().
  4. Extractor runs may slightly increase whole-engine wall time or generator CPU while source sampler p50/p95 stays close to the source-only baseline.

5. Bounded authoring proof

Run one thread × one loop in GUI mode with Debug Sampler. Verify these thread variables:

JSON_ID          = JSON-PRIMARY-100
XML_ID           = XML-PRIMARY-200
HTML_ID          = HTML-PRIMARY-300
TEXT_BOUNDARY_ID = TEXT-PRIMARY-400
TEXT_REGEX_ID    = TEXT-PRIMARY-400

Also test all-match candidate selectors and record their match counts/numbered variables. Remove/disable Debug Sampler and View Results Tree afterward.

6. Run the functional CLI profile

jmeter -n   -t plans/checkpoint-extraction.jmx   -l results/checkpoint/results.jtl   -j results/checkpoint/jmeter.log   -Jjmeter.save.saveservice.print_field_names=true   -Jjmeter.save.saveservice.thread_counts=true

python tools/analyze_extraction.py results/checkpoint/results.jtl
curl --fail --silent http://127.0.0.1:8000/stats

Expected: no failed verification samples. If a verification returns 400, inspect the extractor variable/sentinel and server event before changing the target.

7. Optional two-thread proof

Run at most 2 threads × 2 loops. Extracted variables remain thread-local, so two users can independently parse/use the same synthetic endpoint structure without sharing a mutable JMeter property.

The fixture IDs are static per format for teaching selector behavior, so thread isolation is proven by variable scope/independent executions rather than uniqueness of the returned ID. Chapter 09 covered unique per-thread issued tokens.

8. Deliberately break XPath after a structure change

Create plans/broken-xml.jmx from the normal plan and make only two changes:

  1. XML source path becomes /xml-v2.
  2. XPath2 query becomes brittle /response/primary/id/text().

Run one thread × one loop. Preserve:

  • XML_ID=__NOT_FOUND__ evidence;
  • failed extraction-success assertion;
  • JTL and jmeter.log;
  • server response showing the ID still exists under <data>.

9. Restore the selector contract

If the wrapper is non-contractual, restore //primary/id/text() and rerun into results/repaired-xml/. Expected: extraction and verify request return to success without target modification.

If exact hierarchy is contractual in your real system, do not loosen the XPath: instead classify the response shape as a target contract regression. The checkpoint requires you to state which interpretation applies.

10. Measure small-scale extractor overhead

For each source format, create a source-only baseline and a one-extractor variant, each 1 thread × 20 loops. Time the command externally and record generator CPU/memory:

# Example pair for JSON
/usr/bin/time -p jmeter -n   -t plans/overhead-json-baseline.jmx   -l results/overhead/json-baseline.jtl   -j results/overhead/json-baseline.log

/usr/bin/time -p jmeter -n   -t plans/overhead-json-extractor.jmx   -l results/overhead/json-extractor.jtl   -j results/overhead/json-extractor.log

python tools/compare_runs.py   results/overhead/json-baseline.jtl   results/overhead/json-extractor.jtl

On PowerShell, use Measure-Command around each jmeter.bat -n ... invocation. Repeat the method for XML, HTML, and Text if desired, always comparing each extractor with its own identical source baseline.

11. Interpret overhead carefully

Do not claim “JSONPath is faster than XPath2” from these different payloads. The valid comparison is within one format: same source/workload with and without its extractor. If differences are below noise, report that.

Sampler elapsed largely reflects the HTTP operation. Extraction is Post-Processor work, so use whole-run wall time and generator CPU/GC as primary overhead evidence.

12. Required evidence packet

Artifact Required content
Fixture response samples Bounded JSON/XML/XML-v2/HTML/text examples.
Selector table Extractor, expression/boundaries, match number, default.
Debug variables Synthetic extracted primary IDs and candidate match counts.
Functional JTL + jmeter.log Source/verify sample success and runtime evidence.
Server events/stats Independent downstream supplied/expected IDs and request counts.
Broken XML evidence Sentinel, failed assertion/JTL, unchanged business ID in changed structure.
Repaired XML evidence Successful extraction/verify after selector correction.
Overhead pairs Baseline/extractor JTL/log + wall time + generator CPU/memory.
Validity statement No cross-format speed ranking or production-capacity claim from tiny local runs.

13. Verification checklist

  • All traffic is loopback:8000.
  • JSON uses JSONPath; XML uses XPath2; HTML uses current CSS Selector; text uses Boundary plus regex comparison.
  • Required extractors use explicit sentinel defaults.
  • Downstream /verify independently rejects wrong IDs.
  • Match-all mode is limited to bounded debug/candidate examples.
  • Broken absolute XPath evidence is preserved before repair.
  • JTL and matching jmeter.log exist for meaningful runs.
  • Generator observation accompanies overhead claims.

14. Validity statement

Example: “Using Apache JMeter 5.6.3 with Java 17 against a loopback-only fixture, JSON Extractor/JSONPath, XPath2 Extractor, CSS Selector Extractor, Boundary Extractor, and Regular Expression Extractor each recovered the intended synthetic primary ID from the matching response representation and downstream verification accepted it. Required values used a visible no-match sentinel. An absolute XPath failed after an intentionally introduced wrapper while a selector aligned with the business structure continued to work. Source-only versus extractor runs were compared with JTL, jmeter.log, wall-clock, and generator observations; the small local measurements describe injector extraction cost only and do not rank parsers generally or establish production capacity.”

15. Cleanup

  1. Stop the Python fixture.
  2. Keep normal/broken/repaired evidence until review is complete.
  3. Disable debug listeners and synthetic variable dumps in future load plans.
  4. Do not copy these synthetic defaults into real credential/session extraction.
  5. No production target, credential, recorder certificate, remote engine, container, database, paid service, plugin, or system-wide JVM/OS setting was changed.

16. What Chapter 10 adds to the operating model

The performance-testing operating model now has an extraction contract: expected response representation, selector language, scope, match count, default/failure behavior, produced thread variables, downstream consumers, privacy handling, and generator cost are reviewable inputs.

Chapter 11 builds on this by controlling the other side of dynamic data: CSV Data Set Config, parameterization, unique test data, exhaustion/recycling behavior, and safe multi-user data strategy.

Knowledge check

Why does the checkpoint compare each extractor to its own source-only baseline rather than ranking all extractors directly?

What proves an extracted value is correct beyond seeing it in Debug Sampler?

Why is the broken absolute XPath preserved?

When should //primary/id/text() NOT replace the absolute path?

What is the natural bridge to Chapter 11?

Next chapter

CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy

Chapter 11 turns static fixture inputs into controlled multi-user data streams and teaches file sharing, recycling, stop conditions, uniqueness, and data provenance.

Official references and version notes

  • Component Reference — current Regular Expression, CSS Selector, XPath2/XPath, JSON JMESPath, JSON, and Boundary Extractor semantics.
  • Elements of a Test Plan — Post-Processor scope/execution and thread-local JMeter variables.
  • Regular Expressions — JMeter regular-expression guidance and extractor examples.
  • Best Practices — non-GUI load execution, lean listeners, generator validity, and scripting guidance.
  • Apache JMeter downloads — current production release and Java requirement.
Version and compatibility note

Version-sensitive behavior was rechecked against current Apache JMeter documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK for labs and no third-party plugins; JMeter 5.6.3 requires Java 8+. The current HTML component is named CSS Selector Extractor (formerly CSS/JQuery Extractor) and supports JSoup and Jodd-Lagarto implementations, with JSoup the default when no implementation is selected. For HTML, current JMeter documentation recommends CSS Selector Extractor rather than XPath. Since JMeter 5.0, the documentation recommends XPath2 Extractor over the legacy XPath Extractor because of easier namespace handling, better performance, and XPath 2.0 support. JSON Extractor uses JSONPath syntax; JSON JMESPath Extractor is also a current built-in alternative. Regular Expression and Boundary Extractors can process text/body/header fields and expose match-number/default behavior. For several extractors, match 0 selects a random match, a positive number selects the Nth match, and negative/-1 modes expose all matches through numbered variables and a match-count variable. Defaults are useful during debugging, but a silent permissive default must not be allowed to masquerade as successful correlation. All meaningful load runs preserve both raw JTL and matching jmeter.log.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.