Checkpoint Lab — Regular Expressions, CSS/JQuery, XPath, JSONPath, and Boundary Extraction
The checkpoint tests selector judgment rather than memorization. Four response formats expose synthetic primary IDs; you must select a format-aware extractor, verify each downstream request, quantify small-scale generator overhead, and repair one intentionally brittle selector after XML structure evolves.
Learning objectives
- Extract required primary IDs from JSON, XML, HTML, and text using appropriate built-in components.
- Use both Boundary and Regex Extractors on the same plain-text primary ID and compare maintainability.
- Predict and verify match/default/downstream behavior before running.
- Measure each extractor against its own source-only baseline with small local workloads.
- Break an absolute XPath through a harmless wrapper change and restore the intended selector contract.
- Produce a complete evidence packet without logging real credentials or using paid/remote infrastructure.
1. Assumptions and hard ceilings
| Item | Checkpoint baseline |
|---|---|
| JMeter | Apache JMeter 5.6.3. |
| Java | Java 17 JDK lab baseline; JMeter 5.6.3 requires Java 8+. |
| Plugins | None; all extractors are JMeter core components. |
| Target | http://127.0.0.1:8000 only. |
| Functional profile | 1 thread × 2 loops; source + downstream verification requests. |
| Multi-user proof | Optional 2 threads × 2 loops only. |
| Overhead profiles | 1 thread × 20 loops per source-only/extractor pair. |
| Debug defaults | __NOT_FOUND__ for required values. |
| Debug listeners | View Results Tree/Debug Sampler authoring only. |
2. Start and preflight the fixture
mkdir -p results/checkpoint results/broken-xml results/repaired-xml results/overhead
python fixtures/extraction_fixture.py --log results/checkpoint/server-events.jsonl
curl --fail --silent http://127.0.0.1:8000/health
curl --fail --silent http://127.0.0.1:8000/json
curl --fail --silent http://127.0.0.1:8000/xml
curl --fail --silent http://127.0.0.1:8000/html
curl --fail --silent http://127.0.0.1:8000/text
3. Select the extractor for each response
| Format | Extractor | Selector | Expected primary ID |
|---|---|---|---|
| JSON | JSON Extractor | $.primary.id |
JSON-PRIMARY-100 |
| XML | XPath2 Extractor | //primary/id/text() |
XML-PRIMARY-200 |
| HTML | CSS Selector Extractor |
#primary[data-id] + attribute
data-id
|
HTML-PRIMARY-300 |
| Text | Boundary Extractor | left PRIMARY_ID=, right ; |
TEXT-PRIMARY-400 |
| Text comparison | Regex Extractor |
PRIMARY_ID=([A-Z0-9-]+); → $1$
|
TEXT-PRIMARY-400 |
Set Match Number 1 and default __NOT_FOUND__ for every
required primary extraction.
4. Write predictions before execution
- Each source sampler returns HTTP 200 and creates exactly one expected primary variable.
-
Every
/verifyrequest returns 200 only if the extractor selected the correct ID for its source. -
Changing XML from
/xmlto/xml-v2will break the deliberately absolute XPath/response/primary/id/text()but not//primary/id/text(). - Extractor runs may slightly increase whole-engine wall time or generator CPU while source sampler p50/p95 stays close to the source-only baseline.
6. Run the functional CLI profile
jmeter -n -t plans/checkpoint-extraction.jmx -l results/checkpoint/results.jtl -j results/checkpoint/jmeter.log -Jjmeter.save.saveservice.print_field_names=true -Jjmeter.save.saveservice.thread_counts=true
python tools/analyze_extraction.py results/checkpoint/results.jtl
curl --fail --silent http://127.0.0.1:8000/stats
Expected: no failed verification samples. If a verification returns 400, inspect the extractor variable/sentinel and server event before changing the target.
7. Optional two-thread proof
Run at most 2 threads × 2 loops. Extracted variables remain thread-local, so two users can independently parse/use the same synthetic endpoint structure without sharing a mutable JMeter property.
The fixture IDs are static per format for teaching selector behavior, so thread isolation is proven by variable scope/independent executions rather than uniqueness of the returned ID. Chapter 09 covered unique per-thread issued tokens.
8. Deliberately break XPath after a structure change
Create plans/broken-xml.jmx from the normal plan and
make only two changes:
- XML source path becomes
/xml-v2. -
XPath2 query becomes brittle
/response/primary/id/text().
Run one thread × one loop. Preserve:
XML_ID=__NOT_FOUND__evidence;- failed extraction-success assertion;
- JTL and
jmeter.log; -
server response showing the ID still exists under
<data>.
9. Restore the selector contract
If the wrapper is non-contractual, restore
//primary/id/text() and rerun into
results/repaired-xml/. Expected: extraction and verify
request return to success without target modification.
If exact hierarchy is contractual in your real system, do not loosen the XPath: instead classify the response shape as a target contract regression. The checkpoint requires you to state which interpretation applies.
10. Measure small-scale extractor overhead
For each source format, create a source-only baseline and a one-extractor variant, each 1 thread × 20 loops. Time the command externally and record generator CPU/memory:
# Example pair for JSON
/usr/bin/time -p jmeter -n -t plans/overhead-json-baseline.jmx -l results/overhead/json-baseline.jtl -j results/overhead/json-baseline.log
/usr/bin/time -p jmeter -n -t plans/overhead-json-extractor.jmx -l results/overhead/json-extractor.jtl -j results/overhead/json-extractor.log
python tools/compare_runs.py results/overhead/json-baseline.jtl results/overhead/json-extractor.jtl
On PowerShell, use Measure-Command around each
jmeter.bat -n ... invocation. Repeat the method for
XML, HTML, and Text if desired, always comparing each extractor with
its own identical source baseline.
11. Interpret overhead carefully
Do not claim “JSONPath is faster than XPath2” from these different payloads. The valid comparison is within one format: same source/workload with and without its extractor. If differences are below noise, report that.
Sampler elapsed largely reflects the HTTP operation. Extraction is Post-Processor work, so use whole-run wall time and generator CPU/GC as primary overhead evidence.
12. Required evidence packet
| Artifact | Required content |
|---|---|
| Fixture response samples | Bounded JSON/XML/XML-v2/HTML/text examples. |
| Selector table | Extractor, expression/boundaries, match number, default. |
| Debug variables | Synthetic extracted primary IDs and candidate match counts. |
Functional JTL + jmeter.log |
Source/verify sample success and runtime evidence. |
| Server events/stats | Independent downstream supplied/expected IDs and request counts. |
| Broken XML evidence | Sentinel, failed assertion/JTL, unchanged business ID in changed structure. |
| Repaired XML evidence | Successful extraction/verify after selector correction. |
| Overhead pairs | Baseline/extractor JTL/log + wall time + generator CPU/memory. |
| Validity statement | No cross-format speed ranking or production-capacity claim from tiny local runs. |
13. Verification checklist
- All traffic is loopback:8000.
- JSON uses JSONPath; XML uses XPath2; HTML uses current CSS Selector; text uses Boundary plus regex comparison.
- Required extractors use explicit sentinel defaults.
-
Downstream
/verifyindependently rejects wrong IDs. - Match-all mode is limited to bounded debug/candidate examples.
- Broken absolute XPath evidence is preserved before repair.
-
JTL and matching
jmeter.logexist for meaningful runs. - Generator observation accompanies overhead claims.
14. Validity statement
15. Cleanup
- Stop the Python fixture.
- Keep normal/broken/repaired evidence until review is complete.
- Disable debug listeners and synthetic variable dumps in future load plans.
- Do not copy these synthetic defaults into real credential/session extraction.
- No production target, credential, recorder certificate, remote engine, container, database, paid service, plugin, or system-wide JVM/OS setting was changed.
16. What Chapter 10 adds to the operating model
The performance-testing operating model now has an extraction contract: expected response representation, selector language, scope, match count, default/failure behavior, produced thread variables, downstream consumers, privacy handling, and generator cost are reviewable inputs.
Chapter 11 builds on this by controlling the other side of dynamic data: CSV Data Set Config, parameterization, unique test data, exhaustion/recycling behavior, and safe multi-user data strategy.
Knowledge check
Why does the checkpoint compare each extractor to its own source-only baseline rather than ranking all extractors directly?
The response formats/payloads differ, so cross-format wall-time differences confound parser cost with payload/serialization differences.
What proves an extracted value is correct beyond seeing it in Debug Sampler?
The independent /verify request returns success only when the supplied extracted ID matches the expected ID for that source.
Why is the broken absolute XPath preserved?
It proves the first failure was caused by selector/structure mismatch and prevents the repaired run from erasing causal evidence.
When should //primary/id/text() NOT replace the absolute path?
When exact hierarchy is genuinely part of the contract and the wrapper change should be treated as a target regression.
What is the natural bridge to Chapter 11?
After extracting dynamic response state correctly, the next concern is supplying parameterized, unique, bounded input data to concurrent users.
Official references and version notes
- Component Reference — current Regular Expression, CSS Selector, XPath2/XPath, JSON JMESPath, JSON, and Boundary Extractor semantics.
- Elements of a Test Plan — Post-Processor scope/execution and thread-local JMeter variables.
- Regular Expressions — JMeter regular-expression guidance and extractor examples.
- Best Practices — non-GUI load execution, lean listeners, generator validity, and scripting guidance.
- Apache JMeter downloads — current production release and Java requirement.
Version-sensitive behavior was rechecked against current Apache
JMeter documentation on 2026-09-05. The course baseline remains
Apache JMeter 5.6.3 with a Java 17 JDK for labs
and no third-party plugins; JMeter 5.6.3 requires Java 8+. The
current HTML component is named
CSS Selector Extractor (formerly CSS/JQuery
Extractor) and supports JSoup and Jodd-Lagarto implementations,
with JSoup the default when no implementation is selected. For
HTML, current JMeter documentation recommends CSS Selector
Extractor rather than XPath. Since JMeter 5.0, the documentation
recommends XPath2 Extractor over the legacy XPath
Extractor because of easier namespace handling, better
performance, and XPath 2.0 support.
JSON Extractor uses JSONPath syntax;
JSON JMESPath Extractor is also a current
built-in alternative. Regular Expression and Boundary Extractors
can process text/body/header fields and expose
match-number/default behavior. For several extractors, match
0 selects a random match, a positive number selects
the Nth match, and negative/-1 modes expose all
matches through numbered variables and a match-count variable.
Defaults are useful during debugging, but a silent permissive
default must not be allowed to masquerade as successful
correlation. All meaningful load runs preserve both raw JTL and
matching jmeter.log.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.