Chapter 11Lesson 05~225 minutes

Checkpoint Lab — CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy

The checkpoint treats data allocation as part of the workload model. Workload Class A has reusable accounts that must remain exclusive to one virtual user. Workload Class B consumes one-time tokens exactly once. You will prove both locally, intentionally break each policy, then prepare disjoint remote-engine shards without starting remote RMI.

CheckpointTwo data classesCollision proofEOF proofShard manifest

Learning objectives

  • Allocate one stable reusable account per concurrent local thread.
  • Consume one-time tokens globally without recycling.
  • Predict row demand and achieved sample count before execution.
  • Trigger and diagnose one account collision and one EOF/recycle failure.
  • Generate a deterministic two-engine shard manifest with disjoint values and SHA-256 hashes.
  • Produce a privacy-safe evidence packet covering data schema, allocation, collisions, EOF, sharding, and cleanup.

1. Assumptions and ceilings

Item Checkpoint baseline
JMeter Apache JMeter 5.6.3.
Java Java 17 JDK lab baseline; JMeter 5.6.3 requires Java 8+.
Plugins None.
Target http://127.0.0.1:8000 only.
Class A 2 threads × 3 loops, each thread reuses one exclusive synthetic account.
Class B 2 threads × 3 loops, 6 unique one-time tokens, All-threads shared cursor.
Failure runs Same ceilings; deliberately wrong sharing/recycle settings only.
Remote Simulated sharding only; no RMI/server processes.
Privacy Synthetic example.invalid identities and fake tokens only.
Abort: external target, >2 threads, >3 loops, non-synthetic data, unexpected errors outside the deliberate collision/EOF experiments, or unsafe generator pressure.

2. Generate data and start a clean fixture

python tools/make_data.py
python fixtures/data_fixture.py --log results/checkpoint/server-events.jsonl
curl --fail --silent http://127.0.0.1:8000/health
curl --fail --silent http://127.0.0.1:8000/reset

Archive shards/manifest.json with the run evidence before any load.

3. Workload Class A — stable reusable account per thread

CSV Data Set Config:

Setting Value
Filename data/accounts-thread-${{__threadNum}}.csv
Header first row; Variable Names blank
Recycle true
Stop Thread false
Sharing Current thread

Sampler:

/account?user_id=${user_id}&thread=${__threadNum}

Each file contains one account. Two threads × three outer loops therefore intentionally reread the same account only within the owning thread.

4. Workload Class B — one-time token per iteration

CSV Data Set Config:

Setting Value
Filename data/one-time.csv
Header first row
Recycle false
Stop Thread true
Sharing All threads

Sampler:

/one-time?token=${token}&thread=${__threadNum}

Two threads × three loops demand exactly six rows. The generated file contains six tokens, so the planned data capacity exactly matches the configured outer iterations.

5. Predict before running

  1. Class A: two unique account IDs, each owned by exactly one JMeter thread across all three loops; zero 409 collisions.
  2. Class B: exactly six distinct tokens, each observed once; zero <EOF> values and zero 409 reuse failures.
  3. All-threads Class B does not guarantee which thread gets OT-0001; only global uniqueness is guaranteed.
  4. If Class B data shrinks to four rows while demand remains six and Stop Thread=true, achieved business samples will be data-limited below the configured six.

6. Run Class A

jmeter -n   -t plans/checkpoint-class-a.jmx   -l results/checkpoint/class-a.jtl   -j results/checkpoint/class-a.log   -Jjmeter.save.saveservice.print_field_names=true   -Jjmeter.save.saveservice.thread_counts=true

python tools/analyze_data_events.py results/checkpoint/server-events.jsonl

Verify /stats shows two account owners and no collisions.

7. Reset and run Class B

curl --fail --silent http://127.0.0.1:8000/reset

jmeter -n   -t plans/checkpoint-class-b.jmx   -l results/checkpoint/class-b.jtl   -j results/checkpoint/class-b.log   -Jjmeter.save.saveservice.print_field_names=true   -Jjmeter.save.saveservice.thread_counts=true

python tools/analyze_data_events.py results/checkpoint/server-events.jsonl

Verify exactly six distinct tokens were accepted. Preserve raw JTL and matching logs for both workload classes.

8. Intentionally break Class A sharing

Create a failure copy:

  • Filename becomes common data/accounts.csv.
  • Sharing remains Current thread.
  • 2 threads × 1 loop.

Both private cursors start at row 1. The fixture should record one account owner and at least one 409 collision from the other thread. Preserve the failed sample/event, then restore thread-specific filenames.

9. Intentionally break Class B EOF/recycle policy

Create a three-row one-time file, keep 2 threads × 3 loops, and set Recycle=true. Six iterations consume the three rows twice.

Expected: three first-use successes followed by reuse/collision failures as rows recycle. Restore the full six-row file and Recycle=false/Stop=true.

10. Optional EOF-sentinel diagnostic

With the same short three-row file, set Recycle=false and Stop=false for one bounded run. After valid rows are exhausted, target evidence should show invalid <EOF> data (possibly URL-encoded on the wire). Preserve it, then return to Stop=true for the valid checkpoint profile.

11. Build the distributed uniqueness plan

The generated manifest should prove:

Engine Deployed runtime path Allowed values Rows
engine-a data/users.csv EA-0001 … EA-0006 6
engine-b data/users.csv EB-0001 … EB-0006 6

Both engines can run the same JMX Filename=data/users.csv. Deployment places a different disjoint file at that path on each engine.

Before a real authorized distributed run, verify on every engine: path exists, row count is expected, SHA-256 equals the manifest, JMeter/Java versions match, and no engine received another engine's shard.

12. Distributed row math

If two engines each run 2 threads × 3 one-time iterations, global demand is:

2 engines × 2 threads/engine × 3 iterations/thread = 12 globally unique rows

Two six-row disjoint shards exactly satisfy that example. If an engine fails to start, achieved global load/data consumption is lower; do not quietly redistribute rows mid-run without an explicit revised workload contract.

13. Required evidence packet

Artifact Required content
CSV schema/samples Synthetic account and one-time-token headers/representative rows.
Sharing configuration Filename, Recycle, Stop Thread, Sharing mode for both classes.
Row math Class A ownership model; Class B six-row global demand.
Per-thread account evidence Target owner map showing one account per thread.
One-time evidence Six distinct accepted tokens and no duplicates in valid run.
Collision failure Current-thread same-file account collision JTL/server event.
Recycle/EOF failure Reused one-time token and/or <EOF> evidence.
Shard manifest Engine mapping, relative path, ranges, row counts, SHA-256.
JTL + jmeter.log Matching raw result/runtime evidence for meaningful runs.
Cleanup record Fixture stopped; synthetic files/artifacts retained or removed per policy.

14. Verification checklist

  • All traffic is loopback and bounded.
  • No real PII/password/token appears in CSV/JTL/logs.
  • Class A shows stable exclusive account ownership per thread.
  • Class B consumes six unique rows without recycling.
  • Intentional account collision and recycle/EOF failure are preserved before repair.
  • Configured versus achieved sample counts are explained when EOF stops threads.
  • Engine shards are disjoint and explicitly mapped by manifest/hash.
  • No live distributed engine/RMI setup was needed.

15. Validity statement

Example: “Using Apache JMeter 5.6.3 with Java 17 against a synthetic loopback fixture, reusable accounts were allocated by thread-specific CSV filenames with Current-thread sharing and deliberate recycling, while one-time tokens used one All-threads shared cursor with Recycle=false/Stop Thread=true. The valid profiles produced stable per-thread account ownership and six unique one-time consumptions. Intentionally using Current-thread sharing against one common account file duplicated row 1 across threads; intentionally recycling an undersized one-time file reused tokens. A separate manifest pre-sharded globally disjoint IDs for two simulated engines because remote engines have independent cursors/filesystems. These results validate data allocation and EOF semantics; they do not establish production capacity or live distributed-engine performance.”

16. Cleanup

  1. Stop the local Python fixture.
  2. Preserve valid/broken evidence until review is complete.
  3. Delete disposable generated datasets if run policy requires it; retain generator script/manifest/hash for provenance.
  4. Never archive real credentials/PII with JTL or CSV artifacts.
  5. No production target, real identity, remote RMI engine, paid platform, plugin, container, database, or system-wide JVM/OS setting was changed.

17. What Chapter 11 adds to the operating model

The performance-testing operating model now has a test-data allocation contract: schema, source provenance, row count, sharing/cursor domain, recycle/EOF policy, per-user ownership, privacy classification, file deployment, distributed shard mapping, and cleanup are reviewed alongside workload and results.

Chapter 12 builds on this data contract with Controllers: Simple, Loop, Transaction, If, While, Switch, and Throughput, allowing the test tree to control business-flow repetition/branching without accidentally advancing or reusing data in the wrong scope.

Knowledge check

Why does Class A intentionally recycle a one-row file?

Why does Class B use All threads rather than Current thread?

What does a 409 in the intentionally broken account run prove?

Why can two remote engines not safely use identical one-time CSV content?

How does Chapter 12 relate to CSV consumption?

Next chapter

Controllers: Simple, Loop, Transaction, If, While, Switch, and Throughput

Chapter 12 uses controllers to structure scenario repetition, branching, transactions, and conditional flow while preserving the data ownership and iteration semantics established here.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK for labs and no third-party plugins; JMeter 5.6.3 requires Java 8+. CSV Data Set Config reads one line into variables at the start of each test iteration. With default All threads sharing, one file cursor is shared across threads in that JMeter instance, but which thread receives which row depends on execution order and can vary. Current thread group opens a separate cursor per Thread Group; Current thread opens a separate cursor per thread; an explicit sharing identifier creates a cursor shared by elements using that identifier. If every thread uses Current thread against the same file, every thread starts its own cursor at row 1 unless the filename itself is partitioned (the current docs explicitly show filenames such as test${__threadNum}.csv). At EOF, Recycle=true restarts the file. With Recycle=false and Stop Thread=false, CSV variables become <EOF> (default value, configurable by csvdataset.eofstring). With Recycle=false and Stop Thread=true, the thread stops at EOF. Relative local filenames are resolved against the active test-plan path; for distributed testing, the CSV must already exist on each server host in the correct relative location. Data files are not automatically copied by the distributed controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.