CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy: Configuration, Design Patterns, and Trade-Offs
CSV configuration is not primarily about commas; it is about ownership. The right design follows whether an identity is reusable, exclusive to one virtual user, consumed once per iteration, shared across Thread Groups, unique across engines, or regenerated for every run.
Learning objectives
- Choose a shared CSV or generated data from uniqueness/replay requirements.
- Select All threads, Current thread group, Current thread, or Identifier from allocation scope.
- Choose recycle/stop behavior from whether records are reusable or consumable.
- Compare pre-sharded uniqueness with runtime-generated uniqueness.
- Choose committed fake fixtures or ephemeral generated datasets from governance/privacy needs.
- Connect data design to generator cost, target safety, distributed execution, and measurement validity.
1. Mandatory path remains local and synthetic
http://127.0.0.1:8000,
use fake identifiers, and remain ≤2 threads.
Remote/cloud/CI designs are discussed as deployment patterns only;
no paid service or live RMI setup is required.
4. Recycle versus Stop Thread at EOF
Make reuse policy explicit:
- Reusable account: recycle can be acceptable if the same owner is allowed to reuse it.
- One-time coupon/registration ID: Recycle=false; Stop Thread=true is safer when data exhaustion should end activity.
- Data shortage is a test setup error: preflight row count and fail before load rather than discovering EOF mid-run.
<EOF> is valuable diagnostic evidence but should
not be the normal production workload path.
5. Pre-sharded versus runtime-generated uniqueness
Pre-sharding makes engine ownership obvious and debuggable: shard A contains one range, shard B another, and hashes prove deployment. Runtime-generated UUIDs can avoid CSV exhaustion but make exact replay and collision audit harder and consume generator work.
If runtime uniqueness is used, include an engine/run prefix so independent generators cannot accidentally generate overlapping business keys under a weak custom scheme.
6. Committed fake fixtures versus ephemeral data
| Choice | Strength | Governance caution |
|---|---|---|
| Committed fake CSV | Reviewable schema, deterministic examples, easy local onboarding. | Must be unmistakably synthetic and free of secrets/PII. |
| Generated ephemeral CSV | Can size exactly to run demand; easy to shard per run. | Generator script/version/seed/output hash becomes part of evidence. |
| Real anonymized export | May resemble production distribution. | Still carries privacy/re-identification/legal risk; not required for this course. |
For the Academy, committed/generated synthetic data is sufficient. Real customer records are never a mandatory learning dependency.
7. Stable per-user identity versus iteration-driven CSV consumption
Because CSV rows update at each test iteration, a Thread Group loop can unintentionally rotate accounts. If each virtual user should keep one account while repeating a business action many times, allocate the account once per user and place the repeated action inside a nested Loop Controller. Chapter 12 will deepen this controller pattern.
8. Row-demand arithmetic must match the sharing domain
For one-time rows:
All threads shared cursor:
rows >= sum(iterations across all threads/groups sharing that cursor)
Current thread cursor:
rows in each thread's file >= that thread's outer iterations
Current thread group:
rows per group cursor >= iterations consumed by that group
Distributed:
global rows >= sum(rows consumed by every engine)
and shard sets must be disjoint
Include warm-up, retries only if explicitly modeled, setup calls that consume data, and repeated CI runs in capacity planning.
9. Relative paths versus absolute paths
Use repository/test-plan-relative paths for portability. Absolute paths can work on one workstation but fail on CI/remote engines. In distributed mode, provision the same relative path on every server host but place the correct engine-specific shard there.
10. Deterministic shard mapping versus filesystem ordering
Do not infer engine shard by “sort/list all CSVs and take index N” unless the ordering contract is explicitly guaranteed and tested. Directory enumeration can differ across OS/filesystems/tool versions.
Prefer an explicit manifest:
engine-a -> shards/engine-a/data/users.csv -> rows 1..6 -> sha256 ...
engine-b -> shards/engine-b/data/users.csv -> rows 1..6 -> sha256 ...
11. Privacy and credential boundaries
CSV is plain text. A password/API token stored there is readable by anyone with filesystem/artifact access and may later appear in requests/debug output. If credentials are unavoidable in a real authorized environment, use the platform's secret mechanism and minimize persistence; do not commit them.
Use synthetic emails such as acct-001@example.invalid,
synthetic tokens, and fake IDs in training/repository fixtures.
12. Data strategy affects performance validity
Colliding users may trigger target locks, deduplication, cache locality, rate limits, or error branches. Recycled one-time IDs may artificially increase 4xx/409 rates. Runtime generation may consume injector CPU. Missing remote shards may stop threads. None of those should be misreported as server capacity behavior without attribution.
For every meaningful run, preserve the raw JTL together with its
matching jmeter.log, the exact CSV/shard hash, and the
non-secret run manifest. JTL shows achieved samples and failures;
jmeter.log preserves engine-side file, EOF, and runtime
diagnostics needed to distinguish data exhaustion from target
failure.
13. Keep data allocation separate from surrounding layers
| Layer | Examples | Not a substitute for |
|---|---|---|
| JMeter CSV config | filename, columns, recycle, stop, sharing cursor | Creating valid target business data. |
| Generator filesystem | file presence, permissions, encoding | JMeter remote file distribution. |
| Target/SUT | account uniqueness, one-time semantics, cleanup | CSV cursor correctness. |
| CI/container | workspace, mounted files, artifact retention | Privacy review or shard math. |
| Distributed engine | engine-local file/cursor | Controller-local CSV state. |
| JVM | heap/GC/file IO | Fixing duplicate rows or bad allocation. |
14. Decision table
| Requirement | Preferred data strategy | Reason |
|---|---|---|
| Each iteration needs a fresh one-time ID | All threads shared cursor, Recycle=false, Stop=true | One process consumes each row once. |
| Each thread owns one reusable account | Thread-specific file + Current thread, or one-time allocation outside inner business loop | Stable per-user ownership. |
| Two Thread Groups share one global pool | All threads or explicit shared Identifier | Avoid duplicate group-local cursors. |
| Two remote engines need globally unique data | Disjoint pre-shards deployed per engine | Each engine has independent cursor/process. |
| Large synthetic unique pool | Pre-generate CSV | Lower runtime CPU/memory and better auditability. |
| Tiny run-specific suffix | Derived runtime value may be reasonable | Low cost; record run input. |
Knowledge check
Why can a Thread Group loop rotate accounts unexpectedly?
CSV variables are refreshed at the start of each test iteration, so the cursor advances unless the design isolates a stable per-thread file/value.
When should Recycle=false + Stop Thread=true be preferred?
When records are consumable/one-time and the workload must stop
rather than reuse or send
Why are remote shards needed even with All threads sharing?
All threads sharing is per JMeter process; each remote engine has its own cursor and would otherwise start the same file independently.
Why is an explicit shard manifest better than filesystem ordering?
It deterministically maps engines to paths/ranges/hashes and avoids OS/filesystem enumeration differences.
How can test data invalidate a capacity claim?
Collisions, recycled IDs, missing data, or generator-side generation cost can change errors/latency/throughput for reasons unrelated to target capacity.
Official references and version notes
- Component Reference — CSV Data Set Config — row timing, headers, quoting, relative paths, EOF behavior, sharing modes, and per-thread file example.
- Remote (Distributed) Testing — remote engines execute the plan and have their own local runtime/filesystem state.
- Getting Started — GUI authoring/debugging and CLI execution conventions.
- Best Practices — generator validity, lean listeners, and pre-generated data guidance.
- Apache JMeter downloads — current production release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter documentation on 2026-09-05. The course baseline remains
Apache JMeter 5.6.3 with a Java 17 JDK for labs
and no third-party plugins; JMeter 5.6.3 requires Java 8+. CSV
Data Set Config reads one line into variables at the
start of each test iteration. With default
All threads sharing, one file cursor is shared
across threads in that JMeter instance, but which thread receives
which row depends on execution order and can vary.
Current thread group opens a separate cursor per
Thread Group; Current thread opens a separate
cursor per thread; an explicit sharing identifier creates a cursor
shared by elements using that identifier. If every thread uses
Current thread against the same file, every thread starts its own
cursor at row 1 unless the filename itself is partitioned (the
current docs explicitly show filenames such as
test${__threadNum}.csv). At EOF, Recycle=true
restarts the file. With Recycle=false and Stop Thread=false, CSV
variables become <EOF> (default value,
configurable by csvdataset.eofstring). With
Recycle=false and Stop Thread=true, the thread stops at EOF.
Relative local filenames are resolved against the active test-plan
path; for distributed testing, the CSV must already exist on each
server host in the correct relative location. Data files are not
automatically copied by the distributed controller.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.