Chapter 11Lesson 03~160 minutes

CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy: Configuration, Design Patterns, and Trade-Offs

CSV configuration is not primarily about commas; it is about ownership. The right design follows whether an identity is reusable, exclusive to one virtual user, consumed once per iteration, shared across Thread Groups, unique across engines, or regenerated for every run.

Data ownershipGenerated vs fixtureRecycle policyPre-shardingPrivacy

Learning objectives

  • Choose a shared CSV or generated data from uniqueness/replay requirements.
  • Select All threads, Current thread group, Current thread, or Identifier from allocation scope.
  • Choose recycle/stop behavior from whether records are reusable or consumable.
  • Compare pre-sharded uniqueness with runtime-generated uniqueness.
  • Choose committed fake fixtures or ephemeral generated datasets from governance/privacy needs.
  • Connect data design to generator cost, target safety, distributed execution, and measurement validity.

1. Mandatory path remains local and synthetic

All runnable examples stay on http://127.0.0.1:8000, use fake identifiers, and remain ≤2 threads. Remote/cloud/CI designs are discussed as deployment patterns only; no paid service or live RMI setup is required.

2. Single shared file versus generated data

Need Shared/pre-generated CSV Runtime-generated value
Deterministic replay Strong: exact rows/hash can be archived. Weak unless seed/input is recorded.
Large unique pool Preferred; current docs warn runtime random generation costs CPU/memory. Can burden injector at scale.
Data resembles curated business states Excellent; rows can encode scenario classes. Generation logic may become complex.
Ephemeral run ID suffix Can combine base row + run property. Convenient for small cheap derived values.
Distributed global uniqueness Pre-shard explicitly. Requires engine-aware uniqueness scheme and collision analysis.

A hybrid is often best: pre-generate deterministic base records, then append a non-secret run identifier if the target requires each execution to look new.

3. Choose sharing mode from cursor ownership

Sharing mode Use when Collision risk
All threads One process-wide pool should allocate each row once across all threads. Row-to-thread ownership is not stable; remote engines still have separate pools.
Current thread group Each Thread Group intentionally owns its own cursor/pool. Same file can duplicate data across groups.
Current thread Each thread owns a private cursor, often with thread-specific filename. Same common file makes every thread start at row 1.
Identifier Selected groups/elements intentionally share a named cursor domain. Misused identifiers can join pools that should be separate.

4. Recycle versus Stop Thread at EOF

Make reuse policy explicit:

  • Reusable account: recycle can be acceptable if the same owner is allowed to reuse it.
  • One-time coupon/registration ID: Recycle=false; Stop Thread=true is safer when data exhaustion should end activity.
  • Data shortage is a test setup error: preflight row count and fail before load rather than discovering EOF mid-run.

<EOF> is valuable diagnostic evidence but should not be the normal production workload path.

5. Pre-sharded versus runtime-generated uniqueness

Pre-sharding makes engine ownership obvious and debuggable: shard A contains one range, shard B another, and hashes prove deployment. Runtime-generated UUIDs can avoid CSV exhaustion but make exact replay and collision audit harder and consume generator work.

If runtime uniqueness is used, include an engine/run prefix so independent generators cannot accidentally generate overlapping business keys under a weak custom scheme.

6. Committed fake fixtures versus ephemeral data

Choice Strength Governance caution
Committed fake CSV Reviewable schema, deterministic examples, easy local onboarding. Must be unmistakably synthetic and free of secrets/PII.
Generated ephemeral CSV Can size exactly to run demand; easy to shard per run. Generator script/version/seed/output hash becomes part of evidence.
Real anonymized export May resemble production distribution. Still carries privacy/re-identification/legal risk; not required for this course.

For the Academy, committed/generated synthetic data is sufficient. Real customer records are never a mandatory learning dependency.

7. Stable per-user identity versus iteration-driven CSV consumption

Because CSV rows update at each test iteration, a Thread Group loop can unintentionally rotate accounts. If each virtual user should keep one account while repeating a business action many times, allocate the account once per user and place the repeated action inside a nested Loop Controller. Chapter 12 will deepen this controller pattern.

8. Row-demand arithmetic must match the sharing domain

For one-time rows:

All threads shared cursor:
rows >= sum(iterations across all threads/groups sharing that cursor)

Current thread cursor:
rows in each thread's file >= that thread's outer iterations

Current thread group:
rows per group cursor >= iterations consumed by that group

Distributed:
global rows >= sum(rows consumed by every engine)
and shard sets must be disjoint

Include warm-up, retries only if explicitly modeled, setup calls that consume data, and repeated CI runs in capacity planning.

9. Relative paths versus absolute paths

Use repository/test-plan-relative paths for portability. Absolute paths can work on one workstation but fail on CI/remote engines. In distributed mode, provision the same relative path on every server host but place the correct engine-specific shard there.

10. Deterministic shard mapping versus filesystem ordering

Do not infer engine shard by “sort/list all CSVs and take index N” unless the ordering contract is explicitly guaranteed and tested. Directory enumeration can differ across OS/filesystems/tool versions.

Prefer an explicit manifest:

engine-a -> shards/engine-a/data/users.csv -> rows 1..6 -> sha256 ...
engine-b -> shards/engine-b/data/users.csv -> rows 1..6 -> sha256 ...

11. Privacy and credential boundaries

CSV is plain text. A password/API token stored there is readable by anyone with filesystem/artifact access and may later appear in requests/debug output. If credentials are unavoidable in a real authorized environment, use the platform's secret mechanism and minimize persistence; do not commit them.

Use synthetic emails such as acct-001@example.invalid, synthetic tokens, and fake IDs in training/repository fixtures.

12. Data strategy affects performance validity

Colliding users may trigger target locks, deduplication, cache locality, rate limits, or error branches. Recycled one-time IDs may artificially increase 4xx/409 rates. Runtime generation may consume injector CPU. Missing remote shards may stop threads. None of those should be misreported as server capacity behavior without attribution.

For every meaningful run, preserve the raw JTL together with its matching jmeter.log, the exact CSV/shard hash, and the non-secret run manifest. JTL shows achieved samples and failures; jmeter.log preserves engine-side file, EOF, and runtime diagnostics needed to distinguish data exhaustion from target failure.

13. Keep data allocation separate from surrounding layers

Layer Examples Not a substitute for
JMeter CSV config filename, columns, recycle, stop, sharing cursor Creating valid target business data.
Generator filesystem file presence, permissions, encoding JMeter remote file distribution.
Target/SUT account uniqueness, one-time semantics, cleanup CSV cursor correctness.
CI/container workspace, mounted files, artifact retention Privacy review or shard math.
Distributed engine engine-local file/cursor Controller-local CSV state.
JVM heap/GC/file IO Fixing duplicate rows or bad allocation.

14. Decision table

Requirement Preferred data strategy Reason
Each iteration needs a fresh one-time ID All threads shared cursor, Recycle=false, Stop=true One process consumes each row once.
Each thread owns one reusable account Thread-specific file + Current thread, or one-time allocation outside inner business loop Stable per-user ownership.
Two Thread Groups share one global pool All threads or explicit shared Identifier Avoid duplicate group-local cursors.
Two remote engines need globally unique data Disjoint pre-shards deployed per engine Each engine has independent cursor/process.
Large synthetic unique pool Pre-generate CSV Lower runtime CPU/memory and better auditability.
Tiny run-specific suffix Derived runtime value may be reasonable Low cost; record run input.

Knowledge check

Why can a Thread Group loop rotate accounts unexpectedly?

When should Recycle=false + Stop Thread=true be preferred?

Why are remote shards needed even with All threads sharing?

Why is an explicit shard manifest better than filesystem ordering?

How can test data invalidate a capacity claim?

Next lesson

Diagnose data failures before tuning the system

Lesson 4 engineers account reuse, one-time recycling, remote-file assumptions, PII leakage, nondeterministic shard selection, row-count shortages, and stale-column behavior.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK for labs and no third-party plugins; JMeter 5.6.3 requires Java 8+. CSV Data Set Config reads one line into variables at the start of each test iteration. With default All threads sharing, one file cursor is shared across threads in that JMeter instance, but which thread receives which row depends on execution order and can vary. Current thread group opens a separate cursor per Thread Group; Current thread opens a separate cursor per thread; an explicit sharing identifier creates a cursor shared by elements using that identifier. If every thread uses Current thread against the same file, every thread starts its own cursor at row 1 unless the filename itself is partitioned (the current docs explicitly show filenames such as test${__threadNum}.csv). At EOF, Recycle=true restarts the file. With Recycle=false and Stop Thread=false, CSV variables become <EOF> (default value, configurable by csvdataset.eofstring). With Recycle=false and Stop Thread=true, the thread stops at EOF. Relative local filenames are resolved against the active test-plan path; for distributed testing, the CSV must already exist on each server host in the correct relative location. Data files are not automatically copied by the distributed controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.