Chapter 11Lesson 01~140 minutes

CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy: Core Concepts and Mental Model

Chapter 10 made response extraction deliberate. Chapter 11 controls the opposite data flow: values supplied into virtual users. A performance test can be perfectly paced and correctly correlated yet still be invalid because two users share one account, one-time IDs recycle, a remote engine lacks its CSV, or real customer data was committed to the repository.

CSV Data Set ConfigSharing modeEOFUnique dataDistributed shards

Learning objectives

  • Explain when CSV Data Set Config advances its file cursor and creates per-thread variables.
  • Distinguish All threads, Current thread group, Current thread, and Identifier sharing domains.
  • Predict Recycle/Stop Thread behavior at EOF, including the <EOF> sentinel.
  • Design uniqueness across threads, loops, Thread Groups, engines, and repeated runs.
  • Choose relative file paths and distributed shards deliberately.
  • Keep test data synthetic, auditable, and separate from real PII/credentials.

1. The practical problem: test data is shared mutable infrastructure

Suppose four JMeter threads all log in as the same account. The target may serialize sessions, invalidate older tokens, enforce one active checkout, or rate-limit the account. The resulting 409s and latency spikes say more about data collision than server capacity.

Or suppose a “unique registration” CSV has 1,000 rows but the test consumes 2,000 iterations with Recycle enabled. The second half of the run reuses identities and starts measuring duplicate-data behavior instead of the intended path.

Safety boundary: all executable examples use synthetic example.invalid identities and loopback http://127.0.0.1:8000. Never import production customer exports, authentication secrets, or real PII into this course workflow.

2. Mental model: data source → cursor → thread variables → request

CSV cursor ownership and data-consumption flow

The cursor belongs to a sharing domain inside one JMeter process. Remote engines have independent processes/cursors and therefore need deliberately partitioned data.

flowchart TD
D[Versioned or generated synthetic CSV] --> C[CSV Data Set Config]
C --> S[Sharing domain + cursor]
S --> I[Iteration-start row read]
I --> V[Thread-local fields]
V --> R[Sampler request]
R --> T[Authorized target]
T --> E[JTL + jmeter.log + server data-consumption evidence]
M[Shard manifest] --> D
RE[Remote engine A cursor] -. separate process .-> S
RB[Remote engine B cursor] -. separate process .-> S

The CSV file is not “assigned to a user” by itself. CSV Data Set Config owns a cursor within a sharing domain. At the start of each test iteration, JMeter reads the next row for that domain and writes its columns into the current thread's variable map. The sampler then resolves those variables. The target and evidence log independently show which value was actually consumed.

A remote JMeter engine is a separate process with its own cursor and filesystem. Therefore a globally unique dataset must be partitioned across engines, not merely shared inside one engine.

3. Rows are read at iteration start

Current JMeter documentation states that CSV variables are defined at the start of each test iteration. With one CSV Data Set Config and one row consumed per iteration, approximate row demand is:

rows_required_per_cursor_domain = number_of_iterations_in_that_domain

If a Thread Group has 4 threads × 5 outer loops and all threads share one cursor, that cursor may consume about 20 rows. If each thread has its own cursor, each cursor needs 5 rows unless the filename is itself thread-specific or recycling is intentionally used.

4. All threads: one cursor, nondeterministic row-to-thread ownership

All threads is the default sharing mode. One cursor is shared across all threads that reference that file in the same JMeter process, so different threads normally receive different rows. However, the docs explicitly warn that row assignment depends on execution order and may vary between iterations.

That makes All threads excellent for “each iteration needs the next unused record,” but not for “thread 1 must always own account A.”

5. Current thread group: one cursor per Thread Group

With Current thread group, each Thread Group has an independent cursor. If two groups point at the same one-time file, both can start at row 1 and reuse the same data. This is useful only when duplication across groups is intentional or each group has separate data.

6. Current thread: one cursor per thread

With Current thread, every thread opens its own cursor. If all threads point to accounts.csv, every cursor begins with the first data row—so every thread can receive the same account.

The current documentation gives the safe pattern for per-thread files: use a filename such as accounts-thread-${__threadNum}.csv with Current thread sharing. Each thread then owns a different file/cursor.

7. Identifier sharing: explicit cursor domains

The Identifier mode lets elements/threads that use the same identifier share a cursor. It is useful when several Thread Groups intentionally share one allocation pool while others use a different pool. Treat the identifier as part of the data-allocation contract and record it in the run manifest.

8. EOF is a business decision, not a file accident

Recycle on EOF Stop Thread on EOF Behavior
true irrelevant Cursor restarts at the first row. Safe only when reuse is allowed.
false false Variables become <EOF> by default; downstream requests can accidentally send that literal.
false true Thread stops when no row is available. Preferred for one-time/exhaustible data.

For one-time registration tokens, Recycle should normally be false. If the number of valid rows is the workload limit, Stop Thread=true makes exhaustion explicit rather than silently recycling.

9. Headers, quoting, and stale columns

If Variable Names is empty, JMeter can use the first CSV row as column names. Quoted data can contain delimiters when Allow quoted data is enabled. A subtle edge case: if a later row has fewer values than variable names, the missing variables are not automatically cleared—they can retain their prior values. That is another reason to validate CSV shape before load.

10. File path ownership

For local runs, relative filenames are resolved with respect to the active test plan. Absolute paths couple the JMX to one workstation and are poor portability defaults.

In distributed testing, the CSV must already exist on each JMeter server host at the expected relative location. Sending the JMX to remote engines does not copy the CSV bytes.

11. All threads is not global across remote engines

Imagine two remote engines, each running the full test with the same users.csv. Engine A's shared cursor starts at row 1. Engine B's separate shared cursor also starts at row 1. “All threads” is only global inside one JMeter process—not across machines.

For one-time global data, deploy disjoint shards: Engine A gets IDs A-0001…A-1000; Engine B gets B-0001…B-1000. Keep the JMX path the same if desired, but make each engine's file content distinct and record a shard manifest/hash.

12. Pre-generated data versus runtime uniqueness

Current JMeter documentation specifically notes that generating large volumes of random unique values at runtime costs CPU/memory and recommends creating data in advance. Pre-generation also makes row counts, uniqueness, and replay provenance auditable.

Runtime IDs can still be appropriate for small ephemeral values that do not need deterministic replay. Do not use runtime UUID generation as an excuse to ignore target-side uniqueness constraints or generator cost.

13. Read-only inspection before modifying data

  • Count data rows excluding the header and compare with planned iterations.
  • Check for duplicate primary keys/tokens before load.
  • Record filename, encoding, delimiter, header mode, Recycle, Stop Thread, and Sharing mode.
  • Map each sharing domain to its cursor and expected row demand.
  • Verify no real PII/secrets are present; synthetic examples use example.invalid.
  • For remote design, verify every engine has an explicit shard assignment, row range, relative path, and checksum.
  • Preserve JTL, matching jmeter.log, and target-side consumption evidence.

14. DevOps connection

A reliable performance pipeline treats test data like infrastructure: versioned/generated, capacity-planned, privacy-reviewed, deployed to the correct execution node, and cleaned up after the run. Data provenance belongs beside JMX version, workload parameters, generator version, and result artifacts.

Knowledge check

Does All threads guarantee thread 1 always receives row 1?

Why can Current thread plus one shared accounts.csv cause collisions?

What happens at EOF with Recycle=false and Stop Thread=false?

Does the distributed controller automatically copy CSV files to remote engines?

Why pre-generate large unique datasets?

Next lesson

Watch cursors collide, recycle, stop, and shard

Lesson 2 builds synthetic CSVs, demonstrates All threads versus Current thread, triggers one-time recycling and EOF, assigns stable accounts with per-thread files, and creates deterministic engine shards.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK for labs and no third-party plugins; JMeter 5.6.3 requires Java 8+. CSV Data Set Config reads one line into variables at the start of each test iteration. With default All threads sharing, one file cursor is shared across threads in that JMeter instance, but which thread receives which row depends on execution order and can vary. Current thread group opens a separate cursor per Thread Group; Current thread opens a separate cursor per thread; an explicit sharing identifier creates a cursor shared by elements using that identifier. If every thread uses Current thread against the same file, every thread starts its own cursor at row 1 unless the filename itself is partitioned (the current docs explicitly show filenames such as test${__threadNum}.csv). At EOF, Recycle=true restarts the file. With Recycle=false and Stop Thread=false, CSV variables become <EOF> (default value, configurable by csvdataset.eofstring). With Recycle=false and Stop Thread=true, the thread stops at EOF. Relative local filenames are resolved against the active test-plan path; for distributed testing, the CSV must already exist on each server host in the correct relative location. Data files are not automatically copied by the distributed controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.