CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy: Diagnostics, Failure Modes, and Production Practices
A data failure often appears as a protocol or performance symptom:
409 conflicts, authentication churn, abrupt thread termination,
strange <EOF> requests, remote FileNotFound
errors, or latency from serialized account state. Preserve the first
failure and reconstruct cursor ownership before changing the target
or workload.
Learning objectives
- Diagnose concurrent reuse of one account or one-time token.
- Detect accidental recycle versus intentional reusable data.
- Recognize that distributed JMeter does not automatically transfer CSV files.
- Identify real PII/secrets committed as test data and contain the exposure.
- Reject nondeterministic filesystem-order-based shard selection.
- Diagnose row-demand mismatch and stale CSV variables without deleting evidence.
1. Preserve evidence before changing data
http://127.0.0.1:8000.
Preserve the failed JMX, exact CSV/shard hash, selected
properties/CLI, JTL, matching jmeter.log, target
events/stats, and generator filesystem/resource state before editing
sharing or row counts.
2. Diagnostic sequence
The cursor belongs to a sharing domain inside one JMeter process. Remote engines have independent processes/cursors and therefore need deliberately partitioned data.
flowchart TD E[Preserve JMX + CSV/hash + JTL + jmeter.log + target events] --> V[Confirm JMeter/Java/tool versions] V --> C[Confirm exact CLI, target, file path, properties] C --> D[Count rows + validate schema/duplicates/privacy] D --> S[Map sharing mode to cursor domains] S --> M[Compute row demand + EOF/recycle behavior] M --> R[Inspect thread variables + request IDs] R --> G[Inspect generator filesystem/CPU/GC] G --> T[Inspect target collision/error evidence] T --> X[Remote/CI/container file deployment if relevant] X --> F[Least destructive correction] F --> N[Small bounded rerun]
3. Failure mode: one account reused across threads
Symptom: 409/session invalidation or serialized behavior appears
under concurrency. If Current thread sharing points every thread at
the same accounts.csv, each private cursor starts at
row 1.
Evidence: Debug Sampler shows identical
${user_id} values; server event log shows the same
account claimed by thread 1 and thread 2. Repair the data
mapping—not target locking or JVM heap.
4. Intentionally broken example: one-time data with Recycle=true
Three one-time tokens feed six iterations through an All-threads cursor. Recycle=true restarts after row 3, so tokens OT-1/2/3 are sent again. The fixture returns 409 reuse.
Repair options:
- provide at least six disjoint rows;
- reduce workload demand to three iterations; or
- set Recycle=false + Stop Thread=true when exhaustion should end activity.
Do not add retries. A retry would consume the same wrong/reused data policy again and distort the workload.
5. Failure mode: assuming remote engines receive the CSV automatically
The controller can access data/users.csv, but remote
engine B logs a file-not-found error or reads a stale local copy.
The JMX transfer/remote execution path does not provision arbitrary
CSV data files.
Repair: deploy the exact shard to each authorized engine in the expected relative path, verify checksum/row count, then start the distributed run. Do not weaken RMI/TLS security to solve a filesystem problem.
6. Failure mode: every remote engine gets the same one-time file
Even if the file exists everywhere, uniqueness can still fail. Each remote process owns an independent All-threads cursor and starts at row 1. If every host has identical content, token 0001 can be consumed once per engine.
Repair with disjoint shards or a target-backed allocation service designed for concurrency. This chapter uses pre-shards because they are free, simple, and auditable.
7. Failure mode: real PII committed as fixture data
A developer copies customer email/name rows into
data/users.csv and commits them. This is not merely a
test failure; it is a data-protection incident depending on
organizational policy/law.
8. Failure mode: relying on filesystem ordering for shard selection
A launcher lists shards/*.csv and assigns the first
file to engine A, second to engine B. Another OS/filesystem returns
entries in a different order, so engines receive swapped or
duplicate data.
Repair with an explicit manifest keyed by engine identity. Verify path + SHA-256 before execution.
9. Failure mode: row count does not match demand
Two threads × five outer loops require ten rows in one All-threads one-time cursor, but the file has eight. Depending on settings:
- Recycle=true silently reuses two rows;
-
Recycle=false/Stop=false emits
<EOF>for later iterations; - Recycle=false/Stop=true ends threads before the configured iteration count.
Configured load and achieved load are now different. Report the data-limited achieved sample count rather than claiming the target “couldn't handle” the planned iterations.
10. Failure mode: fewer CSV fields retain previous values
Current component documentation notes that if a row contains fewer values than configured variables, the remaining variables are not updated and can retain earlier values. A malformed row can therefore combine a new user ID with the previous iteration's other field.
Preflight CSV schema/column counts before load. Do not rely on the sampler to reveal malformed records at scale.
11. Failure mode: absolute paths work locally, fail in CI/remote
C:\Users\name\Desktop\users.csv or
/home/me/users.csv couples the plan to one generator.
Use a test-plan-relative path and ensure the file is
packaged/mounted/deployed explicitly in each execution environment.
12. Causal performance separation
| Symptom | Data/generator cause | Target cause to distinguish | Evidence |
|---|---|---|---|
| 409 rate increases | duplicate account/token allocation | real target concurrency rule regression | CSV assignment + server collision log. |
| Achieved samples below configured | Stop Thread at EOF / missing shard | target rejects/blocks | JTL count + jmeter.log EOF/file errors. |
| Latency rises for shared account | target serializes same identity because test collides | general server saturation | per-user IDs + target queues + unique-data rerun. |
| Generator CPU rises | runtime uniqueness generation / file parsing/logging | target slowdown | generator CPU/GC + target service timing. |
| Remote-only failures | missing/stale/wrong shard on engine | network/SUT issue | engine-local filesystem/hash + remote log. |
13. Shortcuts to reject
- Do not retry duplicate one-time data until it passes.
- Do not increase thread count when the dataset is already undersized.
- Do not increase heap before proving a generator memory problem.
- Do not copy the controller's real credentials/PII file to every remote engine.
- Do not disable TLS/RMI verification to solve missing CSV files.
- Do not replace explicit shard mapping with directory-order guesses.
-
Do not delete failed JTL/
jmeter.log/server evidence after repair.
Knowledge check
A six-iteration run has only three one-time rows and Recycle=true. What failure should you expect?
Rows are reused after EOF, so one-time tokens collide/repeat; this is a data-policy failure.
Why can remote engine B fail even when the controller has users.csv?
CSV files are engine-local resources and are not automatically copied to remote servers.
What happens when a CSV row has fewer fields than expected?
Unupdated variables can retain values from the prior row/iteration, creating stale mixed records.
Why can Stop Thread at EOF lower achieved load?
Threads terminate when data is exhausted, so the test cannot execute the originally configured number of iterations.
What should be fixed first when two threads share one account unexpectedly?
The cursor/file allocation strategy and row ownership, not server capacity or JVM tuning.
Official references and version notes
- Component Reference — CSV Data Set Config — row timing, headers, quoting, relative paths, EOF behavior, sharing modes, and per-thread file example.
- Remote (Distributed) Testing — remote engines execute the plan and have their own local runtime/filesystem state.
- Getting Started — GUI authoring/debugging and CLI execution conventions.
- Best Practices — generator validity, lean listeners, and pre-generated data guidance.
- Apache JMeter downloads — current production release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter documentation on 2026-09-05. The course baseline remains
Apache JMeter 5.6.3 with a Java 17 JDK for labs
and no third-party plugins; JMeter 5.6.3 requires Java 8+. CSV
Data Set Config reads one line into variables at the
start of each test iteration. With default
All threads sharing, one file cursor is shared
across threads in that JMeter instance, but which thread receives
which row depends on execution order and can vary.
Current thread group opens a separate cursor per
Thread Group; Current thread opens a separate
cursor per thread; an explicit sharing identifier creates a cursor
shared by elements using that identifier. If every thread uses
Current thread against the same file, every thread starts its own
cursor at row 1 unless the filename itself is partitioned (the
current docs explicitly show filenames such as
test${__threadNum}.csv). At EOF, Recycle=true
restarts the file. With Recycle=false and Stop Thread=false, CSV
variables become <EOF> (default value,
configurable by csvdataset.eofstring). With
Recycle=false and Stop Thread=true, the thread stops at EOF.
Relative local filenames are resolved against the active test-plan
path; for distributed testing, the CSV must already exist on each
server host in the correct relative location. Data files are not
automatically copied by the distributed controller.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.