CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy: Guided Hands-On Workflow
This workflow makes CSV cursor behavior observable against a loopback service. The target remembers which JMeter thread first claimed an account and which one-time tokens were already consumed, so wrong sharing/recycling becomes a visible 409 rather than a vague downstream symptom.
Learning objectives
- Generate deterministic synthetic CSV fixtures and per-engine shards.
- Observe default All-threads cursor allocation.
- Demonstrate why Current-thread sharing with one common file duplicates row 1.
- Use thread-specific filenames for stable account ownership.
-
Trigger Recycle and
<EOF>/Stop Thread behavior deliberately. - Prepare disjoint engine shards without starting remote RMI.
1. Safety envelope
http://127.0.0.1:8000. Maximum 2 threads, maximum 3
loops in functional experiments, and synthetic
example.invalid accounts/tokens only. Remote engines
are simulated by files/manifests; no RMI ports are opened. Reset
fixture state between experiments.
2. Start the local data-consumption fixture
Save as fixtures/data_fixture.py:
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import urlparse, parse_qs
from pathlib import Path
import argparse
import json
import threading
import time
lock = threading.Lock()
total = 0
by_path = {}
account_owner = {}
one_time_used = set()
collisions = []
event_log = None
def record(event):
if event_log is None:
return
with lock:
with event_log.open("a", encoding="utf-8") as handle:
handle.write(json.dumps(event, sort_keys=True) + "\n")
def body(obj):
return json.dumps(obj, sort_keys=True).encode("utf-8")
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def _send(self, status, payload):
raw = body(payload)
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.end_headers()
self.wfile.write(raw)
def do_GET(self):
global total
parsed = urlparse(self.path)
path = parsed.path
q = {k: v[-1] for k, v in parse_qs(parsed.query).items()}
if path == "/health":
self._send(200, {"status": "ok"})
return
if path == "/reset":
with lock:
account_owner.clear()
one_time_used.clear()
collisions.clear()
self._send(200, {"status": "reset"})
return
if path == "/stats":
with lock:
snapshot = {
"total": total,
"by_path": dict(by_path),
"account_owner": dict(account_owner),
"one_time_used": sorted(one_time_used),
"collisions": list(collisions),
}
self._send(200, snapshot)
return
started_ms = int(time.time() * 1000)
with lock:
total += 1
request_no = total
by_path[path] = by_path.get(path, 0) + 1
status = 200
event = {"path": path, "request_no": request_no}
if path == "/account":
user_id = q.get("user_id", "")
thread_id = q.get("thread", "")
event.update({"user_id": user_id, "thread": thread_id})
if not user_id or user_id == "<EOF>":
status = 400
payload = {"status": "invalid_data", "user_id": user_id}
else:
with lock:
owner = account_owner.get(user_id)
if owner is None:
account_owner[user_id] = thread_id
owner = thread_id
collision = owner != thread_id
if collision:
collisions.append({
"kind": "account_owner_collision",
"user_id": user_id,
"first_thread": owner,
"second_thread": thread_id,
})
if collision:
status = 409
payload = {
"status": "collision",
"user_id": user_id,
"owner_thread": owner,
"request_thread": thread_id,
}
else:
time.sleep(0.03)
payload = {"status": "accepted", "user_id": user_id, "thread": thread_id}
elif path == "/one-time":
token = q.get("token", "")
thread_id = q.get("thread", "")
event.update({"token": token, "thread": thread_id})
if not token or token == "<EOF>":
status = 400
payload = {"status": "invalid_data", "token": token}
else:
with lock:
reused = token in one_time_used
if not reused:
one_time_used.add(token)
else:
collisions.append({"kind": "one_time_reuse", "token": token, "thread": thread_id})
if reused:
status = 409
payload = {"status": "reused", "token": token}
else:
time.sleep(0.02)
payload = {"status": "accepted", "token": token, "thread": thread_id}
else:
status = 404
payload = {"status": "not_found", "path": path}
finished_ms = int(time.time() * 1000)
event.update({
"status": status,
"started_ms": started_ms,
"finished_ms": finished_ms,
"service_wall_ms": finished_ms - started_ms,
})
record(event)
self._send(status, payload)
def log_message(self, format, *args):
return
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("--log", default="results/server-events.jsonl")
args = parser.parse_args()
event_log = Path(args.log).resolve()
event_log.parent.mkdir(parents=True, exist_ok=True)
event_log.write_text("", encoding="utf-8")
print("fixture=http://127.0.0.1:8000")
print(f"event_log={event_log}")
ThreadingHTTPServer(("127.0.0.1", 8000), Handler).serve_forever()
Start and preflight:
python fixtures/data_fixture.py --log results/server-events.jsonl
curl --fail --silent http://127.0.0.1:8000/health
curl --fail --silent http://127.0.0.1:8000/reset
3. Generate deterministic synthetic CSVs
Save as tools/make_data.py:
from pathlib import Path
import csv
import hashlib
import json
root = Path("data")
root.mkdir(exist_ok=True)
def write_csv(path, field, values):
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("w", newline="", encoding="utf-8") as handle:
writer = csv.writer(handle)
writer.writerow([field])
for value in values:
writer.writerow([value])
def sha256(path):
return hashlib.sha256(path.read_bytes()).hexdigest()
# Shared account pool: useful for All-threads cursor demonstrations.
write_csv(root / "accounts.csv", "user_id", [
"acct-001@example.invalid",
"acct-002@example.invalid",
"acct-003@example.invalid",
"acct-004@example.invalid",
])
# Stable one-file-per-thread allocation.
write_csv(root / "accounts-thread-1.csv", "user_id", ["acct-thread-001@example.invalid"])
write_csv(root / "accounts-thread-2.csv", "user_id", ["acct-thread-002@example.invalid"])
# One-time records: must not be recycled.
write_csv(root / "one-time.csv", "token", [
"OT-0001", "OT-0002", "OT-0003", "OT-0004", "OT-0005", "OT-0006",
])
# Simulated distributed shards. Deploy each engine's file under the same relative
# path data/users.csv on that engine; content is deliberately disjoint.
shards = {
"engine-a": [f"EA-{n:04d}" for n in range(1, 7)],
"engine-b": [f"EB-{n:04d}" for n in range(1, 7)],
}
manifest = {"schema": 1, "engines": {}}
for engine, values in shards.items():
path = Path("shards") / engine / "data" / "users.csv"
write_csv(path, "token", values)
manifest["engines"][engine] = {
"relative_path": "data/users.csv",
"source_file": path.as_posix(),
"rows": len(values),
"first": values[0],
"last": values[-1],
"sha256": sha256(path),
}
Path("shards").mkdir(exist_ok=True)
(Path("shards") / "manifest.json").write_text(
json.dumps(manifest, indent=2, sort_keys=True),
encoding="utf-8",
)
print(json.dumps(manifest, indent=2, sort_keys=True))
Run:
python tools/make_data.py
Expected files include:
data/accounts.csv
data/accounts-thread-1.csv
data/accounts-thread-2.csv
data/one-time.csv
shards/engine-a/data/users.csv
shards/engine-b/data/users.csv
shards/manifest.json
The addresses end in example.invalid, so they are
explicitly synthetic and cannot accidentally route email to a real
domain.
4. Experiment A — All threads shares one cursor
Configure CSV Data Set Config:
| Setting | Value |
|---|---|
| Filename | data/accounts.csv |
| Variable Names | blank; first row is header |
| Delimiter | comma |
| Recycle | false |
| Stop Thread on EOF | true |
| Sharing mode | All threads |
Thread Group: 2 threads × 1 loop. HTTP Request:
/account?user_id=${user_id}&thread=${__threadNum}
Prediction: the two threads receive two different
rows from the shared cursor, but which thread receives
acct-001 versus acct-002 is not
guaranteed.
Inspect /stats and the event log; both accounts should
have exactly one owner.
5. Experiment B — Current thread + same file duplicates row 1
Reset the fixture. Change only Sharing mode to
Current thread, leaving
Filename=data/accounts.csv. Run 2 threads × 1 loop.
Prediction: each thread opens an independent cursor
and each starts at the first data row. Both therefore try
acct-001@example.invalid. The fixture assigns the first
observed thread as owner and returns 409 when another thread claims
the same account.
This is an intentionally broken configuration. Preserve the failed
JTL, jmeter.log, and server collision event.
6. Experiment C — stable per-thread accounts
Repair the account allocation:
| Setting | Value |
|---|---|
| Filename |
data/accounts-thread-${{__threadNum}}.csv
|
| Variable Names | blank/header |
| Recycle | true |
| Stop Thread on EOF | false |
| Sharing mode | Current thread |
Run 2 threads × 3 loops. Each file has one row, so recycling is intentional: thread 1 repeatedly gets its own synthetic account and thread 2 gets a different one. The fixture should report stable account ownership and no collisions.
This is a case where recycling is correct because the account is deliberately reusable by the same virtual user.
7. Experiment D — accidental recycling of one-time data
Reset the fixture. Create a debug copy of
one-time.csv containing only three tokens. Configure:
- 2 threads × 3 loops → demand = 6 rows;
- Sharing mode = All threads;
- Recycle = true.
Prediction: after three globally consumed rows, the cursor restarts and the next tokens repeat. The fixture returns 409 for reused one-time tokens. This is data-policy failure, not target capacity.
8. Experiment E — visible EOF sentinel
Reset. Keep the same 3-row file and 6-iteration demand, but set:
- Recycle = false;
- Stop Thread on EOF = false.
After the third row, JMeter sets ${token} to
<EOF>. The loopback target rejects that value as
invalid data. Preserve the request/event evidence before repairing
it.
9. Experiment F — stop on EOF for exhaustible data
Reset. Set Recycle=false and Stop Thread=true. The shared cursor contains only three one-time records.
Prediction: exactly three business requests can
consume valid data; when a thread requests another iteration row,
JMeter stops that thread instead of sending
<EOF> or recycling.
Which particular thread stops first is scheduler-dependent. The invariant is total unique consumption, not deterministic thread ownership.
10. Thread-group sharing experiment
Create two one-thread Thread Groups, both pointing to the same two-row one-time file with Sharing mode Current thread group. Each group opens its own cursor and starts at row 1, so the same token can be reused across groups.
Repair by using All threads when both groups should share one allocation pool, an explicit Identifier shared by the intended groups, or separate disjoint files when each group owns separate data.
11. Row-consumption math
| Configuration | Cursor domains | Rows needed without recycle |
|---|---|---|
| 2 threads × 3 loops, All threads | 1 shared cursor | 6 rows total |
| 2 threads × 3 loops, Current thread | 2 cursors | 3 rows per cursor; same file means duplicated row sequence |
| 2 Thread Groups × 3 iterations, Current thread group | 2 cursors | 3 rows per group cursor |
| 2 threads × 3 loops, per-thread one-row file, recycle=true | 2 cursors/files | 1 reusable row per thread |
13. Distributed file boundary
14. Analyze data-consumption evidence
Save as tools/analyze_data_events.py:
import json
import sys
from collections import Counter
from pathlib import Path
path = Path(sys.argv[1] if len(sys.argv) > 1 else "results/server-events.jsonl")
events = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
print(f"events={len(events)}")
print(f"status_counts={dict(Counter(e.get('status') for e in events))}")
print(f"path_counts={dict(Counter(e.get('path') for e in events))}")
account_assignments = {}
one_time = []
for e in events:
if e.get("path") == "/account" and e.get("user_id"):
account_assignments.setdefault(e["user_id"], set()).add(e.get("thread"))
if e.get("path") == "/one-time" and e.get("token"):
one_time.append(e["token"])
print("account_assignment_threads:")
for user, threads in sorted(account_assignments.items()):
print(f" {user}: {sorted(t for t in threads if t is not None)}")
counts = Counter(one_time)
duplicates = {token: n for token, n in counts.items() if n > 1 and token != "<EOF>"}
print(f"one_time_duplicates={duplicates}")
print(f"eof_events={sum(1 for e in events if e.get('user_id') == '<EOF>' or e.get('token') == '<EOF>')}")
print(f"collision_events={sum(1 for e in events if e.get('status') == 409)}")
Run after each isolated experiment:
python tools/analyze_data_events.py results/server-events.jsonl
Pair this with raw JTL and jmeter.log. The server event
log proves which account/token actually reached the target; the JTL
proves sample success/failure and thread timing.
15. Challenge
A 100-thread test needs each thread to keep one reusable account for 20 inner business loops. A teammate proposes All threads with 2,000 account rows. What is the simpler model?
If account ownership must remain stable per virtual user, allocate one account per thread (for example thread-specific files or a one-time startup allocation copied into a stable thread variable), then loop the business actions without advancing the account CSV every business repetition. Do not consume 20 accounts per user merely because the Thread Group outer loop also controls CSV advancement.
Knowledge check
Why does Experiment B collide even though Sharing mode is Current thread?
Each thread has its own cursor into the same file, so both cursors begin at the same first row.
Why is Recycle=true correct for the per-thread account files but wrong for one-time tokens?
The account is intentionally reusable by its owning user; one-time tokens must never be consumed twice.
With 2 threads × 3 loops and All threads sharing, how many rows are needed without recycle?
Six rows total in the shared cursor domain.
What is the correct interpretation when Stop Thread=true and only three rows exist?
Data capacity limits the number of valid iterations; threads stop as they exhaust the shared cursor rather than generating invalid/recycled requests.
How does the shard manifest improve reliability?
It explicitly maps each engine to a disjoint file/range/hash instead of relying on filesystem ordering or implicit copying.
Official references and version notes
- Component Reference — CSV Data Set Config — row timing, headers, quoting, relative paths, EOF behavior, sharing modes, and per-thread file example.
- Remote (Distributed) Testing — remote engines execute the plan and have their own local runtime/filesystem state.
- Getting Started — GUI authoring/debugging and CLI execution conventions.
- Best Practices — generator validity, lean listeners, and pre-generated data guidance.
- Apache JMeter downloads — current production release and Java requirement.
Version-sensitive statements were rechecked against current Apache
JMeter documentation on 2026-09-05. The course baseline remains
Apache JMeter 5.6.3 with a Java 17 JDK for labs
and no third-party plugins; JMeter 5.6.3 requires Java 8+. CSV
Data Set Config reads one line into variables at the
start of each test iteration. With default
All threads sharing, one file cursor is shared
across threads in that JMeter instance, but which thread receives
which row depends on execution order and can vary.
Current thread group opens a separate cursor per
Thread Group; Current thread opens a separate
cursor per thread; an explicit sharing identifier creates a cursor
shared by elements using that identifier. If every thread uses
Current thread against the same file, every thread starts its own
cursor at row 1 unless the filename itself is partitioned (the
current docs explicitly show filenames such as
test${__threadNum}.csv). At EOF, Recycle=true
restarts the file. With Recycle=false and Stop Thread=false, CSV
variables become <EOF> (default value,
configurable by csvdataset.eofstring). With
Recycle=false and Stop Thread=true, the thread stops at EOF.
Relative local filenames are resolved against the active test-plan
path; for distributed testing, the CSV must already exist on each
server host in the correct relative location. Data files are not
automatically copied by the distributed controller.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.