Chapter 11Lesson 04~170 minutes

CSV Data Set Config, Parameterization, Unique Data, and Test Data Strategy: Diagnostics, Failure Modes, and Production Practices

A data failure often appears as a protocol or performance symptom: 409 conflicts, authentication churn, abrupt thread termination, strange <EOF> requests, remote FileNotFound errors, or latency from serialized account state. Preserve the first failure and reconstruct cursor ownership before changing the target or workload.

DiagnosticsCollisionsEOFRemote filesPrivacy

Learning objectives

  • Diagnose concurrent reuse of one account or one-time token.
  • Detect accidental recycle versus intentional reusable data.
  • Recognize that distributed JMeter does not automatically transfer CSV files.
  • Identify real PII/secrets committed as test data and contain the exposure.
  • Reject nondeterministic filesystem-order-based shard selection.
  • Diagnose row-demand mismatch and stale CSV variables without deleting evidence.

1. Preserve evidence before changing data

Diagnostic reruns stay loopback-only and bounded. The only executable target is http://127.0.0.1:8000. Preserve the failed JMX, exact CSV/shard hash, selected properties/CLI, JTL, matching jmeter.log, target events/stats, and generator filesystem/resource state before editing sharing or row counts.

2. Diagnostic sequence

CSV/data failure diagnostic sequence

The cursor belongs to a sharing domain inside one JMeter process. Remote engines have independent processes/cursors and therefore need deliberately partitioned data.

flowchart TD
E[Preserve JMX + CSV/hash + JTL + jmeter.log + target events] --> V[Confirm JMeter/Java/tool versions]
V --> C[Confirm exact CLI, target, file path, properties]
C --> D[Count rows + validate schema/duplicates/privacy]
D --> S[Map sharing mode to cursor domains]
S --> M[Compute row demand + EOF/recycle behavior]
M --> R[Inspect thread variables + request IDs]
R --> G[Inspect generator filesystem/CPU/GC]
G --> T[Inspect target collision/error evidence]
T --> X[Remote/CI/container file deployment if relevant]
X --> F[Least destructive correction]
F --> N[Small bounded rerun]

3. Failure mode: one account reused across threads

Symptom: 409/session invalidation or serialized behavior appears under concurrency. If Current thread sharing points every thread at the same accounts.csv, each private cursor starts at row 1.

Evidence: Debug Sampler shows identical ${user_id} values; server event log shows the same account claimed by thread 1 and thread 2. Repair the data mapping—not target locking or JVM heap.

4. Intentionally broken example: one-time data with Recycle=true

Three one-time tokens feed six iterations through an All-threads cursor. Recycle=true restarts after row 3, so tokens OT-1/2/3 are sent again. The fixture returns 409 reuse.

Repair options:

  • provide at least six disjoint rows;
  • reduce workload demand to three iterations; or
  • set Recycle=false + Stop Thread=true when exhaustion should end activity.

Do not add retries. A retry would consume the same wrong/reused data policy again and distort the workload.

5. Failure mode: assuming remote engines receive the CSV automatically

The controller can access data/users.csv, but remote engine B logs a file-not-found error or reads a stale local copy. The JMX transfer/remote execution path does not provision arbitrary CSV data files.

Repair: deploy the exact shard to each authorized engine in the expected relative path, verify checksum/row count, then start the distributed run. Do not weaken RMI/TLS security to solve a filesystem problem.

6. Failure mode: every remote engine gets the same one-time file

Even if the file exists everywhere, uniqueness can still fail. Each remote process owns an independent All-threads cursor and starts at row 1. If every host has identical content, token 0001 can be consumed once per engine.

Repair with disjoint shards or a target-backed allocation service designed for concurrency. This chapter uses pre-shards because they are free, simple, and auditable.

7. Failure mode: real PII committed as fixture data

A developer copies customer email/name rows into data/users.csv and commits them. This is not merely a test failure; it is a data-protection incident depending on organizational policy/law.

Containment: stop distributing the file, follow your organization's incident/secret/PII response procedure, remove exposure from current working artifacts, and do not pretend a normal Git delete erases previously published history/caches. For course work, use synthetic data from the start.

8. Failure mode: relying on filesystem ordering for shard selection

A launcher lists shards/*.csv and assigns the first file to engine A, second to engine B. Another OS/filesystem returns entries in a different order, so engines receive swapped or duplicate data.

Repair with an explicit manifest keyed by engine identity. Verify path + SHA-256 before execution.

9. Failure mode: row count does not match demand

Two threads × five outer loops require ten rows in one All-threads one-time cursor, but the file has eight. Depending on settings:

  • Recycle=true silently reuses two rows;
  • Recycle=false/Stop=false emits <EOF> for later iterations;
  • Recycle=false/Stop=true ends threads before the configured iteration count.

Configured load and achieved load are now different. Report the data-limited achieved sample count rather than claiming the target “couldn't handle” the planned iterations.

10. Failure mode: fewer CSV fields retain previous values

Current component documentation notes that if a row contains fewer values than configured variables, the remaining variables are not updated and can retain earlier values. A malformed row can therefore combine a new user ID with the previous iteration's other field.

Preflight CSV schema/column counts before load. Do not rely on the sampler to reveal malformed records at scale.

11. Failure mode: absolute paths work locally, fail in CI/remote

C:\Users\name\Desktop\users.csv or /home/me/users.csv couples the plan to one generator. Use a test-plan-relative path and ensure the file is packaged/mounted/deployed explicitly in each execution environment.

12. Causal performance separation

Symptom Data/generator cause Target cause to distinguish Evidence
409 rate increases duplicate account/token allocation real target concurrency rule regression CSV assignment + server collision log.
Achieved samples below configured Stop Thread at EOF / missing shard target rejects/blocks JTL count + jmeter.log EOF/file errors.
Latency rises for shared account target serializes same identity because test collides general server saturation per-user IDs + target queues + unique-data rerun.
Generator CPU rises runtime uniqueness generation / file parsing/logging target slowdown generator CPU/GC + target service timing.
Remote-only failures missing/stale/wrong shard on engine network/SUT issue engine-local filesystem/hash + remote log.

13. Shortcuts to reject

  • Do not retry duplicate one-time data until it passes.
  • Do not increase thread count when the dataset is already undersized.
  • Do not increase heap before proving a generator memory problem.
  • Do not copy the controller's real credentials/PII file to every remote engine.
  • Do not disable TLS/RMI verification to solve missing CSV files.
  • Do not replace explicit shard mapping with directory-order guesses.
  • Do not delete failed JTL/jmeter.log/server evidence after repair.

Knowledge check

A six-iteration run has only three one-time rows and Recycle=true. What failure should you expect?

Why can remote engine B fail even when the controller has users.csv?

What happens when a CSV row has fewer fields than expected?

Why can Stop Thread at EOF lower achieved load?

What should be fixed first when two threads share one account unexpectedly?

Next lesson

Checkpoint: two workload classes and one global uniqueness plan

Lesson 5 combines reusable per-thread accounts with exhaustible one-time tokens, proves local uniqueness, preserves intentional failures, and produces a two-engine shard plan.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK for labs and no third-party plugins; JMeter 5.6.3 requires Java 8+. CSV Data Set Config reads one line into variables at the start of each test iteration. With default All threads sharing, one file cursor is shared across threads in that JMeter instance, but which thread receives which row depends on execution order and can vary. Current thread group opens a separate cursor per Thread Group; Current thread opens a separate cursor per thread; an explicit sharing identifier creates a cursor shared by elements using that identifier. If every thread uses Current thread against the same file, every thread starts its own cursor at row 1 unless the filename itself is partitioned (the current docs explicitly show filenames such as test${__threadNum}.csv). At EOF, Recycle=true restarts the file. With Recycle=false and Stop Thread=false, CSV variables become <EOF> (default value, configurable by csvdataset.eofstring). With Recycle=false and Stop Thread=true, the thread stops at EOF. Relative local filenames are resolved against the active test-plan path; for distributed testing, the CSV must already exist on each server host in the correct relative location. Data files are not automatically copied by the distributed controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.