Chapter 25Lesson 04~225 minutes

Distributed Testing, Remote Engines, Network Topology, and Synchronization: Diagnostics, Failure Modes, and Production Practices

Distributed failures often masquerade as target failures because there are more places for time, CPU, files, connections and results to go wrong. A single local JTL can no longer tell the whole story. Diagnosis starts by preserving controller, every engine and target evidence before changing the plan or “adding more load.”

Load multiplicationVersion driftMissing CSVRMI securityController saturation

Learning objectives

  • Diagnose accidental full-plan load multiplication.
  • Detect JMeter/Java/plugin/version drift before runtime ambiguity.
  • Identify engine-local missing/wrong data files.
  • Reject public jmeter-server and insecure RMI as normal operation.
  • Separate controller/network result bottlenecks from SUT latency.
  • Repair one deliberate engine shard mismatch without deleting original evidence.

1. Preserve first-failure evidence

All runnable diagnosis is loopback/private and ≤12 samples unless explicitly reduced further; the local target remains 127.0.0.1:8025. Preserve aggregate JTL, controller jmeter.log, every engine log, version/file/plugin inventories, exact -R/-G command, listening sockets, engine CPU/heap/network and target JSONL before repair. Do not delete failed evidence or increase threads.

2. Diagnostic sequence

Distributed failure diagnosis

Distributed JMeter is a load-generation system with multiple execution and evidence boundaries. The controller sends one test plan to each engine; every engine executes the full configured workload, reads its own local files/properties, targets the SUT independently, and returns or exports results through a separate network path.

flowchart TD
E[Preserve central/per-engine JTL + logs + target events] --> V[Confirm JMeter / Java / plugin / tool versions on every node]
V --> C[Confirm exact JMX / data / properties / -R/-G command + authorized target]
C --> S[Validate full-plan load math / tree scope / engine-specific properties]
S --> P[Inspect protocol / session / CSV shard / source-IP state]
P --> G[Inspect engine + controller JVM / CPU / GC / network / sender mode]
G --> T[Inspect SUT telemetry / received per-engine work]
T --> D[Inspect RMI ports / SSL / DNS / firewall / clocks / CI-container topology]
D --> F[Least destructive correction]
F --> R[Small controlled rerun]

3. Failure mode: assuming 1,000 threads will be split

Symptoms: target receives roughly N× expected concurrency, rate limits/connection pools collapse, all engines show the full Thread Group and controller report count is multiplied.

Repair: stop immediately, preserve evidence, recalculate per-engine workload and rerun at the smallest safe setting. Do not “fix” the target capacity or increase timeouts for a workload that was never authorized.

4. Failure mode: mismatched JMeter/Java/plugin versions

Possible symptoms: JMX cannot instantiate a component, serialization/remote exceptions, scripts behave differently, SSL/provider differences, or one engine fails to start while another succeeds.

Prevention/repair:

  • compare exact jmeter -v on controller and every engine;
  • compare java -version and approved JVM options;
  • hash every third-party plugin/user JAR and distribute the same set;
  • restart engine after changing JAR/properties;
  • rerun one tiny smoke test before capacity load.

5. Failure mode: CSV file was never deployed

The JMX arrives at engine-b, but ${__P(data.file)} points to a path that does not exist there. Threads stop/CSV variables are missing; aggregate sample count becomes smaller or requests carry wrong data.

Repair the deployment/path on that engine. The controller cannot fix it by resending the JMX because JMeter explicitly does not transfer the external data file.

6. Intentionally broken example: engine-b points to engine-a's shard

Start engine-b with the wrong local property:

# Broken engine-b startup
-Jengine.id=engine-b `
-Jdata.file="$Root\data\engine-a.csv"

Run the same 2×3 remote test. Expected evidence:

  • engine-a uses a001–a006 successfully;
  • engine-b identifies itself as engine-b but sends a-prefixed accounts;
  • fixture returns HTTP 409 wrong_shard for engine-b;
  • aggregate JTL contains engine-b failures with ENGINE_ID=engine-b/ACCOUNT_ID=a...;
  • target JSONL proves the protocol request reached the correct target but carried the wrong engine-local data;
  • RMI/controller may otherwise be completely healthy.

Repair: preserve the failed run, stop/restart only engine-b with data/engine-b.csv, reset target state, use a new run ID, and rerun 12 samples. Do not globally replace CSV variables or disable the assertion.

7. Failure mode: jmeter-server exposed publicly

Remote JMeter engines execute supplied plans and can interact with files/network. Treat them as privileged execution agents. Do not route/port-forward RMI registry/engine/callback ports from the public internet.

Use private networks/VPN/bastion/orchestrated workers, firewall allow-lists and RMI SSL. Public exposure is a security architecture failure, not a testing convenience.

8. Failure mode: disabling RMI SSL to “make it work”

server.rmi.ssl.disable=true can bypass certificate problems, but using it as a normal fix weakens the trust boundary and hides the real issue.

Repair certificate path/password/alias/truststore/clock/hostname or rebuild a valid lab keystore. Do not recommend TLS/RMI verification disablement as the production solution.

9. Failure mode: controller/network saturation

Symptoms: engines/SUT service time stay healthy, but achieved sample rate drops as controller CPU/network rises; remote sender logs show batching/queue/connection pressure. Standard sender can amplify this by returning each sample synchronously.

Compare:

  • engine CPU/GC/network and local completion;
  • controller CPU/heap/network;
  • target event rate/service time;
  • sender mode and result payload fields;
  • central JTL arrival timing.

Potential corrections include default StrippedBatch, lean result fields, backend telemetry/per-engine artifacts, or independent workers—not automatically a larger target timeout.

10. Failure mode: unsynchronized engine clocks

Central result timestamps originate from distributed execution contexts. If engine clocks differ, time-series graphs can interleave events incorrectly and correlations with Chapter 24 telemetry become misleading. Verify NTP/UTC epoch on every host before a distributed run.

11. Failure mode: continue after one engine fails

Setting client.continue_on_fail=true can let a test start with a missing engine. That may be useful for certain resilience workflows, but it silently changes offered load unless your controller explicitly recalculates/records it. Capacity/regression tests should normally fail closed.

12. Causal symptom table

Symptom Generator/distributed cause Target cause to distinguish Evidence
Exactly N× expected requests full-plan replication math mistake target retry/duplicate behavior engine count × Thread Group config + target engine IDs.
One engine contributes zero/few samples version/start/RMI/data/EOF failure target rejects one source engine log + CSV state + target per-engine events.
All engines slow, target service stable controller/result network bottleneck SUT saturation controller CPU/NIC + sender mode + target timing.
Only engine-b fails 409 wrong_shard engine-b data.file mismatch target business regression sample variables + target reason + engine startup manifest.
Remote start fails before traffic RMI SSL/ports/DNS/firewall/version target outage controller/engine RMI logs + zero target events.
Timeline misaligned engine clock skew real asynchronous SUT reaction per-host UTC/NTP + event timestamps.

13. Security-sensitive/disruptive boundaries

Remote engine services, RMI keystores, environment variables, data files, process startup, firewall rules, containers/CI secrets and OS/JVM tuning are privileged operations. Use fake local values and isolated nodes. Never distribute real credentials in a CSV shard for this course lab.

14. Troubleshooting shortcuts to reject

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not allocate giant heaps to mask controller/engine result pressure.
  • Do not mass-disable all listeners/evidence without identifying transfer cost.
  • Do not use global -G hacks for values that must differ by engine.
  • Do not disable TLS/RMI verification.
  • Do not test remote topology fixes against production/public targets.
  • Do not increase workload while an engine is missing/misaligned.
  • Do not delete failed JTL/jmeter.log/engine/target evidence.

Knowledge check

Engine-b sends a-prefixed accounts while ENGINE_ID=engine-b. Which layer failed?

Remote start fails and target receives zero events. What should you inspect first?

Why is a healthy target service time useful when controller throughput collapses?

Why is disabling RMI SSL a bad routine fix?

Why preserve a wrong-shard JTL before repair?

Next lesson

Checkpoint: prove exact distributed workload and repair one engine

Lesson 5 combines topology, inventory, shard hashes, remote/global properties, central/per-engine counts and a deliberate engine-b mismatch into one evidence packet.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. JMeter remote mode sends the test plan to every remote server, but each engine runs the entire test plan; workload is not divided automatically. All controller/server nodes should run exactly the same JMeter version, and JMeter discourages mixing Java versions. External data files are not sent by the controller and must exist on every server in the expected path; plugins/user JARs likewise need an explicit identical deployment. CLI -r starts servers listed in remote_hosts; -Rhost1,host2 explicitly selects/overrides the server list; -Gname=value or -Gpropertyfile sends JMeter properties to remote servers. Since JMeter 4.0, RMI uses SSL by default. JMeter ships create-rmi-keystore, whose generated test certificate is documented as valid for seven days and uses default alias rmi/passphrase changeit; these defaults are suitable only for a disposable private lab. server.rmi.ssl.disable defaults to false and this chapter never disables it. The remote-testing manual describes dynamic server-engine ports conceptually, while the current properties reference lists server.rmi.localport default 4000; production/private labs should set registry/server/callback ports explicitly instead of relying on defaults. The controller reverse callback client.rmi.localport defaults to 0 (random); when set non-zero JMeter can use up to three consecutive ports. Current default remote sample sender mode is StrippedBatch: successful response bodies are stripped and results are batched. Remote mode can consume more resources than equivalent independent CLI workers, and the controller/client or its network can become the bottleneck.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.