Chapter 18Lesson 04~195 minutes

File Uploads, Downloads, Multipart Requests, and Large Payloads: Diagnostics, Failure Modes, and Production Practices

Large-payload failures are often self-inflicted by the load generator: sensitive source files, unbounded file sizes, GUI listeners retaining bodies, XML JTL with response data, duplicated response files, mutable shared inputs, or cleanup scripts that delete outside the lab root. Preserve the evidence and identify the state store that is growing before changing heap, threads, or the server.

DiagnosticsDisk exhaustionFull-body JTLMutable-file raceSafe cleanup

Learning objectives

  • Reject sensitive documents and uncontrolled payload roots.
  • Diagnose generator disk/heap growth caused by result retention.
  • Detect races caused by one mutable file shared across threads.
  • Separate payload generation time from HTTP/server timing.
  • Reject unbounded payload size and unsafe recursive deletion.
  • Use a least-destructive diagnostic sequence and preserve failed evidence.

1. Preserve first-failure evidence

All reproduction stays on 127.0.0.1:8018 with synthetic ≤512 KiB mandatory files. Preserve JTL, jmeter.log, JMX/result-save configuration, payload manifest/hashes, disk snapshot, generator process state, server events/stats, and exact CLI before repair. Do not delete the evidence just because it is large.

2. Diagnostic sequence

Large-payload diagnostic sequence

File-transfer elapsed time can contain several different costs. The experiment is valid only when payload generation, multipart encoding, network transfer, server storage, response buffering/hash work, and result-file storage are kept distinguishable.

flowchart TD
E[Preserve JTL + jmeter.log + manifest + server events] --> V[Confirm JMeter / Java / tool versions]
V --> C[Confirm JMX/data/properties/CLI + authorized target]
C --> S[Inspect HTTP file/multipart/MD5/listener scope + variables]
S --> D[Inspect payload size/hash/path + result/storage directories]
D --> G[Inspect generator heap/GC/CPU/disk/network]
G --> T[Inspect server multipart/hash/storage/download timing]
T --> X[Inspect distributed/CI/container workspace/artifact state if relevant]
X --> F[Least destructive correction]
F --> R[Small controlled rerun + compare]

3. Failure mode: using sensitive documents as “realistic data”

A sensitive business PDF/backup is copied into the repo or CI workspace for load testing. That creates a privacy/retention incident and may expose content through failed response/request capture.

Repair: synthetic structurally representative data, generated manifest/hashes, and explicit classification rules. Real production documents are not required to model byte size/compressibility unless authorized privacy controls explicitly allow them.

4. Failure mode: filling generator disk with artifacts

Every 512 KiB download is saved in JTL plus “Save Responses to a file,” while the server fixture also keeps uploads. The generator may run out of disk before the SUT shows any saturation.

Inspect free disk, JTL/result directory growth, saved-response file count, and target storage independently. Restore lean retention first; do not simply redirect files to another disk without understanding why they are being written.

5. Intentionally broken example: full body in XML JTL

For exactly one medium 512 KiB download, deliberately:

  • uncheck Save response as MD5 hash;
  • run 1 thread ×1 download only;
  • set JTL output to XML;
  • enable jmeter.save.saveservice.response_data=true.

Current JMeter CSV cannot store response data, which is why the demonstration uses XML. Because the fixture returns text/plain, the 512 KiB body can be serialized into the result file.

Expected: JTL size increases by roughly the retained body plus XML escaping/metadata; generator heap/disk work is higher even though server payload processing is unchanged.

Repair: preserve the inefficient JTL, restore CSV + response_data=false + MD5 download mode, rerun the same single download, and compare JTL size/generator state/server service time.

6. Failure mode: one mutable file reused unsafely

Thread A rewrites payload.txt while Thread B reads it for upload. The manifest hash no longer describes a stable object; partial/truncated uploads can appear nondeterministically.

Repair: immutable shared file or pre-generated unique file per writer/thread. Never use a global property as a “lock” for concurrent filesystem mutation.

7. Failure mode: measuring local generation as server latency

A preprocessor/script creates a 100 MiB file immediately before every HTTP sampler and the transaction controller wraps both generation and upload. Reported “upload transaction” time is mostly generator file creation.

Keep generation as a separately labeled setup/action if real users generate content; otherwise pre-generate outside measured transfer timing. Report both only when the business question intentionally includes both.

8. Failure mode: unbounded payload sizes

A CSV contains arbitrary file paths/sizes, or a random-size function has no ceiling. A typo can generate tens of GiB per thread.

Use a manifest with explicit size classes, server hard limits, test-level byte budget, thread/loop ceilings, free-disk abort threshold, and only approved roots. The local fixture rejects files above 1 MiB.

9. Failure mode: deleting outside the guarded temp root

Unsafe cleanup scripts concatenate user variables into rm -rf/Remove-Item -Recurse paths. One empty or .. value can delete unrelated data.

The fixture avoids this by generating upload IDs server-side and checking the stored file's resolved parent before unlink. Manual reset should be performed only after the server stops and after verifying the exact lab directory name/path.

10. Failure mode: listener/JTL retention mistaken for SUT slowdown

View Results Tree holds samples for display, and full-response XML JTL writes large bodies to disk. Generator CPU/GC/disk wait rises; achieved request rate falls; target service time can remain flat.

Compare server service_wall_ms and generator/JTL growth. If server timing is unchanged while JMeter elapsed/throughput changes after result retention is enabled, the evidence points toward generator/result cost.

11. Performance causality table

Symptom Generator/network cause Target cause to distinguish Evidence
Upload p95 rises, server upload wall flat file read/NIC/send/result overhead multipart/hash/storage slowdown JTL + generator disk/NIC + server events.
Download p95 rises only when full-body retention enabled heap/GC/JTL disk/listener server read/storage slowdown lean vs retained-body JTL + server wall.
sentBytes lower than expected file/path/error/request aborted server rejected/truncated manifest + JTL + server request body bytes/status.
Server storage grows between runs missing delete/failed cleanup storage service leak /stats + event upload/delete counts.
CI-only failure workspace quota/artifact upload/container disk SUT regression CI disk/artifacts + same server telemetry.

12. Distributed/CI note

In distributed JMeter, every injector needs the payload files at the expected path; JMeter does not magically distribute arbitrary data files as part of remote load semantics. Large response/result transfer back to a controller can also become a network bottleneck. Keep payload/result ownership explicit per engine.

In CI, workspace quota and artifact upload can dwarf the test itself. Do not upload giant raw response bodies by default.

13. Troubleshooting shortcuts to reject

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not increase heap immediately to accommodate avoidable full-body retention.
  • Do not disable all listeners “because listeners are bad” without evidence; remove/limit the specific expensive path.
  • Do not move file paths/locks into mutable global properties.
  • Do not disable TLS/RMI verification.
  • Do not point tests at production file/object storage.
  • Do not increase payload/thread counts while disk/cleanup limits are unresolved.
  • Do not delete result files before first-failure analysis.

Knowledge check

Why does XML JTL grow dramatically when response_data=true?

What evidence distinguishes result-retention overhead from server slowdown?

Why is one mutable shared payload file unsafe?

What is the safe deletion principle?

Why can a CI run fail even when local/SUT metrics are healthy?

Next lesson

Checkpoint: quantify the cost of retaining one large body

Lesson 5 runs the lean small/medium upload/download path, deliberately turns on one full-body XML JTL pattern for a single medium download, compares artifact/generator/server evidence, restores MD5/CSV, and proves zero storage leftovers.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. For HTTP Request, supplying a File Path causes JMeter to send the request as multipart form data; the file entry also accepts a parameter name and MIME type. “Browser-compatible headers” suppresses per-part Content-Type and Content-Transfer-Encoding headers, leaving Content-Disposition. “Save response as MD5 hash” does not retain the original response in the SampleResult; instead JMeter stores a 32-character MD5 hash and explicitly documents this mode for testing large amounts of data. JMeter 5.6.3 allows ${...} expressions in the “store as MD5” checkbox. Default CSV result settings retain bytes and sentBytes but do not retain response data, sampler data, request headers, or response headers. CSV cannot store response data; demonstrating full-body JTL retention therefore requires a deliberately tiny XML-result run. The HTTP property httpsampler.max_bytes_to_store_per_request defaults to 0 (no truncation), so this chapter uses MD5 mode and lean result retention rather than increasing heap or globally truncating data as a troubleshooting shortcut.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.