Chapter 18Lesson 01~155 minutes

File Uploads, Downloads, Multipart Requests, and Large Payloads: Core Concepts and Mental Model

Chapter 17 showed that protocol connection and message boundaries are part of the measurement. File transfers add a second dimension: the bytes themselves consume generator disk bandwidth, page cache, heap/buffers, network bandwidth, server memory/storage, and result-artifact space. A “slow upload” is not automatically a slow server; the generator may be reading files poorly, the network may be full, multipart encoding may add overhead, the server may hash/write to disk, or the test may be retaining every downloaded body.

Multipart uploadDownloadLarge payloadsMD5 storageDisk/heap validity

Learning objectives

  • Model synthetic payload source → multipart upload → server storage → download → validation → JTL as distinct state/cost boundaries.
  • Separate payload generation, file reads, network bytes, server hashing/storage, response buffering, and result-file retention.
  • Understand JMeter's HTTP file-upload fields and Save response as MD5 hash behavior.
  • Use JTL bytes/sentBytes together with exact server payload metadata instead of guessing transfer volume.
  • Inspect local payloads, generator disk, target storage, and result policy before load.
  • Keep large-payload claims conditional on generator/network/server evidence.

1. The practical problem: payload size can move the bottleneck

A 2 KiB API response mostly exercises request processing. A 500 KiB upload/download can be dominated by file reads, TCP transfer, hashing, object/file storage, response buffering, listener rendering, or JTL output. If the generator disk is full or the JVM retains every body, adding threads can reduce achieved traffic before the server saturates.

Mandatory lab boundary: only synthetic ASCII payloads of 64 KiB and 512 KiB, localhost service at 127.0.0.1:8018, maximum 2 threads, maximum 2 transfer journeys per thread, no payload above 1 MiB, no sensitive documents, and all storage confined to guarded lab directories. Never substitute a public, shared, corporate, or production upload/download endpoint or storage bucket for this fixture.

2. Mental model: bytes move through several independent cost centers

Upload/download cost and evidence flow

File-transfer elapsed time can contain several different costs. The experiment is valid only when payload generation, multipart encoding, network transfer, server storage, response buffering/hash work, and result-file storage are kept distinguishable.

flowchart TD
F[Synthetic payload files + manifest] --> H[JMeter HTTP Request / multipart encoder]
H --> N[Loopback network transfer]
N --> S[Server parses multipart + hashes + writes guarded storage]
S --> M[Upload JSON metadata]
M --> X[Extract upload ID / size / hashes]
X --> D[HTTP download]
D --> B[JMeter response read / MD5 mode or full body]
B --> A[Hash/metadata assertions]
H --> J[JTL sentBytes]
D --> J2[JTL bytes]
S --> E[Server storage/event telemetry]
G[Generator disk / heap / CPU / sockets] --> V[Validity review]
J --> V
J2 --> V
E --> V

The synthetic payload file is generator-side data. HTTP Request reads it and builds a multipart request; that adds boundaries, per-part headers, and form fields around the payload. The server reads/parses the body, hashes the file, and writes it under a guarded storage root. Upload JSON returns only metadata and a server-generated upload ID. Download streams the stored bytes back. JMeter either retains the body or—in the lean path—stores only its MD5 hash. JTL reports client-observed timing and byte counts; the server event log reports exact payload/body/storage state; generator CPU/heap/disk confirms the injector is not the hidden bottleneck.

3. Multipart/form-data is an envelope around the file

Current HTTP Request behavior automatically uses multipart form data when File Path is populated. Each file entry has a file path, form parameter name, and MIME type. Additional form parameters such as run_id and payload_name become other multipart parts.

Therefore sentBytes is expected to exceed the raw payload size: request line/headers, multipart boundaries, Content-Disposition, Content-Type, and metadata fields all add bytes. That overhead is real protocol traffic but is not extra file content.

4. Save response as MD5 hash changes retention, not network transfer

When Save response as MD5 hash is enabled, JMeter still downloads the response bytes so it can calculate the MD5. It then stores only a 32-character MD5 value in the SampleResult rather than the original body. This reduces response-body retention pressure, but does not reduce network bytes or eliminate CPU needed to read/hash the response.

In the chapter, the payload generator computes both SHA-256 and MD5. Upload metadata is validated with SHA-256/size; the lean download path validates JMeter's stored MD5 against the manifest.

5. CSV JTL can record transfer size without storing bodies

JMeter's CSV result format includes bytes and sentBytes when those save-service fields are enabled (they are defaults in the current listener configuration). Default CSV also avoids response data, sampler data, and request/response headers.

This is a strong large-payload baseline: retain timing, success, response code/message, active-thread counts, bytes/sentBytes, and assertions—without duplicating a 512 KiB response into every JTL row.

6. Response-to-disk is another state store

Save Responses to a file writes response data to generated files and can expose filenames through a variable. It is useful for bounded functional/debug cases, but under load it moves the pressure from JTL/heap toward generator disk/inodes and cleanup. The current component also warns that the filename prefix itself must not contain thread variables/functions.

Do not enable response-to-file across thousands of large downloads unless that local-file behavior is explicitly part of the experiment and the disk budget/cleanup is designed.

7. State to define before changing a large-payload test

State Question
Generator JMeter/Java version; free disk; heap/RSS/CPU; file cache; NIC; listener/result policy?
Thread/arrival Threads, loops, pacing, sequential/concurrent transfers, configured bytes/s?
Component scope HTTP Defaults/Header Manager/file upload fields/MD5 option/assertions?
Variables/data Payload path/name/size/hash, RUN_ID, server UPLOAD_ID; shared immutable or unique file?
HTTP/session KeepAlive/TLS/auth/multipart boundary/content type; connection reuse?
Target Storage root, upload count/bytes, hash/write time, download bytes, cleanup state?
Artifacts Manifest, JTL, jmeter.log, server events, optional full-body debug JTL/files?
Credential/trust Synthetic/no-secret local path; authorized HTTPS trust when real environments are used?
Validity Can generator disk/heap/network support achieved transfer volume without distortion?

8. Read-only inspection before the first upload

  • Payload manifest: exact file paths, sizes, SHA-256, MD5; confirm files are synthetic.
  • Generator: free disk, payload directory size, result directory size, Java/JMeter versions.
  • JMX: exact file path/parameter/MIME, multipart mode, MD5 checkbox, result save policy.
  • Target: health/version, current stored upload count/bytes, storage root.
  • Run budget: expected number of uploads/downloads and approximate payload bytes.

Do not test cleanup by deleting arbitrary files. Prove the guarded root/path first.

9. Causal performance questions

Observation Possible cause Independent evidence
Upload elapsed rises; server wall stable generator file read/network/send pressure generator disk/CPU/NIC + server request-body timing.
Server wall rises with size multipart parsing/hash/storage becomes target cost server service_wall_ms + storage/process telemetry.
Download elapsed rises; server wall stable network/client buffering/hash/result cost generator CPU/NIC + MD5/full-body setting.
JTL/result directory grows quickly result retention rather than server workload JTL/file sizes + save-service/listener config.
Configured transfers stay fixed but achieved rate falls generator or target saturation/backpressure active threads, throughput, CPU/GC/NIC/server timing.

10. Configured versus achieved transfer load

Configured load is the intended threads, loops, payload-size mix, and theoretical payload bytes offered to the target. Achieved load is the successfully completed upload/download count and byte rate actually observed after generator disk/heap/network, HTTP connection behavior, server processing/storage, and errors are accounted for. Large-payload results must report both; thread count alone does not prove achieved bytes per second.

11. DevOps connection

A release claim about file transfer must state payload distribution, concurrency, network path, storage backend, validation method, result retention, injector capacity, and cleanup. “512 KiB upload p95=…” is not sufficient if the test accidentally hashes/generated files inside the timed sampler, runs out of disk, or buffers every response in a GUI listener.

Knowledge check

Why can sentBytes be larger than the uploaded file?

Does Save response as MD5 hash reduce network transfer?

Why keep server payload_bytes in addition to JTL bytes?

What is dangerous about Save Responses to a file under large load?

Why must generator disk/heap/network be reported before target-capacity claims?

Next lesson

Build the bounded multipart upload/download workflow

Lesson 2 generates deterministic payloads, runs a localhost-only upload service, configures multipart HTTP Request, validates returned hashes/size, downloads with MD5 storage, records bytes sent/received, and verifies delete/cleanup state.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. For HTTP Request, supplying a File Path causes JMeter to send the request as multipart form data; the file entry also accepts a parameter name and MIME type. “Browser-compatible headers” suppresses per-part Content-Type and Content-Transfer-Encoding headers, leaving Content-Disposition. “Save response as MD5 hash” does not retain the original response in the SampleResult; instead JMeter stores a 32-character MD5 hash and explicitly documents this mode for testing large amounts of data. JMeter 5.6.3 allows ${...} expressions in the “store as MD5” checkbox. Default CSV result settings retain bytes and sentBytes but do not retain response data, sampler data, request headers, or response headers. CSV cannot store response data; demonstrating full-body JTL retention therefore requires a deliberately tiny XML-result run. The HTTP property httpsampler.max_bytes_to_store_per_request defaults to 0 (no truncation), so this chapter uses MD5 mode and lean result retention rather than increasing heap or globally truncating data as a troubleshooting shortcut.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.