Chapter 18Lesson 03~180 minutes

File Uploads, Downloads, Multipart Requests, and Large Payloads: Configuration, Design Patterns, and Trade-Offs

The most maintainable large-payload plan minimizes state that changes during load. Immutable synthetic files, explicit manifests, server-generated upload IDs, bounded result retention, and deterministic cleanup make concurrency safer and make performance evidence easier to reproduce in local and CI environments.

Pre-generated dataImmutable sharingHash validationSequential/concurrentResponse storage

Learning objectives

  • Choose pre-generated versus generated-per-user payloads from the real workload and generator budget.
  • Distinguish immutable shared files from mutable/shared unsafe state.
  • Choose metadata/hash validation versus full response retention.
  • Choose sequential versus concurrent transfer profiles from the performance question.
  • Use JMeter's MD5 and Save Responses to a file options only where their cost/semantics are appropriate.
  • Keep JMeter, JVM, OS/network, SUT/storage, plugin, CI, and container layers separate.

1. Mandatory executable path remains local and bounded

All runnable profiles remain 127.0.0.1:8018, 64/512 KiB synthetic files, ≤2 threads ×2 loops. Public file hosts, production object storage, corporate documents, shared upload endpoints, managed load clouds, and paid telemetry are not required.

2. Pre-generated versus generated-per-user files

Pre-generated bounded files Generated-per-user files
Stable content/hash/size; easy to reproduce and compare. Useful when every real user creates distinct content.
Generation cost is outside transfer timing. Generation consumes generator CPU/disk and may change arrival pacing.
Can be immutable/shared safely when service accepts duplicate content. Needs guarded per-user filenames/storage and cleanup.
Excellent for payload-size experiments. Better for uniqueness/compressibility/content-dependent server behavior when required.

Prefer pre-generated datasets for capacity/regression unless uniqueness/content generation is part of the system behavior. If unique files are required, generate a bounded pool before the timed transfer profile rather than rewriting one file in-place from many threads.

3. Unique versus shared read-only payloads

Many JMeter threads can safely read the same immutable payload file because they do not modify it. The target still creates unique upload IDs, so server-side state remains independently owned.

Unsafe pattern: multiple threads truncate/rewrite payload.txt while other threads upload it. The observed file size/hash becomes scheduler-dependent. Use immutable pre-generated files or one unique path per writer/thread.

4. Full-response retention versus metadata/hash validation

Hash/metadata validation Full body retention
Lean SampleResult/JTL; strong content-integrity signal. Useful for small diagnostic/functional cases.
MD5 mode still downloads/hashes full body. Duplicates large response into heap/result artifact.
Pair with server SHA-256/size for exact file identity. Can explode XML JTL/listener memory/disk.
Best default for large repeated downloads. Use only tiny bounded reproduction with explicit disk budget.

MD5 here is a transfer-integrity comparison, not a password/security design. The manifest also carries SHA-256 for stronger independent content identity.

5. Current MD5 option is version-aware

JMeter 5.6.3 added expression support for the HTTP Request “store as MD5” checkbox. This means a plan can technically toggle it from an expression/property. For maintainability, prefer separate clearly named lean/debug profiles rather than dynamically changing a major result-retention behavior per sample.

6. Save Responses to a file trade-off

This element writes each scoped response to a generated file and can expose the filename to a JMeter variable. It can be better than GUI rendering for a small functional artifact, but can create massive filesystem state under load.

Current behavior also means the filename prefix cannot contain thread variables/functions; JMeter handles numbering/suffixes. If “Don't add number” is used, ensure the path is uniquely safe or responses can overwrite one another.

7. Sequential versus concurrent transfers

A 1-thread small→medium profile isolates payload-size scaling with minimal competition. A 2-thread concurrent profile increases aggregate generator disk reads, socket traffic, server writes/reads, and storage overlap.

Do not compare their p95 as if only “concurrency” changed unless payload mix, loops, timers, server state, result policy, and generator headroom are identical.

8. Payload distribution versus one worst-case size

Real workloads often have a distribution. Model size classes/weights from evidence: for example, 80% 64 KiB, 15% 512 KiB, 5% 1 MiB. A test where every user sends the maximum answers a different stress/worst-case question.

9. Multipart versus raw request body

For POST/PUT/PATCH, a single file with no parameter name can be sent as the entire request body without multipart wrappers. That is correct only when the real API expects raw body upload. If the API expects browser-style multipart with field name file plus metadata fields, use multipart.

Do not optimize away multipart overhead if the production protocol includes it.

10. Browser-compatible multipart headers

The current HTTP Request option suppresses per-part Content-Type and Content-Transfer-Encoding headers and sends only Content-Disposition. Use it only if the target/protocol requires browser-compatible multipart behavior. Changing it can alter request bytes and server parsing.

11. Response retention and JMeter memory properties

httpsampler.max_bytes_to_store_per_request defaults to 0, meaning no truncation. It can cap retained bytes, but changing a global property also changes what post-processors/assertions can see and what evidence is available.

For this chapter's download integrity use case, MD5 mode is clearer: the full response is read and hashed but not retained as body. Do not raise heap or rely on global truncation as the first response to memory pressure.

12. Configuration-layer boundaries

Layer Examples Do not confuse with
JMeter core HTTP multipart fields, MD5 option, assertions, JTL save-service, listeners Server object/file storage behavior.
Java/JVM heap/GC, file buffers, TLS implementation Network bandwidth or disk throughput.
OS/network disk cache, filesystem, NIC, sockets, DNS Server processing/storage latency.
SUT/storage multipart parsing, hash/virus scan, storage write/read, quotas Generator result-file cost.
Plugin/driver None required in mandatory path Core HTTP upload/download semantics.
CI/container workspace quota, artifact upload, ephemeral disk, image startup Target capacity.

13. Worked scenario

Requirement: “95% of uploads are 200 KiB, 5% are 20 MiB; CI runners have 4 GiB RAM and 10 GiB free disk.”

  • Build a representative weighted dataset; do not make every request 20 MiB.
  • Pre-generate synthetic files once and verify hashes.
  • Use lean CSV JTL and metadata/hash validation; avoid full-body listeners.
  • Calculate worst-case concurrent in-flight bytes and result/workspace budget before thread count.
  • Run a separate bounded maximum-size stress scenario if that decision is needed.

14. Decision table

Requirement Preferred pattern Reason/evidence
Reproducible size regression Pre-generated immutable files + manifest Stable size/hash; generation excluded.
Every user requires unique content Pre-generate unique bounded pool No shared mutable-file race.
Repeated large downloads Save response as MD5 + metadata Validates content without body retention.
One failed small response needs inspection Tiny debug profile/full body or Save Response to file Diagnostic artifact explicitly bounded.
Measure max network/storage contention Concurrent representative files with generator headroom Configured/achieved bytes/s + target/generator telemetry.
API expects raw body Single file/no parameter name Avoid multipart only when protocol contract says raw body.

15. Evidence contract

Preserve JMeter/Java versions, exact JMX/upload fields, payload manifest/hash/size distribution, thread/pacing profile, CSV JTL with bytes/sentBytes, matching jmeter.log, generator disk/RSS/heap/CPU/network state, server storage/event telemetry, result/JTL sizes, and cleanup proof. If containers/CI are added later, preserve volume/workspace quotas and artifact-upload behavior too.

Knowledge check

When is sharing one payload file across threads safe?

Why can generated-per-user files reduce achieved load?

When should full response retention be used for a large payload?

What does Browser-compatible multipart headers change?

Why is a maximum-size-only workload not automatically realistic?

Next lesson

Diagnose disk, retention, file-race, and validity failures

Lesson 4 intentionally creates an inefficient full-body XML JTL pattern, then diagnoses why artifact growth is a generator/result problem rather than server latency.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3 with a Java 17 JDK; JMeter 5.6.3 requires Java 8+. For HTTP Request, supplying a File Path causes JMeter to send the request as multipart form data; the file entry also accepts a parameter name and MIME type. “Browser-compatible headers” suppresses per-part Content-Type and Content-Transfer-Encoding headers, leaving Content-Disposition. “Save response as MD5 hash” does not retain the original response in the SampleResult; instead JMeter stores a 32-character MD5 hash and explicitly documents this mode for testing large amounts of data. JMeter 5.6.3 allows ${...} expressions in the “store as MD5” checkbox. Default CSV result settings retain bytes and sentBytes but do not retain response data, sampler data, request headers, or response headers. CSV cannot store response data; demonstrating full-body JTL retention therefore requires a deliberately tiny XML-result run. The HTTP property httpsampler.max_bytes_to_store_per_request defaults to 0 (no truncation), so this chapter uses MD5 mode and lean result retention rather than increasing heap or globally truncating data as a troubleshooting shortcut.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.