File Uploads, Downloads, Cookies, Storage, and Session State: Guided Hands-On Workflow
Run one coherent loopback workflow that makes every file and browser-state transition observable: generated upload source, browser-selected file, transformed download bytes, synthetic cookie, local/session storage, and explicit teardown.
Learning objectives
- Create an isolated project, profile directory, input directory, download directory, and evidence directory.
- Upload a generated text file by sending its absolute path to a real file input.
- Trigger a controlled browser download and verify exact filename and byte content without fixed sleeps.
- Create/read/delete a synthetic cookie and inspect localStorage/sessionStorage on the correct origin.
- Record before/after session, URL, DOM, file, and state evidence to prove causality.
- Clean profile/files/cookies/storage deterministically and explain each owner boundary.
1. Scenario and safety boundary
The AUT is a static loopback page. Selecting a text file lets page
JavaScript read only the user-selected synthetic file, convert its
content to uppercase, append TRANSFORMED , and expose a
normal download link. No network upload leaves the machine; the
exercise is about WebDriver file-input semantics and browser
download/state ownership.
2. Create the isolated project
The following example makes the Create the isolated project behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
mkdir selenium-ch11-workflow
cd selenium-ch11-workflow
python -m venv .venv
# Windows PowerShell: .\.venv\Scripts\Activate.ps1
# Linux/macOS: source .venv/bin/activate
python -m pip install "selenium==4.47.0"
The following example makes the Create the isolated project behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from pathlib import Path
root = Path("site")
root.mkdir(exist_ok=True)
(root / "index.html").write_text("""<!doctype html>
<html lang="en"><head><meta charset="utf-8"><title>Browser State Lab</title>
<style>body{font-family:system-ui,sans-serif;max-width:850px;margin:2rem auto;padding:0 1rem}.panel{border:1px solid #8886;border-radius:.6rem;padding:1rem;margin:1rem 0}button,input{padding:.45rem .65rem;margin:.25rem}</style></head>
<body>
<h1>Browser State Lab</h1>
<p id="upload-state">upload:none</p>
<p id="download-state">download:not-ready</p>
<div class="panel">
<label>Input file <input id="upload" type="file" accept=".txt,text/plain"></label>
<button id="prepare" type="button">Prepare transformed download</button>
<a id="download" download="transformed.txt" hidden>Download transformed.txt</a>
</div>
<p id="storage-view">storage:not-read</p>
<script>
let selectedText = '';
const upload = document.querySelector('#upload');
const prepare = document.querySelector('#prepare');
const download = document.querySelector('#download');
const uploadState = document.querySelector('#upload-state');
const downloadState = document.querySelector('#download-state');
upload.addEventListener('change', async () => {
const file = upload.files[0];
if (!file) return;
selectedText = await file.text();
uploadState.textContent = `upload:${file.name}:${file.size}`;
download.hidden = true;
downloadState.textContent = 'download:not-ready';
});
prepare.addEventListener('click', () => {
if (!selectedText) {
downloadState.textContent = 'download:error:no-file';
return;
}
const transformed = selectedText.toUpperCase() + 'TRANSFORMED\\n';
const blob = new Blob([transformed], {type:'text/plain'});
if (download.dataset.url) URL.revokeObjectURL(download.dataset.url);
const url = URL.createObjectURL(blob);
download.dataset.url = url;
download.href = url;
download.hidden = false;
downloadState.textContent = `download:ready:${transformed.length}`;
});
window.addEventListener('beforeunload', () => {
if (download.dataset.url) URL.revokeObjectURL(download.dataset.url);
});
</script></body></html>""", encoding="utf-8")
print(root.resolve())
Save the Python block as make_fixture.py, run it, then
serve only the generated directory:
python make_fixture.py
python -m http.server 8776 --bind 127.0.0.1 --directory site
3. Preflight and state predictions
Before running Selenium, verify
http://127.0.0.1:8776/ manually. Predict these
transitions:
-
after
send_keys(path), the file input owns one selected file and the page shows its generated name/size; -
after Prepare + Download, exactly one new file named
transformed.txtappears in this test's download directory with deterministic bytes; - after cookie/storage mutation, only the current loopback origin/session contains the synthetic keys;
- after cleanup, the synthetic cookie/storage keys and disposable directories are gone.
4. Run the complete workflow
The following example makes the Run the complete workflow behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from pathlib import Path
import json
import shutil
import tempfile
import selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
URL = "http://127.0.0.1:8776/"
work = Path(tempfile.mkdtemp(prefix="selenium-ch11-"))
input_dir = work / "input"
download_dir = work / "downloads"
profile_dir = work / "profile"
evidence_dir = work / "evidence"
for p in (input_dir, download_dir, profile_dir, evidence_dir):
p.mkdir()
source = input_dir / "academy-input.txt"
source_bytes = b"selenium state lab\n"
source.write_bytes(source_bytes)
expected = source_bytes.upper() + b"TRANSFORMED\n"
downloaded = download_dir / "transformed.txt"
options = webdriver.ChromeOptions()
options.add_argument(f"--user-data-dir={profile_dir.resolve()}")
options.add_experimental_option("prefs", {
"download.default_directory": str(download_dir.resolve()),
"download.prompt_for_download": False,
"download.directory_upgrade": True,
"safebrowsing.enabled": True,
})
driver = webdriver.Chrome(options=options)
try:
driver.get(URL)
caps = driver.capabilities
before = {
"selenium": selenium.__version__,
"session_id": driver.session_id,
"browser": caps.get("browserName"),
"browser_version": caps.get("browserVersion"),
"url": driver.current_url,
"cookies": [c["name"] for c in driver.get_cookies()],
"local_keys": driver.execute_script("return Object.keys(window.localStorage)"),
"session_keys": driver.execute_script("return Object.keys(window.sessionStorage)"),
"download_files": [p.name for p in download_dir.iterdir()],
}
# 1) File input: changes browser-selected-file state and fixture DOM state.
file_input = driver.find_element(By.ID, "upload")
file_input.send_keys(str(source.resolve()))
WebDriverWait(driver, 3).until(
lambda d: d.find_element(By.ID, "upload-state").text.startswith("upload:academy-input.txt:")
)
# 2) Prepare application output, then trigger the browser-managed download.
driver.find_element(By.ID, "prepare").click()
WebDriverWait(driver, 3).until(
lambda d: d.find_element(By.ID, "download-state").text.startswith("download:ready:")
)
driver.find_element(By.ID, "download").click()
WebDriverWait(driver, 5).until(
lambda _: downloaded.is_file() and downloaded.stat().st_size == len(expected)
)
assert downloaded.read_bytes() == expected
# 3) Synthetic cookie scoped to the current origin.
driver.add_cookie({"name": "academy_mode", "value": "synthetic", "sameSite": "Lax"})
assert driver.get_cookie("academy_mode")["value"] == "synthetic"
# 4) Web Storage: narrow page-context APIs, not UI interaction bypasses.
driver.execute_script("window.localStorage.setItem('academy.local', 'synthetic')")
driver.execute_script("window.sessionStorage.setItem('academy.session', 'synthetic')")
assert driver.execute_script("return window.localStorage.getItem('academy.local')") == "synthetic"
assert driver.execute_script("return window.sessionStorage.getItem('academy.session')") == "synthetic"
after = {
"upload_state": driver.find_element(By.ID, "upload-state").text,
"download_state": driver.find_element(By.ID, "download-state").text,
"download_sha_input_bytes": source.read_text(encoding="utf-8"),
"download_text": downloaded.read_text(encoding="utf-8"),
"cookie_names": [c["name"] for c in driver.get_cookies()],
"local_value": driver.execute_script("return window.localStorage.getItem('academy.local')"),
"session_value": driver.execute_script("return window.sessionStorage.getItem('academy.session')"),
}
(evidence_dir / "run.json").write_text(json.dumps({"before": before, "after": after}, indent=2), encoding="utf-8")
driver.save_screenshot(str(evidence_dir / "final.png"))
# Explicit browser-state cleanup while origin/session still exists.
driver.delete_cookie("academy_mode")
driver.execute_script("window.localStorage.removeItem('academy.local')")
driver.execute_script("window.sessionStorage.removeItem('academy.session')")
assert driver.get_cookie("academy_mode") is None
assert driver.execute_script("return window.localStorage.getItem('academy.local')") is None
assert driver.execute_script("return window.sessionStorage.getItem('academy.session')") is None
finally:
driver.quit()
print("work directory:", work)
print("verified bytes:", downloaded.read_bytes())
# Inspect evidence before deleting the disposable directory.
# shutil.rmtree(work) # enable after evidence review
The filesystem wait uses WebDriverWait as a
deadline/poller even though the condition reads a local file rather
than DOM state. There is no fixed sleep and no stale file because
the directory was created uniquely for this test.
5. Verify causality state by state
The following table organizes the key choices and evidence for Verify causality state by state. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Action | Reads | Changes | Independent verification |
|---|---|---|---|
| send file path | runner file + file-input element | browser selected-file list; fixture DOM | upload-state shows exact generated name/size |
| prepare output | selected file text | fixture Blob URL + ready state | download-state becomes ready |
| click download | DOM link / browser download manager | per-test download directory | file exists with exact expected length and bytes |
| add cookie | current origin/session | cookie jar | get_cookie value and metadata |
| set localStorage | current page origin | origin-persistent storage in disposable profile | read same key through page-context storage API |
| set sessionStorage | current origin + context | context-scoped storage | read same key before context/session ends |
6. What changes with RemoteWebDriver/Grid?
The browser may run on a node where the runner path does not exist. Python RemoteWebDriver uses a file detector to recognize a local upload file and can transfer it to the remote end before the browser selects it. Downloads are the reverse problem. If the remote session enables Selenium downloads, RemoteWebDriver can list completed downloadable files, download one back to a runner directory, and delete the remote downloadable files.
from pathlib import Path
from selenium import webdriver
options = webdriver.ChromeOptions()
options.enable_downloads = True
remote = webdriver.Remote("http://127.0.0.1:4444", options=options)
try:
# After the AUT has triggered a completed browser download:
print(remote.capabilities.get("se:downloadsEnabled"))
names = remote.get_downloadable_files()
if "transformed.txt" in names:
remote.download_file("transformed.txt", str(Path("remote-artifacts").resolve()))
remote.delete_downloadable_files()
finally:
remote.quit()
7. Why storage script is intentionally narrow
The test does not use JavaScript to click hidden controls, set input
values, or bypass interactability. It uses
window.localStorage and
window.sessionStorage only because those storage APIs
are not standard WebDriver commands. Keep such scripts centralized
and auditable.
8. Challenge: choose the correct owner/control
A CI runner says /home/runner/input.txt exists, but a
remote browser node reports the upload field still empty. Which
control is relevant: download preferences, a cookie API, a longer
wait, or remote file detection/transfer?
Expected reasoning: this is a runner→remote-browser filesystem boundary. Prove the local file first, inspect whether the RemoteWebDriver file detector is active, then verify the browser-side selected file/application outcome. A longer timeout cannot make a nonexistent remote path appear.
9. Cleanup and rollback
Stop the loopback server. Review evidence/run.json and
final.png if needed, then remove the entire unique work
directory. Never clean a shared Downloads directory with a wildcard.
from pathlib import Path
import shutil
work = Path("/path/printed/by/the/workflow")
# Verify this is the disposable directory created by the lab before removal.
assert work.name.startswith("selenium-ch11-")
shutil.rmtree(work)
assert not work.exists()
10. Summary and next step
The workflow proved every transition using the state store that owns it. Lesson 3 now turns those mechanics into architecture choices: browser UI versus API setup, temporary versus shared profiles, local versus remote file semantics, and when workflow realism is worth session-state cost.
Knowledge check
Why does the workflow create a unique download directory before session creation?
It establishes a clean ownership boundary and prevents stale files from previous tests or retries from satisfying the current assertion.
What proves the download succeeded?
The expected file appears in the unique directory under a bounded deadline and its exact bytes equal the deterministic expected transformation.
Why is the successful cookie added only after navigating to the loopback origin?
Cookie scope is tied to a domain/origin context. Navigating first gives WebDriver the correct host context for a host-scoped synthetic cookie.
What does a RemoteWebDriver file detector solve?
It bridges a local-runner file path to a remote browser host for file-input upload instead of assuming the remote node can see the runner filesystem.
Why is execute_script not considered a UI bypass
here?
It is used only for Web Storage APIs that classic W3C WebDriver does not expose; all user interactions remain normal WebDriver commands.
Official references and version notes
- Selenium 4.47 release notes — stable baseline pinned for this chapter.
- Selenium downloads — current stable bindings and Selenium Server/Grid versions.
- File upload — use a file input and send the full path; do not automate the OS chooser.
- Working with cookies — WebDriver cookie create/read/delete semantics.
- Python RemoteWebDriver API — LocalFileDetector default plus remote downloadable-file methods.
-
Python common Options API
—
enable_downloadscapability surface. - File downloads guidance — browser-triggered downloads do not provide portable progress semantics; prefer lower-layer verification where appropriate.
- Selenium API deprecations — Web Storage — local/session storage are not W3C WebDriver commands; use page-context script when storage inspection is required.
-
Python exceptions
— includes
InvalidCookieDomainException.
Version-sensitive behavior was rechecked against current primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python 4.47.0, require Python 3.10+, use a supported locally
installed Chromium-family browser with Selenium Manager for driver
resolution, and target only 127.0.0.1. Chromium
download preferences are intentionally labeled browser-specific.
Remote/Grid file-transfer and downloadable-file APIs are optional
extensions to the local path. JavaScript execution appears only
for Web Storage access, where Selenium explicitly notes
local/session storage are not W3C WebDriver commands. UI
interactions remain native WebDriver interactions.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.