Chapter 11Lesson 02~205 minutes

File Uploads, Downloads, Cookies, Storage, and Session State: Guided Hands-On Workflow

Run one coherent loopback workflow that makes every file and browser-state transition observable: generated upload source, browser-selected file, transformed download bytes, synthetic cookie, local/session storage, and explicit teardown.

Local labUploadDownloadCookiesStorage cleanup

Learning objectives

  • Create an isolated project, profile directory, input directory, download directory, and evidence directory.
  • Upload a generated text file by sending its absolute path to a real file input.
  • Trigger a controlled browser download and verify exact filename and byte content without fixed sleeps.
  • Create/read/delete a synthetic cookie and inspect localStorage/sessionStorage on the correct origin.
  • Record before/after session, URL, DOM, file, and state evidence to prove causality.
  • Clean profile/files/cookies/storage deterministically and explain each owner boundary.

1. Scenario and safety boundary

The AUT is a static loopback page. Selecting a text file lets page JavaScript read only the user-selected synthetic file, convert its content to uppercase, append TRANSFORMED , and expose a normal download link. No network upload leaves the machine; the exercise is about WebDriver file-input semantics and browser download/state ownership.

Safety: use only the generated text file. Do not point this lab at personal documents, browser profiles, cloud drives, or production downloads.

2. Create the isolated project

The following example makes the Create the isolated project behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

mkdir selenium-ch11-workflow
cd selenium-ch11-workflow
python -m venv .venv
# Windows PowerShell: .\.venv\Scripts\Activate.ps1
# Linux/macOS: source .venv/bin/activate
python -m pip install "selenium==4.47.0"

The following example makes the Create the isolated project behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from pathlib import Path

root = Path("site")
root.mkdir(exist_ok=True)
(root / "index.html").write_text("""<!doctype html>
<html lang="en"><head><meta charset="utf-8"><title>Browser State Lab</title>
<style>body{font-family:system-ui,sans-serif;max-width:850px;margin:2rem auto;padding:0 1rem}.panel{border:1px solid #8886;border-radius:.6rem;padding:1rem;margin:1rem 0}button,input{padding:.45rem .65rem;margin:.25rem}</style></head>
<body>
<h1>Browser State Lab</h1>
<p id="upload-state">upload:none</p>
<p id="download-state">download:not-ready</p>
<div class="panel">
  <label>Input file <input id="upload" type="file" accept=".txt,text/plain"></label>
  <button id="prepare" type="button">Prepare transformed download</button>
  <a id="download" download="transformed.txt" hidden>Download transformed.txt</a>
</div>
<p id="storage-view">storage:not-read</p>
<script>
let selectedText = '';
const upload = document.querySelector('#upload');
const prepare = document.querySelector('#prepare');
const download = document.querySelector('#download');
const uploadState = document.querySelector('#upload-state');
const downloadState = document.querySelector('#download-state');

upload.addEventListener('change', async () => {
  const file = upload.files[0];
  if (!file) return;
  selectedText = await file.text();
  uploadState.textContent = `upload:${file.name}:${file.size}`;
  download.hidden = true;
  downloadState.textContent = 'download:not-ready';
});

prepare.addEventListener('click', () => {
  if (!selectedText) {
    downloadState.textContent = 'download:error:no-file';
    return;
  }
  const transformed = selectedText.toUpperCase() + 'TRANSFORMED\\n';
  const blob = new Blob([transformed], {type:'text/plain'});
  if (download.dataset.url) URL.revokeObjectURL(download.dataset.url);
  const url = URL.createObjectURL(blob);
  download.dataset.url = url;
  download.href = url;
  download.hidden = false;
  downloadState.textContent = `download:ready:${transformed.length}`;
});

window.addEventListener('beforeunload', () => {
  if (download.dataset.url) URL.revokeObjectURL(download.dataset.url);
});
</script></body></html>""", encoding="utf-8")
print(root.resolve())

Save the Python block as make_fixture.py, run it, then serve only the generated directory:

python make_fixture.py
python -m http.server 8776 --bind 127.0.0.1 --directory site

3. Preflight and state predictions

Before running Selenium, verify http://127.0.0.1:8776/ manually. Predict these transitions:

  1. after send_keys(path), the file input owns one selected file and the page shows its generated name/size;
  2. after Prepare + Download, exactly one new file named transformed.txt appears in this test's download directory with deterministic bytes;
  3. after cookie/storage mutation, only the current loopback origin/session contains the synthetic keys;
  4. after cleanup, the synthetic cookie/storage keys and disposable directories are gone.

4. Run the complete workflow

The following example makes the Run the complete workflow behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from pathlib import Path
import json
import shutil
import tempfile

import selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait

URL = "http://127.0.0.1:8776/"
work = Path(tempfile.mkdtemp(prefix="selenium-ch11-"))
input_dir = work / "input"
download_dir = work / "downloads"
profile_dir = work / "profile"
evidence_dir = work / "evidence"
for p in (input_dir, download_dir, profile_dir, evidence_dir):
    p.mkdir()

source = input_dir / "academy-input.txt"
source_bytes = b"selenium state lab\n"
source.write_bytes(source_bytes)
expected = source_bytes.upper() + b"TRANSFORMED\n"
downloaded = download_dir / "transformed.txt"

options = webdriver.ChromeOptions()
options.add_argument(f"--user-data-dir={profile_dir.resolve()}")
options.add_experimental_option("prefs", {
    "download.default_directory": str(download_dir.resolve()),
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
    "safebrowsing.enabled": True,
})

driver = webdriver.Chrome(options=options)
try:
    driver.get(URL)
    caps = driver.capabilities
    before = {
        "selenium": selenium.__version__,
        "session_id": driver.session_id,
        "browser": caps.get("browserName"),
        "browser_version": caps.get("browserVersion"),
        "url": driver.current_url,
        "cookies": [c["name"] for c in driver.get_cookies()],
        "local_keys": driver.execute_script("return Object.keys(window.localStorage)"),
        "session_keys": driver.execute_script("return Object.keys(window.sessionStorage)"),
        "download_files": [p.name for p in download_dir.iterdir()],
    }

    # 1) File input: changes browser-selected-file state and fixture DOM state.
    file_input = driver.find_element(By.ID, "upload")
    file_input.send_keys(str(source.resolve()))
    WebDriverWait(driver, 3).until(
        lambda d: d.find_element(By.ID, "upload-state").text.startswith("upload:academy-input.txt:")
    )

    # 2) Prepare application output, then trigger the browser-managed download.
    driver.find_element(By.ID, "prepare").click()
    WebDriverWait(driver, 3).until(
        lambda d: d.find_element(By.ID, "download-state").text.startswith("download:ready:")
    )
    driver.find_element(By.ID, "download").click()
    WebDriverWait(driver, 5).until(
        lambda _: downloaded.is_file() and downloaded.stat().st_size == len(expected)
    )
    assert downloaded.read_bytes() == expected

    # 3) Synthetic cookie scoped to the current origin.
    driver.add_cookie({"name": "academy_mode", "value": "synthetic", "sameSite": "Lax"})
    assert driver.get_cookie("academy_mode")["value"] == "synthetic"

    # 4) Web Storage: narrow page-context APIs, not UI interaction bypasses.
    driver.execute_script("window.localStorage.setItem('academy.local', 'synthetic')")
    driver.execute_script("window.sessionStorage.setItem('academy.session', 'synthetic')")
    assert driver.execute_script("return window.localStorage.getItem('academy.local')") == "synthetic"
    assert driver.execute_script("return window.sessionStorage.getItem('academy.session')") == "synthetic"

    after = {
        "upload_state": driver.find_element(By.ID, "upload-state").text,
        "download_state": driver.find_element(By.ID, "download-state").text,
        "download_sha_input_bytes": source.read_text(encoding="utf-8"),
        "download_text": downloaded.read_text(encoding="utf-8"),
        "cookie_names": [c["name"] for c in driver.get_cookies()],
        "local_value": driver.execute_script("return window.localStorage.getItem('academy.local')"),
        "session_value": driver.execute_script("return window.sessionStorage.getItem('academy.session')"),
    }
    (evidence_dir / "run.json").write_text(json.dumps({"before": before, "after": after}, indent=2), encoding="utf-8")
    driver.save_screenshot(str(evidence_dir / "final.png"))

    # Explicit browser-state cleanup while origin/session still exists.
    driver.delete_cookie("academy_mode")
    driver.execute_script("window.localStorage.removeItem('academy.local')")
    driver.execute_script("window.sessionStorage.removeItem('academy.session')")
    assert driver.get_cookie("academy_mode") is None
    assert driver.execute_script("return window.localStorage.getItem('academy.local')") is None
    assert driver.execute_script("return window.sessionStorage.getItem('academy.session')") is None
finally:
    driver.quit()

print("work directory:", work)
print("verified bytes:", downloaded.read_bytes())
# Inspect evidence before deleting the disposable directory.
# shutil.rmtree(work)  # enable after evidence review

The filesystem wait uses WebDriverWait as a deadline/poller even though the condition reads a local file rather than DOM state. There is no fixed sleep and no stale file because the directory was created uniquely for this test.

5. Verify causality state by state

The following table organizes the key choices and evidence for Verify causality state by state. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Action Reads Changes Independent verification
send file path runner file + file-input element browser selected-file list; fixture DOM upload-state shows exact generated name/size
prepare output selected file text fixture Blob URL + ready state download-state becomes ready
click download DOM link / browser download manager per-test download directory file exists with exact expected length and bytes
add cookie current origin/session cookie jar get_cookie value and metadata
set localStorage current page origin origin-persistent storage in disposable profile read same key through page-context storage API
set sessionStorage current origin + context context-scoped storage read same key before context/session ends

6. What changes with RemoteWebDriver/Grid?

The browser may run on a node where the runner path does not exist. Python RemoteWebDriver uses a file detector to recognize a local upload file and can transfer it to the remote end before the browser selects it. Downloads are the reverse problem. If the remote session enables Selenium downloads, RemoteWebDriver can list completed downloadable files, download one back to a runner directory, and delete the remote downloadable files.

from pathlib import Path
from selenium import webdriver

options = webdriver.ChromeOptions()
options.enable_downloads = True
remote = webdriver.Remote("http://127.0.0.1:4444", options=options)
try:
    # After the AUT has triggered a completed browser download:
    print(remote.capabilities.get("se:downloadsEnabled"))
    names = remote.get_downloadable_files()
    if "transformed.txt" in names:
        remote.download_file("transformed.txt", str(Path("remote-artifacts").resolve()))
        remote.delete_downloadable_files()
finally:
    remote.quit()
Optional remote path only. The mandatory lesson does not require Grid. Do not assume browser download directories are shared volumes with the test runner.

7. Why storage script is intentionally narrow

The test does not use JavaScript to click hidden controls, set input values, or bypass interactability. It uses window.localStorage and window.sessionStorage only because those storage APIs are not standard WebDriver commands. Keep such scripts centralized and auditable.

8. Challenge: choose the correct owner/control

A CI runner says /home/runner/input.txt exists, but a remote browser node reports the upload field still empty. Which control is relevant: download preferences, a cookie API, a longer wait, or remote file detection/transfer?

Expected reasoning: this is a runner→remote-browser filesystem boundary. Prove the local file first, inspect whether the RemoteWebDriver file detector is active, then verify the browser-side selected file/application outcome. A longer timeout cannot make a nonexistent remote path appear.

9. Cleanup and rollback

Stop the loopback server. Review evidence/run.json and final.png if needed, then remove the entire unique work directory. Never clean a shared Downloads directory with a wildcard.

from pathlib import Path
import shutil

work = Path("/path/printed/by/the/workflow")
# Verify this is the disposable directory created by the lab before removal.
assert work.name.startswith("selenium-ch11-")
shutil.rmtree(work)
assert not work.exists()

10. Summary and next step

The workflow proved every transition using the state store that owns it. Lesson 3 now turns those mechanics into architecture choices: browser UI versus API setup, temporary versus shared profiles, local versus remote file semantics, and when workflow realism is worth session-state cost.

Knowledge check

Why does the workflow create a unique download directory before session creation?

What proves the download succeeded?

Why is the successful cookie added only after navigating to the loopback origin?

What does a RemoteWebDriver file detector solve?

Why is execute_script not considered a UI bypass here?

Next lesson

File Uploads, Downloads, Cookies, Storage, and Session State: Configuration, Design Patterns, and Trade-Offs

Continue with File Uploads, Downloads, Cookies, Storage, and Session State: Configuration, Design Patterns, and Trade-Offs. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current primary documentation on 2026-08-28. Mandatory examples pin Selenium Python 4.47.0, require Python 3.10+, use a supported locally installed Chromium-family browser with Selenium Manager for driver resolution, and target only 127.0.0.1. Chromium download preferences are intentionally labeled browser-specific. Remote/Grid file-transfer and downloadable-file APIs are optional extensions to the local path. JavaScript execution appears only for Web Storage access, where Selenium explicitly notes local/session storage are not W3C WebDriver commands. UI interactions remain native WebDriver interactions.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.