Chapter 21Lesson 04~240 minutes

Parallel Execution, Isolation, Concurrency, and Test Sharding: Diagnostics, Failure Modes, and Production Practices

Concurrency failures are often misdiagnosed as random Selenium flakiness. This lesson deliberately breaks ownership, identity, ports, ordering, capacity, retries, and artifact naming, then applies the same first-failure diagnostic sequence used throughout the course.

Driver raceData collisionPort collisionOverloadArtifact race

Learning objectives

  • Recognize driver sharing and mutable-data sharing as separate race classes.
  • Diagnose port, ordering, capacity, retry-load, and artifact-name collisions from evidence.
  • Reproduce a deterministic shared-identity collision with separate browsers.
  • Preserve first-failure evidence before rerunning or reducing concurrency.
  • Apply the least destructive correction instead of adding retries or giant waits.

1. Broken example: several threads share one WebDriver

This is intentionally wrong. Both tasks mutate the same session. One can navigate while the other locates, one can close a window the other owns, and teardown can race with commands.

from concurrent.futures import ThreadPoolExecutor
from selenium import webdriver

shared_driver = webdriver.Chrome()  # intentionally broken shared mutable session


def broken(case_id):
    shared_driver.get("http://127.0.0.1:8765/work?case=" + case_id)
    return shared_driver.title

try:
    with ThreadPoolExecutor(max_workers=2) as pool:
        print(list(pool.map(broken, ["a", "b"])))
finally:
    shared_driver.quit()
Interpret the symptom, not only the exception

The result can be wrong-title assertions, unexpected navigation, stale elements, closed-session errors, or apparently passing tests that observed another worker’s page. The defect is ownership, not “Selenium needs a retry.”

2. Separate browsers can still collide through shared AUT data

Use the Chapter 21 fixture’s /claim endpoint. Two fresh browser sessions request the same synthetic identity at nearly the same time. One claim succeeds; the other receives a deterministic collision state. This proves that one-driver-per-test is necessary but not sufficient.

from concurrent.futures import ThreadPoolExecutor
from threading import Barrier
from selenium import webdriver
from selenium.webdriver.common.by import By

barrier = Barrier(2)


def claim_same_identity(label):
    driver = webdriver.Chrome()
    try:
        barrier.wait(timeout=10)
        driver.get("http://127.0.0.1:8765/claim?identity=shared@example.test&hold_ms=450")
        return label, driver.find_element(By.CSS_SELECTOR, "[data-testid='status']").text
    finally:
        driver.quit()

with ThreadPoolExecutor(max_workers=2) as pool:
    print(list(pool.map(claim_same_identity, ["A", "B"])))

Expected result: one claimed and one collision. Repair the data contract by giving each test a unique identity such as claim-A@example.test and claim-B@example.test; do not serialize the whole suite unless the resource is truly global.

3. Port and artifact collisions are ordinary shared-resource races

Two workers binding hard-coded port 8765 can produce “address already in use.” Two workers writing failure.png can overwrite one another. Use one fixture server owned by the suite, or let the OS choose an ephemeral port and publish it through fixture configuration. Namespace evidence by run/test/attempt/session.

from pathlib import Path


def evidence_path(root: Path, run_id: str, test_id: str, attempt: int, name: str) -> Path:
    safe_test = test_id.replace("/", "_").replace("::", "__")
    path = root / run_id / safe_test / f"attempt-{attempt}"
    path.mkdir(parents=True, exist_ok=True)
    return path / name

4. Hidden test order becomes nondeterministic under parallel scheduling

If Test B expects Test A to create data first, parallelism merely exposes an existing dependency. Preserve the failing order/evidence, then move setup into B’s own fixture or a lower-layer API. Do not force alphabetical execution to hide the dependency.

5. More workers can create infrastructure incidents

Runner workers above Grid slots create queue pressure. Sessions above safe CPU/RAM create browser instability. Requests above AUT limits create 429/5xx or slow responses. Retries multiply these loads. A two-attempt retry on 20 simultaneously failing tests can turn 20 failing attempts into roughly 40 browser attempts and worsen the incident.

Do not normalize overload with retries

Keep the first attempt’s artifacts. Reduce the experiment to the smallest reproducible concurrency, inspect Grid queue/Node state and host/AUT resources, then adjust capacity or worker limits.

6. Concurrency diagnostic sequence

  1. Preserve first-failure evidence and the original worker/shard count.
  2. Confirm Selenium/binding/browser/driver/Grid versions.
  3. Confirm target environment, shard manifest, synthetic data IDs, and evidence namespace.
  4. Inspect session IDs and prove whether each worker owns a distinct session.
  5. Inspect locator/wait/browser context only after ownership is proven.
  6. Inspect AUT conflicts, rate limits, server logs, and shared records.
  7. If remote, inspect Grid queue, slots, Node health, CPU/RAM and CI executor load.
  8. Apply the least destructive correction: unique state, lower concurrency, or measured capacity change.
  9. Rerun the smallest controlled scenario before reopening the full matrix.

7. Shortcuts that hide the layer

The following table organizes the key choices and evidence for Shortcuts that hide the layer. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Shortcut Why it is wrong Better correction
Blanket retry Multiplies load and can turn races green Preserve first failure; repair ownership/data/capacity
Long fixed sleep Changes timing, not the shared resource Wait on real state or remove the race
Global lock around every test Serializes throughput and hides design flaws Lock only a truly exclusive resource
Restart Grid/browser on failure Destroys useful state/evidence Inspect queue/session/Node first
Reuse one account to reduce fixture cost Creates mutable business-state collisions Generate deterministic per-test identities

8. Performance attribution

Separate runner overhead, browser/session startup, Grid queue time, AUT/network latency, artifact IO, and retry cost. “Parallel run took 30 seconds” is not enough. Compare per-test durations and session creation timing, then identify where concurrency adds waiting or saturation.

Knowledge check

Two threads use different test IDs but the same driver. What layer is broken first?

Two fresh browser sessions still receive one claim collision. What does that prove?

Why can retrying every failed parallel test make a Grid incident worse?

A screenshot from worker B replaces worker A’s failure screenshot. What should change?

Should order-dependent tests be fixed by enforcing alphabetical execution?

Next lesson

Checkpoint Lab — Parallel Execution, Isolation, Concurrency, and Test Sharding

Continue with Checkpoint Lab — Parallel Execution, Isolation, Concurrency, and Test Sharding. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and current-version notes

Version baseline — August 2026

The mandatory examples pin Selenium Python to 4.47.0. Selenium Server/Grid 4.47.0 is the matching stable Grid baseline. Python examples use the standard-library concurrent.futures module rather than a third-party parallel-test plugin so worker ownership is visible.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.