Chapter 21Lesson 03~225 minutes

Parallel Execution, Isolation, Concurrency, and Test Sharding: Configuration, Design Patterns, and Trade-Offs

Parallel suites are design choices, not only worker counts. This lesson compares threads, processes and CI runners; per-test versus per-worker browser lifecycles; deterministic sharding; serial exceptions; and concurrency ceilings that respect Grid, AUT, and evidence capacity.

Threads vs processesLifecycleCapacity ceilingSerial markersTrade-offs

Learning objectives

  • Choose an execution model without assuming one runner mechanism is universally best.
  • Compare per-test and per-worker browser lifecycles with isolation and startup-cost consequences.
  • Design deterministic sharding separately from browser parallelism.
  • Use serial/exclusive lanes only for genuinely non-isolatable scenarios.
  • Justify a concurrency ceiling using observable capacity and saturation evidence.

1. Threads, processes, and CI jobs are scheduling mechanisms

Python threads are convenient for teaching I/O-heavy browser orchestration, but production frameworks may use processes, external workers, or separate CI jobs. Selenium does not require one runner topology. The invariant is simpler: each concurrent test gets exclusive ownership of the WebDriver/session and its mutable test state.

Execution unit Strength Risk Use when
Threads in one process Low setup overhead; easy shared coordination Accidental shared Python objects/driver references Small local experiments and runners designed for it
Worker processes Stronger memory isolation Higher process/bootstrap cost Frameworks that isolate tests in processes
Separate CI jobs/runners Strong filesystem/process boundary; natural sharding More provisioning and artifact aggregation Large suites, browser/OS matrices, horizontal scaling

2. Per-test versus per-worker browser lifecycle

Selenium’s encouraged guidance favors a new WebDriver instance per test because it minimizes shared browser state. Reusing one browser across several tests can reduce startup cost, but it turns cookies, storage, windows, downloads, profile state, and browser crashes into shared worker state.

Lifecycle Isolation Startup cost Diagnostic clarity Default
Fresh browser per test Highest Highest Simple: one session maps to one test Preferred teaching/CI baseline
Browser per worker, many tests Lower Lower Failures can contaminate following tests Only with explicit reset contract and measured need
One browser for whole suite Low Lowest Large blast radius Avoid for independent regression tests

3. Shard by stable test identity, then rebalance intentionally

A simple hash shard is deterministic and cheap, but it may not balance wall time when a few tests are much slower. A larger suite can use historical timing to construct balanced shard manifests. If you do that, version the manifest/algorithm and preserve the original stable test IDs so a failure can still be traced to one owner.

# Example of a versioned explicit shard manifest produced by CI tooling.
SHARDS = {
    0: ["checkout.card", "profile.locale", "search.desktop"],
    1: ["checkout.invoice", "profile.mobile", "search.keyboard"],
}


def tests_for_shard(shard_index: int):
    return list(SHARDS[shard_index])

4. Maximum workers versus safe workers

The Node --max-sessions option limits browser sessions at the Grid layer. The runner’s worker count limits scheduling at the runner layer. Neither knows your AUT’s business-data bottleneck or artifact-store bandwidth. Find the knee of the throughput curve by measuring at a small sequence such as 1 → 2 → 3 workers, stopping when wall time no longer improves or errors/resource pressure appear.

Signal as workers rise Interpretation Action
Wall time falls; CPU/RAM healthy; no queue growth Useful parallel headroom Consider next measured step
Grid queue grows while Nodes are full Runner demand exceeds Grid slots Cap runner or add measured Grid capacity
AUT 429/5xx or shared-data conflicts rise Application/test-data bottleneck Reduce concurrency or improve fixture isolation
Browser crashes / renderer exits / memory pressure Host saturation Reduce sessions; right-size Nodes
Artifact upload dominates completion Evidence pipeline bottleneck Reduce/target artifacts; parallelize storage safely

5. Serial markers are a controlled exception, not a repair tool

Some scenarios are genuinely exclusive: a lab that rotates one global test certificate, validates a singleton maintenance banner, or exercises a device that cannot be virtualized. Put those tests in an explicit serial lane with a named reason. Do not mark tests serial merely because they share a badly designed account or filename; fix the isolation contract instead.

Serial does not mean order dependent

A serial lane should still contain independently runnable tests whenever possible. “Test B requires Test A to have run first” is hidden workflow state, not a concurrency strategy.

6. Keep Selenium, runner, Grid, AUT, and CI configuration in their own layers

Selenium options configure browser sessions. The runner configures workers/fixtures/reports. Grid configures Node slots, queueing, and routing. The AUT controls rate limits and fixture semantics. CI configures job count, matrix, workspace, and artifact aggregation. A worker-count problem is not repaired with a browser capability, and a Grid-slot problem is not repaired with a Selenium explicit wait.

7. Worked decision: eight runner workers, four Grid slots, three safe AUT identities

The initial safe ceiling is three concurrent scenarios because the AUT/test-data boundary is the smallest known capacity. If you can provision six isolated identities but Grid remains four slots, the ceiling becomes four. Only after measuring CPU/RAM and queue behavior should you consider adding Grid capacity or overriding Node defaults.

Knowledge check

Why is per-worker browser reuse faster but riskier than a fresh browser per test?

When should a test be put in a serial lane?

What happens if Grid allows four sessions but the AUT safely supports only two synthetic accounts?

Is a hash shard always balanced by runtime?

Which layer owns --max-sessions?

Next lesson

Parallel Execution, Isolation, Concurrency, and Test Sharding: Diagnostics, Failure Modes, and Production Practices

Continue with Parallel Execution, Isolation, Concurrency, and Test Sharding: Diagnostics, Failure Modes, and Production Practices. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and current-version notes

Version baseline — August 2026

The mandatory examples pin Selenium Python to 4.47.0. Selenium Server/Grid 4.47.0 is the matching stable Grid baseline. Python examples use the standard-library concurrent.futures module rather than a third-party parallel-test plugin so worker ownership is visible.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.