Performance Engineering for Test Suites and Grid Capacity: Core Concepts and Mental Model
Browser suites can become slow for many different reasons. This lesson separates those reasons into measurable components so an optimization can improve feedback without weakening test isolation, evidence, or product coverage.
Learning objectives
- Decompose total pipeline feedback time into setup, session start, queue wait, browser execution, evidence/reporting, and teardown.
- Distinguish latency, throughput, capacity, utilization, saturation, and queueing in Selenium terms.
- Inspect current client/browser/session/Grid/host state before changing performance controls.
- Explain why runner worker count and Grid capacity are independent limits.
- State why Selenium timing is useful for test-system engineering but is not a substitute for application load/performance testing.
1. The practical problem: a “slow Selenium suite” is not one problem
A 20-minute browser suite may spend time downloading dependencies, starting browsers, waiting in a Grid queue, waiting for the AUT, writing screenshots, retrying flakes, or tearing down leaked processes. Calling all of that “test execution” produces bad fixes. Increasing workers can shorten one segment while making queueing, memory pressure, or AUT contention worse.
The operating question is therefore not “How do I make Selenium faster?” It is: Which measured segment dominates feedback time, what resource constrains it, and what reliability cost appears when I change that resource?
Do not gain speed by reusing personal profiles, sharing privileged accounts, disabling security controls, dropping first-failure evidence, or pointing a benchmark at production. Performance engineering inherits all authorization, privacy, and cleanup controls from Chapter 24.
2. Mental model: feedback time and capacity are a pipeline
The following diagram visualizes the relationships described in Mental model: feedback time and capacity are a pipeline. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.
flowchart TD C[Change / CI job] --> S[Setup + build] S --> R[Runner / workers] R --> Q[Grid queue or local session request] Q --> B[Browser session startup] B --> T[Test actions + waits + AUT] T --> E[Evidence + reports] E --> X[Teardown] H[CPU / RAM / browser processes] -. limits .-> B G[Grid slots / Nodes] -. limits .-> Q A[AUT / DB / network capacity] -. limits .-> T I[Artifact IO / storage] -. limits .-> E
The solid arrows are the lifecycle. A change enters setup, the runner schedules work, a local driver or Grid obtains a browser, the scenario executes, evidence is persisted, and resources are torn down. The dotted arrows are capacity constraints: host resources constrain session startup, Grid slots constrain simultaneous sessions, AUT/network capacity constrains actions, and artifact IO constrains evidence.
A useful accounting identity is
T_feedback = T_setup + T_session/queue + T_execution + T_evidence
+ T_teardown. When multiple tests run concurrently, the wall-clock total is not
the sum of every test duration; it is shaped by overlap,
bottlenecks, and queueing.
3. Terms that must not be conflated
The following table organizes the key choices and evidence for Terms that must not be conflated. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Term | Meaning in this course | Observable evidence |
|---|---|---|
| Latency | Time for one operation or scenario. | Session-start ms, one-case duration, request/DOM-ready duration. |
| Throughput | Completed tests or sessions per unit time. | Cases/minute or sessions/minute over a stable window. |
| Capacity | How much concurrent work a layer can sustain while meeting reliability/latency goals. | Grid max sessions, CPU/RAM headroom, AUT service limit. |
| Utilization | Fraction of a resource currently busy. | Active Grid sessions / max sessions; CPU; memory; browser count. |
| Queue wait | Time a session request waits because matching capacity is unavailable. | Grid queue size/traces plus RemoteWebDriver creation latency. |
| Saturation | Point where more offered concurrency causes disproportionate latency, queueing, errors, or instability. | p95 growth, queue growth, CPU/RAM exhaustion, timeouts, flake increase. |
| Speedup | Serial wall time divided by parallel wall time for the same work and evidence policy. | Controlled serial/parallel benchmark. |
4. Read-only inspection before tuning
Record the binding, browser, driver/session, host, and requested options first. A timing result without version and environment context is not a benchmark; it is an anecdote.
from selenium import __version__ as selenium_version
from selenium import webdriver
import os, platform, time
print("selenium", selenium_version)
print("python", platform.python_version())
print("host", platform.system(), platform.release())
print("logical_cpu", os.cpu_count())
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
requested = options.to_capabilities()
print("requested", {k: requested.get(k) for k in ("browserName", "pageLoadStrategy")})
t0 = time.perf_counter()
driver = webdriver.Chrome(options=options) # Selenium Manager resolves the driver
startup_ms = (time.perf_counter() - t0) * 1000
try:
caps = dict(driver.capabilities)
print("session", driver.session_id)
print("browser", caps.get("browserName"), caps.get("browserVersion"))
print("platform", caps.get("platformName"))
print("startup_ms", round(startup_ms, 1))
print("url", driver.current_url)
print("title", driver.title)
finally:
driver.quit()
For a Grid run, take a read-only snapshot before creating sessions.
Grid GraphQL can report maxSession,
sessionCount, node state, slot counts, and queue size.
The query does not change Grid state.
from urllib.request import Request, urlopen
import json
GRID = "http://127.0.0.1:4444"
query = """{
grid { maxSession sessionCount sessionQueueSize }
nodesInfo { nodes { id uri status slotCount sessionCount } }
}"""
req = Request(
GRID + "/graphql",
data=json.dumps({"query": query}).encode(),
headers={"Content-Type": "application/json"},
)
with urlopen(req, timeout=3) as response:
snapshot = json.load(response)["data"]
print(json.dumps(snapshot, indent=2))
For local WebDriver, creation includes driver resolution/launch plus browser startup and session negotiation. For RemoteWebDriver, creation can also include network latency and Grid queue wait. Do not label the whole measurement “browser startup” unless those layers have been separated.
5. Where the time actually goes
The following table organizes the key choices and evidence for Where the time actually goes. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Segment | Common causes | Safe first measurement |
|---|---|---|
| Setup/build | dependency install, image pull, test discovery, fixture generation | CI step durations and dependency/cache state |
| Session start | browser launch, driver startup, profile initialization, extensions | time around driver construction + returned capabilities |
| Queue | all matching Grid slots busy or stereotype unavailable | GraphQL queue size, Grid traces/logs, RemoteWebDriver creation time |
| Execution | AUT/network latency, waits, DOM work, test code | scenario timers + AUT evidence; avoid fixed sleeps |
| Evidence | screenshots, video, page source, logs, upload/compression | bytes written and time spent writing/uploading artifacts |
| Teardown | driver quit, profile deletion, container stop | teardown timer + post-run browser/process check |
| Retries | same work executed again after failure | attempt count × startup/execution/evidence cost |
6. Runner concurrency and Grid capacity are different controls
A runner with eight workers can offer eight session requests. A Grid with four matching free slots can execute at most four of them immediately; the rest queue. Conversely, a Grid with twenty slots does not create parallelism if the runner schedules only one test at a time.
Current Grid guidance uses available processors as the default Node maximum session count and recommends approximately one browser session per processor; overriding that limit can reduce stability when resources are exhausted. Current Grid sizing guidance also expects browser sessions to consume meaningful memory, so CPU count alone is not a sufficient capacity model.
A practical ceiling is the minimum of runner workers, matching Grid slots, safe browser-host capacity, isolated test-data capacity, AUT capacity, and evidence/storage capacity. The smallest of those limits wins.
7. Selenium suite performance is not application load testing
Selenium is excellent for measuring the automation system: how long sessions take to start, how Grid queueing behaves, how many browser workers the host can sustain, and how artifact policy changes feedback time. Those are valid engineering questions because the object under measurement is the browser-test platform.
It is generally a poor load-test generator. Browser startup, WebDriver instrumentation, third-party resources, rendering, OS scheduling, and synchronization logic add variability that makes it hard to attribute results to the application server. If the question is “How many requests/users can our service sustain, what is p99 latency under 5,000 virtual users, or where does the API saturate?”, use JMeter, k6, Gatling, Locust, or another purpose-built performance tool.
This chapter uses a disposable loopback AUT. Never turn Selenium workers into a production load generator. Capacity experiments are scoped to the test system and explicitly authorized disposable environments.
8. DevOps baseline contract
A useful benchmark record contains: commit/build ID; Selenium/binding/Grid/browser versions; runner OS/CPU/RAM class; local versus RemoteWebDriver; Grid topology and max sessions; test inventory; worker count; data/evidence policy; cold/warm cache state; wall time; session-start distribution; queue signal; failures/retries; and cleanup result.
Without that context, a “30% faster” claim cannot be reproduced or trusted. Chapter 25 treats benchmark metadata as evidence, just like Chapter 17 treated failure artifacts as evidence.
Knowledge checks
Answer from the operating model, then reveal the explanation.
Eight runner workers target a Grid with four matching slots. What happens?
At most four sessions run immediately; additional session requests wait in the New Session Queue until matching slots become free or the configured request timeout expires.
Why is RemoteWebDriver construction time not automatically “browser startup time”?
It can include client/network latency, Grid queue wait, slot selection, Node communication, browser launch, and session negotiation.
CPU is only 40%, but p95 case time doubles when concurrency rises. What should you inspect?
Do not assume spare CPU means spare capacity. Check memory/browser processes, Grid queue/slot state, AUT/network latency, test-data contention, and artifact IO.
When should JMeter or another load tool replace Selenium for the question being asked?
When the objective is application/system performance under controlled request/user load rather than browser-test platform feedback time and capacity.
What makes a performance comparison reproducible?
Same work plus recorded versions/hardware/topology/data/evidence/cache conditions, repeated controlled runs, and one intentional variable change at a time.
Summary and next bridge
- Feedback time is a sum of lifecycle segments, not a single “Selenium time.”
- Concurrency is useful only until a constrained layer saturates.
- Runner workers and Grid slots are separate controls.
- Performance claims require version/hardware/topology/evidence context.
- Selenium measures browser-test platform performance; purpose-built tools measure application load performance.
Lesson 2 builds a disposable local measurement workflow and records serial/parallel, session-start, queue, resource, and evidence costs.
Primary references and version notes
- Selenium downloads — stable client and Grid release baseline.
- Getting started with Selenium Grid — CPU/RAM sizing and session-capacity guidance.
-
Grid CLI options
— current
max-sessions, queue timeout, retry interval, and related controls. - Grid GraphQL support — observable max sessions, session count, nodes, slots, and queue size.
- Grid observability — traces, events, and logs for distributed diagnosis.
- Performance testing with Selenium is discouraged — why WebDriver suite timing is not a substitute for load/performance tooling.
The mandatory examples pin selenium==4.47.0 and
Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid
baseline. Selenium Manager remains the normal local
driver-resolution path. Hardware, browser versions, container
images, Grid slot counts, runner sizes, AUT capacity, and network
conditions are recorded as benchmark context rather than assumed
constants.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.