Chapter 25Lesson 01~215 minutes

Performance Engineering for Test Suites and Grid Capacity: Core Concepts and Mental Model

Browser suites can become slow for many different reasons. This lesson separates those reasons into measurable components so an optimization can improve feedback without weakening test isolation, evidence, or product coverage.

feedback timethroughputsession startupGrid queuesaturationbenchmark context

Learning objectives

  • Decompose total pipeline feedback time into setup, session start, queue wait, browser execution, evidence/reporting, and teardown.
  • Distinguish latency, throughput, capacity, utilization, saturation, and queueing in Selenium terms.
  • Inspect current client/browser/session/Grid/host state before changing performance controls.
  • Explain why runner worker count and Grid capacity are independent limits.
  • State why Selenium timing is useful for test-system engineering but is not a substitute for application load/performance testing.

1. The practical problem: a “slow Selenium suite” is not one problem

A 20-minute browser suite may spend time downloading dependencies, starting browsers, waiting in a Grid queue, waiting for the AUT, writing screenshots, retrying flakes, or tearing down leaked processes. Calling all of that “test execution” produces bad fixes. Increasing workers can shorten one segment while making queueing, memory pressure, or AUT contention worse.

The operating question is therefore not “How do I make Selenium faster?” It is: Which measured segment dominates feedback time, what resource constrains it, and what reliability cost appears when I change that resource?

Optimization must preserve Chapter 24 boundaries

Do not gain speed by reusing personal profiles, sharing privileged accounts, disabling security controls, dropping first-failure evidence, or pointing a benchmark at production. Performance engineering inherits all authorization, privacy, and cleanup controls from Chapter 24.

2. Mental model: feedback time and capacity are a pipeline

The following diagram visualizes the relationships described in Mental model: feedback time and capacity are a pipeline. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.

Feedback time and capacity pipeline
flowchart TD
  C[Change / CI job] --> S[Setup + build]
  S --> R[Runner / workers]
  R --> Q[Grid queue or local session request]
  Q --> B[Browser session startup]
  B --> T[Test actions + waits + AUT]
  T --> E[Evidence + reports]
  E --> X[Teardown]
  H[CPU / RAM / browser processes] -. limits .-> B
  G[Grid slots / Nodes] -. limits .-> Q
  A[AUT / DB / network capacity] -. limits .-> T
  I[Artifact IO / storage] -. limits .-> E

The solid arrows are the lifecycle. A change enters setup, the runner schedules work, a local driver or Grid obtains a browser, the scenario executes, evidence is persisted, and resources are torn down. The dotted arrows are capacity constraints: host resources constrain session startup, Grid slots constrain simultaneous sessions, AUT/network capacity constrains actions, and artifact IO constrains evidence.

A useful accounting identity is T_feedback = T_setup + T_session/queue + T_execution + T_evidence + T_teardown. When multiple tests run concurrently, the wall-clock total is not the sum of every test duration; it is shaped by overlap, bottlenecks, and queueing.

3. Terms that must not be conflated

The following table organizes the key choices and evidence for Terms that must not be conflated. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Term Meaning in this course Observable evidence
Latency Time for one operation or scenario. Session-start ms, one-case duration, request/DOM-ready duration.
Throughput Completed tests or sessions per unit time. Cases/minute or sessions/minute over a stable window.
Capacity How much concurrent work a layer can sustain while meeting reliability/latency goals. Grid max sessions, CPU/RAM headroom, AUT service limit.
Utilization Fraction of a resource currently busy. Active Grid sessions / max sessions; CPU; memory; browser count.
Queue wait Time a session request waits because matching capacity is unavailable. Grid queue size/traces plus RemoteWebDriver creation latency.
Saturation Point where more offered concurrency causes disproportionate latency, queueing, errors, or instability. p95 growth, queue growth, CPU/RAM exhaustion, timeouts, flake increase.
Speedup Serial wall time divided by parallel wall time for the same work and evidence policy. Controlled serial/parallel benchmark.

4. Read-only inspection before tuning

Record the binding, browser, driver/session, host, and requested options first. A timing result without version and environment context is not a benchmark; it is an anecdote.

from selenium import __version__ as selenium_version
from selenium import webdriver
import os, platform, time

print("selenium", selenium_version)
print("python", platform.python_version())
print("host", platform.system(), platform.release())
print("logical_cpu", os.cpu_count())

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
requested = options.to_capabilities()
print("requested", {k: requested.get(k) for k in ("browserName", "pageLoadStrategy")})

t0 = time.perf_counter()
driver = webdriver.Chrome(options=options)  # Selenium Manager resolves the driver
startup_ms = (time.perf_counter() - t0) * 1000
try:
    caps = dict(driver.capabilities)
    print("session", driver.session_id)
    print("browser", caps.get("browserName"), caps.get("browserVersion"))
    print("platform", caps.get("platformName"))
    print("startup_ms", round(startup_ms, 1))
    print("url", driver.current_url)
    print("title", driver.title)
finally:
    driver.quit()

For a Grid run, take a read-only snapshot before creating sessions. Grid GraphQL can report maxSession, sessionCount, node state, slot counts, and queue size. The query does not change Grid state.

from urllib.request import Request, urlopen
import json

GRID = "http://127.0.0.1:4444"
query = """{
  grid { maxSession sessionCount sessionQueueSize }
  nodesInfo { nodes { id uri status slotCount sessionCount } }
}"""
req = Request(
    GRID + "/graphql",
    data=json.dumps({"query": query}).encode(),
    headers={"Content-Type": "application/json"},
)
with urlopen(req, timeout=3) as response:
    snapshot = json.load(response)["data"]
print(json.dumps(snapshot, indent=2))
What session-start timing contains

For local WebDriver, creation includes driver resolution/launch plus browser startup and session negotiation. For RemoteWebDriver, creation can also include network latency and Grid queue wait. Do not label the whole measurement “browser startup” unless those layers have been separated.

5. Where the time actually goes

The following table organizes the key choices and evidence for Where the time actually goes. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Segment Common causes Safe first measurement
Setup/build dependency install, image pull, test discovery, fixture generation CI step durations and dependency/cache state
Session start browser launch, driver startup, profile initialization, extensions time around driver construction + returned capabilities
Queue all matching Grid slots busy or stereotype unavailable GraphQL queue size, Grid traces/logs, RemoteWebDriver creation time
Execution AUT/network latency, waits, DOM work, test code scenario timers + AUT evidence; avoid fixed sleeps
Evidence screenshots, video, page source, logs, upload/compression bytes written and time spent writing/uploading artifacts
Teardown driver quit, profile deletion, container stop teardown timer + post-run browser/process check
Retries same work executed again after failure attempt count × startup/execution/evidence cost

6. Runner concurrency and Grid capacity are different controls

A runner with eight workers can offer eight session requests. A Grid with four matching free slots can execute at most four of them immediately; the rest queue. Conversely, a Grid with twenty slots does not create parallelism if the runner schedules only one test at a time.

Current Grid guidance uses available processors as the default Node maximum session count and recommends approximately one browser session per processor; overriding that limit can reduce stability when resources are exhausted. Current Grid sizing guidance also expects browser sessions to consume meaningful memory, so CPU count alone is not a sufficient capacity model.

Effective concurrency

A practical ceiling is the minimum of runner workers, matching Grid slots, safe browser-host capacity, isolated test-data capacity, AUT capacity, and evidence/storage capacity. The smallest of those limits wins.

7. Selenium suite performance is not application load testing

Selenium is excellent for measuring the automation system: how long sessions take to start, how Grid queueing behaves, how many browser workers the host can sustain, and how artifact policy changes feedback time. Those are valid engineering questions because the object under measurement is the browser-test platform.

It is generally a poor load-test generator. Browser startup, WebDriver instrumentation, third-party resources, rendering, OS scheduling, and synchronization logic add variability that makes it hard to attribute results to the application server. If the question is “How many requests/users can our service sustain, what is p99 latency under 5,000 virtual users, or where does the API saturate?”, use JMeter, k6, Gatling, Locust, or another purpose-built performance tool.

No production load experiments

This chapter uses a disposable loopback AUT. Never turn Selenium workers into a production load generator. Capacity experiments are scoped to the test system and explicitly authorized disposable environments.

8. DevOps baseline contract

A useful benchmark record contains: commit/build ID; Selenium/binding/Grid/browser versions; runner OS/CPU/RAM class; local versus RemoteWebDriver; Grid topology and max sessions; test inventory; worker count; data/evidence policy; cold/warm cache state; wall time; session-start distribution; queue signal; failures/retries; and cleanup result.

Without that context, a “30% faster” claim cannot be reproduced or trusted. Chapter 25 treats benchmark metadata as evidence, just like Chapter 17 treated failure artifacts as evidence.

Knowledge checks

Answer from the operating model, then reveal the explanation.

Eight runner workers target a Grid with four matching slots. What happens?

Why is RemoteWebDriver construction time not automatically “browser startup time”?

CPU is only 40%, but p95 case time doubles when concurrency rises. What should you inspect?

When should JMeter or another load tool replace Selenium for the question being asked?

What makes a performance comparison reproducible?

Summary and next bridge

  • Feedback time is a sum of lifecycle segments, not a single “Selenium time.”
  • Concurrency is useful only until a constrained layer saturates.
  • Runner workers and Grid slots are separate controls.
  • Performance claims require version/hardware/topology/evidence context.
  • Selenium measures browser-test platform performance; purpose-built tools measure application load performance.

Lesson 2 builds a disposable local measurement workflow and records serial/parallel, session-start, queue, resource, and evidence costs.

Next lesson

Performance Engineering for Test Suites and Grid Capacity: Guided Hands-On Workflow

Continue with Performance Engineering for Test Suites and Grid Capacity: Guided Hands-On Workflow. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Primary references and version notes

Version baseline — August 2026

The mandatory examples pin selenium==4.47.0 and Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid baseline. Selenium Manager remains the normal local driver-resolution path. Hardware, browser versions, container images, Grid slot counts, runner sizes, AUT capacity, and network conditions are recorded as benchmark context rather than assumed constants.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.