Performance Engineering for Test Suites and Grid Capacity: Diagnostics, Failure Modes, and Production Practices
Bad performance work often makes suites faster on paper while making failures harder to reproduce. This lesson engineers the common traps and repairs them by preserving first-failure evidence and isolating the smallest constrained layer.
Learning objectives
- Diagnose blindly increased worker counts, warm-cache bias, leaked browsers, retry amplification, and resource exhaustion.
- Distinguish Grid/browser saturation from AUT/network throttling using queue and session evidence.
- Interpret an intentionally broken benchmark without hiding the original failure.
- Apply the standard diagnostic sequence before changing performance controls.
- Keep production load generation outside Selenium and within explicit authorization boundaries.
1. The diagnostic sequence does not change because the symptom is “slow”
- Preserve first-failure evidence and timing metadata.
- Confirm Selenium/binding/browser/driver/Grid versions and benchmark host class.
- Confirm target/environment/test data and authorization.
- Inspect session/capabilities/context and browser lifecycle.
- Inspect locator/element/synchronization state so waits are not misread as raw performance.
- Inspect AUT/network/browser evidence.
- If remote, inspect Grid queue/slot/Node/CI resource state.
- Apply the least destructive correction.
- Rerun the smallest controlled scenario with the same evidence policy.
2. Failure mode: blindly increasing workers
Symptoms: lower throughput gain per added worker, rising session-start p95, queue growth, CPU/memory pressure, browser crashes, or more timing flakes. The repair is not “increase timeouts.” Plot wall time, p95 case duration, failures, queue size, and resource state across worker counts. Stop at the first knee where additional concurrency gives little speedup or harms reliability.
3. Failure mode: warm-cache benchmark presented as general speedup
Dependency caches, browser profile caches, image layers, DNS, service caches, and AUT caches all change timing. A warm second run can be useful, but label it as warm. If CI regularly starts from fresh runners, a warm-only result is not representative.
| Benchmark mode | Record | Claim you may make |
|---|---|---|
| cold | fresh runner/image/dependency/profile state | cold-start feedback behavior |
| warm | explicit cache reuse and same commit | steady/warm repeat behavior |
| mixed/unknown | cache state not controlled | no strong comparative speed claim |
4. Failure mode: leaked browsers make later runs slower
A missing quit() is both a correctness and capacity
bug. Each orphan consumes processes, memory, file
descriptors/handles, ports, and possibly Grid slots.
Intentionally broken example:
The following example makes the Failure mode: leaked browsers make later runs slower behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium import webdriver
def bad_benchmark(urls):
durations = []
for url in urls:
driver = webdriver.Chrome()
start = __import__("time").perf_counter()
driver.get(url)
durations.append(__import__("time").perf_counter() - start)
# BUG: no driver.quit(); later iterations inherit host pressure from leaked browsers.
return durations
The first measurement may look fine while later measurements degrade. Repair lifecycle ownership first and repeat from a clean host:
from selenium import webdriver
from time import perf_counter
def measured_navigation(url):
driver = None
total_start = perf_counter()
try:
driver = webdriver.Chrome()
nav_start = perf_counter()
driver.get(url)
return {
"navigation_ms": (perf_counter() - nav_start) * 1000,
"total_ms": (perf_counter() - total_start) * 1000,
}
finally:
if driver is not None:
driver.quit()
5. Failure mode: retries multiply load
A retry repeats session startup, AUT work, evidence capture, and teardown. With broad retries, a 5% flake rate can consume far more than 5% extra capacity because failures are often correlated with saturation—the exact time when extra work is most harmful.
Preserve attempt 1 evidence and count retry cost separately. Chapter 15’s rule still applies: retry is an additional observation, not a cure.
6. Failure mode: AUT throttling mistaken for Grid bottleneck
If Grid queue size is zero, sessions are allocated promptly, browser startup is stable, but DOM-ready/API latency rises only at higher concurrent test counts, investigate the AUT/test-data/backend/network layer. Adding Grid Nodes may make the problem worse by sending even more work to the already constrained AUT.
Conversely, if AUT timings are flat but session creation waits while Grid slots are full, the browser/Grid layer is the stronger suspect.
7. Failure mode: CPU/memory exhaustion creates flakes
Resource exhaustion often appears first as “random Selenium instability”: browser renderer crashes, DevTools/BiDi connection loss, session-creation errors, long garbage-collection pauses in Grid/runner processes, or timeouts. Treat those as resource evidence, not as proof the test needs a longer wait.
Current Grid guidance warns that overriding the CPU-based max-session recommendation can reduce session stability. If you deliberately override it in a disposable benchmark, label it as failure injection and restore the safe setting afterward.
8. Failure mode: timing without context
“Suite now takes 8 minutes instead of 11” is incomplete unless both runs identify versions, machine class, worker count, Grid capacity, browser mode, cache state, evidence policy, test inventory, failures/retries, and target environment. A browser auto-update alone can change startup and rendering behavior enough to invalidate the comparison.
9. Failure mode: using Selenium to generate production load
Do not turn the runner into a load generator, bypass rate limits, or increase browser workers against production. Selenium’s own guidance discourages using WebDriver for performance testing because browser/driver/instrumentation and external resources introduce variability and the tool is not optimized for load analysis.
If the objective is server capacity, move to a purpose-built tool and a separately authorized performance-test environment. Keep Selenium for a small number of functional browser journeys before/after that load test if needed.
Never disable TLS, bypass MFA/CAPTCHA/anti-abuse controls, expose Grid publicly, or use real customer data to “simplify” a performance experiment.
10. Evidence-to-correction map
The correction map below is for the browser-test platform. If the target question is server throughput or sustained virtual-user load, use JMeter or another purpose-built performance tool instead of extending Selenium concurrency.
| Observed evidence | Likely layer | Least destructive next step |
|---|---|---|
| queue grows; max sessions reached; CPU/RAM healthy | Grid matching capacity | reduce offered workers or add matching isolated Node capacity |
| queue empty; browser CPU/RAM near exhaustion | browser host | lower sessions/Node, use smaller Nodes, inspect browser/profile cost |
| session start stable; AUT/API p95 rises | AUT/backend/network | lower test concurrency or scale/isolate test environment |
| artifact stage dominates | evidence IO/storage | tier/compress/upload evidence while preserving first failure |
| browser process count grows after suite | teardown leak | fix lifecycle ownership before any tuning |
| only warm run improves | cache state | publish cold and warm separately; align benchmark with CI reality |
Knowledge checks
Answer from the operating model, then reveal the explanation.
Why is increasing timeout a bad first response to CPU/memory saturation?
It masks the symptom and can increase resource occupancy. Preserve evidence, reduce offered concurrency, and fix the constrained layer.
How can retries amplify a capacity incident?
Each retry adds another session/startup/AUT/evidence cycle precisely when saturation may already be causing failures.
Grid queue is zero but AUT p95 doubles. What should you avoid doing?
Avoid adding more Grid/runner concurrency until the AUT/network/test-data bottleneck is understood.
What makes the intentionally broken benchmark invalid?
It leaks browser sessions, so later samples run under progressively different host resource conditions.
Why must cold and warm cache results be labeled separately?
They represent different system states; presenting warm-only improvement as general CI speedup is misleading.
Summary and next bridge
- Preserve first evidence before changing a slow or flaky system.
- Worker count, retries, and timeouts can worsen saturation rather than fix it.
- Queue/session/resource/AUT evidence identifies the constrained layer.
- Lifecycle leaks invalidate benchmarks and consume capacity.
- Selenium remains a browser-test engineering tool, not a production load generator.
Lesson 5 turns the chapter into a small capacity report with controlled runs, a safe concurrency recommendation, an optimization delta, and a clear handoff boundary to load-testing tools.
Primary references and version notes
- Selenium downloads — stable client and Grid release baseline.
- Getting started with Selenium Grid — CPU/RAM sizing and session-capacity guidance.
-
Grid CLI options
— current
max-sessions, queue timeout, retry interval, and related controls. - Grid GraphQL support — observable max sessions, session count, nodes, slots, and queue size.
- Grid observability — traces, events, and logs for distributed diagnosis.
- Performance testing with Selenium is discouraged — why WebDriver suite timing is not a substitute for load/performance tooling.
The mandatory examples pin selenium==4.47.0 and
Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid
baseline. Selenium Manager remains the normal local
driver-resolution path. Hardware, browser versions, container
images, Grid slot counts, runner sizes, AUT capacity, and network
conditions are recorded as benchmark context rather than assumed
constants.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.