Chapter 25Lesson 04~215 minutes

Performance Engineering for Test Suites and Grid Capacity: Diagnostics, Failure Modes, and Production Practices

Bad performance work often makes suites faster on paper while making failures harder to reproduce. This lesson engineers the common traps and repairs them by preserving first-failure evidence and isolating the smallest constrained layer.

diagnosticsresource leakwarm cache biasretry amplificationAUT throttlingsaturation

Learning objectives

  • Diagnose blindly increased worker counts, warm-cache bias, leaked browsers, retry amplification, and resource exhaustion.
  • Distinguish Grid/browser saturation from AUT/network throttling using queue and session evidence.
  • Interpret an intentionally broken benchmark without hiding the original failure.
  • Apply the standard diagnostic sequence before changing performance controls.
  • Keep production load generation outside Selenium and within explicit authorization boundaries.

1. The diagnostic sequence does not change because the symptom is “slow”

  1. Preserve first-failure evidence and timing metadata.
  2. Confirm Selenium/binding/browser/driver/Grid versions and benchmark host class.
  3. Confirm target/environment/test data and authorization.
  4. Inspect session/capabilities/context and browser lifecycle.
  5. Inspect locator/element/synchronization state so waits are not misread as raw performance.
  6. Inspect AUT/network/browser evidence.
  7. If remote, inspect Grid queue/slot/Node/CI resource state.
  8. Apply the least destructive correction.
  9. Rerun the smallest controlled scenario with the same evidence policy.

2. Failure mode: blindly increasing workers

Symptoms: lower throughput gain per added worker, rising session-start p95, queue growth, CPU/memory pressure, browser crashes, or more timing flakes. The repair is not “increase timeouts.” Plot wall time, p95 case duration, failures, queue size, and resource state across worker counts. Stop at the first knee where additional concurrency gives little speedup or harms reliability.

3. Failure mode: warm-cache benchmark presented as general speedup

Dependency caches, browser profile caches, image layers, DNS, service caches, and AUT caches all change timing. A warm second run can be useful, but label it as warm. If CI regularly starts from fresh runners, a warm-only result is not representative.

Benchmark mode Record Claim you may make
cold fresh runner/image/dependency/profile state cold-start feedback behavior
warm explicit cache reuse and same commit steady/warm repeat behavior
mixed/unknown cache state not controlled no strong comparative speed claim

4. Failure mode: leaked browsers make later runs slower

A missing quit() is both a correctness and capacity bug. Each orphan consumes processes, memory, file descriptors/handles, ports, and possibly Grid slots.

Intentionally broken example:

The following example makes the Failure mode: leaked browsers make later runs slower behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from selenium import webdriver

def bad_benchmark(urls):
    durations = []
    for url in urls:
        driver = webdriver.Chrome()
        start = __import__("time").perf_counter()
        driver.get(url)
        durations.append(__import__("time").perf_counter() - start)
        # BUG: no driver.quit(); later iterations inherit host pressure from leaked browsers.
    return durations

The first measurement may look fine while later measurements degrade. Repair lifecycle ownership first and repeat from a clean host:

from selenium import webdriver
from time import perf_counter

def measured_navigation(url):
    driver = None
    total_start = perf_counter()
    try:
        driver = webdriver.Chrome()
        nav_start = perf_counter()
        driver.get(url)
        return {
            "navigation_ms": (perf_counter() - nav_start) * 1000,
            "total_ms": (perf_counter() - total_start) * 1000,
        }
    finally:
        if driver is not None:
            driver.quit()

5. Failure mode: retries multiply load

A retry repeats session startup, AUT work, evidence capture, and teardown. With broad retries, a 5% flake rate can consume far more than 5% extra capacity because failures are often correlated with saturation—the exact time when extra work is most harmful.

Preserve attempt 1 evidence and count retry cost separately. Chapter 15’s rule still applies: retry is an additional observation, not a cure.

6. Failure mode: AUT throttling mistaken for Grid bottleneck

If Grid queue size is zero, sessions are allocated promptly, browser startup is stable, but DOM-ready/API latency rises only at higher concurrent test counts, investigate the AUT/test-data/backend/network layer. Adding Grid Nodes may make the problem worse by sending even more work to the already constrained AUT.

Conversely, if AUT timings are flat but session creation waits while Grid slots are full, the browser/Grid layer is the stronger suspect.

7. Failure mode: CPU/memory exhaustion creates flakes

Resource exhaustion often appears first as “random Selenium instability”: browser renderer crashes, DevTools/BiDi connection loss, session-creation errors, long garbage-collection pauses in Grid/runner processes, or timeouts. Treat those as resource evidence, not as proof the test needs a longer wait.

Current Grid guidance warns that overriding the CPU-based max-session recommendation can reduce session stability. If you deliberately override it in a disposable benchmark, label it as failure injection and restore the safe setting afterward.

8. Failure mode: timing without context

“Suite now takes 8 minutes instead of 11” is incomplete unless both runs identify versions, machine class, worker count, Grid capacity, browser mode, cache state, evidence policy, test inventory, failures/retries, and target environment. A browser auto-update alone can change startup and rendering behavior enough to invalidate the comparison.

9. Failure mode: using Selenium to generate production load

Do not turn the runner into a load generator, bypass rate limits, or increase browser workers against production. Selenium’s own guidance discourages using WebDriver for performance testing because browser/driver/instrumentation and external resources introduce variability and the tool is not optimized for load analysis.

If the objective is server capacity, move to a purpose-built tool and a separately authorized performance-test environment. Keep Selenium for a small number of functional browser journeys before/after that load test if needed.

Security boundary

Never disable TLS, bypass MFA/CAPTCHA/anti-abuse controls, expose Grid publicly, or use real customer data to “simplify” a performance experiment.

10. Evidence-to-correction map

The correction map below is for the browser-test platform. If the target question is server throughput or sustained virtual-user load, use JMeter or another purpose-built performance tool instead of extending Selenium concurrency.

Observed evidence Likely layer Least destructive next step
queue grows; max sessions reached; CPU/RAM healthy Grid matching capacity reduce offered workers or add matching isolated Node capacity
queue empty; browser CPU/RAM near exhaustion browser host lower sessions/Node, use smaller Nodes, inspect browser/profile cost
session start stable; AUT/API p95 rises AUT/backend/network lower test concurrency or scale/isolate test environment
artifact stage dominates evidence IO/storage tier/compress/upload evidence while preserving first failure
browser process count grows after suite teardown leak fix lifecycle ownership before any tuning
only warm run improves cache state publish cold and warm separately; align benchmark with CI reality

Knowledge checks

Answer from the operating model, then reveal the explanation.

Why is increasing timeout a bad first response to CPU/memory saturation?

How can retries amplify a capacity incident?

Grid queue is zero but AUT p95 doubles. What should you avoid doing?

What makes the intentionally broken benchmark invalid?

Why must cold and warm cache results be labeled separately?

Summary and next bridge

  • Preserve first evidence before changing a slow or flaky system.
  • Worker count, retries, and timeouts can worsen saturation rather than fix it.
  • Queue/session/resource/AUT evidence identifies the constrained layer.
  • Lifecycle leaks invalidate benchmarks and consume capacity.
  • Selenium remains a browser-test engineering tool, not a production load generator.

Lesson 5 turns the chapter into a small capacity report with controlled runs, a safe concurrency recommendation, an optimization delta, and a clear handoff boundary to load-testing tools.

Next lesson

Checkpoint Lab — Performance Engineering for Test Suites and Grid Capacity

Continue with Checkpoint Lab — Performance Engineering for Test Suites and Grid Capacity. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Primary references and version notes

Version baseline — August 2026

The mandatory examples pin selenium==4.47.0 and Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid baseline. Selenium Manager remains the normal local driver-resolution path. Hardware, browser versions, container images, Grid slot counts, runner sizes, AUT capacity, and network conditions are recorded as benchmark context rather than assumed constants.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.