Chapter 25Lesson 03~215 minutes

Performance Engineering for Test Suites and Grid Capacity: Configuration, Design Patterns, and Trade-Offs

Performance controls are architecture choices with reliability costs. This lesson compares the major choices using measurable consequences rather than folklore such as “headless is always faster” or “reuse one browser for everything.”

browser lifecycleheadlessGrid latencyconcurrencyevidence IOfixture reusebrowser matrix

Learning objectives

  • Compare fresh-session isolation with browser reuse and identify the state-risk introduced by reuse.
  • Explain headed/headless and local/Grid trade-offs without assuming one mode is universally faster.
  • Choose concurrency from measured saturation evidence rather than CPU count alone.
  • Balance evidence richness, fixture reuse, browser-matrix breadth, and feedback time.
  • Distinguish Selenium configuration from runner, AUT, browser policy, proxy/TLS, CI, and orchestration configuration.

1. Browser reuse versus isolation

Creating a browser session is expensive, so reuse looks attractive. But a reused session also reuses cookies, storage, history, service workers, open windows, extensions/profile state, and sometimes AUT-side session state. That changes the test contract.

Choice Potential benefit Cost/risk Use when
Fresh session per test strong isolation; clear ownership/cleanup startup cost default for independent production browser tests
Session per worker amortizes startup requires rigorous state reset; failures may poison later cases only after proving reset invariants and documenting ownership
One global browser minimal startup shared mutable state, ordering, thread safety, poor diagnostics not a safe concurrency strategy
Do not hide a semantic change in a benchmark

If Run B reuses a browser while Run A creates fresh sessions, the two runs do not measure only speed. They also change isolation and failure blast radius.

2. Headless versus headed

Headless mode can reduce display/rendering overhead and is operationally convenient in CI, but performance differences vary by browser, version, GPU/display stack, page behavior, and container/VM environment. A headless result is not proof of headed performance or vice versa.

Record the mode as benchmark metadata. If the product risk includes headed-only behavior, keep at least a representative headed lane instead of replacing coverage with a faster but different execution mode.

3. Local WebDriver versus Grid

Local WebDriver removes network and Grid routing/queueing from session creation. Grid adds useful distribution and browser/OS capacity but also adds transport, queue, Node placement, and infrastructure failure domains.

Signal Local interpretation Grid interpretation
slow driver creation driver/browser/profile/host startup client network + queue + Distributor/Node + browser startup
high execution latency browser/AUT/network/test waits same, plus client↔Grid↔Node command path
many waiting tests runner scheduling or local resource contention can also be New Session Queue pressure
orphan process local cleanup failure Node/session lifecycle or container cleanup issue

4. Concurrency versus saturation

The fastest stable point is usually before the maximum theoretical worker count. Once a constrained layer saturates, queueing theory becomes visible operationally: offered work grows faster than service capacity, queues grow, p95/p99 latency rises, and timeouts/flakes appear.

Current Grid defaults Node max sessions to available processors and warns that overriding the recommended one-browser-per-processor model can reduce stability. Treat that as a starting safety guard, not proof that every CPU can sustain one heavyweight test for your AUT/browser/profile/evidence mix.

5. Evidence richness versus IO/storage

Always-on screenshots, video, HAR-like data, page source, console logs, and trace bundles can dominate IO at scale. The wrong optimization is “turn all evidence off.” A better design is tiered capture:

  • small manifest/version/session metadata always;
  • screenshot + targeted DOM/log evidence on failure;
  • richer network/video/tracing on selected suites or diagnostic reruns;
  • bounded retention and per-attempt names so retries never overwrite first failure.

Measure bytes and write/upload time. Chapter 17’s diagnostic value still applies; Chapter 25 merely quantifies the cost.

6. Per-test setup versus fixture reuse

Fast API/database fixture setup can be safely reused only when the reused resource is immutable or correctly partitioned. Browser state is especially risky to share. A practical rule is to amortize expensive infrastructure setup (for example starting a disposable AUT once per worker) while keeping mutable test identity, browser session, and evidence namespace isolated per test unless a different contract is explicitly proven.

7. Browser matrix breadth versus feedback time

Chapter 16 established risk-based compatibility coverage. Performance engineering gives that policy a schedule: fast per-change lanes can cover the highest-risk browser/viewport combinations; broader cross-browser matrices can run on merge or schedule. Do not delete coverage solely because it is slow—move it to a cadence justified by risk and capacity.

8. Dedicated Grid versus shared capacity

A shared Grid improves utilization but introduces workload interference: another team’s browser mix, video settings, or long sessions can change your queue wait. A dedicated Grid costs more infrastructure but makes capacity and benchmark results more reproducible.

If benchmarking on shared capacity, record queue/session/Node state during the run. Otherwise you cannot distinguish your code change from neighbor load.

9. Configuration ownership: put the knob in the correct layer

The following table organizes the key choices and evidence for Configuration ownership: put the knob in the correct layer. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Layer Examples Performance implication
Selenium/WebDriver browser options, page-load strategy, RemoteWebDriver endpoint changes browser/session command behavior
test framework/runner workers, sharding, fixture scope, retries changes offered concurrency and lifecycle
AUT seed/reset method, feature flags, backend limits changes application work/capacity
browser policy/profile extensions, enterprise policy, cache/profile state changes startup/runtime behavior
Grid max sessions, Node stereotypes, queue timeouts, topology changes session scheduling/capacity
CI/container/cloud CPU/RAM class, network, image pull/cache, storage changes host and artifact capacity
proxy/TLS/identity proxy route, trust chain, SSO flow changes network/auth latency and failure domains

10. Decision table: choose by observable behavior

The following table organizes the key choices and evidence for Decision table: choose by observable behavior. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Situation Preferred move Why
Startup dominates, tests remain independent consider more parallel small Nodes before browser reuse preserves isolation while adding session-start capacity
Grid queue grows; Nodes at safe max sessions add matching Node capacity or reduce offered workers queue proves scheduling bottleneck
Queue empty; AUT p95 doubles at concurrency 4 lower test concurrency or isolate/scale test AUT Grid is not the bottleneck
Artifacts consume 35% of job time tier evidence + compress/upload once; preserve failure packet reduces IO without deleting diagnostic value
Shared Grid benchmark varies wildly capture Grid state or use dedicated benchmark window/capacity controls neighbor interference
Need server throughput under thousands of users use JMeter/k6/etc. different objective; Selenium is not optimized for load generation

Knowledge checks

Answer from the operating model, then reveal the explanation.

Why can browser reuse make a benchmark incomparable to a fresh-session baseline?

Is headless always faster?

What does a growing Grid queue with stable AUT latency suggest?

What should remain per-test even when infrastructure setup is reused?

Why can a shared Grid invalidate a performance claim?

Summary and next bridge

  • Performance controls also change isolation, portability, and diagnostics.
  • Fresh sessions remain the reliability baseline; reuse is an explicit trade-off.
  • Headless/local/Grid choices must be measured, not assumed.
  • Concurrency stops being useful at the first constrained layer.
  • Evidence and browser-matrix policies can be tiered without deleting essential coverage.

Lesson 4 diagnoses misleading benchmarks, resource leaks, retry amplification, AUT throttling, and saturation failures using a fixed evidence-first sequence.

Next lesson

Performance Engineering for Test Suites and Grid Capacity: Diagnostics, Failure Modes, and Production Practices

Continue with Performance Engineering for Test Suites and Grid Capacity: Diagnostics, Failure Modes, and Production Practices. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Primary references and version notes

Version baseline — August 2026

The mandatory examples pin selenium==4.47.0 and Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid baseline. Selenium Manager remains the normal local driver-resolution path. Hardware, browser versions, container images, Grid slot counts, runner sizes, AUT capacity, and network conditions are recorded as benchmark context rather than assumed constants.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.