Performance Engineering for Test Suites and Grid Capacity: Guided Hands-On Workflow
Now turn the model into measurements. The workflow uses a loopback AUT, fresh browser sessions, explicit timers, optional Grid snapshots, and a synthetic saturation model so every optimization can be traced to one changed variable.
Learning objectives
- Create a reproducible loopback fixture whose latency rises after a known soft capacity.
- Measure serial and parallel wall time without sharing WebDriver sessions or test identities.
- Record browser session-start, scenario, evidence, and total timings separately.
- Observe Grid max sessions/session count/queue size where a local Grid is used.
- Change one concurrency or evidence variable at a time and compare against a documented baseline.
1. Lab topology and preflight
The mandatory path is free and local. The AUT binds to
127.0.0.1:8815; it does not require real accounts,
external traffic, or a paid browser cloud. Its
/api/work endpoint deliberately adds latency only after
more than two concurrent calls so saturation is visible without
stressing any real service.
- Python 3.10+.
-
selenium==4.47.0for the browser measurement path. - A supported local Chromium-family browser; Selenium Manager resolves the driver.
- Optional: Selenium Server/Grid 4.47.0 on loopback for queue/slot observation.
- Enough local memory for the worker count you choose. Start small.
The fixture is intentionally synthetic. Do not substitute a production base URL or a public Grid.
2. Start a measurable loopback AUT
Save this as aut_server.py. The service reports
current/peak concurrency and applies a deterministic latency penalty
above two active requests.
from http.server import ThreadingHTTPServer, BaseHTTPRequestHandler
from urllib.parse import urlparse, parse_qs
from threading import Lock
import json, time
HOST, PORT = "127.0.0.1", 8815
SOFT_CAPACITY = 2
state = {"active": 0, "peak": 0, "requests": 0}
lock = Lock()
PAGE = b"""<!doctype html><html><head><title>Capacity Lab</title></head><body>
<h1>Capacity Lab</h1><p id='state'>booting</p><script>
const p = new URLSearchParams(location.search);
const id = p.get('id') || 'case';
fetch('/api/work?id=' + encodeURIComponent(id))
.then(r => r.json())
.then(d => { document.querySelector('#state').textContent = 'ready:' + d.id; });
</script></body></html>"""
class Handler(BaseHTTPRequestHandler):
def log_message(self, *args):
pass
def send_json(self, status, obj):
raw = json.dumps(obj).encode()
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.end_headers(); self.wfile.write(raw)
def do_GET(self):
u = urlparse(self.path)
if u.path == "/health":
return self.send_json(200, {"ok": True, **state})
if u.path == "/":
self.send_response(200); self.send_header("Content-Type", "text/html")
self.send_header("Content-Length", str(len(PAGE))); self.end_headers(); self.wfile.write(PAGE); return
if u.path == "/api/work":
case_id = parse_qs(u.query).get("id", ["case"])[0]
with lock:
state["active"] += 1; state["requests"] += 1
state["peak"] = max(state["peak"], state["active"])
active = state["active"]
# The loopback AUT intentionally models a service whose latency rises after two concurrent calls.
delay = 0.12 + 0.18 * max(0, active - SOFT_CAPACITY)
try:
time.sleep(delay)
return self.send_json(200, {"id": case_id, "active": active, "delay_ms": int(delay * 1000)})
finally:
with lock: state["active"] -= 1
self.send_error(404)
if __name__ == "__main__":
ThreadingHTTPServer((HOST, PORT), Handler).serve_forever()
Run python aut_server.py, then verify
http://127.0.0.1:8815/health. That health response is
your pre-mutation state: request count, current active requests, and
observed peak.
3. First measure the fixture without browsers
This control measurement separates AUT-capacity behavior from
browser/session cost. Save as http_capacity.py. It runs
the same six synthetic cases at worker counts 1, 2, and 4 and
optionally writes 64 KiB per-case evidence so artifact IO has an
explicit cost.
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.request import urlopen
from pathlib import Path
import json, statistics, time
BASE = "http://127.0.0.1:8815"
CASES = [f"case-{i}" for i in range(1, 7)]
def percentile(values, p):
ordered = sorted(values)
idx = max(0, min(len(ordered) - 1, round((len(ordered) - 1) * p)))
return ordered[idx]
def one(case_id, artifact_bytes=0):
t0 = time.perf_counter()
with urlopen(f"{BASE}/api/work?id={case_id}", timeout=5) as r:
payload = json.load(r)
latency_ms = (time.perf_counter() - t0) * 1000
artifact_ms = 0.0
if artifact_bytes:
a0 = time.perf_counter()
Path("artifacts").mkdir(exist_ok=True)
(Path("artifacts") / f"{case_id}.bin").write_bytes(b"x" * artifact_bytes)
artifact_ms = (time.perf_counter() - a0) * 1000
return latency_ms, artifact_ms, payload["active"]
def run(workers, artifact_bytes=0):
start = time.perf_counter(); rows = []
with ThreadPoolExecutor(max_workers=workers) as pool:
futures = [pool.submit(one, c, artifact_bytes) for c in CASES]
for f in as_completed(futures): rows.append(f.result())
wall_ms = (time.perf_counter() - start) * 1000
latencies = [x[0] for x in rows]
return {
"workers": workers,
"wall_ms": round(wall_ms, 1),
"p50_ms": round(statistics.median(latencies), 1),
"p95_ms": round(percentile(latencies, .95), 1),
"peak_active": max(x[2] for x in rows),
"artifact_ms": round(sum(x[1] for x in rows), 1),
}
for workers in (1, 2, 4):
print(json.dumps(run(workers, artifact_bytes=64 * 1024)))
print(json.dumps({"optimization": "failure-only evidence simulation", **run(2, artifact_bytes=0)}))
Expected shape—not exact numbers—is: concurrency 2 reduces wall time relative to 1; concurrency 4 raises per-request p95 because the fixture’s soft capacity is 2; the “failure-only evidence simulation” removes artifact-write time. Because the service is local and deterministic, this becomes a useful control when browser results are noisy.
4. Instrument Selenium session startup, execution, evidence, and teardown
Now measure the browser-test system. Each concurrent case creates its own temporary profile and its own driver. No driver, cookie store, or profile is shared across threads.
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import json, tempfile, shutil, time
BASE = "http://127.0.0.1:8815"
CASES = [f"case-{i}" for i in range(1, 7)]
def run_case(case_id, capture=False):
profile = Path(tempfile.mkdtemp(prefix=f"sel-{case_id}-"))
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1280,900")
options.add_argument(f"--user-data-dir={profile}")
driver = None
t0 = time.perf_counter()
try:
s0 = time.perf_counter()
driver = webdriver.Chrome(options=options)
startup_ms = (time.perf_counter() - s0) * 1000
e0 = time.perf_counter()
driver.get(f"{BASE}/?id={case_id}")
WebDriverWait(driver, 5).until(
lambda d: d.find_element(By.ID, "state").text == f"ready:{case_id}"
)
execution_ms = (time.perf_counter() - e0) * 1000
a0 = time.perf_counter()
if capture:
Path("evidence").mkdir(exist_ok=True)
driver.save_screenshot(str(Path("evidence") / f"{case_id}.png"))
artifact_ms = (time.perf_counter() - a0) * 1000
caps = dict(driver.capabilities)
return {
"case": case_id,
"session": driver.session_id,
"browser": caps.get("browserName"),
"browserVersion": caps.get("browserVersion"),
"startup_ms": round(startup_ms, 1),
"execution_ms": round(execution_ms, 1),
"artifact_ms": round(artifact_ms, 1),
"total_ms": round((time.perf_counter() - t0) * 1000, 1),
}
finally:
if driver is not None:
driver.quit()
shutil.rmtree(profile, ignore_errors=True)
def run_suite(workers, capture=False):
t0 = time.perf_counter(); results = []
with ThreadPoolExecutor(max_workers=workers) as pool:
futures = [pool.submit(run_case, c, capture) for c in CASES]
for f in as_completed(futures): results.append(f.result())
return {"workers": workers, "wall_ms": round((time.perf_counter()-t0)*1000,1), "results": results}
if __name__ == "__main__":
print(json.dumps(run_suite(1, capture=True), indent=2))
print(json.dumps(run_suite(2, capture=True), indent=2))
The returned data separates startup_ms,
execution_ms, artifact_ms, and total case
time. The suite-level wall_ms measures feedback time
for the same six cases. Compare one worker and two workers first.
Only increase further if host and AUT evidence remain healthy.
This baseline preserves test independence from Chapters 14 and 21. Later you may evaluate browser reuse as an explicit trade-off, but do not silently change isolation and call the result a pure speed improvement.
5. Observe queue and slot state on Grid
If using a local Grid, record a snapshot before, during, and after
the parallel run. GraphQL is useful because it can expose
maxSession, sessionCount,
sessionQueueSize, and Node/slot state without mutating
the Grid.
from urllib.request import Request, urlopen
import json
GRID = "http://127.0.0.1:4444"
query = """{
grid { maxSession sessionCount sessionQueueSize }
nodesInfo { nodes { id uri status slotCount sessionCount } }
}"""
req = Request(
GRID + "/graphql",
data=json.dumps({"query": query}).encode(),
headers={"Content-Type": "application/json"},
)
with urlopen(req, timeout=3) as response:
snapshot = json.load(response)["data"]
print(json.dumps(snapshot, indent=2))
Interpretation matters: a growing queue with all matching slots busy points to Grid/browser capacity. An empty queue with long DOM-ready time points elsewhere—often the AUT, network, waits, or browser host.
6. Capture host resource context without confusing it with Selenium state
Selenium does not own CPU/RAM telemetry. Capture host evidence from the operating system and correlate it by timestamp with the test run.
# Linux examples — read-only
nproc
free -h
ps -eo pid,comm,rss,%cpu --sort=-%cpu | grep -E 'chrome|chromium|firefox|java' | head -30
The following example makes the Capture host resource context without confusing it with Selenium state behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
# Windows PowerShell examples — read-only
Get-CimInstance Win32_ComputerSystem | Select-Object NumberOfLogicalProcessors,TotalPhysicalMemory
Get-Process chrome,msedge,firefox,java -ErrorAction SilentlyContinue |
Select-Object Id,ProcessName,CPU,WorkingSet64
Record process counts too. A flat wall-time improvement that leaves orphaned browsers is not an optimization—it is a leak.
7. Change one variable at a time
A valid experiment keeps test cases, synthetic data, browser version, evidence policy, and machine class constant while changing exactly one chosen control.
| Run | Changed variable | Everything else held constant | Question |
|---|---|---|---|
| A | workers=1 | same six cases, fresh sessions, screenshots on | serial baseline |
| B | workers=2 | same cases/evidence/browser/host | does throughput improve without p95/resource regression? |
| C | workers=4 | same cases/evidence/browser/host | where does first saturation evidence appear? |
| D | screenshots failure-only | workers=2 and same cases | how much evidence IO is avoided? |
8. Before/after evidence packet
For each controlled run, store one compact JSON summary with: run ID, timestamp, Selenium/browser/Grid versions, worker count, Grid max/current sessions and queue size if applicable, host CPU count/memory class, six session IDs, startup/execution/artifact distributions, wall time, failures/retries, and cleanup result.
Do not store whole personal profiles, cookies, or unrelated page sources just to make the packet “richer.” Performance evidence is still subject to Chapter 24 minimization and retention policy.
9. Challenge: choose the control, not the syntax
Scenario: two workers cut wall time from 120 s to 72 s. Four workers cut it only to 69 s, Grid queue remains zero, CPU reaches 95%, memory pressure rises, and one browser times out. What should you change next?
Do not add retries or six more workers. The mental model points to browser-host saturation. Keep the safe ceiling near two workers, inspect browser/resource cost, and evaluate a different runner/Node size or more isolated Nodes before increasing offered concurrency.
If the real question changes from test-suite feedback time to service throughput under controlled virtual-user/request load, stop scaling Selenium workers and move that experiment to JMeter, k6, or another purpose-built load-testing tool.
Knowledge checks
Answer from the operating model, then reveal the explanation.
Why measure the HTTP fixture before Selenium?
It provides a control for AUT capacity. If browser timings change while direct fixture behavior is stable, the browser/test/Grid layer becomes a stronger suspect.
Why must every parallel case own a fresh driver/profile in the baseline?
To keep session state and browser runtime isolated so speed changes are not purchased by hidden state sharing.
A Grid queue grows while all matching slots are occupied. What layer is constrained?
Grid/browser capacity for that stereotype is constrained; the runner is offering work faster than matching slots become free.
Wall time improves but p95 case time and failures worsen sharply. Is that automatically a win?
No. Throughput, latency, and reliability must be considered together. The system may already be past its safe saturation point.
What is the purpose of the artifact-on/off comparison?
To quantify evidence IO/storage cost while keeping test work and concurrency constant, not to justify deleting necessary failure evidence.
Summary and next bridge
- The fixture provides a controlled, local saturation signal.
- Browser timing is decomposed into session startup, execution, evidence, and total time.
- Grid queue/slot state distinguishes scheduling pressure from AUT latency.
- OS resource telemetry complements but does not replace Selenium/Grid evidence.
- One-variable experiments make causal claims defensible.
Lesson 3 evaluates the design choices behind lifecycle reuse, headless mode, Grid latency, concurrency, evidence, fixtures, browser matrices, and shared versus dedicated capacity.
Primary references and version notes
- Selenium downloads — stable client and Grid release baseline.
- Getting started with Selenium Grid — CPU/RAM sizing and session-capacity guidance.
-
Grid CLI options
— current
max-sessions, queue timeout, retry interval, and related controls. - Grid GraphQL support — observable max sessions, session count, nodes, slots, and queue size.
- Grid observability — traces, events, and logs for distributed diagnosis.
- Performance testing with Selenium is discouraged — why WebDriver suite timing is not a substitute for load/performance tooling.
The mandatory examples pin selenium==4.47.0 and
Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid
baseline. Selenium Manager remains the normal local
driver-resolution path. Hardware, browser versions, container
images, Grid slot counts, runner sizes, AUT capacity, and network
conditions are recorded as benchmark context rather than assumed
constants.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.