Chapter 25Lesson 02~215 minutes

Performance Engineering for Test Suites and Grid Capacity: Guided Hands-On Workflow

Now turn the model into measurements. The workflow uses a loopback AUT, fresh browser sessions, explicit timers, optional Grid snapshots, and a synthetic saturation model so every optimization can be traced to one changed variable.

instrumentationserial vs parallelGrid GraphQLresource snapshotsartifact overheadcontrolled experiment

Learning objectives

  • Create a reproducible loopback fixture whose latency rises after a known soft capacity.
  • Measure serial and parallel wall time without sharing WebDriver sessions or test identities.
  • Record browser session-start, scenario, evidence, and total timings separately.
  • Observe Grid max sessions/session count/queue size where a local Grid is used.
  • Change one concurrency or evidence variable at a time and compare against a documented baseline.

1. Lab topology and preflight

The mandatory path is free and local. The AUT binds to 127.0.0.1:8815; it does not require real accounts, external traffic, or a paid browser cloud. Its /api/work endpoint deliberately adds latency only after more than two concurrent calls so saturation is visible without stressing any real service.

  • Python 3.10+.
  • selenium==4.47.0 for the browser measurement path.
  • A supported local Chromium-family browser; Selenium Manager resolves the driver.
  • Optional: Selenium Server/Grid 4.47.0 on loopback for queue/slot observation.
  • Enough local memory for the worker count you choose. Start small.
Disposable target only

The fixture is intentionally synthetic. Do not substitute a production base URL or a public Grid.

2. Start a measurable loopback AUT

Save this as aut_server.py. The service reports current/peak concurrency and applies a deterministic latency penalty above two active requests.

from http.server import ThreadingHTTPServer, BaseHTTPRequestHandler
from urllib.parse import urlparse, parse_qs
from threading import Lock
import json, time

HOST, PORT = "127.0.0.1", 8815
SOFT_CAPACITY = 2
state = {"active": 0, "peak": 0, "requests": 0}
lock = Lock()

PAGE = b"""<!doctype html><html><head><title>Capacity Lab</title></head><body>
<h1>Capacity Lab</h1><p id='state'>booting</p><script>
const p = new URLSearchParams(location.search);
const id = p.get('id') || 'case';
fetch('/api/work?id=' + encodeURIComponent(id))
 .then(r => r.json())
 .then(d => { document.querySelector('#state').textContent = 'ready:' + d.id; });
</script></body></html>"""

class Handler(BaseHTTPRequestHandler):
    def log_message(self, *args):
        pass
    def send_json(self, status, obj):
        raw = json.dumps(obj).encode()
        self.send_response(status)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(raw)))
        self.end_headers(); self.wfile.write(raw)
    def do_GET(self):
        u = urlparse(self.path)
        if u.path == "/health":
            return self.send_json(200, {"ok": True, **state})
        if u.path == "/":
            self.send_response(200); self.send_header("Content-Type", "text/html")
            self.send_header("Content-Length", str(len(PAGE))); self.end_headers(); self.wfile.write(PAGE); return
        if u.path == "/api/work":
            case_id = parse_qs(u.query).get("id", ["case"])[0]
            with lock:
                state["active"] += 1; state["requests"] += 1
                state["peak"] = max(state["peak"], state["active"])
                active = state["active"]
            # The loopback AUT intentionally models a service whose latency rises after two concurrent calls.
            delay = 0.12 + 0.18 * max(0, active - SOFT_CAPACITY)
            try:
                time.sleep(delay)
                return self.send_json(200, {"id": case_id, "active": active, "delay_ms": int(delay * 1000)})
            finally:
                with lock: state["active"] -= 1
        self.send_error(404)

if __name__ == "__main__":
    ThreadingHTTPServer((HOST, PORT), Handler).serve_forever()

Run python aut_server.py, then verify http://127.0.0.1:8815/health. That health response is your pre-mutation state: request count, current active requests, and observed peak.

3. First measure the fixture without browsers

This control measurement separates AUT-capacity behavior from browser/session cost. Save as http_capacity.py. It runs the same six synthetic cases at worker counts 1, 2, and 4 and optionally writes 64 KiB per-case evidence so artifact IO has an explicit cost.

from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.request import urlopen
from pathlib import Path
import json, statistics, time

BASE = "http://127.0.0.1:8815"
CASES = [f"case-{i}" for i in range(1, 7)]

def percentile(values, p):
    ordered = sorted(values)
    idx = max(0, min(len(ordered) - 1, round((len(ordered) - 1) * p)))
    return ordered[idx]

def one(case_id, artifact_bytes=0):
    t0 = time.perf_counter()
    with urlopen(f"{BASE}/api/work?id={case_id}", timeout=5) as r:
        payload = json.load(r)
    latency_ms = (time.perf_counter() - t0) * 1000
    artifact_ms = 0.0
    if artifact_bytes:
        a0 = time.perf_counter()
        Path("artifacts").mkdir(exist_ok=True)
        (Path("artifacts") / f"{case_id}.bin").write_bytes(b"x" * artifact_bytes)
        artifact_ms = (time.perf_counter() - a0) * 1000
    return latency_ms, artifact_ms, payload["active"]

def run(workers, artifact_bytes=0):
    start = time.perf_counter(); rows = []
    with ThreadPoolExecutor(max_workers=workers) as pool:
        futures = [pool.submit(one, c, artifact_bytes) for c in CASES]
        for f in as_completed(futures): rows.append(f.result())
    wall_ms = (time.perf_counter() - start) * 1000
    latencies = [x[0] for x in rows]
    return {
        "workers": workers,
        "wall_ms": round(wall_ms, 1),
        "p50_ms": round(statistics.median(latencies), 1),
        "p95_ms": round(percentile(latencies, .95), 1),
        "peak_active": max(x[2] for x in rows),
        "artifact_ms": round(sum(x[1] for x in rows), 1),
    }

for workers in (1, 2, 4):
    print(json.dumps(run(workers, artifact_bytes=64 * 1024)))
print(json.dumps({"optimization": "failure-only evidence simulation", **run(2, artifact_bytes=0)}))

Expected shape—not exact numbers—is: concurrency 2 reduces wall time relative to 1; concurrency 4 raises per-request p95 because the fixture’s soft capacity is 2; the “failure-only evidence simulation” removes artifact-write time. Because the service is local and deterministic, this becomes a useful control when browser results are noisy.

4. Instrument Selenium session startup, execution, evidence, and teardown

Now measure the browser-test system. Each concurrent case creates its own temporary profile and its own driver. No driver, cookie store, or profile is shared across threads.

from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import json, tempfile, shutil, time

BASE = "http://127.0.0.1:8815"
CASES = [f"case-{i}" for i in range(1, 7)]

def run_case(case_id, capture=False):
    profile = Path(tempfile.mkdtemp(prefix=f"sel-{case_id}-"))
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")
    options.add_argument("--window-size=1280,900")
    options.add_argument(f"--user-data-dir={profile}")
    driver = None
    t0 = time.perf_counter()
    try:
        s0 = time.perf_counter()
        driver = webdriver.Chrome(options=options)
        startup_ms = (time.perf_counter() - s0) * 1000
        e0 = time.perf_counter()
        driver.get(f"{BASE}/?id={case_id}")
        WebDriverWait(driver, 5).until(
            lambda d: d.find_element(By.ID, "state").text == f"ready:{case_id}"
        )
        execution_ms = (time.perf_counter() - e0) * 1000
        a0 = time.perf_counter()
        if capture:
            Path("evidence").mkdir(exist_ok=True)
            driver.save_screenshot(str(Path("evidence") / f"{case_id}.png"))
        artifact_ms = (time.perf_counter() - a0) * 1000
        caps = dict(driver.capabilities)
        return {
            "case": case_id,
            "session": driver.session_id,
            "browser": caps.get("browserName"),
            "browserVersion": caps.get("browserVersion"),
            "startup_ms": round(startup_ms, 1),
            "execution_ms": round(execution_ms, 1),
            "artifact_ms": round(artifact_ms, 1),
            "total_ms": round((time.perf_counter() - t0) * 1000, 1),
        }
    finally:
        if driver is not None:
            driver.quit()
        shutil.rmtree(profile, ignore_errors=True)

def run_suite(workers, capture=False):
    t0 = time.perf_counter(); results = []
    with ThreadPoolExecutor(max_workers=workers) as pool:
        futures = [pool.submit(run_case, c, capture) for c in CASES]
        for f in as_completed(futures): results.append(f.result())
    return {"workers": workers, "wall_ms": round((time.perf_counter()-t0)*1000,1), "results": results}

if __name__ == "__main__":
    print(json.dumps(run_suite(1, capture=True), indent=2))
    print(json.dumps(run_suite(2, capture=True), indent=2))

The returned data separates startup_ms, execution_ms, artifact_ms, and total case time. The suite-level wall_ms measures feedback time for the same six cases. Compare one worker and two workers first. Only increase further if host and AUT evidence remain healthy.

Why fresh sessions even though reuse is faster

This baseline preserves test independence from Chapters 14 and 21. Later you may evaluate browser reuse as an explicit trade-off, but do not silently change isolation and call the result a pure speed improvement.

5. Observe queue and slot state on Grid

If using a local Grid, record a snapshot before, during, and after the parallel run. GraphQL is useful because it can expose maxSession, sessionCount, sessionQueueSize, and Node/slot state without mutating the Grid.

from urllib.request import Request, urlopen
import json

GRID = "http://127.0.0.1:4444"
query = """{
  grid { maxSession sessionCount sessionQueueSize }
  nodesInfo { nodes { id uri status slotCount sessionCount } }
}"""
req = Request(
    GRID + "/graphql",
    data=json.dumps({"query": query}).encode(),
    headers={"Content-Type": "application/json"},
)
with urlopen(req, timeout=3) as response:
    snapshot = json.load(response)["data"]
print(json.dumps(snapshot, indent=2))

Interpretation matters: a growing queue with all matching slots busy points to Grid/browser capacity. An empty queue with long DOM-ready time points elsewhere—often the AUT, network, waits, or browser host.

6. Capture host resource context without confusing it with Selenium state

Selenium does not own CPU/RAM telemetry. Capture host evidence from the operating system and correlate it by timestamp with the test run.

# Linux examples — read-only
nproc
free -h
ps -eo pid,comm,rss,%cpu --sort=-%cpu | grep -E 'chrome|chromium|firefox|java' | head -30

The following example makes the Capture host resource context without confusing it with Selenium state behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

# Windows PowerShell examples — read-only
Get-CimInstance Win32_ComputerSystem | Select-Object NumberOfLogicalProcessors,TotalPhysicalMemory
Get-Process chrome,msedge,firefox,java -ErrorAction SilentlyContinue |
  Select-Object Id,ProcessName,CPU,WorkingSet64

Record process counts too. A flat wall-time improvement that leaves orphaned browsers is not an optimization—it is a leak.

7. Change one variable at a time

A valid experiment keeps test cases, synthetic data, browser version, evidence policy, and machine class constant while changing exactly one chosen control.

Run Changed variable Everything else held constant Question
A workers=1 same six cases, fresh sessions, screenshots on serial baseline
B workers=2 same cases/evidence/browser/host does throughput improve without p95/resource regression?
C workers=4 same cases/evidence/browser/host where does first saturation evidence appear?
D screenshots failure-only workers=2 and same cases how much evidence IO is avoided?

8. Before/after evidence packet

For each controlled run, store one compact JSON summary with: run ID, timestamp, Selenium/browser/Grid versions, worker count, Grid max/current sessions and queue size if applicable, host CPU count/memory class, six session IDs, startup/execution/artifact distributions, wall time, failures/retries, and cleanup result.

Do not store whole personal profiles, cookies, or unrelated page sources just to make the packet “richer.” Performance evidence is still subject to Chapter 24 minimization and retention policy.

9. Challenge: choose the control, not the syntax

Scenario: two workers cut wall time from 120 s to 72 s. Four workers cut it only to 69 s, Grid queue remains zero, CPU reaches 95%, memory pressure rises, and one browser times out. What should you change next?

Do not add retries or six more workers. The mental model points to browser-host saturation. Keep the safe ceiling near two workers, inspect browser/resource cost, and evaluate a different runner/Node size or more isolated Nodes before increasing offered concurrency.

If the real question changes from test-suite feedback time to service throughput under controlled virtual-user/request load, stop scaling Selenium workers and move that experiment to JMeter, k6, or another purpose-built load-testing tool.

Knowledge checks

Answer from the operating model, then reveal the explanation.

Why measure the HTTP fixture before Selenium?

Why must every parallel case own a fresh driver/profile in the baseline?

A Grid queue grows while all matching slots are occupied. What layer is constrained?

Wall time improves but p95 case time and failures worsen sharply. Is that automatically a win?

What is the purpose of the artifact-on/off comparison?

Summary and next bridge

  • The fixture provides a controlled, local saturation signal.
  • Browser timing is decomposed into session startup, execution, evidence, and total time.
  • Grid queue/slot state distinguishes scheduling pressure from AUT latency.
  • OS resource telemetry complements but does not replace Selenium/Grid evidence.
  • One-variable experiments make causal claims defensible.

Lesson 3 evaluates the design choices behind lifecycle reuse, headless mode, Grid latency, concurrency, evidence, fixtures, browser matrices, and shared versus dedicated capacity.

Next lesson

Performance Engineering for Test Suites and Grid Capacity: Configuration, Design Patterns, and Trade-Offs

Continue with Performance Engineering for Test Suites and Grid Capacity: Configuration, Design Patterns, and Trade-Offs. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Primary references and version notes

Version baseline — August 2026

The mandatory examples pin selenium==4.47.0 and Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid baseline. Selenium Manager remains the normal local driver-resolution path. Hardware, browser versions, container images, Grid slot counts, runner sizes, AUT capacity, and network conditions are recorded as benchmark context rather than assumed constants.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.