Chapter 25Lesson 05~245 minutes

Checkpoint Lab — Performance Engineering for Test Suites and Grid Capacity

The checkpoint produces a small but auditable capacity report. You will run the same six-case fixture at multiple concurrency levels, measure session and wall-clock cost, find the first saturation signal, quantify one evidence optimization, and document the point where a load-testing tool becomes the correct instrument.

checkpointcapacity reportsafe concurrencysaturationoptimization deltaJMeter boundary

Learning objectives

  • Execute at least three controlled runs over the same six-case inventory and evidence policy.
  • Predict session, AUT, and evidence changes before each run and verify them independently.
  • Identify the first saturation signal from wall time, p95 latency, queue/resource evidence, or failures.
  • Choose and justify a safe concurrency target rather than maximizing worker count.
  • Quantify one optimization and state which questions require JMeter or another load-testing tool.

1. Checkpoint architecture and exact assumptions

Use the same aut_server.py from Lesson 2 on 127.0.0.1:8815. The primary Selenium path uses Python 3.10+, selenium==4.47.0, a locally installed supported Chromium-family browser, Selenium Manager, headless mode, and one fresh browser/profile per concurrent case. Grid is optional; if used, pin Selenium Server/Grid 4.47.0 and record maxSession, sessionCount, and queue size.

The six synthetic cases contain no credentials or personal data. Evidence paths are per case/run. The lab is intentionally local and disposable.

2. Preflight: predict before executing

Write predictions before running:

  1. Moving from one to two workers should reduce wall time because the fixture’s soft capacity is two, while per-case p95 should remain near baseline.
  2. Moving beyond two workers should raise /api/work latency; browser sessions may still start quickly, proving AUT saturation can occur even when Grid/browser capacity exists.
  3. Changing from always-on screenshots to failure-only evidence should reduce artifact bytes/write time without changing business assertions.

Also predict resource state: browser process count should roughly follow active sessions and return to baseline after teardown. If it does not, stop and repair cleanup before interpreting timings.

3. Run the control capacity model first

Start the fixture, then run http_capacity.py from Lesson 2. It prints three controlled concurrency runs plus one evidence-policy optimization run. Save the JSON lines as http-capacity.jsonl.

from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.request import urlopen
from pathlib import Path
import json, statistics, time

BASE = "http://127.0.0.1:8815"
CASES = [f"case-{i}" for i in range(1, 7)]

def percentile(values, p):
    ordered = sorted(values)
    idx = max(0, min(len(ordered) - 1, round((len(ordered) - 1) * p)))
    return ordered[idx]

def one(case_id, artifact_bytes=0):
    t0 = time.perf_counter()
    with urlopen(f"{BASE}/api/work?id={case_id}", timeout=5) as r:
        payload = json.load(r)
    latency_ms = (time.perf_counter() - t0) * 1000
    artifact_ms = 0.0
    if artifact_bytes:
        a0 = time.perf_counter()
        Path("artifacts").mkdir(exist_ok=True)
        (Path("artifacts") / f"{case_id}.bin").write_bytes(b"x" * artifact_bytes)
        artifact_ms = (time.perf_counter() - a0) * 1000
    return latency_ms, artifact_ms, payload["active"]

def run(workers, artifact_bytes=0):
    start = time.perf_counter(); rows = []
    with ThreadPoolExecutor(max_workers=workers) as pool:
        futures = [pool.submit(one, c, artifact_bytes) for c in CASES]
        for f in as_completed(futures): rows.append(f.result())
    wall_ms = (time.perf_counter() - start) * 1000
    latencies = [x[0] for x in rows]
    return {
        "workers": workers,
        "wall_ms": round(wall_ms, 1),
        "p50_ms": round(statistics.median(latencies), 1),
        "p95_ms": round(percentile(latencies, .95), 1),
        "peak_active": max(x[2] for x in rows),
        "artifact_ms": round(sum(x[1] for x in rows), 1),
    }

for workers in (1, 2, 4):
    print(json.dumps(run(workers, artifact_bytes=64 * 1024)))
print(json.dumps({"optimization": "failure-only evidence simulation", **run(2, artifact_bytes=0)}))

This direct-HTTP result is not a load test; it is a tiny deterministic control for this lab fixture. Its purpose is to tell you where the synthetic AUT begins adding latency so you can interpret the browser results.

4. Run the browser capacity experiment

Use the Selenium measurement harness from Lesson 2, but execute three candidate worker levels. Start with 1 and 2. Add a third level only within your local resource budget; the sample below uses 3. If your host cannot safely launch three isolated browsers, use a local Grid with enough isolated capacity or use the HTTP control as the saturation demonstration and document the browser ceiling as 2.

import json
from pathlib import Path
from selenium_measure import run_suite

runs = []
for workers in (1, 2, 3):
    result = run_suite(workers, capture=True)
    runs.append(result)
    Path(f"run-{workers}.json").write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps({"runs": runs}, indent=2))

Do not change browser version, case inventory, fixture data, or screenshot policy between those three runs. Otherwise you no longer have a one-variable concurrency experiment.

5. Collect independent evidence

Your evidence packet must include:

  • Selenium/binding/browser/driver versions and each session ID;
  • host OS, logical CPUs, memory class, and browser process observations;
  • one-, two-, and three-worker wall times plus startup/execution/artifact distributions;
  • fixture /health state before/after, including peak active requests;
  • Grid maxSession, sessionCount, queue size, and Node/slot evidence if RemoteWebDriver is used;
  • failure/retry count (expected zero);
  • artifact byte count and cleanup verification.

Grid evidence belongs in the packet only when Grid actually participates. Local WebDriver runs should not invent a Grid layer.

6. Identify the first saturation signal and choose the ceiling

Choose the highest concurrency that still meets your reliability and latency criteria before the first clear knee. A simple lab rule can be:

  • zero failures/retries;
  • no growing Grid queue at steady state;
  • no orphan browsers;
  • p95 execution latency no more than 1.5× the one-worker baseline;
  • host memory/CPU remains below the point where browser/session instability appears.

Do not treat 1.5× as a universal production SLO; it is only a transparent checkpoint rule. Your real threshold should come from the team’s feedback-time and reliability objectives.

7. Quantify one optimization without changing semantics

Repeat the selected safe worker count with failure-only screenshot policy. Business assertions, browser lifecycle, data, and concurrency remain unchanged. Report:

artifact_saving_ms = baseline_artifact_ms - optimized_artifact_ms and feedback_delta_% = (baseline_wall - optimized_wall) / baseline_wall × 100.

Keep a small always-on manifest plus complete failure evidence. The optimization is “write less redundant success evidence,” not “erase diagnostic evidence.”

8. Capacity report template

The following table organizes the key choices and evidence for Capacity report template. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Field Run A Run B Run C Optimized
workers 1 2 3 safe target
wall time measure measure measure measure
session-start p50/p95 measure measure measure measure
execution p50/p95 measure measure measure measure
Grid queue peak N/A or measure N/A or measure N/A or measure N/A or measure
host/browser pressure record record record record
failures / retries 0 expected 0 expected 0 expected 0 expected
artifact bytes / write ms measure measure measure reduced policy
decision baseline candidate saturation? adopt/reject

9. What belongs in JMeter or another performance tool instead

Close the report with a scope statement. Examples that belong outside Selenium:

  • maximum sustainable API requests/second;
  • p95/p99 service latency under hundreds or thousands of virtual users;
  • backend throughput, database saturation, or autoscaling thresholds under controlled load;
  • long-duration soak/stress tests;
  • protocol-level workload models where browser rendering is irrelevant.

A sensible combined strategy is: use JMeter/k6/etc. for controlled service load; use a small Selenium smoke journey before/during/after in a separately designed environment if you need end-user functional confirmation. Do not make Selenium generate the load.

10. Verification and cleanup checklist

  • Three controlled runs exist and use the same six cases.
  • Predictions are written before results and compared afterward.
  • Safe concurrency is justified by measured evidence, not CPU count alone.
  • One optimization has a quantified delta and does not weaken assertions/isolation.
  • No production target, real account, public Grid, or security bypass is used.
  • Every WebDriver session is quit; temporary profiles and evidence beyond the retained packet are removed.
  • Fixture server is stopped and browser process count returns to baseline.
  • The report names questions that require a purpose-built load/performance tool.

Knowledge checks

Answer from the operating model, then reveal the explanation.

Four workers reduce wall time by only 2% versus two workers while p95 doubles. Which worker count is the better ceiling?

Why must the optimization run keep browser lifecycle and case inventory unchanged?

The Grid queue remains empty but fixture peak active and p95 rise sharply. What saturated first?

What is the strongest evidence that cleanup failed?

Which question should be handed to JMeter/k6 rather than answered with more Selenium workers?

Summary and next bridge

  • A capacity report needs repeated controlled runs, context, and explicit predictions.
  • The safe worker ceiling is selected before saturation, not at the theoretical maximum.
  • Optimization deltas are meaningful only when test semantics and evidence requirements remain controlled.
  • Browser/Grid/AUT/resource signals must be interpreted together.
  • Load/performance questions beyond the browser-test platform belong to purpose-built tools.

Chapter 26 moves to Selenium IDE, record/playback, and migration from captured interactions to maintainable code. The performance baseline from this chapter helps distinguish a maintainable migration from one that merely replays steps faster.

Next chapter

Selenium IDE, Record/Playback, and Migration to Maintainable Code: Core Concepts and Mental Model

Continue with Selenium IDE, Record/Playback, and Migration to Maintainable Code: Core Concepts and Mental Model. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Primary references and version notes

Version baseline — August 2026

The mandatory examples pin selenium==4.47.0 and Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid baseline. Selenium Manager remains the normal local driver-resolution path. Hardware, browser versions, container images, Grid slot counts, runner sizes, AUT capacity, and network conditions are recorded as benchmark context rather than assumed constants.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.