Checkpoint Lab — Performance Engineering for Test Suites and Grid Capacity
The checkpoint produces a small but auditable capacity report. You will run the same six-case fixture at multiple concurrency levels, measure session and wall-clock cost, find the first saturation signal, quantify one evidence optimization, and document the point where a load-testing tool becomes the correct instrument.
Learning objectives
- Execute at least three controlled runs over the same six-case inventory and evidence policy.
- Predict session, AUT, and evidence changes before each run and verify them independently.
- Identify the first saturation signal from wall time, p95 latency, queue/resource evidence, or failures.
- Choose and justify a safe concurrency target rather than maximizing worker count.
- Quantify one optimization and state which questions require JMeter or another load-testing tool.
1. Checkpoint architecture and exact assumptions
Use the same aut_server.py from Lesson 2 on
127.0.0.1:8815. The primary Selenium path uses Python
3.10+, selenium==4.47.0, a locally installed supported
Chromium-family browser, Selenium Manager, headless mode, and one
fresh browser/profile per concurrent case. Grid is optional; if
used, pin Selenium Server/Grid 4.47.0 and record
maxSession, sessionCount, and queue size.
The six synthetic cases contain no credentials or personal data. Evidence paths are per case/run. The lab is intentionally local and disposable.
2. Preflight: predict before executing
Write predictions before running:
- Moving from one to two workers should reduce wall time because the fixture’s soft capacity is two, while per-case p95 should remain near baseline.
-
Moving beyond two workers should raise
/api/worklatency; browser sessions may still start quickly, proving AUT saturation can occur even when Grid/browser capacity exists. - Changing from always-on screenshots to failure-only evidence should reduce artifact bytes/write time without changing business assertions.
Also predict resource state: browser process count should roughly follow active sessions and return to baseline after teardown. If it does not, stop and repair cleanup before interpreting timings.
3. Run the control capacity model first
Start the fixture, then run http_capacity.py from
Lesson 2. It prints three controlled concurrency runs plus one
evidence-policy optimization run. Save the JSON lines as
http-capacity.jsonl.
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.request import urlopen
from pathlib import Path
import json, statistics, time
BASE = "http://127.0.0.1:8815"
CASES = [f"case-{i}" for i in range(1, 7)]
def percentile(values, p):
ordered = sorted(values)
idx = max(0, min(len(ordered) - 1, round((len(ordered) - 1) * p)))
return ordered[idx]
def one(case_id, artifact_bytes=0):
t0 = time.perf_counter()
with urlopen(f"{BASE}/api/work?id={case_id}", timeout=5) as r:
payload = json.load(r)
latency_ms = (time.perf_counter() - t0) * 1000
artifact_ms = 0.0
if artifact_bytes:
a0 = time.perf_counter()
Path("artifacts").mkdir(exist_ok=True)
(Path("artifacts") / f"{case_id}.bin").write_bytes(b"x" * artifact_bytes)
artifact_ms = (time.perf_counter() - a0) * 1000
return latency_ms, artifact_ms, payload["active"]
def run(workers, artifact_bytes=0):
start = time.perf_counter(); rows = []
with ThreadPoolExecutor(max_workers=workers) as pool:
futures = [pool.submit(one, c, artifact_bytes) for c in CASES]
for f in as_completed(futures): rows.append(f.result())
wall_ms = (time.perf_counter() - start) * 1000
latencies = [x[0] for x in rows]
return {
"workers": workers,
"wall_ms": round(wall_ms, 1),
"p50_ms": round(statistics.median(latencies), 1),
"p95_ms": round(percentile(latencies, .95), 1),
"peak_active": max(x[2] for x in rows),
"artifact_ms": round(sum(x[1] for x in rows), 1),
}
for workers in (1, 2, 4):
print(json.dumps(run(workers, artifact_bytes=64 * 1024)))
print(json.dumps({"optimization": "failure-only evidence simulation", **run(2, artifact_bytes=0)}))
This direct-HTTP result is not a load test; it is a tiny deterministic control for this lab fixture. Its purpose is to tell you where the synthetic AUT begins adding latency so you can interpret the browser results.
4. Run the browser capacity experiment
Use the Selenium measurement harness from Lesson 2, but execute three candidate worker levels. Start with 1 and 2. Add a third level only within your local resource budget; the sample below uses 3. If your host cannot safely launch three isolated browsers, use a local Grid with enough isolated capacity or use the HTTP control as the saturation demonstration and document the browser ceiling as 2.
import json
from pathlib import Path
from selenium_measure import run_suite
runs = []
for workers in (1, 2, 3):
result = run_suite(workers, capture=True)
runs.append(result)
Path(f"run-{workers}.json").write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps({"runs": runs}, indent=2))
Do not change browser version, case inventory, fixture data, or screenshot policy between those three runs. Otherwise you no longer have a one-variable concurrency experiment.
5. Collect independent evidence
Your evidence packet must include:
- Selenium/binding/browser/driver versions and each session ID;
- host OS, logical CPUs, memory class, and browser process observations;
- one-, two-, and three-worker wall times plus startup/execution/artifact distributions;
-
fixture
/healthstate before/after, including peak active requests; -
Grid
maxSession,sessionCount, queue size, and Node/slot evidence if RemoteWebDriver is used; - failure/retry count (expected zero);
- artifact byte count and cleanup verification.
Grid evidence belongs in the packet only when Grid actually participates. Local WebDriver runs should not invent a Grid layer.
6. Identify the first saturation signal and choose the ceiling
Choose the highest concurrency that still meets your reliability and latency criteria before the first clear knee. A simple lab rule can be:
- zero failures/retries;
- no growing Grid queue at steady state;
- no orphan browsers;
- p95 execution latency no more than 1.5× the one-worker baseline;
- host memory/CPU remains below the point where browser/session instability appears.
Do not treat 1.5× as a universal production SLO; it is only a transparent checkpoint rule. Your real threshold should come from the team’s feedback-time and reliability objectives.
7. Quantify one optimization without changing semantics
Repeat the selected safe worker count with failure-only screenshot policy. Business assertions, browser lifecycle, data, and concurrency remain unchanged. Report:
artifact_saving_ms = baseline_artifact_ms -
optimized_artifact_ms
and
feedback_delta_% = (baseline_wall - optimized_wall) /
baseline_wall × 100.
Keep a small always-on manifest plus complete failure evidence. The optimization is “write less redundant success evidence,” not “erase diagnostic evidence.”
8. Capacity report template
The following table organizes the key choices and evidence for Capacity report template. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Field | Run A | Run B | Run C | Optimized |
|---|---|---|---|---|
| workers | 1 | 2 | 3 | safe target |
| wall time | measure | measure | measure | measure |
| session-start p50/p95 | measure | measure | measure | measure |
| execution p50/p95 | measure | measure | measure | measure |
| Grid queue peak | N/A or measure | N/A or measure | N/A or measure | N/A or measure |
| host/browser pressure | record | record | record | record |
| failures / retries | 0 expected | 0 expected | 0 expected | 0 expected |
| artifact bytes / write ms | measure | measure | measure | reduced policy |
| decision | baseline | candidate | saturation? | adopt/reject |
9. What belongs in JMeter or another performance tool instead
Close the report with a scope statement. Examples that belong outside Selenium:
- maximum sustainable API requests/second;
- p95/p99 service latency under hundreds or thousands of virtual users;
- backend throughput, database saturation, or autoscaling thresholds under controlled load;
- long-duration soak/stress tests;
- protocol-level workload models where browser rendering is irrelevant.
A sensible combined strategy is: use JMeter/k6/etc. for controlled service load; use a small Selenium smoke journey before/during/after in a separately designed environment if you need end-user functional confirmation. Do not make Selenium generate the load.
10. Verification and cleanup checklist
- Three controlled runs exist and use the same six cases.
- Predictions are written before results and compared afterward.
- Safe concurrency is justified by measured evidence, not CPU count alone.
- One optimization has a quantified delta and does not weaken assertions/isolation.
- No production target, real account, public Grid, or security bypass is used.
- Every WebDriver session is quit; temporary profiles and evidence beyond the retained packet are removed.
- Fixture server is stopped and browser process count returns to baseline.
- The report names questions that require a purpose-built load/performance tool.
Knowledge checks
Answer from the operating model, then reveal the explanation.
Four workers reduce wall time by only 2% versus two workers while p95 doubles. Which worker count is the better ceiling?
Two is the stronger safe target unless other evidence justifies otherwise; four is already past the useful concurrency knee.
Why must the optimization run keep browser lifecycle and case inventory unchanged?
Otherwise the measured delta mixes multiple semantic/performance changes and cannot be attributed to the evidence policy.
The Grid queue remains empty but fixture peak active and p95 rise sharply. What saturated first?
The AUT/test-work layer, not Grid scheduling capacity.
What is the strongest evidence that cleanup failed?
Browser/driver processes, temporary profiles, Grid sessions, or other owned state remain after the run instead of returning to baseline.
Which question should be handed to JMeter/k6 rather than answered with more Selenium workers?
A question about controlled application/server throughput or latency under large virtual-user/request load.
Summary and next bridge
- A capacity report needs repeated controlled runs, context, and explicit predictions.
- The safe worker ceiling is selected before saturation, not at the theoretical maximum.
- Optimization deltas are meaningful only when test semantics and evidence requirements remain controlled.
- Browser/Grid/AUT/resource signals must be interpreted together.
- Load/performance questions beyond the browser-test platform belong to purpose-built tools.
Chapter 26 moves to Selenium IDE, record/playback, and migration from captured interactions to maintainable code. The performance baseline from this chapter helps distinguish a maintainable migration from one that merely replays steps faster.
Primary references and version notes
- Selenium downloads — stable client and Grid release baseline.
- Getting started with Selenium Grid — CPU/RAM sizing and session-capacity guidance.
-
Grid CLI options
— current
max-sessions, queue timeout, retry interval, and related controls. - Grid GraphQL support — observable max sessions, session count, nodes, slots, and queue size.
- Grid observability — traces, events, and logs for distributed diagnosis.
- Performance testing with Selenium is discouraged — why WebDriver suite timing is not a substitute for load/performance tooling.
The mandatory examples pin selenium==4.47.0 and
Python 3.10+. Selenium Server/Grid 4.47.0 is the stable Grid
baseline. Selenium Manager remains the normal local
driver-resolution path. Hardware, browser versions, container
images, Grid slot counts, runner sizes, AUT capacity, and network
conditions are recorded as benchmark context rather than assumed
constants.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.