Checkpoint Lab — Parallel Execution, Isolation, Concurrency, and Test Sharding
The checkpoint combines the chapter into one operational decision: run a six-test suite with deterministic shards, fresh sessions and isolated state, measure one-versus-two-worker behavior, inject a shared-identity race, and write down a concurrency ceiling that is justified by evidence rather than habit.
Learning objectives
- Execute six synthetic cases with deterministic two-shard ownership.
- Prove per-test isolation through distinct session IDs, identities, download directories, and evidence paths.
- Measure serial/one-worker and two-worker wall time and report observed speedup.
- Inject a deterministic shared-state race while keeping browser sessions independent.
- Document a safe concurrency ceiling and the evidence required before increasing it.
1. Checkpoint operating model
The following diagram visualizes the relationships described in Checkpoint operating model. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.
flowchart TD C[6 test IDs] --> H[Stable 2-shard mapping] H --> R1[Run with 1 worker] H --> R2[Run with 2 workers] R1 --> E[Timing + session evidence] R2 --> E X[Injected shared identity] --> F[Deterministic collision] F --> D[Diagnosis + isolation rule] E --> L[Candidate safe ceiling]
The checkpoint does not promise that two workers are faster on every workstation. It requires you to measure the result and explain it. Browser startup is deliberately part of the cost because production suites pay that cost unless their framework consciously chooses another lifecycle.
2. Setup and preflight
Reuse server.py from Lesson 2 and start it on loopback.
Create a fresh virtual environment and install exactly Selenium
4.47.0. Confirm /health, record CPU count, and ensure
no previous lab browsers remain. If you use Grid instead of local
Chrome, record Grid 4.47.0 status/slots and start with a worker
count no higher than the intended available slots.
python -m venv .venv
# activate the virtual environment
python -m pip install selenium==4.47.0
python server.py
# In another terminal, verify http://127.0.0.1:8765/health
3. Write predictions before execution
Record at least these predictions in a text file before running:
- Each of the six cases will receive a stable shard number that does not depend on completion order.
- Each browser execution will have a distinct Selenium session ID and distinct synthetic identity.
- The two-worker run may reduce wall time, but the amount is environment-dependent.
- The deliberately shared identity race will produce one successful claim and one collision while both browsers remain independent.
4. Run the six-test checkpoint
Save as checkpoint.py. It writes immutable evidence
under a random run ID. The candidate safe ceiling starts at two only
for this small lab; it is not a universal Selenium recommendation.
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from tempfile import TemporaryDirectory
from threading import Barrier
from time import perf_counter
import hashlib
import json
import os
import uuid
from selenium import webdriver
from selenium.webdriver.common.by import By
BASE_URL = "http://127.0.0.1:8765"
CASES = [f"case-{n}" for n in range(1, 7)]
def shard_for(test_id, shard_count=2):
d = hashlib.sha256(test_id.encode()).digest()
return int.from_bytes(d[:8], "big") % shard_count
def isolated_case(case_id, root):
identity = f"{case_id}@example.test"
case_dir = root / case_id
case_dir.mkdir(parents=True, exist_ok=False)
with TemporaryDirectory(prefix=f"downloads-{case_id}-") as download_dir:
options = webdriver.ChromeOptions()
options.add_experimental_option("prefs", {"download.default_directory": str(Path(download_dir).resolve())})
driver = webdriver.Chrome(options=options)
start = perf_counter()
try:
driver.get(f"{BASE_URL}/work?case={case_id}&identity={identity}&delay_ms=300")
status = driver.find_element(By.CSS_SELECTOR, "[data-testid='status']").text
observed_identity = driver.find_element(By.CSS_SELECTOR, "[data-testid='identity']").text
assert status == "done"
assert observed_identity == identity
row = {
"case": case_id,
"shard": shard_for(case_id),
"session_id": driver.session_id,
"identity": identity,
"duration_s": round(perf_counter() - start, 3),
"browser": driver.capabilities.get("browserName"),
"browserVersion": driver.capabilities.get("browserVersion"),
}
(case_dir / "session.json").write_text(json.dumps(row, indent=2), encoding="utf-8")
driver.save_screenshot(str(case_dir / "viewport.png"))
return row
finally:
driver.quit()
def run_suite(workers, root):
start = perf_counter()
rows = []
if workers == 1:
for case in CASES:
rows.append(isolated_case(case, root))
else:
with ThreadPoolExecutor(max_workers=workers, thread_name_prefix="checkpoint") as pool:
future_map = {pool.submit(isolated_case, case, root): case for case in CASES}
for future in as_completed(future_map):
rows.append(future.result())
return rows, perf_counter() - start
def inject_shared_identity_race():
barrier = Barrier(2)
def one(label):
driver = webdriver.Chrome()
try:
barrier.wait(timeout=10)
driver.get(f"{BASE_URL}/claim?identity=shared@example.test&hold_ms=450")
return label, driver.find_element(By.CSS_SELECTOR, "[data-testid='status']").text
finally:
driver.quit()
with ThreadPoolExecutor(max_workers=2) as pool:
return list(pool.map(one, ["A", "B"]))
if __name__ == "__main__":
run_id = uuid.uuid4().hex[:10]
root = Path("checkpoint-evidence") / run_id
root.mkdir(parents=True)
print("shards", {c: shard_for(c) for c in CASES})
rows_1, t1 = run_suite(1, root / "workers-1")
rows_2, t2 = run_suite(2, root / "workers-2")
race = inject_shared_identity_race()
summary = {
"run_id": run_id,
"cpu_count": os.cpu_count(),
"workers_1_s": round(t1, 3),
"workers_2_s": round(t2, 3),
"observed_speedup_2_vs_1": round(t1 / t2, 3) if t2 else None,
"workers_1_sessions": [x["session_id"] for x in rows_1],
"workers_2_sessions": [x["session_id"] for x in rows_2],
"race_result": race,
"candidate_safe_ceiling": 2,
"ceiling_rationale": "Start at 2 for this lab; raise only after CPU/RAM/Grid/AUT evidence remains healthy and wall time improves.",
}
(root / "summary.json").write_text(json.dumps(summary, indent=2), encoding="utf-8")
print(json.dumps(summary, indent=2))
5. Execute and inspect the evidence packet
The following example makes the Execute and inspect the evidence packet behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
python checkpoint.py
# Inspect checkpoint-evidence/<run-id>/summary.json
# Inspect workers-1/<case>/session.json and screenshots
# Inspect workers-2/<case>/session.json and screenshots
Verify that each run contains six unique session IDs. Verify that
the two shard assignments cover the six case IDs without overlap or
omission. The observed_speedup_2_vs_1 value is
evidence, not a pass criterion by itself.
| Evidence | Question it answers |
|---|---|
summary.json |
What worker counts, timings, speedup, CPU count, and race result were observed? |
Per-case session.json |
Did each test own a distinct session/identity and what browser version ran? |
| Per-case screenshot | What visible state did that isolated browser observe? |
| Stable shard map | Which shard owns each test independently of finish order? |
| Race result | Can shared AUT identity still collide even with separate browsers? |
6. Diagnose the injected race instead of retrying it
The expected race result contains one claimed and one
collision. That is not a Selenium synchronization
failure. The first correction is to generate a distinct identity per
test. Repeating the same shared identity with retries would only
make the business-state race intermittent.
# Stable repair pattern: identity belongs to the test, not to the suite.
def identity_for(test_id: str) -> str:
return f"{test_id}@example.test"
7. Document a safe concurrency ceiling
Start by documenting 2 workers as the checkpoint candidate because that is the measured configuration. Raise it only after another controlled run shows useful wall-time improvement with healthy browser startup, no Grid queue timeout, no AUT collision/rate-limit increase, no browser resource failures, and no artifact loss. If two workers are slower or unstable, the safe ceiling may be one on that machine.
Write the machine/runner identity, Selenium/browser/Grid versions, worker count, wall time, and observed bottleneck next to the ceiling. “We always use 8 workers” is not a capacity argument.
8. Optional Grid translation
For Grid, replace local driver creation with
webdriver.Remote() but keep the same per-test ownership
and evidence namespaces. Capture Grid /status or
GraphQL before/during the run. If two runner workers create two
active sessions and a third request queues, the queue/slot state
becomes part of your concurrency evidence.
9. Verification checklist
- Exactly six case IDs are present in each measured suite run.
- Every case owns a distinct WebDriver session and synthetic identity.
- No parallel case shares an evidence directory or download directory.
- The deterministic shard function covers every test exactly once for the selected shard count.
- One-worker and two-worker wall times are recorded, not guessed.
- The injected race is classified as shared AUT/test-data state, not browser-session failure.
- The concurrency ceiling is justified by measured environment evidence and can be revised.
10. Cleanup and bridge to Chapter 22
Stop the loopback fixture, remove only
checkpoint-evidence/ and the disposable virtual
environment if desired, and verify all browser sessions quit.
Chapter 21 adds a capacity/ownership model to the production
browser-automation platform. Chapter 22 moves that model into GitHub
Actions, GitLab CI, and Jenkins, where shards, runners, secrets,
artifacts, and cancellation behavior become CI/CD concerns.
Knowledge check
What must remain true when moving from one worker to two?
Each test still owns a fresh WebDriver session and isolated data/download/evidence state; only scheduling concurrency changes.
Why is the shared-identity race useful in the checkpoint?
It proves that separate browser sessions do not automatically isolate application/test-data state.
Is a measured speedup below 1 automatically a failed checkpoint?
No. It is valid capacity evidence; the operator should explain startup/contention and may keep a lower concurrency ceiling.
What should happen before raising the documented ceiling above two?
Run another controlled measurement and verify useful wall-time improvement plus healthy browser/Grid/AUT/resource/artifact behavior.
A third Grid request queues while two slots are busy. Which layer should be investigated before changing Selenium waits?
Grid capacity/runner demand. Explicit waits control browser/AUT synchronization, not slot availability.
Official references and current-version notes
- Selenium 4.47 release notes
- Selenium downloads — current stable client and Grid versions
- Avoid sharing state — Selenium test practices
- Test independency — Selenium test practices
- Getting started with Selenium Grid — capacity guidance
- Grid CLI options — max sessions and queue controls
- Python concurrent.futures — ThreadPoolExecutor
The mandatory examples pin Selenium Python to 4.47.0.
Selenium Server/Grid 4.47.0 is the matching stable
Grid baseline. Python examples use the standard-library
concurrent.futures module rather than a third-party
parallel-test plugin so worker ownership is visible.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.