Checkpoint Lab — Debugging Stale Elements, Timing Races, Driver Failures, and Grid Incidents
Run a five-incident troubleshooting drill, preserve one evidence packet per incident, write root-cause statements, apply narrow repairs, and leave behind a reusable diagnostic helper and runbook.
Learning objectives
- Execute or simulate five incident classes from one disposable fixture.
- Produce correlated evidence and one root-cause statement per incident.
- Predict at least two state changes before injecting failures.
- Repair each incident without blanket sleeps/retries/restarts.
- Produce a reusable diagnostic helper and runbook decision tree.
1. Checkpoint mission
Operate a five-incident drill covering stale reference, timing race, browser/session creation, invalid session lifecycle, and Grid distribution. For each incident, predict the state change, preserve one evidence packet, classify the layer, write a one-sentence root cause, apply one narrow repair, and define a regression guard.
2. Setup and exact assumptions
- Python 3.10+ and
selenium==4.47.0. - Selenium Manager is the default local driver resolver.
- Supported local Chromium-family browser for live WebDriver portions.
- Loopback fixture only:
127.0.0.1:8788. - Grid is optional; deterministic simulation is mandatory fallback.
- No production URL, real account, public Grid, TLS bypass, or personal profile.
3. Generate and serve the fixture
The following example makes the Generate and serve the fixture behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
# file: make_incident_fixture.py
from pathlib import Path
root = Path("incident-lab")
root.mkdir(exist_ok=True)
(root / "index.html").write_text("""<!doctype html>
<html lang="en"><head><meta charset="utf-8"><title>Incident Lab</title></head>
<body data-generation="1">
<main>
<h1>Incident Lab</h1>
<p id="target" data-testid="target">generation:1</p>
<button id="rerender" data-testid="rerender">Rerender target</button>
<button id="delayed" data-testid="delayed">Create delayed result</button>
<p id="health" data-testid="health">healthy</p>
</main>
<script>
let generation = 1;
document.querySelector('#rerender').addEventListener('click', () => {
generation += 1;
const old = document.querySelector('#target');
const fresh = document.createElement('p');
fresh.id = 'target'; fresh.dataset.testid = 'target';
fresh.textContent = `generation:${generation}`;
old.replaceWith(fresh);
document.body.dataset.generation = String(generation);
});
document.querySelector('#delayed').addEventListener('click', () => {
document.querySelector('#delayed-result')?.remove();
setTimeout(() => {
const p = document.createElement('p');
p.id = 'delayed-result'; p.dataset.testid = 'delayed-result';
p.textContent = 'ready'; document.querySelector('main').appendChild(p);
}, 300);
});
</script></body></html>""", encoding="utf-8")
print(root.resolve())
print("Serve with: python -m http.server 8788 --bind 127.0.0.1 --directory incident-lab")
Verify the page title and health=healthy before
injecting anything. Record the installed Selenium/browser versions.
If the fixture is unhealthy, stop: the baseline environment is
already the incident.
4. Prediction worksheet
The following table organizes the key choices and evidence for Prediction worksheet. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Incident | Prediction before action | Independent verification |
|---|---|---|
| I1 stale | body generation changes 1→2; old reference dies | read body data-generation and fresh target |
| I2 timing | delayed-result absent immediately, then appears | DOM query before + explicit wait after |
| I3 missing browser | no usable session ID returned | exception + missing path exists=False |
| I4 invalid session | session ID becomes unusable after quit | command after quit raises session/WebDriver exception |
| I5 Grid | Firefox request has no matching free UP slot | simulated/live status shows matching Node down/unavailable |
5. Reusable evidence helper
The following example makes the Reusable evidence helper behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from __future__ import annotations
from datetime import datetime, timezone
from pathlib import Path
import json, re, uuid
SECRET_PATTERNS = [re.compile(r"(?i)(token|password|secret)=([^&\s]+)")]
def redact(text: str) -> str:
for pattern in SECRET_PATTERNS:
text = pattern.sub(lambda m: f"{m.group(1)}=[REDACTED]", text)
return text
def capture_failure(driver, root: Path, incident: str, error: BaseException):
packet = root / f"{incident}-{uuid.uuid4().hex[:8]}"
packet.mkdir(parents=True, exist_ok=False)
metadata = {
"incident": incident,
"utc": datetime.now(timezone.utc).isoformat(),
"exception": type(error).__name__,
"message": redact(str(error)),
"session_id": getattr(driver, "session_id", None),
}
try:
metadata["url"] = redact(driver.current_url)
metadata["title"] = driver.title
metadata["capabilities"] = {
k: driver.capabilities.get(k)
for k in ("browserName", "browserVersion", "platformName", "pageLoadStrategy")
}
driver.save_screenshot(str(packet / "failure.png"))
except Exception as collector_error:
metadata["collector_error"] = type(collector_error).__name__
(packet / "metadata.json").write_text(json.dumps(metadata, indent=2), encoding="utf-8")
return packet
Extend the packet in your own environment with test ID, worker/shard, attempt, Grid endpoint ID, CI job URL/reference, and server/network correlation ID. Do not add raw secrets or full personal profiles.
6. Run I1 and I2: DOM/reference and synchronization
The following example makes the Run I1 and I2: DOM/reference and synchronization behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium import webdriver
from selenium.common.exceptions import StaleElementReferenceException
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
BASE_URL = "http://127.0.0.1:8788/"
driver = webdriver.Chrome()
try:
driver.get(BASE_URL)
old_target = driver.find_element(By.CSS_SELECTOR, '[data-testid="target"]')
driver.find_element(By.CSS_SELECTOR, '[data-testid="rerender"]').click()
try:
print(old_target.text) # intentionally uses the dead DOM reference
except StaleElementReferenceException as exc:
print("classified: stale-reference", type(exc).__name__)
fresh = WebDriverWait(driver, 3).until(
EC.text_to_be_present_in_element((By.CSS_SELECTOR, '[data-testid="target"]'), "generation:2")
)
assert fresh is True
finally:
driver.quit()
The following example makes the Run I1 and I2: DOM/reference and synchronization behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
driver = webdriver.Chrome()
try:
driver.get("http://127.0.0.1:8788/")
driver.find_element(By.CSS_SELECTOR, '[data-testid="delayed"]').click()
try:
driver.find_element(By.CSS_SELECTOR, '[data-testid="delayed-result"]')
except NoSuchElementException as exc:
print("classified: too-early-observation", type(exc).__name__)
result = WebDriverWait(driver, 3).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, '[data-testid="delayed-result"]'))
)
assert result.text == "ready"
finally:
driver.quit()
Root-cause examples: I1 — test retained a DOM reference across a known node replacement; repair reacquires after rerender. I2 — test observed the DOM before the AUT readiness condition; repair waits for delayed-result visibility.
7. Run I3 and I4: creation and lifecycle
The following example makes the Run I3 and I4: creation and lifecycle behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from pathlib import Path
from selenium import webdriver
from selenium.common.exceptions import SessionNotCreatedException, WebDriverException
from selenium.webdriver.chrome.options import Options
missing = Path("/definitely/not/a/browser/chrome")
print("configured browser exists:", missing.exists())
options = Options()
options.binary_location = str(missing)
try:
driver = webdriver.Chrome(options=options)
except (SessionNotCreatedException, WebDriverException) as exc:
print("classified: browser/session-start configuration")
print(type(exc).__name__, str(exc).splitlines()[0][:180])
else:
driver.quit()
raise AssertionError("The deliberately missing browser path unexpectedly launched")
The following example makes the Run I3 and I4: creation and lifecycle behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium import webdriver
from selenium.common.exceptions import InvalidSessionIdException, WebDriverException
driver = webdriver.Chrome()
session_id = driver.session_id
driver.quit()
try:
_ = driver.current_url
except (InvalidSessionIdException, WebDriverException) as exc:
print("closed session:", session_id)
print("classified:", type(exc).__name__)
Keep these diagnoses separate. I3 never owns a valid browser session. I4 owns one and then deletes it. Their repairs therefore belong to different layers: browser/session creation configuration versus test lifecycle ownership.
8. Run I5: Grid request distribution
The following example makes the Run I5: Grid request distribution behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
# Deterministic queue/slot explanation when a live Grid is not available.
status = {
"nodes": [
{"id": "node-chrome", "availability": "up", "slots": [
{"stereotype": {"browserName": "chrome"}, "session": None}
]},
{"id": "node-firefox", "availability": "down", "slots": [
{"stereotype": {"browserName": "firefox"}, "session": None}
]},
]
}
requested = {"browserName": "firefox"}
matching_up = [
(node["id"], slot)
for node in status["nodes"] if node["availability"] == "up"
for slot in node["slots"]
if slot["stereotype"].get("browserName") == requested["browserName"] and slot["session"] is None
]
print("request:", requested)
print("matching free slots on UP nodes:", len(matching_up))
print("classification:", "distributable" if matching_up else "queue-unsatisfied / node-capability incident")
If a private local Grid is available, capture
/status beside the simulator and compare Node
availability/stereotypes. Do not deliberately saturate or disable
shared Grid Nodes for this lab.
9. Generate the five root-cause packets
The following example makes the Generate the five root-cause packets behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
# file: checkpoint_drill.py
from dataclasses import dataclass
from pathlib import Path
import json
@dataclass(frozen=True)
class Incident:
id: str
symptom: str
expected_layer: str
repair: str
INCIDENTS = [
Incident("I1", "old element reference after DOM replacement", "DOM/reference", "re-locate after the known rerender boundary"),
Incident("I2", "result queried before readiness condition", "synchronization", "explicit wait for the real ready state"),
Incident("I3", "configured browser binary is missing", "browser/session creation", "repair installation/path; keep Selenium Manager default where possible"),
Incident("I4", "command sent after driver.quit()", "session lifecycle", "fix lifecycle ownership; do not reuse deleted session"),
Incident("I5", "requested stereotype has no free UP node", "Grid distribution/capacity", "restore matching node/slot or request supported capability"),
]
root = Path("incident-evidence")
root.mkdir(exist_ok=True)
for item in INCIDENTS:
packet = root / item.id
packet.mkdir(exist_ok=True)
(packet / "root-cause.json").write_text(json.dumps(item.__dict__, indent=2), encoding="utf-8")
(root / "runbook.txt").write_text(
"preserve evidence -> versions/target -> session/context -> DOM/waits -> AUT/network -> Grid/CI/resources -> narrow repair -> smallest rerun -> regression guard\n",
encoding="utf-8",
)
print("incidents:", len(INCIDENTS))
print("evidence packets:", len([p for p in root.iterdir() if p.is_dir()]))
The deterministic script creates one directory per incident and a shared runbook. Replace the synthetic root-cause files with your actual captured metadata if you ran the live browser portions.
10. Required evidence packet
The following table organizes the key choices and evidence for Required evidence packet. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| File/evidence | Purpose |
|---|---|
| root-cause.json or incident.md | symptom, layer, causal statement, repair |
| metadata.json | UTC, Selenium/browser/session/capabilities where available |
| failure.png | visual browser state when a live session still exists |
| grid-status.json / simulator output | queue/node/slot evidence for I5 |
| test-output.txt | original exception/traceback without secret values |
| runbook.txt | repeatable decision sequence |
11. Verification checklist
- Five incident IDs exist and are not merged.
- Every root-cause statement names the failing layer and causal state.
- No repair is “sleep longer,” “retry everything,” “restart everything,” “disable TLS,” or “download random driver.”
- I1 uses a fresh element after state change; I2 uses an explicit readiness condition.
- I3 and I4 are not conflated.
- I5 includes requested capabilities plus Node/slot evidence.
- Evidence is redacted and disposable.
12. Safe concurrency and CI note
If you repeat the drill in CI, each worker needs its own test ID/evidence directory and independent browser session. Grid queue/capacity evidence must be correlated to the exact attempt. Do not run multiple failure-injection jobs against the same mutable fixture unless the fixture is designed for isolation.
13. Cleanup and rollback
Quit surviving sessions, stop the loopback server, remove the artificial browser binary override, restore any local Grid configuration you changed, and delete disposable profiles/downloads. Retain only the redacted incident packet for the agreed diagnostic window, then remove it.
14. Production operating model and Chapter 29 bridge
Chapter 28 adds a consistent incident operating model: immutable first evidence, layer classification, smallest reproduction, least destructive repair, and regression guard. Chapter 29 turns those mechanics into architecture/governance—coding standards, ownership rules, review gates, exception policies, and suite evolution practices that keep diagnostic quality from degrading as the test estate grows.
Official references and current-version notes
- Selenium downloads — Stable clients and Selenium Server/Grid 4.47.0, released August 10, 2026.
- Understanding Common Errors — Official current guidance for stale element, invalid session ID, session-not-created, and related WebDriver errors.
- Python exceptions API 4.47.0 — Current Python exception classes and stale-reference semantics.
- Selenium Manager — Official default driver/browser management path shipped with Selenium.
- Grid Components — Router, New Session Queue, Distributor, Session Map, Event Bus, Node, and slot behavior.
- Grid architecture — Node heartbeat/status, slots/stereotypes, session routing, and Grid model concepts.
- Grid CLI options — Current component configuration and diagnostic option reference.
Version-sensitive statements in this lesson retain the pinned baseline used when the lesson was authored. Before changing Selenium, browser, driver, Grid, BiDi, container, or framework dependencies, compare that baseline with current primary documentation instead of silently substituting an unverified “latest” environment.
Knowledge checks
What are the five checkpoint incident layers?
DOM/reference, synchronization, browser/session creation, session lifecycle, and Grid distribution/capacity.
Why are I3 and I4 deliberately separate?
I3 never creates a usable session; I4 creates one and later deletes it, so their evidence and repairs differ.
What is the minimum acceptable I5 evidence?
Requested capability plus status/simulator evidence showing no matching free UP slot/Node.
If the evidence collector fails after the test exception, which failure remains primary?
The original test/WebDriver incident; collector failure is secondary evidence.
What does Chapter 28 add to production operation?
A repeatable evidence-first diagnosis workflow that reduces MTTR and turns each repair into a regression guard/runbook improvement.
Summary and next bridge
Debugging is an evidence workflow, not a bag of retries: preserve the first state, classify the failing layer, reproduce narrowly, repair minimally, and encode a regression guard.
Next: Chapter 29 — Test Architecture, Governance, Coding Standards, and Suite Evolution
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.