Assertions, Failure Semantics, Retries, and Flaky-Test Control: Guided Hands-On Workflow
Now you will observe the difference between deterministic assertion failure and timing flakiness. A local fixture varies only its asynchronous readiness delay across otherwise equivalent loads. The workflow preserves attempt-1 evidence, performs one narrowly scoped retry experiment, and records that pass-after-retry is a distinct state.
Learning objectives
- Create deterministic pass/fail assertions against a disposable loopback AUT.
- Generate and measure a controlled timing race without fixed sleeps.
- Capture first-attempt screenshot, page source, session provenance, and exception before any retry.
- Run one bounded retry experiment using materially equivalent inputs and keep both attempts.
- Build a small evidence-backed failure taxonomy and flake counter.
1. Create the disposable lab
The following example makes the Create the disposable lab behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
mkdir selenium-ch15-workflow
cd selenium-ch15-workflow
python -m venv .venv
# PowerShell: .\.venv\Scripts\Activate.ps1
# POSIX: source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "selenium==4.47.0"
mkdir evidence
2. app_fixture.py — reproducible variable readiness
The following example makes the app_fixture.py — reproducible variable readiness behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from contextlib import contextmanager
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from urllib.parse import urlparse
DELAYS_MS = [70, 310, 110, 360, 85, 270, 95, 330, 75, 290, 105, 350]
class Handler(BaseHTTPRequestHandler):
request_index = 0
def log_message(self, fmt, *args):
return
def do_GET(self):
path = urlparse(self.path).path
if path == "/health":
raw = b'{"ok":true}'
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.end_headers(); self.wfile.write(raw); return
if path != "/":
self.send_response(404); self.end_headers(); return
i = type(self).request_index
type(self).request_index += 1
delay = DELAYS_MS[i % len(DELAYS_MS)]
page = f"""<!doctype html><html lang='en'><head><meta charset='utf-8'>
<title>Flake Control Lab</title><style>body{{font-family:system-ui,sans-serif;max-width:760px;margin:2rem auto}}#result{{font-weight:700}}</style></head>
<body><h1>Flake Control Lab</h1><p data-testid='attempt'>fixture-load:{i + 1}</p>
<button data-testid='start'>Start calculation</button><p id='result' data-testid='result'>idle</p>
<p data-testid='delay'>configured-delay-ms:{delay}</p>
<script>
const button=document.querySelector('[data-testid=start]');
const result=document.querySelector('[data-testid=result]');
button.addEventListener('click',()=>{{
result.textContent='working';
setTimeout(()=>{{ result.textContent='ready:42'; }}, {delay});
}});
</script></body></html>"""
raw = page.encode()
self.send_response(200)
self.send_header("Content-Type", "text/html; charset=utf-8")
self.send_header("Cache-Control", "no-store")
self.send_header("Content-Length", str(len(raw)))
self.end_headers(); self.wfile.write(raw)
@contextmanager
def running_app(start_index=0):
Handler.request_index = start_index
server = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
thread = Thread(target=server.serve_forever, daemon=True)
thread.start()
host, port = server.server_address
try:
yield f"http://{host}:{port}/"
finally:
server.shutdown(); server.server_close(); thread.join(timeout=2)
The delay sequence is deterministic and resets when the server fixture starts. The browser test does not choose a different business case: the same page and expected result are used. Only application readiness timing varies, simulating a race in a way that can be replayed.
3. evidence.py — preserve every attempt
The following example makes the evidence.py — preserve every attempt behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from pathlib import Path
import json, time
class Evidence:
def __init__(self, root="evidence"):
self.root = Path(root)
self.root.mkdir(parents=True, exist_ok=True)
def capture(self, driver, attempt, classification, exc=None):
folder = self.root / f"attempt-{attempt:02d}"
folder.mkdir(parents=True, exist_ok=True)
record = {
"attempt": attempt,
"classification": classification,
"session_id": driver.session_id,
"browser": driver.capabilities.get("browserName"),
"browser_version": driver.capabilities.get("browserVersion"),
"url": driver.current_url,
"title": driver.title,
"exception_type": type(exc).__name__ if exc else None,
"exception_message": str(exc)[:500] if exc else None,
"captured_at_monotonic": time.monotonic(),
}
(folder / "result.json").write_text(json.dumps(record, indent=2), encoding="utf-8")
(folder / "page.html").write_text(driver.page_source, encoding="utf-8")
driver.save_screenshot(str(folder / "page.png"))
return record
The evidence is synthetic and safe to retain. Capture happens before teardown and before another attempt. Session IDs are diagnostic identifiers, not credentials.
4. Deterministic assertion: wait for the right state, then assert meaning
The following example makes the Deterministic assertion: wait for the right state, then assert meaning behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
def clean_run(base_url):
driver = webdriver.Chrome()
try:
driver.get(base_url)
driver.find_element(By.CSS_SELECTOR, "[data-testid='start']").click()
observed = WebDriverWait(driver, 1.5, poll_frequency=0.02).until(
lambda d: d.find_element(By.CSS_SELECTOR, "[data-testid='result']").text
if d.find_element(By.CSS_SELECTOR, "[data-testid='result']").text.startswith("ready:")
else False
)
assert observed == "ready:42", f"stable result mismatch: {observed!r}"
return observed
finally:
driver.quit()
The wait asks for a domain-relevant terminal prefix, then the assertion checks the exact result. The timeout is a bounded safety budget, not the business contract.
5. Intentionally flaky version: a wrong timeout budget
The following example makes the Intentionally flaky version: a wrong timeout budget behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium.common.exceptions import TimeoutException
def flaky_attempt(base_url, attempt, evidence):
driver = webdriver.Chrome()
try:
driver.get(base_url)
driver.find_element(By.CSS_SELECTOR, "[data-testid='start']").click()
try:
observed = WebDriverWait(driver, 0.18, poll_frequency=0.02).until(
lambda d: d.find_element(By.CSS_SELECTOR, "[data-testid='result']").text == "ready:42"
)
evidence.capture(driver, attempt, "first-pass" if attempt == 1 else "retry-pass")
return True, observed
except TimeoutException as exc:
evidence.capture(driver, attempt, "synchronization-suspect", exc)
return False, exc
finally:
driver.quit()
There is no fixed sleep. Some fixture loads finish within 180 ms;
others intentionally do not. Because Selenium Python 4.47.0 defaults
to a 0.5 s wait poll interval, this deliberately short experiment
sets poll_frequency=0.02 explicitly so the 180 ms
deadline is actually observed at useful resolution. The test
therefore alternates between pass and timeout even though the
expected business outcome is always ready:42.
6. One retry as an experiment—not a pass eraser
The following example makes the One retry as an experiment—not a pass eraser behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from app_fixture import running_app
from evidence import Evidence
def run_experiment():
evidence = Evidence("evidence/retry-experiment")
with running_app(start_index=1) as base_url:
first_ok, first = flaky_attempt(base_url, 1, evidence)
if first_ok:
return {"disposition": "clean-first-pass", "attempts": 1}
# Exactly one additional observation with the same code/data/environment.
second_ok, second = flaky_attempt(base_url, 2, evidence)
return {
"disposition": "pass-after-retry" if second_ok else "reproduced-failure",
"attempts": 2,
"first_exception": type(first).__name__,
"second_exception": None if second_ok else type(second).__name__,
}
if __name__ == "__main__":
print(run_experiment())
The retry experiment starts the deterministic fixture at index 1, so
attempt 1 receives 310 ms and should time out under the 180 ms
budget; attempt 2 receives 110 ms and should pass. Preserve both
evidence directories. This creates a reproducible
fail → pass observation without changing the business
case, test data, browser policy, or expected result.
7. Small failure taxonomy with evidence before classification
The following example makes the Small failure taxonomy with evidence before classification behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium.common.exceptions import TimeoutException, NoSuchElementException, WebDriverException
def initial_classification(exc):
if isinstance(exc, TimeoutException):
return "synchronization-or-aut-readiness"
if isinstance(exc, NoSuchElementException):
return "locator-context-or-timing"
if isinstance(exc, WebDriverException):
return "webdriver-browser-grid-or-host"
if isinstance(exc, AssertionError):
return "product-expectation-or-test-data"
return "unknown-investigate"
# This is triage routing, not final root cause.
Record the initial category, then refine it after examining the current DOM, expected state, browser/session identity, AUT evidence, and Grid/CI logs if remote.
8. flake_counter.py — count outcome histories without hiding them
The following example makes the flake_counter.py — count outcome histories without hiding them behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from collections import Counter
histories = [
("pass",),
("timeout", "pass"),
("pass",),
("timeout", "timeout"),
("pass",),
]
counts = Counter()
for h in histories:
if h == ("pass",):
counts["clean_first_pass"] += 1
elif h[-1] == "pass":
counts["pass_after_retry"] += 1
else:
counts["reproduced_failure"] += 1
print(dict(counts))
print("controlled executions:", len(histories))
Do not report only “eventual pass rate.” The first-pass and pass-after-retry counts answer different operational questions.
9. What each action reads or changes
The following table organizes the key choices and evidence for What each action reads or changes. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Operation | State affected | Evidence/meaning |
|---|---|---|
| Start loopback app | ephemeral server + request counter | base URL; deterministic timing sequence |
| Create WebDriver | new browser session/profile | session ID + returned capabilities |
| Click Start | DOM event + AUT client state | result changes idle → working |
| WebDriverWait | reads DOM repeatedly; no AUT mutation | condition/timeout + observed text |
| Assertion | test-runner result only | expected vs observed |
| Evidence capture | writes synthetic artifacts | attempt identity retained before teardown |
| Retry experiment | new browser observation | does not mutate attempt-1 record |
| quit() | browser/session state | deterministic cleanup |
10. Challenge: choose the control, not the sequence
The result element appears immediately with text
working, then later becomes ready:42.
Which control is correct: presence wait, visibility wait, fixed
sleep, or wait for exact terminal text? Explain why. Then change the
fixture so it becomes error:backend on one load. Your
wait should surface that terminal error quickly instead of waiting
the full timeout.
11. Cleanup
The following example makes the Cleanup behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
# Ensure no browser from the lab remains open.
cd ..
# Preserve synthetic evidence only if you want to review it.
# Then remove the disposable lab directory.
# PowerShell: Remove-Item -Recurse -Force selenium-ch15-workflow
# POSIX: rm -rf selenium-ch15-workflow
12. Summary and bridge
You have separated clean first-pass, timing timeout, retry-pass, and reproduced failure while retaining attempt evidence. Lesson 3 turns those mechanics into design choices for hard/assert-all behavior, retry scope, quarantine policy, and stability thresholds.
Knowledge check
Why is the 180 ms timeout intentionally wrong?
It is shorter than some controlled AUT readiness delays, so it creates an observable race without fixed sleeps and demonstrates why timeout number alone is not the synchronization contract.
What must happen before retry attempt 2?
Attempt-1 evidence must already be persisted, including exception, session/browser provenance, DOM/page source, and screenshot where safe.
Why create a fresh browser for each retry attempt?
It avoids carrying transient browser state from the failed attempt and keeps the additional observation easier to interpret.
Should a pass-after-retry be counted as a clean pass?
No. It is a distinct outcome that signals instability even though the later attempt passed.
What would make this retry experiment invalid?
Changing data, environment, expected result, browser policy, or test code between attempts without recording that change would make it a different experiment.
Official references and version notes
- Selenium 4.47 release notes — stable binding/Grid baseline pinned for this chapter.
- Selenium downloads — current stable Selenium client and Server/Grid versions.
- Waiting strategies — race conditions between test and application readiness are a primary source of flaky browser tests.
- Overview of Test Automation — browser tests should keep setup, actions, and evaluation compact and intentional.
- Avoid sharing state — isolate data and create a new WebDriver instance per test where practical.
- Fresh browser per test — begin from a clean, known browser state.
- Test independency — scenarios should not depend on another test's success or state.
- Python unittest — hard assertion semantics, subtests, setup/cleanup, and failure reporting used in the mandatory path.
Version-sensitive behavior was rechecked against Selenium primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python 4.47.0 and Python 3.10+, use Python standard-library
unittest, a supported locally installed
Chromium-family browser with Selenium Manager, and only loopback
synthetic applications. Retry/quarantine policy is deliberately
modeled as test-governance logic rather than a Selenium
capability. The mandatory path deliberately avoids automatic retry
plugins so attempt semantics stay visible. A diagnostic retry is
shown only as an explicitly coded experiment with fresh-session
isolation and immutable first-attempt evidence. Quarantine
thresholds/counts in examples are illustrative governance values,
not Selenium defaults.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.