Checkpoint Lab — Assertions, Failure Semantics, Retries, and Flaky-Test Control
The checkpoint treats flakiness as measurable evidence, not folklore. You will run a deliberately unstable test repeatedly against a deterministic timing fixture, preserve the first failure, classify the cause, replace the wrong synchronization contract, prove stability without blanket retries, and write a governance rule for retry/quarantine.
Learning objectives
- Measure repeated first-pass outcomes under equivalent controlled inputs and record a flake history.
- Preserve the first failing attempt before any additional observation.
- Classify the root cause from DOM/session/timing evidence rather than exception name alone.
- Stabilize the test with a domain-specific terminal-state wait and prove improvement across repeated runs.
- Write a bounded retry/quarantine rule with ownership, evidence retention, and an exit condition.
1. Checkpoint scenario and invariants
The same synthetic calculation always ends at ready:42,
but readiness delay varies by a fixed replayable sequence. The
broken test uses a too-short wait and therefore flakes. Your job is
to prove the flake, retain attempt evidence, fix the synchronization
contract, and show that the corrected test is first-pass stable over
the same timing sequence.
- No production or external site.
- No account, credential, or personal profile.
- No fixed sleeps.
- No blanket retry in the corrected suite.
- No JavaScript interaction bypass.
- Every attempt uses a fresh WebDriver session.
2. Preflight and exact assumptions
The following example makes the Preflight and exact assumptions behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
mkdir selenium-ch15-checkpoint
cd selenium-ch15-checkpoint
python -m venv .venv
# PowerShell: .\.venv\Scripts\Activate.ps1
# POSIX: source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "selenium==4.47.0"
python -c "import selenium; print(selenium.__version__)"
Assumptions: Python 3.10+, Selenium Python 4.47.0, a supported locally installed Chromium-family browser, Selenium Manager for normal driver resolution, and enough local resources for one browser at a time. Grid/BiDi are not required; if you repeat the experiment remotely later, retain Grid queue/node/session evidence separately.
3. app_fixture.py — use the same controlled timing fixture
The following example makes the app_fixture.py — use the same controlled timing fixture behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from contextlib import contextmanager
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from urllib.parse import urlparse
DELAYS_MS = [70, 310, 110, 360, 85, 270, 95, 330, 75, 290, 105, 350]
class Handler(BaseHTTPRequestHandler):
request_index = 0
def log_message(self, fmt, *args):
return
def do_GET(self):
path = urlparse(self.path).path
if path == "/health":
raw = b'{"ok":true}'
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.end_headers(); self.wfile.write(raw); return
if path != "/":
self.send_response(404); self.end_headers(); return
i = type(self).request_index
type(self).request_index += 1
delay = DELAYS_MS[i % len(DELAYS_MS)]
page = f"""<!doctype html><html lang='en'><head><meta charset='utf-8'>
<title>Flake Control Lab</title><style>body{{font-family:system-ui,sans-serif;max-width:760px;margin:2rem auto}}#result{{font-weight:700}}</style></head>
<body><h1>Flake Control Lab</h1><p data-testid='attempt'>fixture-load:{i + 1}</p>
<button data-testid='start'>Start calculation</button><p id='result' data-testid='result'>idle</p>
<p data-testid='delay'>configured-delay-ms:{delay}</p>
<script>
const button=document.querySelector('[data-testid=start]');
const result=document.querySelector('[data-testid=result]');
button.addEventListener('click',()=>{{
result.textContent='working';
setTimeout(()=>{{ result.textContent='ready:42'; }}, {delay});
}});
</script></body></html>"""
raw = page.encode()
self.send_response(200)
self.send_header("Content-Type", "text/html; charset=utf-8")
self.send_header("Cache-Control", "no-store")
self.send_header("Content-Length", str(len(raw)))
self.end_headers(); self.wfile.write(raw)
@contextmanager
def running_app(start_index=0):
Handler.request_index = start_index
server = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
thread = Thread(target=server.serve_forever, daemon=True)
thread.start()
host, port = server.server_address
try:
yield f"http://{host}:{port}/"
finally:
server.shutdown(); server.server_close(); thread.join(timeout=2)
4. evidence.py — immutable attempt evidence
The following example makes the evidence.py — immutable attempt evidence behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from pathlib import Path
import json, time
class Evidence:
def __init__(self, root="evidence"):
self.root = Path(root)
self.root.mkdir(parents=True, exist_ok=True)
def capture(self, driver, attempt, classification, exc=None):
folder = self.root / f"attempt-{attempt:02d}"
folder.mkdir(parents=True, exist_ok=True)
record = {
"attempt": attempt,
"classification": classification,
"session_id": driver.session_id,
"browser": driver.capabilities.get("browserName"),
"browser_version": driver.capabilities.get("browserVersion"),
"url": driver.current_url,
"title": driver.title,
"exception_type": type(exc).__name__ if exc else None,
"exception_message": str(exc)[:500] if exc else None,
"captured_at_monotonic": time.monotonic(),
}
(folder / "result.json").write_text(json.dumps(record, indent=2), encoding="utf-8")
(folder / "page.html").write_text(driver.page_source, encoding="utf-8")
driver.save_screenshot(str(folder / "page.png"))
return record
5. Predict before running
Write these predictions in predictions.txt before
executing:
- With the 180 ms broken wait, loads configured below 180 ms should usually pass while longer loads should time out; exact browser scheduling can shift boundaries, but the short budget should produce mixed first-pass outcomes.
- Every attempt creates a new session ID, but the same URL path and business expectation remain.
- The corrected terminal-state wait with a 1.5 s budget should pass all delays in the fixture sequence because the maximum configured delay is 360 ms.
-
Evidence from the first timeout will show
workingrather than a stable wrong result, supporting a synchronization root cause.
6. checkpoint.py — measure the broken test, then the corrected test
The following example makes the checkpoint.py — measure the broken test, then the corrected test behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from pathlib import Path
import json
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import TimeoutException
from app_fixture import running_app, DELAYS_MS
from evidence import Evidence
RESULT = (By.CSS_SELECTOR, "[data-testid='result']")
START = (By.CSS_SELECTOR, "[data-testid='start']")
DELAY = (By.CSS_SELECTOR, "[data-testid='delay']")
def run_one(base_url, attempt, timeout, evidence_root):
driver = webdriver.Chrome()
evidence = Evidence(evidence_root)
try:
driver.get(base_url)
configured_delay = driver.find_element(*DELAY).text
driver.find_element(*START).click()
try:
observed = WebDriverWait(driver, timeout, poll_frequency=0.02).until(
lambda d: d.find_element(*RESULT).text
if d.find_element(*RESULT).text.startswith("ready:")
or d.find_element(*RESULT).text.startswith("error:")
else False
)
if observed != "ready:42":
exc = AssertionError(f"stable terminal result={observed!r}")
evidence.capture(driver, attempt, "stable-assertion-failure", exc)
return {"attempt": attempt, "outcome": "assertion-fail", "delay": configured_delay}
evidence.capture(driver, attempt, "clean-first-pass")
return {"attempt": attempt, "outcome": "pass", "delay": configured_delay}
except TimeoutException as exc:
current = driver.find_element(*RESULT).text
evidence.capture(driver, attempt, "readiness-timeout", exc)
return {"attempt": attempt, "outcome": "timeout", "delay": configured_delay, "current": current}
finally:
driver.quit()
def run_series(base_url, timeout, label):
rows = []
for attempt in range(1, len(DELAYS_MS) + 1):
rows.append(run_one(base_url, attempt, timeout, f"evidence/{label}"))
Path(f"{label}.json").write_text(json.dumps(rows, indent=2), encoding="utf-8")
return rows
def summarize(rows):
counts = {}
for row in rows:
counts[row["outcome"]] = counts.get(row["outcome"], 0) + 1
return counts
def main():
with running_app() as base_url:
broken = run_series(base_url, 0.18, "broken")
with running_app() as base_url:
stable = run_series(base_url, 1.5, "stable")
print("broken:", summarize(broken))
print("stable:", summarize(stable))
assert summarize(stable) == {"pass": len(DELAYS_MS)}, summarize(stable)
if __name__ == "__main__":
main()
The corrected test does not add a retry. Both broken and stable waits use an explicit 20 ms poll interval; only the timeout contract changes from “be ready within an arbitrary 180 ms” to “reach a terminal business state within a bounded 1.5 s safety budget that comfortably covers the controlled fixture.”
7. Run and preserve the first failure
The following example makes the Run and preserve the first failure behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
python checkpoint.py
# Review the first evidence/broken/attempt-XX whose result.json says readiness-timeout.
# Do not delete it after the stable series passes.
Expected pattern: the broken series contains both
pass and timeout; the stable series
reports all passes. Exact short-timeout boundary can vary slightly
with host scheduling, so your classification must use the recorded
configured delay plus DOM evidence—not merely the count.
8. Classify the root cause from evidence
Open the first timeout’s result.json and
page.html. If the page is still working,
the session is healthy, the locator exists, and the same scenario
later reaches ready:42 under a suitable wait, the
supported root cause is a
test synchronization defect: the test sampled
before the observable transition was complete.
If instead the terminal state were error:dependency,
increasing the timeout would be the wrong fix. That would be
application/dependency evidence.
9. Compute first-pass stability and retain history
The following example makes the Compute first-pass stability and retain history behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
import json
from pathlib import Path
for label in ("broken", "stable"):
rows = json.loads(Path(f"{label}.json").read_text())
total = len(rows)
passes = sum(r["outcome"] == "pass" for r in rows)
timeouts = sum(r["outcome"] == "timeout" for r in rows)
print(label, {
"controlled_runs": total,
"first_pass_rate": passes / total,
"timeouts": timeouts,
})
Do not call the broken series “eventually passing” by rerunning timed-out cases. Its first-pass rate is the useful signal for this experiment.
10. Write a retry/quarantine rule
The following example makes the Write a retry/quarantine rule behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
retry / quarantine policy for browser-functional gate
1. Default: no automatic blanket retries for product-facing Selenium scenarios.
2. On first failure: preserve attempt evidence before any rerun.
3. Diagnostic rerun: at most one fresh-session rerun may be requested by explicit policy
for a classified transient hypothesis; it must use the same revision, case data,
environment, browser family/version policy, and expected result.
4. Result semantics: fail->pass is "pass-after-retry", never clean first-pass.
5. Non-idempotent actions: no step replay unless AUT idempotency and current state are proven.
6. Quarantine: allowed only for a known unstable test with owner, tracking issue,
bounded scope, continued execution/reporting, and review/exit condition.
7. Artifact retention: attempt-1 evidence is immutable even after later pass.
8. Exit from quarantine: fix root cause, then satisfy an organization-defined controlled
first-pass stability window on the relevant browser/environment matrix.
9. Infrastructure incidents: tracked separately from product assertion failures.
10. Forbidden shortcuts: giant sleeps, hidden exception swallowing, production experiments,
TLS disablement, or changing data/environment between retry attempts.
11. Example quarantine record for the broken version
The following example makes the Example quarantine record for the broken version behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
test_id: flake-control.ready-result
status: quarantine-candidate (teaching example only)
classification: test synchronization defect
first_failure_evidence: evidence/broken/attempt-02/
owner: browser-automation-team
tracking_issue: LAB-CH15-001
blocking: no (only if organizational policy explicitly approves quarantine)
continues_to_run: yes
exit_condition: replace 180ms contract with terminal-state wait and demonstrate
organization-defined clean first-pass stability window
forbidden: blanket retry / artifact deletion / longer fixed sleep
In this checkpoint, you already know the correct fix, so real quarantine is unnecessary. The record exists to practice the governance shape.
12. Verification checklist
- The broken series contains mixed first-pass outcomes under the controlled timing sequence.
- The first timeout evidence remains present after all later executions.
- Every attempt has a fresh WebDriver session ID.
-
The failure classification cites DOM/session/timing evidence, not
only
TimeoutException. - The stable series uses no retry and no fixed sleep.
-
The stable series reaches
ready:42on all fixture delays. - The retry/quarantine rule distinguishes clean pass, pass-after-retry, reproduced failure, quarantine, and infrastructure incident.
- No production target, credential, personal profile, TLS bypass, or external Grid is used.
13. Evidence packet
predictions.txt.broken.jsonandstable.json.- Immutable per-attempt JSON, screenshot, and page source.
- Console summary showing mixed broken outcomes and all-pass stable outcomes.
- The first timeout’s classification note.
- The retry/quarantine policy and example record.
All artifacts come from the synthetic loopback fixture. In real pipelines, apply privacy/security retention rules before collecting comparable browser evidence.
14. Cleanup and rollback
Confirm all WebDriver sessions are closed. Preserve only the
synthetic checkpoint evidence needed for review, then delete
selenium-ch15-checkpoint/. The server uses memory only
and shuts down at context exit; no repository, account, remote Grid,
database, or browser profile is modified.
15. What Chapter 15 adds to the operating model
The automation platform now has explicit result semantics: stable business assertions, immutable first-failure evidence, layered classification, first-pass stability metrics, bounded diagnostic rerun semantics, visible quarantine governance, and capacity/security awareness around retries. A pipeline can distinguish product risk from test or infrastructure unreliability instead of hiding all three behind eventual green.
Chapter 16 expands that trusted signal across browsers, responsive layouts, localization, and compatibility matrices. Cross-browser comparison is meaningful only when each individual test already has deterministic failure semantics.
16. Summary
The checkpoint demonstrated the correct order of operations: measure → preserve → classify → fix → prove. The broken test flaked because its synchronization contract was wrong. The corrected test became stable by waiting for the right observable state, not by retrying until green.
Knowledge check
What proves the checkpoint failure is a synchronization defect?
The first failure retains a healthy session and correct element/context while DOM evidence is still working; the same AUT reaches ready:42 when the wait targets the terminal state with an adequate bounded budget.
Why does the corrected series use zero retries?
The root cause is repaired directly. Retrying would add cost and obscure whether the synchronization contract is actually stable.
What must remain after the stable series passes?
The first broken-attempt evidence and its classification must remain immutable; later success must not erase it.
When is a diagnostic rerun valid under the sample policy?
Only when explicitly allowed for a classified transient hypothesis, in a fresh session, with the same controlled revision/data/environment/browser policy, and with attempt 1 already preserved.
Why is Chapter 15 necessary before cross-browser matrices?
Without explicit first-pass/failure/retry semantics, adding browsers multiplies ambiguous noise rather than producing comparable compatibility evidence.
Official references and version notes
- Selenium 4.47 release notes — stable binding/Grid baseline pinned for this chapter.
- Selenium downloads — current stable Selenium client and Server/Grid versions.
- Waiting strategies — race conditions between test and application readiness are a primary source of flaky browser tests.
- Overview of Test Automation — browser tests should keep setup, actions, and evaluation compact and intentional.
- Avoid sharing state — isolate data and create a new WebDriver instance per test where practical.
- Fresh browser per test — begin from a clean, known browser state.
- Test independency — scenarios should not depend on another test's success or state.
- Python unittest — hard assertion semantics, subtests, setup/cleanup, and failure reporting used in the mandatory path.
Version-sensitive behavior was rechecked against Selenium primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python 4.47.0 and Python 3.10+, use Python standard-library
unittest, a supported locally installed
Chromium-family browser with Selenium Manager, and only loopback
synthetic applications. Retry/quarantine policy is deliberately
modeled as test-governance logic rather than a Selenium
capability. The mandatory path deliberately avoids automatic retry
plugins so attempt semantics stay visible. A diagnostic retry is
shown only as an explicitly coded experiment with fresh-session
isolation and immutable first-attempt evidence. Quarantine
thresholds/counts in examples are illustrative governance values,
not Selenium defaults.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.