Chapter 15Lesson 05~245 minutes

Checkpoint Lab — Assertions, Failure Semantics, Retries, and Flaky-Test Control

The checkpoint treats flakiness as measurable evidence, not folklore. You will run a deliberately unstable test repeatedly against a deterministic timing fixture, preserve the first failure, classify the cause, replace the wrong synchronization contract, prove stability without blanket retries, and write a governance rule for retry/quarantine.

Checkpoint labRepeated runsStability proofGovernance policyChapter 16 bridge

Learning objectives

  • Measure repeated first-pass outcomes under equivalent controlled inputs and record a flake history.
  • Preserve the first failing attempt before any additional observation.
  • Classify the root cause from DOM/session/timing evidence rather than exception name alone.
  • Stabilize the test with a domain-specific terminal-state wait and prove improvement across repeated runs.
  • Write a bounded retry/quarantine rule with ownership, evidence retention, and an exit condition.

1. Checkpoint scenario and invariants

The same synthetic calculation always ends at ready:42, but readiness delay varies by a fixed replayable sequence. The broken test uses a too-short wait and therefore flakes. Your job is to prove the flake, retain attempt evidence, fix the synchronization contract, and show that the corrected test is first-pass stable over the same timing sequence.

  • No production or external site.
  • No account, credential, or personal profile.
  • No fixed sleeps.
  • No blanket retry in the corrected suite.
  • No JavaScript interaction bypass.
  • Every attempt uses a fresh WebDriver session.

2. Preflight and exact assumptions

The following example makes the Preflight and exact assumptions behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

mkdir selenium-ch15-checkpoint
cd selenium-ch15-checkpoint
python -m venv .venv
# PowerShell: .\.venv\Scripts\Activate.ps1
# POSIX: source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "selenium==4.47.0"
python -c "import selenium; print(selenium.__version__)"

Assumptions: Python 3.10+, Selenium Python 4.47.0, a supported locally installed Chromium-family browser, Selenium Manager for normal driver resolution, and enough local resources for one browser at a time. Grid/BiDi are not required; if you repeat the experiment remotely later, retain Grid queue/node/session evidence separately.

3. app_fixture.py — use the same controlled timing fixture

The following example makes the app_fixture.py — use the same controlled timing fixture behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from contextlib import contextmanager
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from urllib.parse import urlparse

DELAYS_MS = [70, 310, 110, 360, 85, 270, 95, 330, 75, 290, 105, 350]

class Handler(BaseHTTPRequestHandler):
    request_index = 0

    def log_message(self, fmt, *args):
        return

    def do_GET(self):
        path = urlparse(self.path).path
        if path == "/health":
            raw = b'{"ok":true}'
            self.send_response(200)
            self.send_header("Content-Type", "application/json")
            self.send_header("Content-Length", str(len(raw)))
            self.end_headers(); self.wfile.write(raw); return
        if path != "/":
            self.send_response(404); self.end_headers(); return

        i = type(self).request_index
        type(self).request_index += 1
        delay = DELAYS_MS[i % len(DELAYS_MS)]
        page = f"""<!doctype html><html lang='en'><head><meta charset='utf-8'>
<title>Flake Control Lab</title><style>body{{font-family:system-ui,sans-serif;max-width:760px;margin:2rem auto}}#result{{font-weight:700}}</style></head>
<body><h1>Flake Control Lab</h1><p data-testid='attempt'>fixture-load:{i + 1}</p>
<button data-testid='start'>Start calculation</button><p id='result' data-testid='result'>idle</p>
<p data-testid='delay'>configured-delay-ms:{delay}</p>
<script>
const button=document.querySelector('[data-testid=start]');
const result=document.querySelector('[data-testid=result]');
button.addEventListener('click',()=>{{
  result.textContent='working';
  setTimeout(()=>{{ result.textContent='ready:42'; }}, {delay});
}});
</script></body></html>"""
        raw = page.encode()
        self.send_response(200)
        self.send_header("Content-Type", "text/html; charset=utf-8")
        self.send_header("Cache-Control", "no-store")
        self.send_header("Content-Length", str(len(raw)))
        self.end_headers(); self.wfile.write(raw)

@contextmanager
def running_app(start_index=0):
    Handler.request_index = start_index
    server = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
    thread = Thread(target=server.serve_forever, daemon=True)
    thread.start()
    host, port = server.server_address
    try:
        yield f"http://{host}:{port}/"
    finally:
        server.shutdown(); server.server_close(); thread.join(timeout=2)

4. evidence.py — immutable attempt evidence

The following example makes the evidence.py — immutable attempt evidence behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from pathlib import Path
import json, time

class Evidence:
    def __init__(self, root="evidence"):
        self.root = Path(root)
        self.root.mkdir(parents=True, exist_ok=True)

    def capture(self, driver, attempt, classification, exc=None):
        folder = self.root / f"attempt-{attempt:02d}"
        folder.mkdir(parents=True, exist_ok=True)
        record = {
            "attempt": attempt,
            "classification": classification,
            "session_id": driver.session_id,
            "browser": driver.capabilities.get("browserName"),
            "browser_version": driver.capabilities.get("browserVersion"),
            "url": driver.current_url,
            "title": driver.title,
            "exception_type": type(exc).__name__ if exc else None,
            "exception_message": str(exc)[:500] if exc else None,
            "captured_at_monotonic": time.monotonic(),
        }
        (folder / "result.json").write_text(json.dumps(record, indent=2), encoding="utf-8")
        (folder / "page.html").write_text(driver.page_source, encoding="utf-8")
        driver.save_screenshot(str(folder / "page.png"))
        return record

5. Predict before running

Write these predictions in predictions.txt before executing:

  1. With the 180 ms broken wait, loads configured below 180 ms should usually pass while longer loads should time out; exact browser scheduling can shift boundaries, but the short budget should produce mixed first-pass outcomes.
  2. Every attempt creates a new session ID, but the same URL path and business expectation remain.
  3. The corrected terminal-state wait with a 1.5 s budget should pass all delays in the fixture sequence because the maximum configured delay is 360 ms.
  4. Evidence from the first timeout will show working rather than a stable wrong result, supporting a synchronization root cause.

6. checkpoint.py — measure the broken test, then the corrected test

The following example makes the checkpoint.py — measure the broken test, then the corrected test behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from pathlib import Path
import json
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import TimeoutException

from app_fixture import running_app, DELAYS_MS
from evidence import Evidence

RESULT = (By.CSS_SELECTOR, "[data-testid='result']")
START = (By.CSS_SELECTOR, "[data-testid='start']")
DELAY = (By.CSS_SELECTOR, "[data-testid='delay']")


def run_one(base_url, attempt, timeout, evidence_root):
    driver = webdriver.Chrome()
    evidence = Evidence(evidence_root)
    try:
        driver.get(base_url)
        configured_delay = driver.find_element(*DELAY).text
        driver.find_element(*START).click()
        try:
            observed = WebDriverWait(driver, timeout, poll_frequency=0.02).until(
                lambda d: d.find_element(*RESULT).text
                if d.find_element(*RESULT).text.startswith("ready:")
                or d.find_element(*RESULT).text.startswith("error:")
                else False
            )
            if observed != "ready:42":
                exc = AssertionError(f"stable terminal result={observed!r}")
                evidence.capture(driver, attempt, "stable-assertion-failure", exc)
                return {"attempt": attempt, "outcome": "assertion-fail", "delay": configured_delay}
            evidence.capture(driver, attempt, "clean-first-pass")
            return {"attempt": attempt, "outcome": "pass", "delay": configured_delay}
        except TimeoutException as exc:
            current = driver.find_element(*RESULT).text
            evidence.capture(driver, attempt, "readiness-timeout", exc)
            return {"attempt": attempt, "outcome": "timeout", "delay": configured_delay, "current": current}
    finally:
        driver.quit()


def run_series(base_url, timeout, label):
    rows = []
    for attempt in range(1, len(DELAYS_MS) + 1):
        rows.append(run_one(base_url, attempt, timeout, f"evidence/{label}"))
    Path(f"{label}.json").write_text(json.dumps(rows, indent=2), encoding="utf-8")
    return rows


def summarize(rows):
    counts = {}
    for row in rows:
        counts[row["outcome"]] = counts.get(row["outcome"], 0) + 1
    return counts


def main():
    with running_app() as base_url:
        broken = run_series(base_url, 0.18, "broken")
    with running_app() as base_url:
        stable = run_series(base_url, 1.5, "stable")

    print("broken:", summarize(broken))
    print("stable:", summarize(stable))
    assert summarize(stable) == {"pass": len(DELAYS_MS)}, summarize(stable)

if __name__ == "__main__":
    main()

The corrected test does not add a retry. Both broken and stable waits use an explicit 20 ms poll interval; only the timeout contract changes from “be ready within an arbitrary 180 ms” to “reach a terminal business state within a bounded 1.5 s safety budget that comfortably covers the controlled fixture.”

7. Run and preserve the first failure

The following example makes the Run and preserve the first failure behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

python checkpoint.py
# Review the first evidence/broken/attempt-XX whose result.json says readiness-timeout.
# Do not delete it after the stable series passes.

Expected pattern: the broken series contains both pass and timeout; the stable series reports all passes. Exact short-timeout boundary can vary slightly with host scheduling, so your classification must use the recorded configured delay plus DOM evidence—not merely the count.

8. Classify the root cause from evidence

Open the first timeout’s result.json and page.html. If the page is still working, the session is healthy, the locator exists, and the same scenario later reaches ready:42 under a suitable wait, the supported root cause is a test synchronization defect: the test sampled before the observable transition was complete.

If instead the terminal state were error:dependency, increasing the timeout would be the wrong fix. That would be application/dependency evidence.

9. Compute first-pass stability and retain history

The following example makes the Compute first-pass stability and retain history behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

import json
from pathlib import Path

for label in ("broken", "stable"):
    rows = json.loads(Path(f"{label}.json").read_text())
    total = len(rows)
    passes = sum(r["outcome"] == "pass" for r in rows)
    timeouts = sum(r["outcome"] == "timeout" for r in rows)
    print(label, {
        "controlled_runs": total,
        "first_pass_rate": passes / total,
        "timeouts": timeouts,
    })

Do not call the broken series “eventually passing” by rerunning timed-out cases. Its first-pass rate is the useful signal for this experiment.

10. Write a retry/quarantine rule

The following example makes the Write a retry/quarantine rule behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

retry / quarantine policy for browser-functional gate

1. Default: no automatic blanket retries for product-facing Selenium scenarios.
2. On first failure: preserve attempt evidence before any rerun.
3. Diagnostic rerun: at most one fresh-session rerun may be requested by explicit policy
   for a classified transient hypothesis; it must use the same revision, case data,
   environment, browser family/version policy, and expected result.
4. Result semantics: fail->pass is "pass-after-retry", never clean first-pass.
5. Non-idempotent actions: no step replay unless AUT idempotency and current state are proven.
6. Quarantine: allowed only for a known unstable test with owner, tracking issue,
   bounded scope, continued execution/reporting, and review/exit condition.
7. Artifact retention: attempt-1 evidence is immutable even after later pass.
8. Exit from quarantine: fix root cause, then satisfy an organization-defined controlled
   first-pass stability window on the relevant browser/environment matrix.
9. Infrastructure incidents: tracked separately from product assertion failures.
10. Forbidden shortcuts: giant sleeps, hidden exception swallowing, production experiments,
    TLS disablement, or changing data/environment between retry attempts.

11. Example quarantine record for the broken version

The following example makes the Example quarantine record for the broken version behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

test_id: flake-control.ready-result
status: quarantine-candidate (teaching example only)
classification: test synchronization defect
first_failure_evidence: evidence/broken/attempt-02/
owner: browser-automation-team
tracking_issue: LAB-CH15-001
blocking: no (only if organizational policy explicitly approves quarantine)
continues_to_run: yes
exit_condition: replace 180ms contract with terminal-state wait and demonstrate
                organization-defined clean first-pass stability window
forbidden: blanket retry / artifact deletion / longer fixed sleep

In this checkpoint, you already know the correct fix, so real quarantine is unnecessary. The record exists to practice the governance shape.

12. Verification checklist

  • The broken series contains mixed first-pass outcomes under the controlled timing sequence.
  • The first timeout evidence remains present after all later executions.
  • Every attempt has a fresh WebDriver session ID.
  • The failure classification cites DOM/session/timing evidence, not only TimeoutException.
  • The stable series uses no retry and no fixed sleep.
  • The stable series reaches ready:42 on all fixture delays.
  • The retry/quarantine rule distinguishes clean pass, pass-after-retry, reproduced failure, quarantine, and infrastructure incident.
  • No production target, credential, personal profile, TLS bypass, or external Grid is used.

13. Evidence packet

  • predictions.txt.
  • broken.json and stable.json.
  • Immutable per-attempt JSON, screenshot, and page source.
  • Console summary showing mixed broken outcomes and all-pass stable outcomes.
  • The first timeout’s classification note.
  • The retry/quarantine policy and example record.

All artifacts come from the synthetic loopback fixture. In real pipelines, apply privacy/security retention rules before collecting comparable browser evidence.

14. Cleanup and rollback

Confirm all WebDriver sessions are closed. Preserve only the synthetic checkpoint evidence needed for review, then delete selenium-ch15-checkpoint/. The server uses memory only and shuts down at context exit; no repository, account, remote Grid, database, or browser profile is modified.

15. What Chapter 15 adds to the operating model

The automation platform now has explicit result semantics: stable business assertions, immutable first-failure evidence, layered classification, first-pass stability metrics, bounded diagnostic rerun semantics, visible quarantine governance, and capacity/security awareness around retries. A pipeline can distinguish product risk from test or infrastructure unreliability instead of hiding all three behind eventual green.

Chapter 16 expands that trusted signal across browsers, responsive layouts, localization, and compatibility matrices. Cross-browser comparison is meaningful only when each individual test already has deterministic failure semantics.

16. Summary

The checkpoint demonstrated the correct order of operations: measure → preserve → classify → fix → prove. The broken test flaked because its synchronization contract was wrong. The corrected test became stable by waiting for the right observable state, not by retrying until green.

Knowledge check

What proves the checkpoint failure is a synchronization defect?

Why does the corrected series use zero retries?

What must remain after the stable series passes?

When is a diagnostic rerun valid under the sample policy?

Why is Chapter 15 necessary before cross-browser matrices?

Next chapter

Cross-Browser, Responsive, Localization, and Compatibility Testing: Core Concepts and Mental Model

Continue with Cross-Browser, Responsive, Localization, and Compatibility Testing: Core Concepts and Mental Model. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against Selenium primary documentation on 2026-08-28. Mandatory examples pin Selenium Python 4.47.0 and Python 3.10+, use Python standard-library unittest, a supported locally installed Chromium-family browser with Selenium Manager, and only loopback synthetic applications. Retry/quarantine policy is deliberately modeled as test-governance logic rather than a Selenium capability. The mandatory path deliberately avoids automatic retry plugins so attempt semantics stay visible. A diagnostic retry is shown only as an explicitly coded experiment with fresh-session isolation and immutable first-attempt evidence. Quarantine thresholds/counts in examples are illustrative governance values, not Selenium defaults.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.