Assertions, Failure Semantics, Retries, and Flaky-Test Control: Diagnostics, Failure Modes, and Production Practices
A flaky suite becomes expensive when diagnosis begins with “rerun it.” This lesson starts with preserved evidence, then separates transient-state assertions, synchronization defects, swallowed exceptions, changed retry inputs, deleted artifacts, infrastructure timeouts, and sleep-based suppression before applying the smallest correction.
Learning objectives
- Apply the chapter diagnostic sequence before any retry or timeout change.
- Diagnose blanket retries, transient assertions, swallowed exceptions, changed retry data, and deleted first-failure artifacts.
- Separate infrastructure timeout from application assertion failure using layered evidence.
- Repair an intentionally broken test without JavaScript bypass, giant sleep, or blanket retry.
- Relate retry cost to browser startup, Grid queue/capacity, AUT latency, and artifact IO.
1. Diagnostic sequence: preserve before you perturb
The following diagram visualizes the relationships described in Diagnostic sequence: preserve before you perturb. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.
flowchart TB F[First failure] --> E[Freeze evidence before rerun] E --> V[Versions / target / case identity] V --> S[Session / context / locator / wait state] S --> A[AUT / browser / network evidence] A --> G[Grid / CI / resource state if remote] G --> H[Smallest supported hypothesis] H --> C[Least destructive correction] C --> R[Smallest controlled rerun] R --> D[Record disposition: fixed / product bug / infra / quarantine]
- Preserve first-failure evidence.
- Confirm Selenium/binding/browser/driver/Grid versions.
- Confirm target, environment, case data, and revision.
- Inspect session, browsing context, locator, element, and synchronization state.
- Inspect AUT/browser/network evidence.
- If remote, inspect Grid/CI queue, node, host, and resource state.
- State the smallest supported hypothesis.
- Apply the least destructive correction.
- Rerun the smallest controlled scenario with the same inputs.
2. Realistic failure modes and what they destroy
The following table organizes the key choices and evidence for Realistic failure modes and what they destroy. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Failure mode | Why it is dangerous | Correction direction |
|---|---|---|
| Blanket retry turns red green | erases first-pass reliability from gate meaning | retain attempt history; scope retry only as explicit experiment |
| Assert transient state | test samples before business outcome stabilizes | wait on terminal condition then assert exact meaning |
| Swallow exception | removes stack/error classification and may produce false pass | capture evidence then re-raise/fail |
| Retry with different data | second attempt is not equivalent | freeze/record controlled inputs |
| Delete failed artifacts after retry pass | destroys causal evidence | immutable per-attempt artifact namespace |
| Treat infra timeout as product assertion | wrong team/churn; product signal corrupted | separate session/runner/Grid health from AUT outcome |
| Longer sleep suppresses flake | adds latency and still races at tail | wait for observable state with bounded deadline |
3. Intentionally broken example: transient assertion + swallowed retry
The following example makes the Intentionally broken example: transient assertion + swallowed retry behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
def broken_test(driver, base_url):
for _ in range(3):
try:
driver.get(base_url)
driver.find_element(By.CSS_SELECTOR, "[data-testid='start']").click()
# Broken: element is present immediately, but state is still "working".
text = driver.find_element(By.CSS_SELECTOR, "[data-testid='result']").text
assert text == "ready:42"
return
except Exception:
# Broken: hides assertion/driver/locator differences and destroys evidence.
continue
# Broken: caller cannot tell what failed on which attempt.
return None
This code mixes all failure categories, retries in the same browser
state, preserves no evidence, and can replay after unknown partial
state. The final None can even be ignored by a test
runner.
4. Repair: terminal-state wait, evidence, then fail with original cause
The following example makes the Repair: terminal-state wait, evidence, then fail with original cause behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.support.ui import WebDriverWait
def wait_for_terminal(driver):
def terminal(d):
text = d.find_element(By.CSS_SELECTOR, "[data-testid='result']").text
if text.startswith("ready:") or text.startswith("error:"):
return text
return False
return WebDriverWait(driver, 1.5, poll_frequency=0.02).until(terminal)
def one_controlled_attempt(driver, base_url, evidence):
driver.get(base_url)
driver.find_element(By.CSS_SELECTOR, "[data-testid='start']").click()
try:
observed = wait_for_terminal(driver)
if observed != "ready:42":
exc = AssertionError(f"terminal business result: {observed!r}")
evidence.capture(driver, 1, "stable-assertion-failure", exc)
raise exc
evidence.capture(driver, 1, "clean-first-pass")
return observed
except TimeoutException as exc:
evidence.capture(driver, 1, "readiness-timeout", exc)
raise
The terminal condition can surface an application error before the deadline; the exact assertion remains separate. No retry occurs until the first attempt has a complete record and a human/policy layer decides another observation is justified.
5. Never catch broad exceptions just to keep the suite green
The following example makes the Never catch broad exceptions just to keep the suite green behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
# Better pattern when you must attach evidence around any failure:
try:
run_scenario(driver)
except Exception as exc:
evidence.capture(driver, 1, "unclassified-first-failure", exc)
raise # preserve original traceback and test-runner failure semantics
Evidence capture must itself be defensive: if screenshot capture fails because the browser crashed, preserve the original exception and record the evidence-capture failure separately rather than replacing the cause.
6. A rerun with changed data is not the same experiment
If attempt 1 used case=editor-01 and attempt 2 silently
generates editor-02, a later pass does not show that
attempt 1 was flaky. It shows that a different case passed. Carry
forward Chapter 14 case ID, seed, revision, environment, browser,
and expected result into every attempt record.
7. Artifact namespaces must be append-only per attempt
The following example makes the Artifact namespaces must be append-only per attempt behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
artifacts/
test-checkout-approval/
run-4812/
attempt-01/
result.json
page.png
page.html
browser.log
attempt-02/
result.json
page.png
page.html
disposition.json
Do not overwrite attempt-01 with a successful retry. A
disposition record may summarize the run, but it links to immutable
attempt evidence.
8. Infrastructure timeout versus application failure
The following table organizes the key choices and evidence for Infrastructure timeout versus application failure. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Evidence | Likely layer to inspect |
|---|---|
| No session ID; driver start failed | driver/browser/host/Selenium Manager/Grid infrastructure |
| Grid session queued, no node capacity | Grid capacity/scheduling |
| Session healthy; AUT shows error:dependency | application/dependency path |
| Session healthy; result remains working; network request pending | AUT/network or synchronization contract |
| Stable wrong business value after terminal state | product behavior, expected value, or test data |
9. Performance only where causal
Retries multiply browser startup cost, Grid queue time, AUT/network load, and artifact IO. Longer timeouts also hold a node slot while doing no productive work. Measure those layers separately; do not infer that “tests got slower” means Selenium itself regressed.
10. Security-sensitive failure evidence
Failure screenshots/page sources can capture account names, tokens, personal data, downloads, or proxy pages. Use synthetic accounts and local fixtures in this chapter. In real pipelines, define redaction and retention policy before enabling broad artifact capture. Never disable TLS, bypass identity controls, or test failure injection against production merely to reproduce a flaky test.
11. Rerun the smallest controlled scenario
After correcting the wait contract, rerun only the affected scenario with the same case/environment/browser first. If it is stable, expand to the relevant browser matrix. Starting with the whole suite increases noise and can consume capacity before you know whether the hypothesis was correct.
12. Summary and bridge
Diagnosis begins with immutable first-failure evidence and ends with a controlled rerun after the least destructive correction. Lesson 5 turns that discipline into a checkpoint: measure a deliberately flaky test, stabilize it without blanket retry, and write an explicit quarantine/retry rule.
Knowledge check
What is the first action after a test fails?
Preserve first-failure evidence before retrying, changing timeouts, restarting browsers/Grid, changing data, or otherwise perturbing the state.
Why is catch Exception + continue dangerous?
It collapses different failure classes, hides tracebacks, may replay unknown state, and can let the test finish without a meaningful failure.
Why should evidence directories be attempt-specific?
So a later pass cannot overwrite the exact DOM/screenshot/logs that explain the first failure.
How do you distinguish infra from product timeout?
Correlate session/Grid/runner health with AUT/browser/network state; an exception name alone is insufficient.
What should you rerun after a targeted correction?
The smallest controlled scenario with the same inputs first, then expand scope only after the hypothesis is supported.
Official references and version notes
- Selenium 4.47 release notes — stable binding/Grid baseline pinned for this chapter.
- Selenium downloads — current stable Selenium client and Server/Grid versions.
- Waiting strategies — race conditions between test and application readiness are a primary source of flaky browser tests.
- Overview of Test Automation — browser tests should keep setup, actions, and evaluation compact and intentional.
- Avoid sharing state — isolate data and create a new WebDriver instance per test where practical.
- Fresh browser per test — begin from a clean, known browser state.
- Test independency — scenarios should not depend on another test's success or state.
- Python unittest — hard assertion semantics, subtests, setup/cleanup, and failure reporting used in the mandatory path.
Version-sensitive behavior was rechecked against Selenium primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python 4.47.0 and Python 3.10+, use Python standard-library
unittest, a supported locally installed
Chromium-family browser with Selenium Manager, and only loopback
synthetic applications. Retry/quarantine policy is deliberately
modeled as test-governance logic rather than a Selenium
capability. The mandatory path deliberately avoids automatic retry
plugins so attempt semantics stay visible. A diagnostic retry is
shown only as an explicitly coded experiment with fresh-session
isolation and immutable first-attempt evidence. Quarantine
thresholds/counts in examples are illustrative governance values,
not Selenium defaults.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.