Chapter 15Lesson 04~195 minutes

Assertions, Failure Semantics, Retries, and Flaky-Test Control: Diagnostics, Failure Modes, and Production Practices

A flaky suite becomes expensive when diagnosis begins with “rerun it.” This lesson starts with preserved evidence, then separates transient-state assertions, synchronization defects, swallowed exceptions, changed retry inputs, deleted artifacts, infrastructure timeouts, and sleep-based suppression before applying the smallest correction.

DiagnosticsArtifact retentionSynchronizationInfrastructureProduction practice

Learning objectives

  • Apply the chapter diagnostic sequence before any retry or timeout change.
  • Diagnose blanket retries, transient assertions, swallowed exceptions, changed retry data, and deleted first-failure artifacts.
  • Separate infrastructure timeout from application assertion failure using layered evidence.
  • Repair an intentionally broken test without JavaScript bypass, giant sleep, or blanket retry.
  • Relate retry cost to browser startup, Grid queue/capacity, AUT latency, and artifact IO.

1. Diagnostic sequence: preserve before you perturb

Failure triage preserves causality

The following diagram visualizes the relationships described in Diagnostic sequence: preserve before you perturb. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.

flowchart TB
  F[First failure] --> E[Freeze evidence before rerun]
  E --> V[Versions / target / case identity]
  V --> S[Session / context / locator / wait state]
  S --> A[AUT / browser / network evidence]
  A --> G[Grid / CI / resource state if remote]
  G --> H[Smallest supported hypothesis]
  H --> C[Least destructive correction]
  C --> R[Smallest controlled rerun]
  R --> D[Record disposition: fixed / product bug / infra / quarantine]
  1. Preserve first-failure evidence.
  2. Confirm Selenium/binding/browser/driver/Grid versions.
  3. Confirm target, environment, case data, and revision.
  4. Inspect session, browsing context, locator, element, and synchronization state.
  5. Inspect AUT/browser/network evidence.
  6. If remote, inspect Grid/CI queue, node, host, and resource state.
  7. State the smallest supported hypothesis.
  8. Apply the least destructive correction.
  9. Rerun the smallest controlled scenario with the same inputs.

2. Realistic failure modes and what they destroy

The following table organizes the key choices and evidence for Realistic failure modes and what they destroy. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Failure mode Why it is dangerous Correction direction
Blanket retry turns red green erases first-pass reliability from gate meaning retain attempt history; scope retry only as explicit experiment
Assert transient state test samples before business outcome stabilizes wait on terminal condition then assert exact meaning
Swallow exception removes stack/error classification and may produce false pass capture evidence then re-raise/fail
Retry with different data second attempt is not equivalent freeze/record controlled inputs
Delete failed artifacts after retry pass destroys causal evidence immutable per-attempt artifact namespace
Treat infra timeout as product assertion wrong team/churn; product signal corrupted separate session/runner/Grid health from AUT outcome
Longer sleep suppresses flake adds latency and still races at tail wait for observable state with bounded deadline

3. Intentionally broken example: transient assertion + swallowed retry

The following example makes the Intentionally broken example: transient assertion + swallowed retry behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

def broken_test(driver, base_url):
    for _ in range(3):
        try:
            driver.get(base_url)
            driver.find_element(By.CSS_SELECTOR, "[data-testid='start']").click()
            # Broken: element is present immediately, but state is still "working".
            text = driver.find_element(By.CSS_SELECTOR, "[data-testid='result']").text
            assert text == "ready:42"
            return
        except Exception:
            # Broken: hides assertion/driver/locator differences and destroys evidence.
            continue
    # Broken: caller cannot tell what failed on which attempt.
    return None

This code mixes all failure categories, retries in the same browser state, preserves no evidence, and can replay after unknown partial state. The final None can even be ignored by a test runner.

4. Repair: terminal-state wait, evidence, then fail with original cause

The following example makes the Repair: terminal-state wait, evidence, then fail with original cause behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from selenium.common.exceptions import TimeoutException
from selenium.webdriver.support.ui import WebDriverWait


def wait_for_terminal(driver):
    def terminal(d):
        text = d.find_element(By.CSS_SELECTOR, "[data-testid='result']").text
        if text.startswith("ready:") or text.startswith("error:"):
            return text
        return False
    return WebDriverWait(driver, 1.5, poll_frequency=0.02).until(terminal)


def one_controlled_attempt(driver, base_url, evidence):
    driver.get(base_url)
    driver.find_element(By.CSS_SELECTOR, "[data-testid='start']").click()
    try:
        observed = wait_for_terminal(driver)
        if observed != "ready:42":
            exc = AssertionError(f"terminal business result: {observed!r}")
            evidence.capture(driver, 1, "stable-assertion-failure", exc)
            raise exc
        evidence.capture(driver, 1, "clean-first-pass")
        return observed
    except TimeoutException as exc:
        evidence.capture(driver, 1, "readiness-timeout", exc)
        raise

The terminal condition can surface an application error before the deadline; the exact assertion remains separate. No retry occurs until the first attempt has a complete record and a human/policy layer decides another observation is justified.

5. Never catch broad exceptions just to keep the suite green

The following example makes the Never catch broad exceptions just to keep the suite green behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

# Better pattern when you must attach evidence around any failure:
try:
    run_scenario(driver)
except Exception as exc:
    evidence.capture(driver, 1, "unclassified-first-failure", exc)
    raise  # preserve original traceback and test-runner failure semantics

Evidence capture must itself be defensive: if screenshot capture fails because the browser crashed, preserve the original exception and record the evidence-capture failure separately rather than replacing the cause.

6. A rerun with changed data is not the same experiment

If attempt 1 used case=editor-01 and attempt 2 silently generates editor-02, a later pass does not show that attempt 1 was flaky. It shows that a different case passed. Carry forward Chapter 14 case ID, seed, revision, environment, browser, and expected result into every attempt record.

7. Artifact namespaces must be append-only per attempt

The following example makes the Artifact namespaces must be append-only per attempt behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

artifacts/
  test-checkout-approval/
    run-4812/
      attempt-01/
        result.json
        page.png
        page.html
        browser.log
      attempt-02/
        result.json
        page.png
        page.html
      disposition.json

Do not overwrite attempt-01 with a successful retry. A disposition record may summarize the run, but it links to immutable attempt evidence.

8. Infrastructure timeout versus application failure

The following table organizes the key choices and evidence for Infrastructure timeout versus application failure. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Evidence Likely layer to inspect
No session ID; driver start failed driver/browser/host/Selenium Manager/Grid infrastructure
Grid session queued, no node capacity Grid capacity/scheduling
Session healthy; AUT shows error:dependency application/dependency path
Session healthy; result remains working; network request pending AUT/network or synchronization contract
Stable wrong business value after terminal state product behavior, expected value, or test data

9. Performance only where causal

Retries multiply browser startup cost, Grid queue time, AUT/network load, and artifact IO. Longer timeouts also hold a node slot while doing no productive work. Measure those layers separately; do not infer that “tests got slower” means Selenium itself regressed.

10. Security-sensitive failure evidence

Failure screenshots/page sources can capture account names, tokens, personal data, downloads, or proxy pages. Use synthetic accounts and local fixtures in this chapter. In real pipelines, define redaction and retention policy before enabling broad artifact capture. Never disable TLS, bypass identity controls, or test failure injection against production merely to reproduce a flaky test.

11. Rerun the smallest controlled scenario

After correcting the wait contract, rerun only the affected scenario with the same case/environment/browser first. If it is stable, expand to the relevant browser matrix. Starting with the whole suite increases noise and can consume capacity before you know whether the hypothesis was correct.

12. Summary and bridge

Diagnosis begins with immutable first-failure evidence and ends with a controlled rerun after the least destructive correction. Lesson 5 turns that discipline into a checkpoint: measure a deliberately flaky test, stabilize it without blanket retry, and write an explicit quarantine/retry rule.

Knowledge check

What is the first action after a test fails?

Why is catch Exception + continue dangerous?

Why should evidence directories be attempt-specific?

How do you distinguish infra from product timeout?

What should you rerun after a targeted correction?

Next lesson

Checkpoint Lab — Assertions, Failure Semantics, Retries, and Flaky-Test Control

Continue with Checkpoint Lab — Assertions, Failure Semantics, Retries, and Flaky-Test Control. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against Selenium primary documentation on 2026-08-28. Mandatory examples pin Selenium Python 4.47.0 and Python 3.10+, use Python standard-library unittest, a supported locally installed Chromium-family browser with Selenium Manager, and only loopback synthetic applications. Retry/quarantine policy is deliberately modeled as test-governance logic rather than a Selenium capability. The mandatory path deliberately avoids automatic retry plugins so attempt semantics stay visible. A diagnostic retry is shown only as an explicitly coded experiment with fresh-session isolation and immutable first-attempt evidence. Quarantine thresholds/counts in examples are illustrative governance values, not Selenium defaults.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.