Chapter 01Lesson 04~120 minutes

Browser Automation and Test Engineering Foundations: Diagnostics, Failure Modes, and Production Practices

Diagnose browser failures by preserving the first symptom, locating the failing layer, correcting the smallest cause, and proving the fix without retries, giant sleeps, or destructive shortcuts.

DiagnosticsFlakinessEvidenceSynchronizationProduction safety

Learning objectives

  • Apply a repeatable first-failure diagnostic sequence.
  • Distinguish assertion, locator/context, synchronization, session/driver, AUT/network, and resource failures.
  • Reproduce and repair a deterministic timing race with an explicit condition.
  • Capture screenshots/page source without leaking sensitive state.
  • Explain why blanket retries and giant sleeps mask rather than solve root causes.
  • Separate runner/browser/Grid/AUT performance symptoms before tuning.

1. Diagnostic sequence

  1. Preserve the first failure. Keep the original exception, assertion, timestamp, screenshot/log/DOM evidence that is safe to retain.
  2. Confirm versions. Selenium binding, browser, driver/remote end, Grid (if used), platform.
  3. Confirm target and data. Correct environment, base URL, synthetic account/record, feature state.
  4. Inspect session/context. Session ID, returned capabilities, current URL, window/frame/shadow context.
  5. Inspect locator/element/synchronization. Does the element exist now? Is it visible/interactable? Was an old reference cached?
  6. Inspect AUT/browser/network evidence. DOM, console/network events where supported, application logs.
  7. Inspect Grid/CI/resource state when remote. Queue, node, CPU/memory, browser process, workspace.
  8. Apply the least destructive correction. Change the smallest causal layer.
  9. Rerun the smallest controlled scenario. Verify cause and fix before broad suite execution.

2. Failure families are clues, not verdicts

The following table organizes the key choices and evidence for Failure families are clues, not verdicts. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Symptom Likely layers to inspect Wrong shortcut
NoSuchElementException Wrong locator/context, element not created yet, wrong page/environment. Add a giant sleep everywhere.
StaleElementReferenceException DOM replaced/element reference lifetime. Cache the same stale object and retry blindly.
TimeoutException Condition never became true, wrong condition, slow AUT, context mismatch. Raise timeout to minutes without evidence.
InvalidSessionIdException Session was quit/closed or lost. Restart everything without finding who destroyed it.
Assertion failure Product behavior, test expectation, test data/environment. Replace assertion with a looser one.
Intermittent pass/fail Timing, shared state, environment capacity, nondeterministic AUT. Mark as “flaky” and retry until green.

3. Intentionally broken race

Create app/index.html with a result that appears 700 ms after the click:

<!doctype html>
<html lang="en">
<head><meta charset="utf-8"><title>Timing Lab</title></head>
<body>
  <button id="load" type="button">Load result</button>
  <div id="result-host"></div>
  <script>
    document.querySelector('#load').addEventListener('click', () => {
      setTimeout(() => {
        const p = document.createElement('p');
        p.id = 'result';
        p.textContent = 'ready';
        document.querySelector('#result-host').replaceChildren(p);
      }, 700);
    });
  </script>
</body>
</html>

Serve it locally:

The following example makes the Intentionally broken race behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

python -m http.server 8766 --bind 127.0.0.1 --directory app

This test is deliberately wrong because it looks for #result immediately after the click:

from selenium import webdriver
from selenium.webdriver.common.by import By

URL = "http://127.0.0.1:8766/"
driver = webdriver.Chrome()
try:
    driver.get(URL)
    driver.find_element(By.ID, "load").click()
    # Broken: #result does not exist yet.
    assert driver.find_element(By.ID, "result").text == "ready"
finally:
    driver.quit()

Predict the result: on a normal run, Selenium is likely to raise NoSuchElementException because the element does not exist yet. The AUT is behaving exactly as written. The failure is a synchronization defect in the test.

4. Why common “fixes” are not fixes

  • time.sleep(10) guesses at time, slows every run, and still fails if the environment exceeds the guess.
  • Retrying the entire test mutates state again and can hide the first causal symptom.
  • Using JavaScript to create/read the result directly bypasses the user-visible interaction being tested.
  • Restarting the browser/Grid destroys evidence and changes the environment before diagnosis.
  • Disabling TLS/certificate checks to “get through” a network error weakens the trust boundary and may hide real deployment problems.

5. Repair the cause with an intended-state condition

The corrected test waits for the result element itself, then asserts its text. The timeout is bounded and the failure path captures evidence before teardown.

from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import TimeoutException

URL = "http://127.0.0.1:8766/"
EVIDENCE = Path("evidence")
EVIDENCE.mkdir(exist_ok=True)

driver = webdriver.Chrome()
try:
    driver.get(URL)
    driver.find_element(By.ID, "load").click()

    result = WebDriverWait(driver, 3).until(
        lambda d: d.find_element(By.ID, "result")
    )
    assert result.text == "ready"
except Exception:
    driver.save_screenshot(str(EVIDENCE / "first-failure.png"))
    (EVIDENCE / "first-failure.html").write_text(
        driver.page_source, encoding="utf-8"
    )
    raise
finally:
    driver.quit()

The wait does not make the test “more patient” in the abstract; it encodes the precise readiness contract. If #result never appears within three seconds, the timeout is useful evidence that either the AUT did not reach the intended state or the test is observing the wrong context.

6. Preserve evidence before mutation

Useful first-failure evidence can include the exception/stack trace, current URL/title, session ID, selected capability fields, screenshot, relevant DOM/page source, browser console/network events, and Grid/CI resource state. Capture only what helps diagnose the failure and is safe to retain.

Privacy rule: screenshots and page source can contain names, tokens, account numbers, private messages, or hidden form values. Mandatory labs use synthetic data. In real systems, redact and apply retention/access controls before publishing diagnostics as CI artifacts.

7. Teardown defects are production defects in the test platform

Leaked browser sessions consume CPU/memory and can exhaust Grid slots. Conversely, calling quit() too early causes invalid-session failures. Put session ownership in one clear lifecycle boundary and make teardown unconditional. Selenium’s current guidance favors fresh WebDriver instances per test, which also makes ownership easier to reason about.

8. Performance symptoms: separate the layers

The following table organizes the key choices and evidence for Performance symptoms: separate the layers. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Time/capacity component Example symptom Evidence
Runner/setup Tests spend time preparing fixtures before browser starts. Runner timings.
Session startup Browser takes several seconds to launch. Timestamp before/after constructor, driver logs.
Grid queue Remote session request waits before allocation. Grid queue/session metrics.
AUT/network Page/API resources respond slowly. Application/network timing evidence.
Evidence I/O Video/screenshots/logging dominate suite time. Artifact sizes and capture timings.
Retries Suite looks “stable” but runtime doubles. Retry count and first-failure rate.

Do not call Selenium a load-testing tool and do not tune Grid capacity before measuring where time is actually spent.

9. Foundation failure modes and their real corrections

  • Using Selenium for every test: move low-level logic/contracts down where browser behavior is irrelevant.
  • Asserting before readiness: wait for the business-relevant condition.
  • Leaving sessions open: centralize ownership and unconditional quit.
  • Coupling to incidental markup: use stable test-owned/user-semantic selectors.
  • Equating interaction with correctness: assert the intended observable result.
  • Accepting intermittence as noise: preserve first-failure evidence and eliminate uncontrolled state/timing.

10. Break-and-recover lab

  1. Run the broken race once and preserve its exact exception.
  2. Before changing code, state whether the AUT, locator, session, or synchronization layer is most likely responsible.
  3. Verify manually/through page source that #result appears after the delay.
  4. Replace only the synchronization defect with the bounded explicit condition.
  5. Run the fixed scenario three times without retries.
  6. Verify each run quits the browser and that no unexpected evidence contains sensitive data.

Knowledge check

Why is increasing a timeout not automatically a valid fix for TimeoutException?

What should be preserved before a retry or browser restart?

What is the causal defect in the intentionally broken example?

Why are orphan browser sessions a Grid/capacity issue?

When should JavaScript execution be used as a troubleshooting shortcut?

Next lesson

Integrate the chapter into a checkpoint

Lesson 5 combines scope selection, two independent browser tests, environment guards, evidence, failure semantics, and cleanup into a small production-style acceptance packet.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Selenium primary documentation on 2026-08-27. Mandatory examples pin the Python Selenium binding to 4.47.0, require Python 3.10+, use an installed supported local browser with Selenium Manager as the default driver-management path, and do not require Selenium Grid, a paid browser cloud, enterprise identity, or a production website. Record the browser and driver versions returned by the actual session because those remain environment-specific.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.