Chapter 01Lesson 01~95 minutes

Browser Automation and Test Engineering Foundations: Core Concepts and Mental Model

Learn what Selenium controls, what it does not control, and which pieces of state must line up before a browser interaction can become trustworthy release evidence.

WebDriverTest intentSession stateAUTEvidence

Learning objectives

  • Separate the application under test, test runner, Selenium binding, driver/remote end, browser, assertions, and evidence.
  • Trace one test intent through a WebDriver command to observable application state.
  • Distinguish browser automation from unit, component, API, and performance testing.
  • Inspect Selenium, browser, driver, capability, session, URL, and title state before deeper automation.
  • Explain why a passing browser interaction is not automatically proof of business correctness.
  • Use explicit lifecycle and evidence language suitable for CI/CD diagnostics.

1. Current baseline and scope

This chapter is written against Selenium 4.47.0, released on 2026-08-10. The Python binding requires Python 3.10 or newer. Selenium Manager ships with Selenium and is invoked by the bindings when no usable driver is supplied, so modern beginner code can normally create webdriver.Chrome() without hard-coding a driver executable path. Your installed browser and the matching driver remain separate pieces of software; record what the created session actually reports instead of assuming versions from memory.

Safety boundary: every mandatory exercise targets loopback-only pages or synthetic fixtures. Do not point these examples at employer/customer production sites, real accounts, commerce actions, MFA/CAPTCHA flows, or publicly exposed Grid/browser-debug endpoints.

2. The problem browser automation actually solves

A unit test can prove a pure function returns the expected value, and an API test can prove a server endpoint returns the expected response. Neither proves that a real browser can render the page, locate the intended control, accept input, execute browser-side JavaScript, navigate through the user-visible flow, and expose the expected result. Selenium exists at that browser boundary.

The cost of that realism is additional moving parts. A browser test can fail because the application is wrong, because the test asks the wrong question, because the browser is not ready, because the selector points at incidental markup, because the driver/browser pair cannot start, because the test data is dirty, or because the execution environment is starved. Treating every red test as “Selenium failed” destroys diagnostic value.

Layer Good question Typical cost What Selenium adds
Unit Does a small function/class behave correctly? Very low Usually nothing; keep it below the browser.
Component/API Does a service or component contract behave correctly? Low–medium Usually nothing unless browser behavior is part of the contract.
Browser/E2E Can a user-visible journey work through a real browser? High Native browser control and observable UI state.
Load/performance What happens under sustained/concurrent traffic? Specialized Selenium is not the load generator; use dedicated tools.

3. Mental model: intent to evidence

One browser-test command path

The following diagram visualizes the relationships described in Mental model: intent to evidence. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.

flowchart TD
I[Test intent] --> R[Test runner]
R --> B[Selenium language binding]
B --> W[WebDriver command]
W --> D[Driver or Grid remote end]
D --> C[Browser context]
C --> A[Application under test]
A --> O[Observable browser/AUT state]
O --> X[Assertion]
O --> E[Evidence]
X --> V[Release signal]
E --> V

Test intent is the behavior you mean to verify: for example, “an invalid quantity is rejected.” The test runner decides which test function runs and whether it passed or failed. The Selenium binding is the language library your code imports. A WebDriver command is a standardized request such as navigate, find element, click, or get text. The driver/remote end translates that request for a browser. The browser context is the window/tab/frame in which the browser acts. The AUT is the application under test. An assertion converts observed state into pass/fail semantics. Evidence preserves enough context—version data, screenshot, logs, DOM state, timing—to explain the result later.

4. Objects and state stores you must not blur together

The following table organizes the key choices and evidence for Objects and state stores you must not blur together. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Object/state Owned by Examples Why it matters
Test process Runner/runtime Python process, selected test, environment variables A crash here can end a test before WebDriver reports anything.
WebDriver session Driver/remote end + browser Session ID, capabilities, timeouts Commands are scoped to a live session; quitting destroys it.
Browser profile/context Browser Cookies, cache, local storage, tabs Dirty/shared state can make tests order-dependent.
AUT state Application/services Database records, server session, feature flags Selenium cannot make bad test data deterministic by itself.
Evidence Test/CI workspace Screenshots, page source, logs, capability JSON Useful only if correlated to the right test/session and redacted.
CI orchestration CI platform Job, workspace, secret store, artifact upload It schedules and retains results; it does not define browser semantics.

5. Session lifecycle before API memorization

  1. Create: the binding requests a new session. For local Chrome, Selenium may use Selenium Manager to locate/manage the driver and then starts a browser.
  2. Negotiate: requested options and browser capabilities become the session’s returned capabilities. Returned values are evidence of what you actually got.
  3. Operate: navigation, location, interaction, JavaScript/browser behavior, and the AUT change state over time.
  4. Observe: read URL/title/element state, collect diagnostics, and make assertions against stable intended state.
  5. Quit: end the entire WebDriver session. close() only closes a window/tab and is not a substitute for deterministic session teardown.

A test that opens a browser and clicks a button but never asserts the intended effect is automation, not a complete test. A test that asserts the effect but leaks the browser session is incomplete test engineering because it leaves resources and state behind.

6. Read-only inspection: prove what session exists

Create an isolated environment, pin the binding, and inspect the new session before doing any application mutation. This code intentionally starts at the browser’s initial page and prints returned session evidence.

python -m venv .venv
# Linux/macOS
. .venv/bin/activate
# Windows PowerShell equivalent: .\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install "selenium==4.47.0"
python -c "import selenium,sys; print(sys.version); print(selenium.__version__)"

The following example makes the Read-only inspection: prove what session exists behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from selenium import __version__ as selenium_version
from selenium import webdriver

print(f"selenium={selenium_version}")

driver = None
try:
    driver = webdriver.Chrome()
    caps = driver.capabilities
    print(f"session_id={driver.session_id}")
    print(f"browser={caps.get('browserName')}")
    print(f"browser_version={caps.get('browserVersion')}")
    print(f"platform={caps.get('platformName')}")
    print(f"current_url={driver.current_url}")
    print(f"title={driver.title!r}")

    chrome = caps.get('chrome', {})
    if chrome:
        print(f"chromedriver_version={chrome.get('chromedriverVersion', 'not reported')}")
finally:
    if driver is not None:
        driver.quit()

Expected shape: selenium=4.47.0, a non-empty session ID while the session is alive, a browser name/version, a platform, and an initial URL such as data:, or another browser-specific blank page. Chrome-family capabilities may also report a ChromeDriver version. Exact browser/driver values are environment evidence, not constants the lesson should invent.

7. Local and remote execution change routing, not test intent

With local WebDriver, the driver service and browser run on the same host as the test process. With Selenium Grid or a browser cloud, the test binding sends commands to a remote endpoint that allocates a compatible browser session elsewhere. The test still needs the same fundamental contract: a clean session, known AUT/test data, deterministic waits, meaningful assertions, and evidence. Grid does not make a poor locator stable and does not create application readiness.

8. Passing interaction versus trustworthy test signal

Suppose Selenium clicks “Submit” and no exception occurs. That proves only that the command path completed. It does not prove the server accepted the intended payload, the visible confirmation is correct, the right user was charged, or the database state is valid. A trustworthy test states the intended behavior and selects observable assertions that are strong enough for that intent.

Useful rule: ask “what evidence would distinguish the intended behavior from a coincidental successful click?” That question usually identifies the correct assertion and the minimum evidence to preserve.

9. Common foundation mistakes

  • “Selenium is the test framework.” Selenium provides browser automation APIs; a runner/framework supplies test discovery, fixtures, reporting, and assertion organization.
  • “A browser pass proves the whole system.” Browser evidence is one layer of confidence, not proof of every backend rule.
  • “Every check should be UI-driven.” Repetitive setup and low-level rules are often faster and more stable below the browser.
  • “The driver is the browser.” They are separate processes/components connected by the WebDriver implementation.
  • “An intermittent red test is acceptable noise.” Intermittence is evidence of uncontrolled state, timing, environment, or product behavior and deserves diagnosis.

10. Hands-on lab: create an environment evidence card

  1. Run the inspection script above on a disposable workstation or lab VM.
  2. Record Python, Selenium, browser, driver (if reported), platform, and session ID.
  3. Predict what changes after driver.quit(): the browser process/session should end and the previous session ID must no longer accept commands.
  4. Verify the prediction by keeping no orphan browser and by never issuing commands after teardown in normal test code.
  5. Save only non-sensitive version/capability fields. Do not dump full profiles, cookies, authorization headers, or personal browsing data.

11. Why this matters in DevOps

CI/CD treats automated tests as release gates. A gate is useful only if engineers can explain its scope, reproduce its environment, distinguish application failure from automation failure, and recover evidence from the first failure. The mental model in this lesson is therefore operational: every later Selenium technique will be tied to session state, application state, assertions, and evidence rather than to a list of methods.

Knowledge check

What is the difference between the Selenium binding and the browser driver?

Why is a successful click not enough to call a browser script a test?

Which state should a new WebDriver session isolate by default?

What should you record instead of assuming browser and driver versions?

Does moving a test to Grid fix a flaky locator or synchronization race?

Next lesson

From mental model to one controlled browser workflow

Lesson 2 builds a loopback-only demo application, drives one complete user-visible path, captures evidence, and compares the browser check with a lower-layer HTTP check.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Selenium primary documentation on 2026-08-27. Mandatory examples pin the Python Selenium binding to 4.47.0, require Python 3.10+, use an installed supported local browser with Selenium Manager as the default driver-management path, and do not require Selenium Grid, a paid browser cloud, enterprise identity, or a production website. Record the browser and driver versions returned by the actual session because those remain environment-specific.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.