Screenshots, Logs, Network Evidence, and Failure Diagnostics: Core Concepts and Mental Model
A red browser test is only the beginning of diagnosis. The useful question is whether the evidence lets another engineer distinguish an application defect, browser/client error, WebDriver failure, network problem, or test-code mistake without rerunning the job. This lesson builds that evidence model before any collector mutates files or logging configuration.
Learning objectives
- Trace a failing test from test-step timestamp through browser, application, network, WebDriver/Grid, and CI evidence.
- Distinguish screenshots, targeted DOM state, browser logs, network evidence, driver/Grid logs, and CI metadata by what each can and cannot prove.
- Inspect session identity, capabilities, URL/title, DOM state, and available log types before enabling or collecting extra evidence.
- Apply correlation, redaction, retention, and trust-boundary rules before artifacts leave a disposable workspace.
- Explain why evidence must preserve the first failure rather than be overwritten by retries or cleanup.
1. The practical problem: a screenshot is not a diagnosis
Chapters 06 and 15 already established two constraints: synchronization must describe a real condition, and a retry does not erase the first failure. Chapter 17 adds the operational consequence. A screenshot can show a button that appears disabled, but it cannot by itself tell you whether the AUT rejected a request, JavaScript crashed, a request never reached the server, WebDriver lost its session, or the test asserted the wrong state.
Capture enough independent evidence to test competing hypotheses. Do not collect everything merely because the browser can expose it.
2. Mental model: one failure, correlated evidence layers
The following diagram visualizes the relationships described in Mental model: one failure, correlated evidence layers. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.
flowchart TD T[Test step + UTC timestamp + test_id] --> F[Assertion / exception] F --> V[Viewport screenshot] F --> D[URL + title + targeted DOM state] F --> B[Browser console / JS errors] F --> N[Network / server evidence] F --> W[WebDriver + session + capabilities] W --> G[Grid / node logs if remote] T --> C[CI job / attempt / shard metadata] V --> P[Privacy-aware evidence bundle] D --> P B --> P N --> P W --> P G --> P C --> P P --> R[Triage conclusion]
The test step creates the correlation anchor: a stable test ID, attempt number, and UTC timestamp. The assertion or exception is the primary symptom. Visual, DOM, browser, network, WebDriver/Grid, and CI evidence are separate observations. They become useful only when they refer to the same session and time window. The bundle then supports a conclusion; it does not replace engineering judgment.
3. Objects, state stores, and trust boundaries
The following table organizes the key choices and evidence for Objects, state stores, and trust boundaries. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Object / store | What it owns | Diagnostic value | Boundary / risk |
|---|---|---|---|
| Test runner | Scenario ID, attempt, assertion, fixture lifecycle | Explains intent and where failure surfaced | May contain parameter values or CI metadata |
| WebDriver session | Session ID, capabilities, current browsing context | Correlates browser commands and remote Grid logs | Session ID is operational metadata; avoid exposing remote endpoints/tokens |
| Browser | Rendered viewport, console, JS runtime, resource timing | Shows client-side symptoms and rendering state | Screenshots/logs can expose PII, tokens, URLs, or form content |
| AUT | DOM state and business outcome | Separates application state from test assumptions | Page source can contain hidden values or user data |
| Network boundary | Requests, responses, timing, transport errors | Shows whether a request happened and how it failed | Headers/bodies are often the most sensitive evidence |
| Grid / driver | Node/session routing, driver process, queue state | Explains remote infrastructure failures | Grid endpoints and logs must not be publicly exposed |
| CI artifact store | Bundle retention and access policy | Enables remote triage after ephemeral runner exits | Retention and permissions become part of the security model |
4. Read-only inspection before changing evidence settings
Prove what session you actually have. Do not assume a log type, browser version, or remote/local topology from the test name.
import selenium
from selenium import webdriver
browser = webdriver.Chrome()
try:
browser.get("http://127.0.0.1:8781/?test_id=preflight")
caps = browser.capabilities
print("selenium", selenium.__version__)
print("session_id", browser.session_id)
print("browser", caps.get("browserName"), caps.get("browserVersion"))
print("platform", caps.get("platformName"))
print("url", browser.current_url)
print("title", browser.title)
print("log_types", browser.log_types)
finally:
browser.quit()
driver.log_types is the discovery mechanism: if a
browser does not expose a requested classic log type, the collector
must record that limitation instead of manufacturing an empty
“proof.” Returned capabilities are evidence of the negotiated
session, not a copy of what the test hoped to request.
5. What each evidence type can actually prove
The following table organizes the key choices and evidence for What each evidence type can actually prove. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Evidence | Good question | What it does not prove alone |
|---|---|---|
| Viewport screenshot | What was visible at capture time? | Why the state occurred, hidden DOM state, or network cause |
| Targeted DOM state | What stable element state/text/attribute did the AUT expose? | The server accepted a request or the browser had no JS error |
| Page source | What serialized document source is available now? | Current JS object state; also may over-collect sensitive data |
| Browser console / JS error | Did the client runtime log or throw? | That the application business assertion is wrong |
| Resource timing | Which resources were observed by the page and approximately how long? | Complete request/response headers/bodies or transport history |
| Server/network log | Did the request reach a controlled server boundary? | What the user visually saw afterward |
| WebDriver/Grid logs | Was the session routed/served correctly? | That the AUT business logic is correct |
| CI metadata | Which commit/job/shard/attempt produced the bundle? | The browser-side root cause |
6. Screenshots: evidence of pixels, not causality
driver.save_screenshot(path) stores a PNG of the
current window. Capture it before navigation,
teardown, retry, or browser closure can destroy the state. Use
unique paths per test ID and attempt. For privacy-sensitive pages, a
screenshot may need masking or may be prohibited entirely; that is a
test-data and artifact-policy decision, not a Selenium limitation.
from pathlib import Path
from datetime import datetime, timezone
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S.%fZ")
root = Path("evidence") / "checkout-guest" / f"{stamp}-attempt-1"
root.mkdir(parents=True, exist_ok=False)
driver.save_screenshot(str(root / "viewport.png"))
7. Prefer targeted state over indiscriminate page-source dumps
Page source can be useful, but collecting it by default expands privacy and storage cost. A targeted evidence schema is usually easier to review and less likely to contain hidden form fields or embedded data.
import json
from selenium.webdriver.common.by import By
status = driver.find_element(By.CSS_SELECTOR, '[data-testid="status"]')
state = {
"text": status.text,
"data_state": status.get_attribute("data-state"),
"url": driver.current_url,
"title": driver.title,
}
(root / "dom.json").write_text(json.dumps(state, indent=2), encoding="utf-8")
8. Browser logs, BiDi events, and network evidence are different layers
Python 4.47 exposes log_types and
get_log(type), but available classic logs are
browser/driver dependent. WebDriver BiDi is the standards-based
cross-browser direction for event streams such as console messages,
JavaScript errors, and network events. Because Chapter 18 is
dedicated to BiDi, this chapter uses classic logs only when
explicitly available and introduces BiDi as a preview rather than
making it a hidden prerequisite.
Headers, query strings, request bodies, response bodies, and HAR-like traces can contain credentials, session cookies, PII, and proprietary payloads. Capture the smallest field set needed for the hypothesis and redact before persistence.
9. Correlation is an evidence schema, not a filename afterthought
A production bundle should be attributable to a unique scenario and
attempt. Useful correlation fields are test_id, UTC
capture time, attempt, browser session ID, browser name/version,
platform, CI job/shard ID, and—when remote—the Grid session/node
correlation available to operators. Never assume wall-clock time
alone is unique under parallelism.
{
"test_id": "checkout-guest",
"attempt": 1,
"captured_at": "2026-08-28T04:30:15.123456Z",
"session_id": "<webdriver-session-id>",
"browser": "chrome",
"browser_version": "<returned capability>",
"ci_job": "job-4312",
"shard": "3-of-8"
}
10. Sensitivity, redaction, retention, and access
The following table organizes the key choices and evidence for Sensitivity, redaction, retention, and access. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Control | Minimum practice | Reason |
|---|---|---|
| Synthetic accounts/data | Use disposable identities and fake secrets | Reduces impact if an artifact is mishandled |
| Redaction | Sanitize URLs, headers, logs, and selected DOM fields before writing | Prevents known secrets from entering artifact storage |
| Retention | Define expiry by evidence class and incident need | Unlimited retention converts diagnostics into a data lake of risk |
| Access | Restrict CI artifacts to the people/systems that need them | Screenshots and traces can reveal internal application state |
| Manifest | Record files, hashes, session/test correlation | Makes missing/overwritten evidence detectable |
11. Local versus Grid/remote evidence ownership
Locally, the test runner, browser, fixture server, and artifact directory may share one machine. In Grid, the browser and driver may run on another node. A local path cannot be assumed to exist on the node; node logs may live outside the test workspace; and CI must preserve the runner-side bundle before the ephemeral job is destroyed. Correlate remote evidence by session ID rather than by “latest driver log.”
12. DevOps connection: make failure bundles triageable remotely
A useful pipeline does not simply upload a directory named
screenshots. It uploads a structured bundle after a
failed scenario, records its correlation metadata, preserves the
first attempt even if a controlled rerun occurs, applies a retention
policy, and links the artifact to the exact commit/job/shard. That
turns an intermittent red CI job into an incident another engineer
can analyze without recreating the runner.
13. Anti-patterns that destroy evidence
The following table organizes the key choices and evidence for Anti-patterns that destroy evidence. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Anti-pattern | Why it fails | Repair |
|---|---|---|
| Capture after cleanup | The original DOM/context may already be gone | Capture inside the failure boundary before teardown |
Write failure.png for every attempt |
Parallel jobs/retries overwrite evidence | Use test ID + attempt + timestamp/unique directory |
| Dump all logs/headers/bodies | High privacy/storage cost and noisy triage | Collect hypothesis-relevant fields with redaction |
| Treat console noise as root cause | Unrelated warnings are common | Correlate message, timestamp, test action, and observed failure |
| Retry then keep only the green attempt | Erases the original incident | Preserve attempt 1 immutably; label later observations separately |
14. Summary and bridge
Evidence engineering is the disciplined correlation of independent observations around one failure. Lesson 2 builds a loopback fixture and a privacy-aware collector that captures screenshot, URL/title, selected DOM state, supported browser logs, resource timing, and synthetic server evidence before teardown.
Knowledge check
Why is a screenshot insufficient to classify a failure?
It shows visible pixels at one moment but cannot independently distinguish application logic, client JavaScript, transport, WebDriver/Grid, or test-code causes.
What should you inspect before calling
get_log("browser")?
Inspect driver.log_types and the returned
session/browser capabilities; classic log availability is
browser/driver dependent.
Why include session ID in evidence metadata?
It correlates test-runner artifacts with the exact WebDriver session and, for remote execution, with Grid/node-side operational evidence.
What is the safer default: full page source or targeted DOM state?
Targeted DOM state, because it is more diagnostic and usually collects less hidden/sensitive data. Page source is an explicit escalation when needed.
A retry passes after attempt 1 failed. Which evidence is authoritative for the original incident?
Attempt 1 must remain preserved. The retry is an additional observation, not a replacement for the first failure.
Official references and version notes
- Selenium 4.47 release notes — stable binding/Grid baseline pinned for this chapter.
- Selenium downloads — current stable clients and Selenium Server/Grid.
- Selenium Python 4.47 API — supported Python versions and WebDriver API surface.
-
Chromium WebDriver API
—
get_log,log_types, screenshot APIs, and Chromium-specific capabilities. - WebDriver BiDi — W3C bidirectional event direction and enablement.
- WebDriver BiDi logging — current console-message and JavaScript-error handlers.
- WebDriver BiDi network — current network handler concepts; Chapter 18 covers these APIs in depth.
- Selenium Grid — remote session and node boundary.
- Grid getting started — local/free Grid execution and version alignment.
- Docker Selenium — official containerized Grid distribution; optional in this chapter.
Version-sensitive behavior was rechecked against Selenium primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python 4.47.0 and Python 3.10+, use a supported local
Chromium-family browser and Selenium Manager, and target only a
loopback synthetic AUT. Classic get_log() examples
explicitly discover log_types and do not claim
browser parity. WebDriver BiDi console/JavaScript-error APIs are
labeled as a preview because Chapter 18 is dedicated to BiDi; paid
browser-cloud videos/HARs, enterprise log platforms, and remote
Grid observability are optional architecture only. No production
target, real account, personal profile, TLS bypass, or real secret
is required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.