Screenshots, Logs, Network Evidence, and Failure Diagnostics: Diagnostics, Failure Modes, and Production Practices
Diagnostic evidence can fail just like test code. A collector that runs after teardown, overwrites the first attempt, leaks tokens, or detaches driver logs from the session can actively mislead incident response. This lesson engineers those failure modes deliberately and repairs the evidence system without hiding the original cause.
Learning objectives
- Apply a repeatable first-failure diagnostic sequence before changing retries, waits, browsers, or Grid state.
- Recognize and repair late capture, artifact overwrite, secret leakage, console-noise, incomplete-network, and parallel-collision failures.
- Interpret an intentionally broken evidence collector without swallowing its WebDriver exception.
- Correlate driver/Grid/CI evidence with the exact session ID and attempt.
- Separate evidence I/O cost from browser startup, Grid queueing, AUT latency, and retry cost.
1. Use the same diagnostic sequence every time
- Preserve first-failure evidence.
- Confirm Selenium/binding/browser/driver/Grid versions.
- Confirm target environment and synthetic test data.
- Inspect session ID, returned capabilities, window/frame context, URL/title.
- Inspect locator, element lifetime, and synchronization state.
- Inspect AUT state plus browser/console/network evidence.
- If remote, correlate Grid queue/node/session evidence and CI resource state.
- Apply the least destructive correction.
- Rerun the smallest controlled scenario as a new observation.
2. Failure mode: screenshot captured after cleanup
This intentionally broken code destroys the browsing context before evidence capture:
from selenium import webdriver
from selenium.common.exceptions import InvalidSessionIdException, WebDriverException
browser = webdriver.Chrome()
browser.get("http://127.0.0.1:8781/?test_id=late-capture")
browser.quit()
try:
browser.save_screenshot("evidence/late.png")
except (InvalidSessionIdException, WebDriverException) as exc:
print(type(exc).__name__, str(exc))
raise # preserve the collector failure; do not pretend evidence exists
The repair is lifecycle ordering: capture inside the
assertion/exception boundary, then teardown in finally.
Do not restart the browser and take a screenshot of a new session;
that would be evidence from the wrong state.
3. Failure mode: retry overwrites the first failure
A single filename such as failure.png is not retry-safe
or parallel-safe. A second attempt may replace the only evidence of
the original incident.
from pathlib import Path
# Broken: both attempts target the same file.
path = Path("evidence/failure.png")
for attempt in (1, 2):
driver.save_screenshot(str(path))
# Repair: immutable attempt directories.
for attempt in (1, 2):
root = Path("evidence") / "checkout" / f"attempt-{attempt}"
root.mkdir(parents=True, exist_ok=False)
driver.save_screenshot(str(root / "viewport.png"))
4. Failure mode: tokens or PII stored in logs/URLs
A collector that serializes driver.current_url, browser
logs, headers, or page source verbatim can move a secret from
ephemeral browser state into a long-lived artifact store. The
correction is
minimize first, redact second, verify third.
from urllib.parse import urlsplit, parse_qsl, urlencode, urlunsplit
def safe_url(url):
parts = urlsplit(url)
clean = []
for key, value in parse_qsl(parts.query, keep_blank_values=True):
clean.append((key, "<redacted>" if key.lower() in {"token", "secret", "key"} else value))
return urlunsplit((parts.scheme, parts.netloc, parts.path, urlencode(clean), parts.fragment))
5. Failure mode: console noise mistaken for root cause
A console can contain extension noise, deprecation warnings, failed
analytics, or previous unrelated messages. A
SEVERE record is a hypothesis, not automatically the
root cause. Correlate its timestamp/message to the action and to an
observable AUT failure. In BiDi, scope handlers to the
scenario/context and remove them when done.
6. Failure mode: network trace lacks the context you need
A URL/status-only trace may prove that a request occurred but not why the server rejected it. Conversely, capturing all request/response bodies can be unjustifiably sensitive. Escalate deliberately: start with route, correlation ID, status/outcome, timing, and server-side synthetic logs; add safe header/body fields only when the hypothesis requires them.
7. Failure mode: artifact names collide under parallelism
Parallel shards often execute the same test name simultaneously. Add shard/worker/attempt plus a UTC timestamp or nonce, and create directories with collision failure enabled. Do not let “last writer wins” become your evidence policy.
8. Failure mode: driver/Grid logs detached from the session
“The node log has an error around 12:00” is weak evidence on a busy
Grid. Store the WebDriver session_id in runner
artifacts and use it to query/locate Grid or vendor-side logs
according to your platform. Preserve node identity/job ID when
available. The test runner should never expose a public Grid
endpoint or store its authentication token in the bundle.
metadata = {
"session_id": driver.session_id,
"browser": driver.capabilities.get("browserName"),
"browser_version": driver.capabilities.get("browserVersion"),
"platform": driver.capabilities.get("platformName"),
"ci_job": "${CI_JOB_ID or local}",
"attempt": 1,
}
# Serialize only sanitized metadata; do not include remote credentials/endpoints.
9. Do not use troubleshooting shortcuts that destroy evidence
The following table organizes the key choices and evidence for Do not use troubleshooting shortcuts that destroy evidence. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Shortcut | Why it is destructive | Least-destructive correction |
|---|---|---|
| Blanket retry | Changes the observation before diagnosis | Capture attempt 1, classify, rerun only as an explicit experiment |
| Giant sleep/timeout | Hides timing symptom and wastes capacity | Wait on the actual state contract |
| JavaScript force-click | Bypasses interactability and user semantics | Diagnose obstruction/readiness |
| Browser/Grid restart | Destroys session/node state | Preserve logs/session evidence first; restart only after layer isolation |
| Disable TLS | Changes security semantics | Fix trust/certificate test environment |
| Production experiment | Risks real data/users | Use the disposable local fixture or an approved synthetic environment |
10. Separate performance costs by layer
The following table organizes the key choices and evidence for Separate performance costs by layer. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Cost | Measure separately | Why |
|---|---|---|
| Test-runner overhead | Serialization, hashing, redaction CPU | Can grow with bundle size |
| Browser/session startup | Driver/browser creation time | Unrelated to screenshot write time |
| Grid queue/capacity | Queue wait and node saturation | Can dominate remote latency |
| AUT/network latency | Request/response timing | Product/environment signal |
| Evidence I/O | PNG/log write and artifact upload duration | Evidence policy can slow CI |
| Retry cost | Additional full/partial attempts | Multiplies all layers and can normalize flakiness |
11. A failure-safe capture boundary
The following example makes the A failure-safe capture boundary behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from pathlib import Path
root = Path("evidence") / "case-42" / "attempt-1"
root.mkdir(parents=True, exist_ok=False)
try:
# scenario actions
assert driver.find_element("css selector", '[data-testid="status"]').text == "healthy"
except Exception as original:
try:
driver.save_screenshot(str(root / "viewport.png"))
(root / "url.txt").write_text(driver.current_url, encoding="utf-8")
except Exception as capture_error:
(root / "capture-error.txt").write_text(
f"{type(capture_error).__name__}: {capture_error}", encoding="utf-8"
)
raise original
finally:
driver.quit()
The collector may fail too. Record that secondary failure without replacing the original exception. In production code, sanitize URLs/text before writing them.
12. Security-sensitive evidence operations
Screenshots, downloads, browser profiles, cookies, proxy/TLS data, request headers, CI variables, and Grid logs can all contain secrets or PII. This course uses fake credentials and loopback targets only.
Do not copy personal browser profiles into CI, do not upload raw profile directories as artifacts, and do not retain browser-cloud video/network traces without reviewing provider retention and access controls.
13. Minimal incident runbook
- Freeze attempt-1 bundle and record test/session/job correlation.
- Classify symptom layer: assertion, browser, network, WebDriver/Grid, or test code.
- Check whether evidence is complete and sanitized.
- Form the smallest falsifiable hypothesis.
- Rerun only the smallest controlled scenario, preserving it as attempt 2.
- Repair the failing layer; do not weaken the oracle.
- Update evidence policy only if the incident revealed a systematic blind spot.
14. Summary and bridge
A diagnostic system is trustworthy only when its own failure modes are visible. Lesson 5 combines the chapter into a three-case checkpoint: application error, JavaScript error, and network-like failure, each with correlated, redacted evidence and a written conclusion that does not rely on screenshots alone.
Knowledge check
Why is restarting the browser before capture a bad repair for a lost screenshot?
The new browser is a different session/state. Its screenshot is not evidence of the original failure.
A retry passes and failure.png now shows the green
state. What evidence-system bug occurred?
The retry overwrote first-failure evidence because filenames/attempt directories were not immutable.
Why is a severe console message not automatically the root cause?
Console logs can contain unrelated/noisy entries. Correlate the message to the action, time window, context, and observed AUT failure.
What identifier should connect runner artifacts to Grid/node logs?
The exact WebDriver session ID, plus CI job/attempt and node/vendor correlation when available.
If evidence collection throws after the original assertion failed, which exception should drive the test result?
The original scenario failure. Record the capture error as secondary evidence without replacing or swallowing the first failure.
Official references and version notes
- Selenium 4.47 release notes — stable binding/Grid baseline pinned for this chapter.
- Selenium downloads — current stable clients and Selenium Server/Grid.
- Selenium Python 4.47 API — supported Python versions and WebDriver API surface.
-
Chromium WebDriver API
—
get_log,log_types, screenshot APIs, and Chromium-specific capabilities. - WebDriver BiDi — W3C bidirectional event direction and enablement.
- WebDriver BiDi logging — current console-message and JavaScript-error handlers.
- WebDriver BiDi network — current network handler concepts; Chapter 18 covers these APIs in depth.
- Selenium Grid — remote session and node boundary.
- Grid getting started — local/free Grid execution and version alignment.
- Docker Selenium — official containerized Grid distribution; optional in this chapter.
Version-sensitive behavior was rechecked against Selenium primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python 4.47.0 and Python 3.10+, use a supported local
Chromium-family browser and Selenium Manager, and target only a
loopback synthetic AUT. Classic get_log() examples
explicitly discover log_types and do not claim
browser parity. WebDriver BiDi console/JavaScript-error APIs are
labeled as a preview because Chapter 18 is dedicated to BiDi; paid
browser-cloud videos/HARs, enterprise log platforms, and remote
Grid observability are optional architecture only. No production
target, real account, personal profile, TLS bypass, or real secret
is required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.