Chapter 17Lesson 04~215 minutes

Screenshots, Logs, Network Evidence, and Failure Diagnostics: Diagnostics, Failure Modes, and Production Practices

Diagnostic evidence can fail just like test code. A collector that runs after teardown, overwrites the first attempt, leaks tokens, or detaches driver logs from the session can actively mislead incident response. This lesson engineers those failure modes deliberately and repairs the evidence system without hiding the original cause.

DiagnosticsFailure modesGrid correlationEvidence integritySecurity

Learning objectives

  • Apply a repeatable first-failure diagnostic sequence before changing retries, waits, browsers, or Grid state.
  • Recognize and repair late capture, artifact overwrite, secret leakage, console-noise, incomplete-network, and parallel-collision failures.
  • Interpret an intentionally broken evidence collector without swallowing its WebDriver exception.
  • Correlate driver/Grid/CI evidence with the exact session ID and attempt.
  • Separate evidence I/O cost from browser startup, Grid queueing, AUT latency, and retry cost.

1. Use the same diagnostic sequence every time

  1. Preserve first-failure evidence.
  2. Confirm Selenium/binding/browser/driver/Grid versions.
  3. Confirm target environment and synthetic test data.
  4. Inspect session ID, returned capabilities, window/frame context, URL/title.
  5. Inspect locator, element lifetime, and synchronization state.
  6. Inspect AUT state plus browser/console/network evidence.
  7. If remote, correlate Grid queue/node/session evidence and CI resource state.
  8. Apply the least destructive correction.
  9. Rerun the smallest controlled scenario as a new observation.

2. Failure mode: screenshot captured after cleanup

This intentionally broken code destroys the browsing context before evidence capture:

from selenium import webdriver
from selenium.common.exceptions import InvalidSessionIdException, WebDriverException

browser = webdriver.Chrome()
browser.get("http://127.0.0.1:8781/?test_id=late-capture")
browser.quit()
try:
    browser.save_screenshot("evidence/late.png")
except (InvalidSessionIdException, WebDriverException) as exc:
    print(type(exc).__name__, str(exc))
    raise  # preserve the collector failure; do not pretend evidence exists

The repair is lifecycle ordering: capture inside the assertion/exception boundary, then teardown in finally. Do not restart the browser and take a screenshot of a new session; that would be evidence from the wrong state.

3. Failure mode: retry overwrites the first failure

A single filename such as failure.png is not retry-safe or parallel-safe. A second attempt may replace the only evidence of the original incident.

from pathlib import Path

# Broken: both attempts target the same file.
path = Path("evidence/failure.png")
for attempt in (1, 2):
    driver.save_screenshot(str(path))

# Repair: immutable attempt directories.
for attempt in (1, 2):
    root = Path("evidence") / "checkout" / f"attempt-{attempt}"
    root.mkdir(parents=True, exist_ok=False)
    driver.save_screenshot(str(root / "viewport.png"))

4. Failure mode: tokens or PII stored in logs/URLs

A collector that serializes driver.current_url, browser logs, headers, or page source verbatim can move a secret from ephemeral browser state into a long-lived artifact store. The correction is minimize first, redact second, verify third.

from urllib.parse import urlsplit, parse_qsl, urlencode, urlunsplit

def safe_url(url):
    parts = urlsplit(url)
    clean = []
    for key, value in parse_qsl(parts.query, keep_blank_values=True):
        clean.append((key, "<redacted>" if key.lower() in {"token", "secret", "key"} else value))
    return urlunsplit((parts.scheme, parts.netloc, parts.path, urlencode(clean), parts.fragment))

5. Failure mode: console noise mistaken for root cause

A console can contain extension noise, deprecation warnings, failed analytics, or previous unrelated messages. A SEVERE record is a hypothesis, not automatically the root cause. Correlate its timestamp/message to the action and to an observable AUT failure. In BiDi, scope handlers to the scenario/context and remove them when done.

6. Failure mode: network trace lacks the context you need

A URL/status-only trace may prove that a request occurred but not why the server rejected it. Conversely, capturing all request/response bodies can be unjustifiably sensitive. Escalate deliberately: start with route, correlation ID, status/outcome, timing, and server-side synthetic logs; add safe header/body fields only when the hypothesis requires them.

7. Failure mode: artifact names collide under parallelism

Parallel shards often execute the same test name simultaneously. Add shard/worker/attempt plus a UTC timestamp or nonce, and create directories with collision failure enabled. Do not let “last writer wins” become your evidence policy.

8. Failure mode: driver/Grid logs detached from the session

“The node log has an error around 12:00” is weak evidence on a busy Grid. Store the WebDriver session_id in runner artifacts and use it to query/locate Grid or vendor-side logs according to your platform. Preserve node identity/job ID when available. The test runner should never expose a public Grid endpoint or store its authentication token in the bundle.

metadata = {
    "session_id": driver.session_id,
    "browser": driver.capabilities.get("browserName"),
    "browser_version": driver.capabilities.get("browserVersion"),
    "platform": driver.capabilities.get("platformName"),
    "ci_job": "${CI_JOB_ID or local}",
    "attempt": 1,
}
# Serialize only sanitized metadata; do not include remote credentials/endpoints.

9. Do not use troubleshooting shortcuts that destroy evidence

The following table organizes the key choices and evidence for Do not use troubleshooting shortcuts that destroy evidence. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Shortcut Why it is destructive Least-destructive correction
Blanket retry Changes the observation before diagnosis Capture attempt 1, classify, rerun only as an explicit experiment
Giant sleep/timeout Hides timing symptom and wastes capacity Wait on the actual state contract
JavaScript force-click Bypasses interactability and user semantics Diagnose obstruction/readiness
Browser/Grid restart Destroys session/node state Preserve logs/session evidence first; restart only after layer isolation
Disable TLS Changes security semantics Fix trust/certificate test environment
Production experiment Risks real data/users Use the disposable local fixture or an approved synthetic environment

10. Separate performance costs by layer

The following table organizes the key choices and evidence for Separate performance costs by layer. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Cost Measure separately Why
Test-runner overhead Serialization, hashing, redaction CPU Can grow with bundle size
Browser/session startup Driver/browser creation time Unrelated to screenshot write time
Grid queue/capacity Queue wait and node saturation Can dominate remote latency
AUT/network latency Request/response timing Product/environment signal
Evidence I/O PNG/log write and artifact upload duration Evidence policy can slow CI
Retry cost Additional full/partial attempts Multiplies all layers and can normalize flakiness

11. A failure-safe capture boundary

The following example makes the A failure-safe capture boundary behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from pathlib import Path

root = Path("evidence") / "case-42" / "attempt-1"
root.mkdir(parents=True, exist_ok=False)
try:
    # scenario actions
    assert driver.find_element("css selector", '[data-testid="status"]').text == "healthy"
except Exception as original:
    try:
        driver.save_screenshot(str(root / "viewport.png"))
        (root / "url.txt").write_text(driver.current_url, encoding="utf-8")
    except Exception as capture_error:
        (root / "capture-error.txt").write_text(
            f"{type(capture_error).__name__}: {capture_error}", encoding="utf-8"
        )
    raise original
finally:
    driver.quit()

The collector may fail too. Record that secondary failure without replacing the original exception. In production code, sanitize URLs/text before writing them.

12. Security-sensitive evidence operations

Treat evidence like application data

Screenshots, downloads, browser profiles, cookies, proxy/TLS data, request headers, CI variables, and Grid logs can all contain secrets or PII. This course uses fake credentials and loopback targets only.

Do not copy personal browser profiles into CI, do not upload raw profile directories as artifacts, and do not retain browser-cloud video/network traces without reviewing provider retention and access controls.

13. Minimal incident runbook

  1. Freeze attempt-1 bundle and record test/session/job correlation.
  2. Classify symptom layer: assertion, browser, network, WebDriver/Grid, or test code.
  3. Check whether evidence is complete and sanitized.
  4. Form the smallest falsifiable hypothesis.
  5. Rerun only the smallest controlled scenario, preserving it as attempt 2.
  6. Repair the failing layer; do not weaken the oracle.
  7. Update evidence policy only if the incident revealed a systematic blind spot.

14. Summary and bridge

A diagnostic system is trustworthy only when its own failure modes are visible. Lesson 5 combines the chapter into a three-case checkpoint: application error, JavaScript error, and network-like failure, each with correlated, redacted evidence and a written conclusion that does not rely on screenshots alone.

Knowledge check

Why is restarting the browser before capture a bad repair for a lost screenshot?

A retry passes and failure.png now shows the green state. What evidence-system bug occurred?

Why is a severe console message not automatically the root cause?

What identifier should connect runner artifacts to Grid/node logs?

If evidence collection throws after the original assertion failed, which exception should drive the test result?

Next lesson

Checkpoint Lab — Screenshots, Logs, Network Evidence, and Failure Diagnostics

Continue with Checkpoint Lab — Screenshots, Logs, Network Evidence, and Failure Diagnostics. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against Selenium primary documentation on 2026-08-28. Mandatory examples pin Selenium Python 4.47.0 and Python 3.10+, use a supported local Chromium-family browser and Selenium Manager, and target only a loopback synthetic AUT. Classic get_log() examples explicitly discover log_types and do not claim browser parity. WebDriver BiDi console/JavaScript-error APIs are labeled as a preview because Chapter 18 is dedicated to BiDi; paid browser-cloud videos/HARs, enterprise log platforms, and remote Grid observability are optional architecture only. No production target, real account, personal profile, TLS bypass, or real secret is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.