Chapter 17Lesson 03~200 minutes

Screenshots, Logs, Network Evidence, and Failure Diagnostics: Configuration, Design Patterns, and Trade-Offs

An evidence collector that captures everything on every test can become slower, noisier, more expensive, and less safe than the failures it is supposed to diagnose. This lesson treats evidence as a configurable production subsystem with explicit trade-offs in diagnostic value, privacy, portability, runtime, and retention.

Evidence policyBiDi vs classic logsRetentionCI artifactsTrade-offs

Learning objectives

  • Select capture-on-failure, sampled, or always-capture policies based on diagnostic need and cost.
  • Choose viewport, element, or browser-specific full-page screenshots without assuming pixel identity across environments.
  • Compare classic browser logs with WebDriver BiDi event streams and state the portability boundary.
  • Choose targeted DOM state, page source, resource timing, server logs, or HAR-like network data proportionally.
  • Design an evidence retention and naming policy that remains safe under CI parallelism and retries.

1. Evidence is a policy surface

The capture API is the easy part. The hard part is deciding when, how much, and for how long. Evidence policy belongs in test infrastructure and CI governance, not scattered across page objects or hidden inside random assertions.

2. Capture-on-failure versus always capture

The following table organizes the key choices and evidence for Capture-on-failure versus always capture. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Policy Best use Cost / risk Failure semantics
Failure only Stable suites with good first-failure hooks Low storage/runtime Must capture before teardown and before retry
Always Short critical flows, auditing of controlled synthetic labs High artifact volume and privacy exposure Pass artifacts can establish baseline context
Sampled Large suites where full pass evidence is unnecessary Requires explicit sampling metadata Never sample away first failure
Escalating Start targeted; capture deeper network/source only on specific failure classes More collector logic Best balance when hypotheses are classified

3. Viewport, element, and full-page screenshots

The standard portable baseline is a current-window screenshot through save_screenshot(). Element screenshots can isolate a component. Some bindings/drivers expose browser-specific full-page screenshot functions, but do not make them a cross-browser contract unless the matrix verifies support. Full-page images also increase storage and privacy exposure.

from pathlib import Path
from selenium.webdriver.common.by import By

root = Path("evidence/case-17")
root.mkdir(parents=True, exist_ok=True)
driver.save_screenshot(str(root / "viewport.png"))
card = driver.find_element(By.CSS_SELECTOR, '[data-testid="status"]')
card.screenshot(str(root / "status-element.png"))

4. Targeted DOM snippets versus page source

The following table organizes the key choices and evidence for Targeted DOM snippets versus page source. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Approach Diagnostic strength Privacy/storage Recommendation
Targeted fields/attributes High for known state contracts Low Default
Selected element outerHTML Useful for locator/state investigation Medium Escalate locally
Entire page_source Broad snapshot of serialized document High Explicit failure class only
JavaScript object dump Can expose app internals Potentially very high Only a documented testability contract; never arbitrary secret-bearing globals

5. Classic logs versus WebDriver BiDi

get_log() remains useful when the selected browser/driver exposes the needed type. It is easy to integrate into existing failure hooks, but log-type parity is not guaranteed. Current WebDriver BiDi provides event-oriented console and JavaScript-error handlers and is the standards-based cross-browser direction. That stream model is better when timing and subscription scope matter, but it introduces listener lifecycle and browser-support prerequisites.

Need Classic log path BiDi path
Post-failure console snapshot get_log("browser") if exposed Event list retained by handler
Precise event timing Depends on returned log records Natural event stream
Cross-browser direction Varies by log type/driver W3C WebDriver BiDi
Lifecycle complexity Low Must enable BiDi and manage handler/subscription IDs
Chapter use Mandatory local fallback Preview here; full treatment in Chapter 18

6. HAR-like network evidence versus privacy/storage cost

Network evidence has a steep sensitivity curve. URL and status may be enough to explain a 503. Request headers can contain bearer tokens/cookies. Bodies can contain credentials, user content, and proprietary records. A “capture all HAR” default is therefore a security and storage decision, not merely a debugging convenience.

Layer Example fields Use when Redaction burden
Resource timing URL, initiator type, duration You need client-observed resource timing Query strings
Server request log Route, correlation ID, outcome You control the fixture/service boundary Headers/query/test data
BiDi network events Request/response event metadata You need browser-level timing/events Headers, cookies, URLs
HAR/vendor trace Broad request/response timeline Incident requires deep network reconstruction Highest; often contains sensitive payloads

7. Retention and artifact classes

Classify artifacts by sensitivity and operational value. A synthetic screenshot from a training fixture can have a different retention period from an authenticated enterprise network trace. CI artifact expiration should be explicit, and incident preservation should be an auditable exception rather than “keep forever.”

8. Parallel-safe naming and immutable attempts

The following example makes the Parallel-safe naming and immutable attempts behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from datetime import datetime, timezone
from pathlib import Path
import uuid

def artifact_root(test_id, attempt, shard):
    stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S.%fZ")
    nonce = uuid.uuid4().hex[:8]
    return Path("evidence") / test_id / f"a{attempt}-{shard}-{stamp}-{nonce}"

root = artifact_root("cart-tax", 1, "s03")
root.mkdir(parents=True, exist_ok=False)

exist_ok=False turns a collision into a visible infrastructure error instead of silently merging two bundles.

9. Worked decision table

The following table organizes the key choices and evidence for Worked decision table. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Failure hypothesis Capture now Do not default to Reason
Wrong visible state Viewport + targeted DOM + metadata Full HAR Start at the layer where symptom lives
Client JS exception Targeted DOM + console/BiDi JS error + screenshot Full page source on every pass Correlate runtime exception with UI state
Request never completes DOM state + server/network event + session metadata Retrying before capture Preserve original transport evidence
Remote session lost Exception + session/capabilities + Grid/node/CI correlation Screenshot-only diagnosis Browsing context may already be unavailable
Locator regression Exception + locator/DOM snippet + screenshot All network bodies Network payloads are unrelated and risky

10. Keep configuration layers separate

Selenium browser options determine session capabilities such as BiDi enablement or browser-specific logging. The test framework decides failure hooks and assertion lifecycle. The AUT controls application logging/testability hooks. Grid controls remote routing/node logs. CI controls artifact upload/retention. A browser preference cannot fix a CI retention policy, and a test retry cannot fix a Grid node capacity problem.

11. Performance and capacity are causal

Evidence I/O adds real runtime: PNG encoding, filesystem writes, network trace processing, and CI upload consume CPU, disk, and bandwidth. On a large parallel Grid, “capture everything always” can itself extend test time and increase queue pressure. Measure artifact volume and upload duration separately from browser session startup and AUT latency.

12. Security/privacy design patterns

Never treat redaction as permission to collect arbitrary data

Minimize collection first, then redact. A regex cannot reliably recognize every secret, PII field, binary body, or proprietary payload.

Use synthetic accounts for automated flows, prevent CI secrets from entering URLs, avoid personal browser profiles, and review screenshots/download directories before artifact upload. If an enterprise proxy or authenticated browser cloud is optional, describe its evidence access/retention separately from the free local path.

13. Summary and bridge

A mature evidence policy is proportional, correlated, immutable per attempt, and privacy-aware. Lesson 4 stress-tests that policy against realistic failures: late capture, retry overwrites, secret leakage, noisy consoles, missing network context, parallel collisions, and Grid logs detached from session identity.

Knowledge check

When is “always capture everything” a poor default?

Why is WebDriver BiDi preferable as a long-term cross-browser event direction?

Why can page_source be riskier than targeted DOM fields?

What should happen if two parallel tests compute the same artifact directory?

A remote session disappears before a screenshot can be captured. What evidence becomes primary?

Next lesson

Screenshots, Logs, Network Evidence, and Failure Diagnostics: Diagnostics, Failure Modes, and Production Practices

Continue with Screenshots, Logs, Network Evidence, and Failure Diagnostics: Diagnostics, Failure Modes, and Production Practices. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against Selenium primary documentation on 2026-08-28. Mandatory examples pin Selenium Python 4.47.0 and Python 3.10+, use a supported local Chromium-family browser and Selenium Manager, and target only a loopback synthetic AUT. Classic get_log() examples explicitly discover log_types and do not claim browser parity. WebDriver BiDi console/JavaScript-error APIs are labeled as a preview because Chapter 18 is dedicated to BiDi; paid browser-cloud videos/HARs, enterprise log platforms, and remote Grid observability are optional architecture only. No production target, real account, personal profile, TLS bypass, or real secret is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.