Screenshots, Logs, Network Evidence, and Failure Diagnostics: Configuration, Design Patterns, and Trade-Offs
An evidence collector that captures everything on every test can become slower, noisier, more expensive, and less safe than the failures it is supposed to diagnose. This lesson treats evidence as a configurable production subsystem with explicit trade-offs in diagnostic value, privacy, portability, runtime, and retention.
Learning objectives
- Select capture-on-failure, sampled, or always-capture policies based on diagnostic need and cost.
- Choose viewport, element, or browser-specific full-page screenshots without assuming pixel identity across environments.
- Compare classic browser logs with WebDriver BiDi event streams and state the portability boundary.
- Choose targeted DOM state, page source, resource timing, server logs, or HAR-like network data proportionally.
- Design an evidence retention and naming policy that remains safe under CI parallelism and retries.
1. Evidence is a policy surface
The capture API is the easy part. The hard part is deciding when, how much, and for how long. Evidence policy belongs in test infrastructure and CI governance, not scattered across page objects or hidden inside random assertions.
2. Capture-on-failure versus always capture
The following table organizes the key choices and evidence for Capture-on-failure versus always capture. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Policy | Best use | Cost / risk | Failure semantics |
|---|---|---|---|
| Failure only | Stable suites with good first-failure hooks | Low storage/runtime | Must capture before teardown and before retry |
| Always | Short critical flows, auditing of controlled synthetic labs | High artifact volume and privacy exposure | Pass artifacts can establish baseline context |
| Sampled | Large suites where full pass evidence is unnecessary | Requires explicit sampling metadata | Never sample away first failure |
| Escalating | Start targeted; capture deeper network/source only on specific failure classes | More collector logic | Best balance when hypotheses are classified |
3. Viewport, element, and full-page screenshots
The standard portable baseline is a current-window screenshot
through save_screenshot(). Element screenshots can
isolate a component. Some bindings/drivers expose browser-specific
full-page screenshot functions, but do not make them a cross-browser
contract unless the matrix verifies support. Full-page images also
increase storage and privacy exposure.
from pathlib import Path
from selenium.webdriver.common.by import By
root = Path("evidence/case-17")
root.mkdir(parents=True, exist_ok=True)
driver.save_screenshot(str(root / "viewport.png"))
card = driver.find_element(By.CSS_SELECTOR, '[data-testid="status"]')
card.screenshot(str(root / "status-element.png"))
4. Targeted DOM snippets versus page source
The following table organizes the key choices and evidence for Targeted DOM snippets versus page source. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Approach | Diagnostic strength | Privacy/storage | Recommendation |
|---|---|---|---|
| Targeted fields/attributes | High for known state contracts | Low | Default |
Selected element outerHTML |
Useful for locator/state investigation | Medium | Escalate locally |
Entire page_source |
Broad snapshot of serialized document | High | Explicit failure class only |
| JavaScript object dump | Can expose app internals | Potentially very high | Only a documented testability contract; never arbitrary secret-bearing globals |
5. Classic logs versus WebDriver BiDi
get_log() remains useful when the selected
browser/driver exposes the needed type. It is easy to integrate into
existing failure hooks, but log-type parity is not guaranteed.
Current WebDriver BiDi provides event-oriented console and
JavaScript-error handlers and is the standards-based cross-browser
direction. That stream model is better when timing and subscription
scope matter, but it introduces listener lifecycle and
browser-support prerequisites.
| Need | Classic log path | BiDi path |
|---|---|---|
| Post-failure console snapshot | get_log("browser") if exposed |
Event list retained by handler |
| Precise event timing | Depends on returned log records | Natural event stream |
| Cross-browser direction | Varies by log type/driver | W3C WebDriver BiDi |
| Lifecycle complexity | Low | Must enable BiDi and manage handler/subscription IDs |
| Chapter use | Mandatory local fallback | Preview here; full treatment in Chapter 18 |
6. HAR-like network evidence versus privacy/storage cost
Network evidence has a steep sensitivity curve. URL and status may be enough to explain a 503. Request headers can contain bearer tokens/cookies. Bodies can contain credentials, user content, and proprietary records. A “capture all HAR” default is therefore a security and storage decision, not merely a debugging convenience.
| Layer | Example fields | Use when | Redaction burden |
|---|---|---|---|
| Resource timing | URL, initiator type, duration | You need client-observed resource timing | Query strings |
| Server request log | Route, correlation ID, outcome | You control the fixture/service boundary | Headers/query/test data |
| BiDi network events | Request/response event metadata | You need browser-level timing/events | Headers, cookies, URLs |
| HAR/vendor trace | Broad request/response timeline | Incident requires deep network reconstruction | Highest; often contains sensitive payloads |
7. Retention and artifact classes
Classify artifacts by sensitivity and operational value. A synthetic screenshot from a training fixture can have a different retention period from an authenticated enterprise network trace. CI artifact expiration should be explicit, and incident preservation should be an auditable exception rather than “keep forever.”
8. Parallel-safe naming and immutable attempts
The following example makes the Parallel-safe naming and immutable attempts behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from datetime import datetime, timezone
from pathlib import Path
import uuid
def artifact_root(test_id, attempt, shard):
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S.%fZ")
nonce = uuid.uuid4().hex[:8]
return Path("evidence") / test_id / f"a{attempt}-{shard}-{stamp}-{nonce}"
root = artifact_root("cart-tax", 1, "s03")
root.mkdir(parents=True, exist_ok=False)
exist_ok=False turns a collision into a visible
infrastructure error instead of silently merging two bundles.
9. Worked decision table
The following table organizes the key choices and evidence for Worked decision table. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Failure hypothesis | Capture now | Do not default to | Reason |
|---|---|---|---|
| Wrong visible state | Viewport + targeted DOM + metadata | Full HAR | Start at the layer where symptom lives |
| Client JS exception | Targeted DOM + console/BiDi JS error + screenshot | Full page source on every pass | Correlate runtime exception with UI state |
| Request never completes | DOM state + server/network event + session metadata | Retrying before capture | Preserve original transport evidence |
| Remote session lost | Exception + session/capabilities + Grid/node/CI correlation | Screenshot-only diagnosis | Browsing context may already be unavailable |
| Locator regression | Exception + locator/DOM snippet + screenshot | All network bodies | Network payloads are unrelated and risky |
10. Keep configuration layers separate
Selenium browser options determine session capabilities such as BiDi enablement or browser-specific logging. The test framework decides failure hooks and assertion lifecycle. The AUT controls application logging/testability hooks. Grid controls remote routing/node logs. CI controls artifact upload/retention. A browser preference cannot fix a CI retention policy, and a test retry cannot fix a Grid node capacity problem.
11. Performance and capacity are causal
Evidence I/O adds real runtime: PNG encoding, filesystem writes, network trace processing, and CI upload consume CPU, disk, and bandwidth. On a large parallel Grid, “capture everything always” can itself extend test time and increase queue pressure. Measure artifact volume and upload duration separately from browser session startup and AUT latency.
12. Security/privacy design patterns
Minimize collection first, then redact. A regex cannot reliably recognize every secret, PII field, binary body, or proprietary payload.
Use synthetic accounts for automated flows, prevent CI secrets from entering URLs, avoid personal browser profiles, and review screenshots/download directories before artifact upload. If an enterprise proxy or authenticated browser cloud is optional, describe its evidence access/retention separately from the free local path.
13. Summary and bridge
A mature evidence policy is proportional, correlated, immutable per attempt, and privacy-aware. Lesson 4 stress-tests that policy against realistic failures: late capture, retry overwrites, secret leakage, noisy consoles, missing network context, parallel collisions, and Grid logs detached from session identity.
Knowledge check
When is “always capture everything” a poor default?
When pass volume, privacy exposure, I/O, artifact storage, or CI upload cost exceeds the diagnostic value. Failure-only or escalating capture is often better.
Why is WebDriver BiDi preferable as a long-term cross-browser event direction?
It is the W3C bidirectional standard designed for cross-browser event streams such as console, JavaScript errors, and network events; classic log availability varies by driver.
Why can page_source be riskier than targeted DOM
fields?
It can serialize far more page content—including hidden or unrelated sensitive values—than the hypothesis needs.
What should happen if two parallel tests compute the same artifact directory?
Fail visibly rather than merge/overwrite bundles; unique test/attempt/shard/time/nonce naming prevents ambiguity.
A remote session disappears before a screenshot can be captured. What evidence becomes primary?
The WebDriver exception plus session ID/capabilities, Grid/node correlation, and CI infrastructure evidence; a missing screenshot is itself consistent with a lost context.
Official references and version notes
- Selenium 4.47 release notes — stable binding/Grid baseline pinned for this chapter.
- Selenium downloads — current stable clients and Selenium Server/Grid.
- Selenium Python 4.47 API — supported Python versions and WebDriver API surface.
-
Chromium WebDriver API
—
get_log,log_types, screenshot APIs, and Chromium-specific capabilities. - WebDriver BiDi — W3C bidirectional event direction and enablement.
- WebDriver BiDi logging — current console-message and JavaScript-error handlers.
- WebDriver BiDi network — current network handler concepts; Chapter 18 covers these APIs in depth.
- Selenium Grid — remote session and node boundary.
- Grid getting started — local/free Grid execution and version alignment.
- Docker Selenium — official containerized Grid distribution; optional in this chapter.
Version-sensitive behavior was rechecked against Selenium primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python 4.47.0 and Python 3.10+, use a supported local
Chromium-family browser and Selenium Manager, and target only a
loopback synthetic AUT. Classic get_log() examples
explicitly discover log_types and do not claim
browser parity. WebDriver BiDi console/JavaScript-error APIs are
labeled as a preview because Chapter 18 is dedicated to BiDi; paid
browser-cloud videos/HARs, enterprise log platforms, and remote
Grid observability are optional architecture only. No production
target, real account, personal profile, TLS bypass, or real secret
is required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.