Chapter 14Lesson 03~185 minutes

Test Data, Parameterization, Fixtures, and Environment Configuration: Configuration, Design Patterns, and Trade-Offs

Once data and fixtures are explicit, the next question is where each input should live and how broad each matrix should be. This lesson compares design choices by maintenance cost, isolation, portability, diagnostic value, runtime, and security rather than by convenience alone.

Trade-offsFixture scopeSeeded dataConfig precedenceMatrix design

Learning objectives

  • Choose between inline data, external fixture files, and builders using change frequency and diagnostic needs.
  • Decide when UI setup is the test and when lower-layer setup is a better prerequisite mechanism.
  • Choose fixture scope deliberately instead of sharing mutable state for speed.
  • Use deterministic generation with recorded seeds/identities and separate non-secret config from secret injection.
  • Bound environment/browser/data matrix breadth using risk and CI capacity.

1. Start with the state that must be reproducible

Configuration design is not a preference contest. Identify the state that must be identical for a rerun, the state that may vary intentionally, and the state that must never be logged. Then choose the smallest mechanism that makes those boundaries explicit.

2. Inline rows, fixture files, or builders?

The following table organizes the key choices and evidence for Inline rows, fixture files, or builders?. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Approach Good fit Trade-off
inline rows small behavior tables with a few obvious fields easy to read; noisy when payloads become large
JSON/CSV fixture file large static datasets maintained separately portable but IDs/schema validation become important
builder/factory many related synthetic variants and defaults expressive but can hide surprising generated values unless evidence is recorded

The following example makes the Inline rows, fixture files, or builders? behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from dataclasses import dataclass, replace

@dataclass(frozen=True)
class UserCase:
    id: str
    name: str = "Synthetic User"
    role: str = "viewer"

BASE = UserCase(id="viewer-valid")
ADMIN = replace(BASE, id="admin-valid", role="admin")

Immutable builders make intentional variation visible. Avoid a factory that silently chooses a random role, locale, or identity unless the seed and generated values are part of evidence.

3. UI setup versus API-assisted setup

The following table organizes the key choices and evidence for UI setup versus API-assisted setup. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Question UI setup API/helper setup
is setup behavior itself under test? yes usually no
speed/stability for repeated prerequisites lower higher
real user-path coverage high not the goal
failure surface browser + DOM + network + AUT usually API + AUT
cleanup suitability often fragile usually better if idempotent

4. Per-test isolation versus shared expensive setup

A static loopback server can be shared if its mutable records are reset per case; a browser session should normally remain per-test. If a real Grid startup is expensive, sharing the Grid is different from sharing a browser session. Infrastructure can be broad-scoped while mutable test state stays narrow-scoped.

Resource Reasonable scope Required control
Grid service suite/session infrastructure capacity/health monitoring; no shared browser session
static app server module/suite case-specific data reset
WebDriver session test/case fresh profile/session
mutable synthetic record test/case unique identity + teardown
read-only reference dataset suite immutable/versioned

5. Random data versus seeded deterministic data

The following example makes the Random data versus seeded deterministic data behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

import random

seed = 20260828
rng = random.Random(seed)
case = {
    "seed": seed,
    "quantity": rng.randint(1, 5),
    "tier": rng.choice(["basic", "pro"]),
}
print(case)

Seeded generation is useful when exploring a bounded data space, but deterministic explicit rows are often clearer for core regression cases. If a generated case fails, preserve both the seed and generated inputs.

6. Environment variables versus config files

The following table organizes the key choices and evidence for Environment variables versus config files. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Mechanism Use Avoid
environment variable CI-injected target/browser/feature toggle/secret reference printing whole environment
versioned config file non-secret named environment defaults embedding passwords/tokens
CLI argument one-off local override with visible invocation making every test parse ad-hoc flags
secret store/injected secret credentials committing value to test-data files

A useful precedence rule is explicit and documented, for example: CLI override → environment variable → non-secret config file → safe local default. Do not create hidden fallback from “staging missing” to production.

7. Matrix breadth versus runtime/capacity

Parameterization multiplies quickly. Three browsers × two environments × five roles × four locales is 120 executions before retries. More rows are not automatically more confidence. Use risk to decide which dimensions belong in every pull request and which belong in scheduled or pre-release suites.

Layer Example matrix Purpose
PR smoke 1 browser × local/staging simulator × critical roles fast regression signal
cross-browser gate Chrome/Firefox/Edge × critical cases browser compatibility
scheduled breadth more locales/data variants broader exploration without blocking every commit

8. Keep Selenium/browser config separate from AUT/test data

SELENIUM_BROWSER=firefox changes the browser implementation. TEST_ENV=local selects an environment. role=admin changes a test parameter. A proxy or certificate setting changes browser/network policy. A Grid URL changes transport infrastructure. They may all be present in one CI job, but they are different control planes and should not be collapsed into one “config” dictionary passed everywhere.

9. Worked decision table

The following table organizes the key choices and evidence for Worked decision table. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Scenario Recommended design Why
registration validation for 3 roles inline named rows + per-case browser + UI registration small table; behavior itself under test
checkout requires existing catalog items API seed catalog + browser checkout catalog creation is prerequisite, not checkout signal
large immutable country list versioned fixture file shared read-only reference data
explore boundary quantities nightly seeded generator + seed evidence reproducible exploration
production password never ordinary fixture/config data secret boundary

10. CI reliability and evidence

Store the case ID, environment name, Selenium version, returned browser/version, session ID, seed if any, and sanitized target URL. Do not persist credentials, personal profiles, or unrelated environment variables. If a matrix case fails, rerun the smallest identical tuple before broadening the experiment.

11. Summary and next step

Choose data/configuration patterns by state ownership and reproducibility: inline rows for small tables, files for large static data, builders for structured variation, APIs for repetitive prerequisites, narrow fixture scope for mutable state, and explicit environment precedence.

Knowledge check

When is a shared suite-level fixture acceptable?

Why can a builder be worse than inline data?

What is wrong with falling back to production when TEST_ENV is absent?

Why not run every data variant in every browser on every commit?

What should be recorded for seeded generated data?

Next lesson

Diagnose configuration and data failures

Lesson 4 engineers shared accounts, order dependence, unseeded randomness, skipped teardown, secret leakage, hard-coded targets, unsafe target selection, and opaque parameter IDs.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against Selenium primary documentation on 2026-08-28. Mandatory examples pin Selenium Python 4.47.0 and Python 3.10+, use Python standard-library unittest for parameterization/fixture examples, use a supported locally installed Chromium-family browser with Selenium Manager, and target only loopback synthetic applications. Parameterization, fixture scope, environment parsing, and secret injection are test-runner/application-infrastructure concerns rather than WebDriver capabilities. The mandatory path deliberately uses Python standard-library unittest rather than requiring pytest. Selenium guidance is applied by keeping tests independent, using fresh browser sessions, and moving repetitive prerequisite/reset work to a disposable lower-layer API when that setup is not the UI behavior under test.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.