Chapter 11Lesson 04200–280 min

Data-Driven Testing, Templates, Embedded Arguments, and Variants: Diagnostics, Failure Modes, and Production Practices

Diagnose matrix explosion, heterogeneous templates, vague row identity, hidden conversion, false-green expected-failure patterns, non-reproducible generated data, and secret leakage without masking the first failure.

DiagnosticsMatrix explosionFailure identityReproducibilitySecurity

Learning objectives

  • Preserve the first failing data row and result artifacts before editing the matrix.
  • Separate parser/import/argument-conversion failures from template assertion failures.
  • Recognize false-green expected-failure and broad status-swallowing patterns.
  • Reduce non-causal matrix dimensions and make generated data reproducible.
  • Keep secrets and production data out of template source and diagnostic artifacts.

Current compatibility baseline. Verified 2026-08-31: Robot Framework 7.4.2 is the current stable release and requires Python 3.8+; 7.5b1 is a pre-release and is not required in this chapter. In 7.4.2, templates can use normal, named, and embedded arguments. A templated test containing multiple data rows executes all rows even when one fails or is skipped, then aggregates the test status. User-keyword argument type conversion is stable from Robot Framework 7.3.

1. Diagnostic sequence: isolate the failing layer

  1. Preserve output.xml, log.html, report.html, console output, and the exact variant source.
  2. Confirm Python/Robot versions and the executed suite path.
  3. Confirm test selection and the test/template name that actually ran.
  4. Run --dryrun to separate parse/import/argument-contract errors from runtime behavior.
  5. Inspect the failing template call and concrete argument values.
  6. Inspect variable scope and any library/external-system state the template touched.
  7. Only then inspect timing/parallel/CI issues if relevant.
  8. Apply the smallest correction and rerun the smallest named test/controlled grouped slice.

2. Failure map

Symptom Likely layer Evidence to inspect
Template keyword not found Import/keyword resolution --dryrun, import graph, exact template name
Expected int but got text Argument conversion / data contract Failing row + conversion message
Wrong classification Template/business assertion Concrete arguments + keyword trace
Only one report failure hides many bad rows Grouped template identity log.html iterations + source row names
CI runtime grows quadratically Matrix design Dimension counts, setup cost, row runtime
Secret appears in log Data/privacy boundary Source row, variable source, log level/output retention

3. Failure mode: Cartesian expansion without causal coverage

A matrix that combines every browser × environment × locale × role × input boundary can grow faster than the defects it finds. Start by marking which dimensions influence the actual code path. Pairwise or risk-based representative selection can be appropriate, but the rationale must be reviewable and deterministic.

Do not solve a slow matrix by blindly adding Pabot workers. Parallelism cannot repair redundant coverage and may create shared-state contention.

4. Failure mode: one template covers unrelated workflows

If rows require columns such as mode, action, expected_error, cleanup_strategy, and many blanks to steer different paths, the template has become a hidden dispatcher. Split it into purpose-specific keywords/tests. The report should tell the operator which contract failed without reconstructing a mini-program from a row.

5. Failure mode: vague row/test names

Names such as case_17 or row_42 are poor release evidence. Prefer names that encode the risk class, for example Minimum valid length or Dash character rejected. Keep large raw payloads in arguments, not in test names.

6. Failure mode: hidden type conversion changes the meaning of a row

*** Settings ***
Test Template    Compare Numbers

*** Test Cases ***     LEFT    RIGHT
Looks numeric           03      3

*** Keywords ***
Compare Numbers
    [Arguments]    ${left: int}    ${right: int}
    Should Be Equal    ${left}    ${right}

This test passes because both values are converted to integers. That may be correct for a numeric rule but wrong for a lexical identifier rule where 03 and 3 are distinct. Type annotations are a contract, not decoration.

7. Failure mode: “expected failure” passes for the wrong reason

An unsafe anti-pattern is to wrap an entire variant in a broad “any error is expected” check. Then a typo, missing keyword, parser-adjacent runtime error, or wrong dependency can satisfy the row even though the intended validation was never reached.

# INTENTIONALLY UNSAFE — diagnostic anti-pattern, do not copy as a default.
Invalid case
    ${status}=    Run Keyword And Return Status    Validate Complex Workflow    bad-input
    Should Be Equal    ${status}    ${False}

The Boolean only proves that something failed. Repair the design by giving the validator an explicit classification/result contract, or—when failure itself is the contract—assert a narrow, intentional error at the smallest keyword boundary. Chapter 13 covers expected-failure patterns in depth.

8. Failure mode: template failure has no input context

If the assertion message says only ACCEPT != REJECT, operators must search the log to find the row. Include non-sensitive context in the assertion message or embedded call name:

Should Be Equal    ${actual}    ${expected}
...    msg=classification mismatch for candidate=${candidate}

Do not add secrets or entire personal records merely to improve diagnostics.

9. Failure mode: generated data cannot be reproduced

Random values without a recorded seed/source make a CI failure impossible to recreate. If generation is necessary, store the seed, generator version, generation parameters, and failing concrete value with the result artifacts. For core business boundaries, prefer explicit deterministic rows.

10. Failure mode: secrets in data tables

A template row is not a credential channel. Source control and result files may expose it. Use fake values in this course. In production, inject secrets from an approved boundary and keep the test row limited to a symbolic case such as valid credential/expired credential rather than the credential itself.

11. Intentionally broken diagnostic exercise

*** Settings ***
Test Template    Identifier Should Have Classification

*** Test Cases ***          CANDIDATE    EXPECTED
case_1                       abc          ACCEPT
case_2                       ab           ACCEPT
case_3                       abc          ACCEPT
case_4                       abc          ACCEPT

*** Keywords ***
Identifier Should Have Classification
    [Arguments]    ${candidate}    ${expected}
    ${actual}=    Classify Identifier    ${candidate}
    Should Be Equal    ${actual}    ${expected}

Three problems exist: the failing row name is vague, ab has the wrong expectation, and three rows duplicate the same case. Preserve the failure first. Then rename case_2 to Too short rejected, repair the expected value to REJECT, and remove redundant duplicates with a recorded rationale.

12. Troubleshooting shortcuts to reject

  • Blanket retries around a failing matrix.
  • Giant sleeps/timeouts to make generated cases “stable.”
  • Broad EXCEPT or status conversion around all template logic.
  • Arbitrary PYTHONPATH changes instead of fixing imports.
  • Global variables to share mutable row state.
  • Deleting the original failing result directory before comparing the repair.
  • Adding workers before measuring whether rows are actually independent.
  • Copying production customer data into a “realistic” test table.

13. Knowledge check

Why can a broad “status should be False” negative test be false-green?

What should you inspect first when a typed template row fails before any assertion?

What is the safest first response to a 1,000-row matrix that is too slow?

What evidence makes generated failing data reproducible?

14. Summary and next step

Data-driven failures are diagnosable when test identity, variant values, input contracts, and result artifacts remain explicit. Preserve the first failure, reduce redundant dimensions, reject false-green error suppression, and keep sensitive data out of rows. Lesson 5 applies these rules in a compact checkpoint matrix and a deliberate matrix-reduction exercise.

Next lesson

Checkpoint Lab — Data-Driven Testing, Templates, Embedded Arguments, and Variants

Continue with Checkpoint Lab — Data-Driven Testing, Templates, Embedded Arguments, and Variants. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.