Chapter 13Lesson 03180–240 min

Assertions, Error Handling, Expected Failures, and Recovery Patterns: Configuration, Design Patterns, and Trade-Offs

Choose assertion, error-matching, continuation, skip, and retry patterns by their observable status/evidence consequences, not by convenience or how easily they make a suite green.

Trade-offsError matchingFail-fastRetriesNegative testing

Learning objectives

  • Choose assertion-specific keywords when they improve intent and diagnostics.
  • Select exact or pattern error matching based on stable observable contracts.
  • Compare native TRY/EXCEPT with BuiltIn status wrappers and explain false-green risk.
  • Choose fail-fast or continue-on-failure based on safety and evidence independence.
  • Distinguish deterministic recovery from retry and negative testing from error suppression.

Current compatibility baseline. Verified 2026-08-31: Robot Framework 7.4.2 is the current stable release and requires Python 3.8+; 7.5b1 is a pre-release and is not required in this chapter. Native TRY/EXCEPT is the preferred modern error-handling surface. EXCEPT uses exact matching by default and supports type=GLOB, REGEXP, START, or LITERAL. BuiltIn Run Keyword And Expect Error defaults to glob matching. Continuable failures continue execution but still make the test fail. Skip/Skip If produce SKIP, and Wait Until Keyword Succeeds is a bounded retry helper for normal failures—not syntax errors, timeouts, or fatal execution-stopping errors.

1. Design from the result contract backward

Start with the statement the report must be able to make: “the value was correct,” “the invalid input was rejected for reason X,” “a known transient condition was recovered,” “additional diagnostics ran but the scenario failed,” or “the scenario could not be executed and was skipped.” Then select the mechanism that naturally yields that result. Starting from a convenient helper and deciding later what its Boolean/status means is how false greens are created.

2. Assertion-specific keyword versus generic equality

Need Prefer Why
Exact object/value equality Should Be Equal Direct actual-versus-expected contract
Numeric equality after conversion Should Be Equal As Integers/Numbers Makes conversion/comparison intent explicit
Membership Should Contain Failure message describes missing containment
Pattern contract Should Match / regexp assertion Pattern is first-class, not hidden in generic string logic
Complex domain invariant Named user keyword with focused assertions Business intent remains visible while low-level checks stay diagnosable

3. Exact versus pattern error matching

Native EXCEPT is exact by default. That is the safest default when the message is a stable public contract. Use type=START when a stable prefix is followed by volatile context; GLOB for a deliberately simple wildcard contract; REGEXP only when the structure truly needs it; and LITERAL when you want to make exactness explicit. Run Keyword And Expect Error differs by using glob matching by default, so use the EQUALS: prefix for exact negative tests.

Matcher Good use Risk
Exact / LITERAL / EQUALS Stable domain error Can be brittle if message includes irrelevant volatile detail
START Stable category prefix + variable context Too-short prefix can catch unrelated errors
GLOB Small controlled wildcard A lone * proves almost nothing
REGEXP Structured variable error contract Complex patterns can be unreadable and over-match

4. TRY/EXCEPT versus status-returning BuiltIn helpers

Native TRY/EXCEPT keeps the decision close to the error message and makes unmatched failures propagate automatically. Run Keyword And Return Status is useful when a Boolean is truly the desired data value, but it removes the original failure from normal propagation. Run Keyword And Ignore Error similarly returns PASS/FAIL plus a value/message. Both require the caller to do something meaningful with the returned state.

Approach Strength Failure mode
Narrow TRY/EXCEPT Unmatched errors naturally fail Broad/message-less EXCEPT can hide defects
Return Status Convenient Boolean probe Ignored Boolean creates false green
Ignore Error Captures status + message/value Caller may log and continue without asserting
Expect Error Concise negative test Broad glob can pass for wrong error

5. Continue-on-failure versus fail-fast

Fail-fast is the default because later steps may be invalid or destructive after a failed prerequisite. Continue-on-failure is appropriate when later observations are independent and safe: for example, collect three read-only configuration assertions in one diagnostic scenario while still failing the scenario at the end.

Situation Preferred behavior Reason
Authentication prerequisite failed Fail-fast Later actions do not have valid identity/state
Three independent read-only assertions Continuable may be justified More evidence without unsafe mutation
Teardown cleanup Continue cleanup while preserving failures Cleanup ownership requires best-effort completion
Destructive remediation after failed guard Fail-fast Do not mutate after an unmet safety condition

6. Retry versus deterministic wait/recovery

A retry is not a diagnosis. Use it only when the system contract explicitly allows eventual readiness and the operation is idempotent/safe to repeat. Prefer a condition-specific library wait when one exists because it observes the real condition and can provide better diagnostics. Wait Until Keyword Succeeds is generic and can amplify output; it should have a small count/time and meaningful interval.

Do not retry syntax/import errors, deterministic validation failures, credential denial, or data corruption. Those conditions do not become correct by repetition.

7. Negative testing versus suppressing errors

*** Test Cases ***
Good negative test
    Run Keyword And Expect Error
    ...    EQUALS:amount must be non-negative
    ...    Reject Negative Amount    ${-1}

Bad suppression
    ${status}    ${message}=    Run Keyword And Ignore Error
    ...    Reject Negative Amount    ${-1}
    Log    ${status}: ${message}
    # No assertion: any ordinary failure is effectively accepted.

Negative testing asserts both that failure occurs and that it is the correct failure. Suppression merely prevents propagation.

8. Skip is a governance signal, not recovery

Skip a scenario when the precondition for meaningful verification is absent and that absence is an accepted execution state. Examples include an intentionally unprovisioned optional capability or a platform-specific test on the wrong platform. Do not skip because an assertion failed. robot:skip-on-failure/--skiponfailure exist for known not-ready cases, but they should be governed carefully because they transform failed tests into SKIP and can conceal a growing backlog if applied broadly.

9. Worked decision table

Scenario Decision Observable result
Invalid amount must be rejected Exact expected error PASS only for intended rejection
Optional capability absent by design Skip with reason SKIP
Cache cold but local fallback is approved START-matched TRY recovery + fallback assertion PASS if fallback works
Checksum corrupt No recovery/retry FAIL
Independent read-only checks Continue-on-failure FAIL if any check fails, with more evidence
Eventually-ready synthetic probe Bounded retry PASS only when observed condition succeeds within bound

10. Production design checklist

  • Every caught error has a documented reason it is safe to recover.
  • Every pattern matcher is as narrow as practical and has a test proving an unrelated error is rejected.
  • Every returned status is asserted or explicitly used in a decision.
  • Continuable failures are followed only by independent, safe observations/cleanup.
  • Retries are bounded, safe to repeat, and preserve attempt evidence.
  • Skip reasons explain why verification was not performed.
  • Real secrets/PII are not embedded in assertion messages, expected patterns, or logs.
  • Output directories are retained long enough to diagnose the first failure.

11. Knowledge check

Why can EQUALS: be safer than the default Run Keyword And Expect Error matcher?

When should a continuable failure be rejected?

What is the key difference between a negative test and Run Keyword And Ignore Error?

Why should checksum corruption not be retried?

12. Summary and next step

Good failure handling is policy expressed in executable form. Lesson 4 applies that policy to production failure modes: catch-all recovery, weak expected patterns, blanket retry, weak messages, continuable destructive flows, teardown masking, and normalized flakiness.

Next lesson

Assertions, Error Handling, Expected Failures, and Recovery Patterns: Diagnostics, Failure Modes, and Production Practices

Continue with Assertions, Error Handling, Expected Failures, and Recovery Patterns: Diagnostics, Failure Modes, and Production Practices. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.