Chapter 13Lesson 04240–330 min

Assertions, Error Handling, Expected Failures, and Recovery Patterns: Diagnostics, Failure Modes, and Production Practices

Diagnose false-green and failure-masking patterns by preserving original evidence, isolating the failing layer, and applying the least destructive correction without blanket catch/retry behavior.

DiagnosticsFalse greenFailure injectionRetry hazardsProduction practices

Learning objectives

  • Apply the academy diagnostic sequence before changing error-handling code.
  • Detect broad catch/expected-error patterns that turn unrelated failures into success.
  • Recognize unsafe continuable flows and teardown behavior that obscures primary evidence.
  • Separate deterministic failure from transient retry conditions and performance cost.
  • Repair an intentionally false-green test without deleting first-failure artifacts.

Current compatibility baseline. Verified 2026-08-31: Robot Framework 7.4.2 is the current stable release and requires Python 3.8+; 7.5b1 is a pre-release and is not required in this chapter. Native TRY/EXCEPT is the preferred modern error-handling surface. EXCEPT uses exact matching by default and supports type=GLOB, REGEXP, START, or LITERAL. BuiltIn Run Keyword And Expect Error defaults to glob matching. Continuable failures continue execution but still make the test fail. Skip/Skip If produce SKIP, and Wait Until Keyword Succeeds is a bounded retry helper for normal failures—not syntax errors, timeouts, or fatal execution-stopping errors.

1. Diagnostic sequence: preserve, identify, isolate, then repair

  1. Preserve first-failure artifacts: console output, output.xml, log.html, report.html, relevant synthetic input.
  2. Confirm versions: Python, Robot Framework, imported libraries/tools.
  3. Confirm executed path/selection/environment: exact suite/test, tags, variables, output directory.
  4. Validate parse/import graph: syntax and imports must be healthy before runtime recovery matters.
  5. Inspect variable scope and keyword resolution: confirm the keyword and data you think ran actually ran.
  6. Inspect library/external state: partial side effects, sessions, files, processes, APIs, databases, SSH, RPA state as applicable.
  7. Inspect timing/parallel/CI/container state: only if causally relevant.
  8. Apply the least destructive correction.
  9. Rerun the smallest controlled slice into a new evidence directory.

2. Failure mode: catch every error and declare recovery

*** Test Cases ***
False green — catch-all recovery
    TRY
        Should Be Equal    expected    actual
    EXCEPT    AS    ${error}
        Log    Ignoring: ${error}
    END
    Log    Test reaches PASS even though validation failed.

The message-less EXCEPT catches any ordinary failure from the TRY block. If the handler only logs, the test can PASS. Repair by removing the catch when the assertion must fail, or match only the approved technical condition and assert the recovery outcome.

3. Failure mode: expected-error pattern matches the wrong failure

*** Keywords ***
Broken Validator
    Fail    internal parser bug

*** Test Cases ***
Bad negative test
    Run Keyword And Expect Error    *    Broken Validator

Good negative test
    Run Keyword And Expect Error
    ...    EQUALS:amount must be non-negative
    ...    Reject Negative Amount    ${-1}

The bad test passes for any ordinary error. It proves only “something failed.” In a production negative test, failure identity is part of the specification.

4. Failure mode: blanket retry as a substitute for diagnosis

*** Test Cases ***
Bad retry
    Wait Until Keyword Succeeds    20x    5s    Validate Checksum

If checksum validation is deterministic, 20 retries produce the same failure, increase runtime, and enlarge output. Retry only if the contract is truly transient and safe to repeat. For browser/API/database synchronization, prefer condition-specific waits from the relevant domain library when available. Keep the first direct failure before adding any retry.

5. Failure mode: assertion messages without actual/expected context

Fail bad tells the operator almost nothing. Prefer assertions that already provide actual/expected values, and add a custom message that contributes domain context rather than replacing evidence. Do not include credentials, tokens, private payloads, or PII in the message.

Should Be Equal
...    ${actual_state}
...    ready
...    msg=Release state must match the approved readiness fixture

6. Failure mode: continuable failure followed by unsafe mutation

*** Test Cases ***
Unsafe pattern — do not copy
    Run Keyword And Continue On Failure
    ...    Should Be Equal    ${approved}    ${True}
    Delete Production Resource    ${id}

The guard can fail but execution continues into deletion. The correct design is fail-fast before any destructive action. This course does not execute such targets; the example is intentionally non-runnable and demonstrates why continuation should be limited to independent read-only evidence/cleanup.

7. Failure mode: cleanup failure masks or complicates the root cause

Robot can report multiple failures, including teardown failures. That is useful evidence, but cleanup should be written so it does not introduce unrelated noise. Teardown must target only resources owned by the current run, avoid killing unrelated processes, and record enough context to distinguish “test failed” from “cleanup also failed.” Never delete the original result files as part of cleanup.

8. Failure mode: treating flaky recovery as normal success

If a test passes only after retries on a regular basis, “green” is not the complete operational truth. Track retry counts/first failure messages and investigate the underlying synchronization, resource contention, environment instability, or external-system behavior. A retry that routinely succeeds on attempt 3 is a reliability signal, not a solved problem.

9. Performance: attribute cost to the correct layer

Cost Failure/recovery relevance
Parsing/import/setup Occurs before domain assertions; recovery wrappers cannot fix import errors
Robot keyword execution Repeated by continuations/retries
External-system latency Often dominates browser/API/database/SSH waits
Logging/output Retries and continuable diagnostics can substantially grow output.xml/log.html
Pabot scheduling/shared resources Can create contention that looks flaky; Chapter 24 covers it
Container/CI startup Not a reason to retry a deterministic test assertion
Rerun/retry cost Multiplies resource usage and can delay feedback

10. Security and trust boundaries during failure diagnostics

  • Use synthetic credentials/values in demonstrations.
  • Do not log environment secrets just to prove a failure path.
  • Do not disable TLS/SSH verification to make a failing integration pass.
  • Do not test recovery against production browser/API/database/SSH/RPA targets.
  • Do not use Remote libraries or shared CI resources for failure injection in this chapter.
  • Preserve evidence, but redact secret-bearing data at the source rather than deleting all logs after the fact.

11. Intentionally broken false-green exercise

*** Test Cases ***
Broken release validation
    ${status}    ${message}=    Run Keyword And Ignore Error
    ...    Should Be Equal    pending    ready
    Log    status=${status}; message=${message}
    Log    Continuing to publish readiness evidence.

Observe: the assertion fails internally, ${status} becomes FAIL, but the test can PASS because no later keyword fails. Preserve that run in results/false-green.

Repair A (preferred): remove the wrapper and call the assertion directly. Repair B (when status is genuinely data): assert Should Be Equal ${status} PASS and include the captured message if it is safe. Run the repair into results/false-green-repaired and confirm final test FAILs.

12. Troubleshooting shortcuts to reject

  • Catch-all EXCEPT with “continue” as the default response.
  • Run Keyword And Ignore Error without asserting/using its result.
  • Run Keyword And Expect Error * for negative tests.
  • Blanket Wait Until Keyword Succeeds around deterministic assertions.
  • Giant sleeps/timeouts instead of observing the relevant condition.
  • Arbitrary PYTHONPATH changes to work around import failures.
  • Global variables to smuggle recovery state across unrelated tests.
  • Deleting result artifacts before root cause is understood.
  • Disabling security verification or experimenting on production targets.

13. Knowledge check

Why can a message-less EXCEPT create a false green?

A negative test uses expected pattern *. What does it really prove?

Why is a continuable safety guard dangerous before deletion?

What should be preserved before experimenting with a retry?

14. Summary and next step

The diagnostic rule is simple: do not change status semantics until you can name the original failure and its owner. Lesson 5 turns that rule into a checkpoint matrix containing hard, expected, recovered, continuable, skipped, retried, and false-green cases with predicted and verified final statuses.

Next lesson

Checkpoint Lab — Assertions, Error Handling, Expected Failures, and Recovery Patterns

Continue with Checkpoint Lab — Assertions, Error Handling, Expected Failures, and Recovery Patterns. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.