Chapter 01Lesson 0495–125 min

Robot Framework Foundations, ATDD, BDD, RPA, and Use Cases: Diagnostics, Failure Modes, and Production Practices

Diagnose realistic Robot Framework failures by preserving first-failure evidence and locating the failing parser, import, scope, keyword, assertion, library, external-system, or execution-environment layer.

DiagnosticsFailure semanticsKeyword resolutionFirst-failure evidenceSafe troubleshooting

Learning objectives

  • Apply a layered diagnostic sequence instead of guessing at fixes.
  • Preserve and compare first-failure and corrected result artifacts.
  • Diagnose BDD keyword resolution, assertion mismatches, and missing domain libraries.
  • Recognize retry, giant-timeout, broad-EXCEPT, PYTHONPATH, and global-state anti-patterns.
  • Keep troubleshooting inside disposable local boundaries without exposing secrets.

Failure-lab safety. Every failure in this lesson is local and synthetic. Do not reproduce troubleshooting techniques against production accounts, public systems, real databases, SSH targets, browser profiles, or CI secrets. Preserve first-failure evidence before changing anything.

1. Failure is a location problem before it is a fix problem

"Robot failed" is too coarse to diagnose. A failure can occur while parsing source, importing a resource or library, resolving a keyword, converting arguments, executing a library keyword, asserting a result, interacting with an external system, serializing outputs, or running inside a constrained CI/container/parallel environment. The first job is to identify the failing layer.

Diagnostic sequence: preserve evidence, then narrow the failing layer
flowchart TD
A[Preserve first-failure artifacts] --> B[Confirm Robot / Python / library versions]
B --> C[Confirm executed path, selection, environment and data]
C --> D[Validate parse and import graph]
D --> E[Inspect variable scope and keyword resolution]
E --> F[Inspect library and external-system state]
F --> G[Inspect timing / parallel / CI / container state]
G --> H[Apply least-destructive correction]
H --> I[Rerun smallest controlled slice]

2. Preserve the original evidence

Before retrying, copying files around, or editing the suite, save the console output and the original result directory. A rerun can make the visible outcome greener while destroying the temporal clue that explains the first failure. When external systems are involved, also preserve their relevant logs or synthetic fixture state.

For this chapter, create separate output directories such as results/fail-01 and results/fixed-01. That makes the before/after comparison explicit.

3. Failure mode: using Robot Framework for every test layer

Suppose a tiny pure Python function has a defect. Wrapping it in Robot keywords, starting a full suite, and routing it through CI can make a simple failure slower and less localized. The correction is architectural: keep unit-level checks close to the code and reserve Robot for the layers where readable orchestration, acceptance intent, system integration, or task automation adds value.

A useful symptom is an automation estate full of one-line Robot tests that merely call one Python function and assert one primitive result. That is not automatically wrong, but it should trigger a design review: what value is the Robot layer adding?

4. Failure mode: treating BDD prefixes as a separate engine

This suite fails because the underlying keyword does not exist:

*** Test Cases ***
Candidate Can Proceed
    Given release candidate exists
    Then candidate may proceed

Robot can remove the Given prefix during matching, but it still needs a keyword matching release candidate exists (or a full-name match). Installing a "BDD plugin" is not the appropriate first response. Inspect keyword resolution and define/import the correct keyword.

*** Keywords ***
Release candidate exists
    Log    Synthetic candidate exists

Candidate may proceed
    Should Be Equal    ready    ready

The failure is at the definition/resolution layer, not in a browser, API, or CI system.

5. Failure mode: long procedural tests instead of readable keywords

A test body with dozens of low-level steps creates brittle ownership. When one selector, API route, or file layout changes, many tests change. When a failure occurs, business intent is buried in mechanics.

The correction is not "hide everything." Extract stable, meaningful user keywords that express domain intent, keep lower-level calls visible in nested logs, return data rather than mutating global state where practical, and place assertions at the level that can explain the business outcome.

6. Failure mode: confusing RPA task completion with test correctness

An RPA task can complete its keyword sequence and still produce the wrong business outcome if the workflow lacks independent verification. Conversely, a task can legitimately stop because a precondition fails even though the target system is healthy. Define what PASS/FAIL means for the task and, where necessary, verify durable external state separately.

For long-running or human-in-the-loop automation, do not rely on a suite variable as the only checkpoint. A process restart destroys in-memory state.

7. Failure mode: assuming Robot core includes domain drivers

Consider:

*** Test Cases ***
Open Example Page
    Open Browser    https://example.invalid

With no browser library imported, Robot cannot resolve Open Browser. The right diagnosis is "missing/unresolved keyword provider," not "browser driver broken." Only after the correct library is installed and imported would browser-runtime compatibility become the next layer to inspect.

Symptom Likely first layer Do not jump directly to
No keyword with expected name Imports / keyword resolution Network or browser debugging
Variable not found Variable definition/scope Retry loops
Assertion mismatch Observed value vs expected value Increasing timeout
Library keyword raises connection error Library/external-system boundary Changing Robot parser settings
Only parallel CI fails Shared state/capacity/environment Marking test flaky and ignoring it

8. Intentionally broken example — diagnose without hiding the cause

Create this safe suite:

*** Variables ***
${STATUS}    blocked

*** Test Cases ***
Candidate Is Ready
    Log    Observed status=${STATUS}
    Should Be Equal    ${STATUS}    ready

Run it once into results/fail-01. The expected failure is a direct assertion mismatch: blocked is not equal to ready. Do not add a retry, sleep, broad TRY/EXCEPT, or fake expected value. Preserve the failed artifacts, decide whether the fixture or expectation is wrong, then make one intentional correction.

For this exercise, change the synthetic fixture to ready and rerun into results/fixed-01. Compare the two logs. Your evidence should show that the first failure remains visible and the second run proves the controlled correction.

9. Troubleshooting shortcuts that create false confidence

  • Blanket retries: can turn intermittent unknown failures into green results without fixing causality.
  • Giant sleeps/timeouts: increase runtime and can still race; later chapters teach condition-based synchronization for external systems.
  • Broad EXCEPT blocks: can convert unexpected failures into misleading success paths.
  • Arbitrary PYTHONPATH changes: can make imports work on one machine while hiding packaging/project-layout defects.
  • Global variable mutation: increases order dependence and makes parallel execution harder.
  • Deleting result files before investigation: destroys the evidence needed to explain the incident.

10. Diagnostic lab — classify three failures

For each scenario, write the first layer you would inspect and the smallest evidence packet you would preserve:

  1. A BDD-style step says no keyword was found.
  2. A BuiltIn assertion reports blocked != ready.
  3. A future browser-library keyword is resolved correctly but reports that its browser runtime cannot start.

Correct ordering is resolution/imports for the first, test data/expectation for the second, and library/runtime compatibility for the third. Robot core is not the first suspect in all three.

11. Security-sensitive diagnosis

Do not paste real tokens, passwords, SSH keys, HTTP Authorization headers, database connection strings, or personal file paths into logs merely to "see what Robot is using." Use fake values in course labs and secret-safe diagnostics in production. If a future library logs a sensitive payload, the Robot result artifact can become sensitive even when the source file contains no secret.

12. Knowledge check

A BDD step is unresolved. What should you inspect first?

Why preserve results/fail-01 after a successful rerun?

Does increasing a timeout fix a deterministic assertion mismatch?

If Open Browser is unresolved, should you debug browser-driver compatibility first?

What is wrong with catching every error and logging "continuing"?

13. Summary and next step

Reliable diagnosis starts by preserving first-failure evidence and locating the failing layer: parser/import, keyword resolution, variable scope, assertion, library runtime, external system, or execution environment. Avoid fixes that merely make the run green. Lesson 5 combines the chapter into a checkpoint lab with an automation charter, one tiny test, one tiny task, explicit predictions, a controlled failure, and an evidence packet.

Next lesson

Checkpoint Lab — Robot Framework Foundations, ATDD, BDD, RPA, and Use Cases

Continue with Checkpoint Lab — Robot Framework Foundations, ATDD, BDD, RPA, and Use Cases. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading and current primary references

Version-sensitive statements in this lesson were checked on 2026-08-30. The mandatory path uses Robot Framework 7.4.2, the current stable release at the time of authoring. Robot Framework 7.5b1 is a pre-release and is not required here. Re-check these sources before pinning versions in a new environment.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.