Chapter 29Lesson 04200–280 min

Troubleshooting Imports, Scope, Timing, Encoding, and Flaky Automation: Diagnostics, Failure Modes, and Production Practices

Production incidents are often prolonged by symptom patches. This lesson diagnoses sleeps, path hacks, globals, silent decoding, blanket retries, evidence deletion, blind reinstalls, and parallel collisions using a fixed sequence.

Robot Framework 7.4.2Failure analysisPabot 5.2.2First-failure evidenceFlakiness

Learning objectives

  • Apply a fixed evidence-first diagnostic sequence before changing Robot data, environments, retries, or timing.
  • Explain why common “fixes” such as sleeps, broad search paths, global variables, silent decoding, and blanket retries hide root causes.
  • Differentiate deterministic Robot defects from external timing and from parallel shared-state collisions.
  • Preserve first-failure output/debug evidence and avoid repairing only generated reports or derived artifacts.
  • Write production-quality root-cause statements and verify the smallest repaired slice before expanding execution.

Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 is a pre-release and is not required here. Pabot 5.2.2 is the stable parallel-runner baseline; parallel-only diagnostics are optional in this chapter. Mandatory labs use local files, local Python processes, synthetic data, loopback/private state only, and disposable result directories.

1. Diagnostic sequence: preserve first, then move outward

Production incident sequence
flowchart TD
A[Preserve first-failure artifacts] --> B[Record Python, Robot, library, Pabot versions]
B --> C[Confirm command, selection, environment, data]
C --> D[Validate parse and import graph]
D --> E[Inspect variable scope and keyword resolution]
E --> F[Inspect library and external state]
F --> G[Inspect timing, parallel, CI, container state]
G --> H[Apply least destructive correction]
H --> I[Rerun smallest controlled slice]
I --> J[Expand only after root cause is supported]

This ordering makes deterministic, cheap checks precede expensive environment changes. It also prevents a passing retry from erasing the question “why did the first attempt fail?” The sequence is a default, not a rigid law: if a CI worker clearly ran out of disk, inspect that evidence immediately. But do not skip inward layers merely because “flaky usually means timing.”

2. Anti-pattern: adding sleeps to timing races

A fixed sleep changes elapsed time, not readiness semantics. It can make a race rarer while increasing every run's latency. Worse, when the system becomes slower, the same race returns.

# Symptom patch
Sleep    3 s
Element Should Be Visible    ${locator}

# Design intent
Wait Until Element Is Visible    ${locator}    timeout=10 s
Element Should Be Visible         ${locator}

The exact wait keyword depends on the external library. The production principle is library-neutral: observe the condition that proves readiness, bound the wait, and retain the assertion that proves the business outcome.

3. Anti-pattern: arbitrary PYTHONPATH accumulation

When an import fails, teams sometimes append the repository root, parent, grandparent, home directory, and tool directories until Python can find something with the requested name. This can make a different module shadow the intended one and creates a local/CI mismatch.

# Avoid folklore like:
PYTHONPATH=.:..:../..:/opt/random-libs:$PYTHONPATH

# Prefer an explicit project command when a search path is truly required:
python -m robot --pythonpath libraries -d results tests

For Robot resource files, first verify the path relative to the importing file. For Python libraries, verify interpreter ownership, installed package/module, and intentional search path. The repair should reduce ambiguity, not increase it.

4. Anti-pattern: globalizing a missing variable

A missing or unexpected variable often reflects a provenance or lifecycle defect. Set Global Variable can hide it by changing scope for everything that executes later. This creates order dependence and makes parallel execution harder.

# Risky incident patch
Set Global Variable    ${SESSION_ID}    ${value}

# Prefer explicit data flow when ownership is local
${session_id}=    Create Disposable Session
Use Session    ${session_id}

If configuration is genuinely global, inject it explicitly at execution start and document its priority. If mutable runtime state belongs to one test, keep it local/test-scoped or in a test-scoped library instance.

5. Anti-pattern: decoding with ignore or replace to “stop errors”

# Dangerous unless lossy decoding is explicitly required:
text = raw.decode("utf-8", errors="ignore")

# Evidence-preserving failure:
text = raw.decode("utf-8")

Strict decoding makes the mismatch visible. If bytes are corrupt or use another codec, determine the source contract. Silent loss can turn “customer name differs” into a false pass because the distinguishing character was discarded.

6. Anti-pattern: blanket catch/retry

Broad TRY/EXCEPT, Run Keyword And Ignore Error, repeated Wait Until Keyword Succeeds, or test-level reruns can each be legitimate in narrow designs. They become dangerous when applied to unknown failures because they collapse different semantics into “try again.”

Failure Will retry make it correct? What to inspect instead
Missing resource No Import path and repository contents.
Wrong variable owner No Priority/scope/provenance.
ASCII decoding UTF-8 No Codec boundary and bytes.
Shared Pabot output filename No Worker isolation/resource naming.
Eventually consistent local fixture Possibly, but only with explicit condition/state semantics Readiness condition, timeout, idempotency.

7. Anti-pattern: deleting or overwriting first-failure outputs

Robot result artifacts are evidence. If a cleanup job removes output.xml, log, screenshots, process output, or worker artifacts before diagnosis, a later passing rerun cannot reconstruct the initial state. Use distinct output directories or timestamped incident IDs.

python -m robot -d results/incident-001-fail tests/incidents.robot
# after repair:
python -m robot -d results/incident-001-repaired --test "Import incident" tests/incidents.robot

Generated log.html and report.html are presentations of result data. Fixing text in a generated report does not repair the source suite, library, or external state. Always repair the owning source/configuration layer and regenerate results.

8. Anti-pattern: reinstalling everything before recording versions

“Delete the venv and reinstall latest” destroys a critical comparison point. Before dependency changes, record Python, Robot Framework, external library, Pabot, browser/runtime, container image, and lockfile state. If the reinstall fixes the incident, you need a diff to know which dependency changed.

python --version
python -m robot --version
python -m pip freeze > results/incident-001-freeze.txt

Do not publish a freeze file that contains private package indexes or credentials. Course labs use public packages only.

9. “Flaky means external” is a false assumption

Intermittence can originate inside the automation estate: unordered data, mutable globals, stale library instances, nondeterministic test selection, file collisions, random ports, locale/encoding defaults, filesystem case sensitivity, or teardown races. External systems are only one category.

Observed pattern Likely hypotheses to test
Fails only after another test Shared suite/global/library/external state; cleanup/lifecycle defect.
Fails only on Linux CI Case-sensitive paths, permissions, locale/encoding, container networking, dependency drift.
Fails only in parallel Shared file/port/account/record, Pabot worker isolation, target capacity.
Fails at variable-dependent branch Priority/scope/input injection, not necessarily timing.
Fails with non-ASCII data Codec boundary or platform default.

10. Intentionally broken example: parallel-only file collision

Consider two Pabot workers executing tests that both write the same filename:

*** Settings ***
Library    OperatingSystem

*** Test Cases ***
Worker A
    Create File    ${EXECDIR}/state/result.txt    A
    ${text}=    Get File    ${EXECDIR}/state/result.txt
    Should Be Equal    ${text}    A

Worker B
    Create File    ${EXECDIR}/state/result.txt    B
    ${text}=    Get File    ${EXECDIR}/state/result.txt
    Should Be Equal    ${text}    B

Serial execution may pass. Parallel execution can fail depending on interleaving. Increasing sleeps only changes which worker wins. The root cause is shared mutable artifact identity.

Repair by giving each worker/test a unique disposable path. Pabot exposes worker-related variables/options in its current documentation; alternatively derive uniqueness from test name plus a generated run ID in a safe local helper. The key is to eliminate shared ownership rather than serialize all workers with a lock unless sharing is genuinely required.

Concurrency boundary. Pabot processes, CI jobs, browser sessions, database rows, API accounts, files, and ports are separate resources. A lock can hide a design that should instead use independent resources.

11. Write a root-cause statement that can be reviewed

A useful statement names four things:

  1. Layer: where the first incorrect state occurred.
  2. Trigger: the condition that exposed it.
  3. Evidence: the artifact or observation supporting the claim.
  4. Repair/proof: the minimal change and targeted rerun demonstrating correction.

Example: “Under two Pabot workers, both tests wrote state/result.txt; worker artifacts showed cross-overwritten content; unique per-test paths removed the collision and the same parallel selection passed repeatedly with unchanged assertions.”

12. Production troubleshooting practices

  • Preserve failing and repaired artifacts separately.
  • Record exact commands and version/environment context before upgrades.
  • Use allowlisted environment diagnostics rather than dumping all secrets.
  • Keep deep debug/TRACE instrumentation time-bounded and access-controlled.
  • Never disable TLS/SSH verification, authorization, MFA, or anti-abuse controls to “rule them out.”
  • Never reproduce incidents against uncontrolled production targets. Use authorized staging or disposable local fixtures.
  • Expand from one failing test to the full suite only after the root-cause hypothesis requires broader interaction.

Knowledge check

Why can a Pabot-only failure be a state-isolation defect rather than a timeout problem?

What is lost if you reinstall dependencies before recording versions?

Why is editing log.html not a repair?

When can retries be justified?

A test fails after silent decoding removes a character. Which troubleshooting principle was violated?

Summary and bridge

Production troubleshooting rejects symptom patches in favor of layer ownership, preserved evidence, and minimal verified repairs. Lesson 5 combines the chapter into a checkpoint: four incidents, predictions, root-cause statements, smallest-slice reruns, an evidence packet, and a reusable incident checklist ready for the governed platform capstone.

Next lesson

Checkpoint Lab — Troubleshooting Imports, Scope, Timing, Encoding, and Flaky Automation

Continue with Checkpoint Lab — Troubleshooting Imports, Scope, Timing, Encoding, and Flaky Automation. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Troubleshooting commands are intentionally conservative and local. Re-check primary documentation when Robot Framework, Python, external libraries, Pabot, containers, or CI runners change.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.