Troubleshooting Imports, Scope, Timing, Encoding, and Flaky Automation: Configuration, Design Patterns, and Trade-Offs
Troubleshooting tools have trade-offs. This lesson turns dry-run, TRACE/debug files, waits, search paths, UTF-8 normalization, targeted reruns, and CI instrumentation into explicit design decisions instead of reflexes.
Learning objectives
- Choose between dry-run and full execution based on whether the suspected defect is structural or runtime-dependent.
- Use TRACE/debug files surgically while preserving privacy and manageable evidence volume.
- Prefer explicit condition waits and isolated state over retries and fixed sleeps for timing-related incidents.
- Standardize project-relative paths, documented Python search paths, and explicit UTF-8 boundaries across local and CI execution.
- Decide when to reproduce locally, when to instrument CI, and when to reduce the failing selection before changing code.
Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 is a pre-release and is not required here. Pabot 5.2.2 is the stable parallel-runner baseline; parallel-only diagnostics are optional in this chapter. Mandatory labs use local files, local Python processes, synthetic data, loopback/private state only, and disposable result directories.
1. Decision model: diagnostic power must match the suspected layer
More logging, more retries, and a larger rerun are not automatically better diagnostics. Each tool has a cost and blind spot. Dry-run is safe and fast for model/import/keyword/signature problems but intentionally skips test-library keyword execution and does not validate variables. TRACE can expose detailed messages but may increase noise and sensitive-data risk. A CI rerun reproduces the runner environment but can destroy the exact first-failure state if artifacts are overwritten.
The design goal is minimum sufficient instrumentation: collect enough evidence to distinguish competing root causes, then stop increasing diagnostic surface.
2. Troubleshooting trade-off matrix
| Decision | Prefer left side when… | Prefer right side when… | Key risk |
|---|---|---|---|
| Dry-run vs full execution | Suspect syntax/import/keyword/signature issues. | Need variable values, external state, timing, return values, teardown behavior. | Treating dry-run as runtime proof. |
| INFO/DEBUG vs TRACE/debugfile | Normal incident evidence is sufficient. | Need detailed library messages or execution sequencing in a controlled scope. | PII/secrets/noise and larger artifacts. |
| Explicit condition wait vs retry | A known external condition becomes ready asynchronously. | A narrowly defined transient operation is explicitly safe/idempotent and retry semantics are intentional. | Masking deterministic defects with retries. |
| Project-relative/importable layout vs cwd/search-path changes | Resources/libraries belong to the project/package. | A documented external extension directory is a real requirement. | Machine-specific path folklore. |
| Explicit UTF-8 vs platform default | Course files, JSON/text fixtures, logs, and custom files have defined UTF-8 policy. | Only when an external protocol explicitly mandates another codec. | Mojibake or silent corruption. |
| Smallest failing slice vs whole suite | Failure can be reproduced independently. | Cross-suite lifecycle/order/shared-state interaction is itself the hypothesis. | Losing causality in unrelated noise. |
| Local reproduction vs CI-only instrumentation | Environment can be recreated faithfully. | Failure depends on runner image, env injection, service DNS, permissions, worker concurrency, or ephemeral state. | Changing CI while evidence is incomplete. |
3. Dry-run versus full execution
Use dry-run first when the console says “resource not found,” “keyword not found,” “expected N arguments,” or the suite appears malformed. Its safety advantage is important: library keywords are not executed. This makes it appropriate before touching a target system when you only need model validation.
python -m robot --dryrun --test "Import incident" -d results/dry tests/incidents.robot
Switch to a real targeted run when your hypothesis involves variable provenance, library instance state, returned values, external timing, process/network state, encoding performed by a Python library, or teardown. Document why the transition from dry-run to execution is necessary.
Common misunderstanding. A dry-run PASS means the dry-run checks passed. It does not mean the system under automation is healthy, credentials work, variables have intended runtime values, or timing is correct.
4. TRACE and debug files versus privacy/noise
Robot's default log threshold is INFO. DEBUG and TRACE can reveal
more detail, and --debugfile writes execution
events/messages to a plain-text file. These are powerful but should
be scoped to the smallest reproducible slice because arguments and
library messages may contain paths, payloads, IDs, or secrets in
real projects.
python -m robot --test "Encoding incident" --loglevel TRACE:INFO --debugfile debug.txt -d results/encoding-trace tests/incident_encoding.robot
TRACE:INFO means the execution retains TRACE-level
messages while the HTML log initially displays INFO and above. In
real environments, that richer retained data can itself be
sensitive. Synthetic course values make the mandatory lab safe, but
production teams should define who can enable deep traces, where
they are stored, and how long they are retained.
5. Explicit condition wait versus retry
Three mechanisms are often conflated:
| Mechanism | What it means | Appropriate use | Bad use |
|---|---|---|---|
| Fixed sleep | Wait a duration regardless of state. | Rare pacing constraints where elapsed time itself is the requirement. | Readiness synchronization. |
| Condition wait/poll | Observe a known condition until true or timeout. | File/process/browser/API readiness. | Hiding an unknown deterministic failure. |
| Retry/rerun | Repeat an operation/test after failure. | Narrow, idempotent, understood transient errors with preserved first-failure evidence. | Making the pipeline green without root cause. |
If a file is created asynchronously, poll file existence. If an API returns a job state, poll that state with a bounded budget. If a browser element becomes interactable, use the browser library's explicit wait. Retrying an entire test repeats setup, mutations, and external side effects and can create a false pass if the first failure is discarded.
6. Project-root path discipline versus cwd hacks
Robot exposes several different path concepts:
${CURDIR} is the directory containing the current data
file; ${EXECDIR} is the directory where execution
started; the OS cwd belongs to the process;
--pythonpath/PYTHONPATH affect Python's
module search path. These mechanisms are not interchangeable.
# Resource anchored to the importing file:
Resource ../resources/common.resource
# Explicit custom-library search path in a project command:
python -m robot --pythonpath libraries -d results tests
# Avoid relying on "cd some/magic/folder" as the only reason imports work.
A maintainable project chooses a canonical execution command, package/layout boundary, and explicit resource paths. If CI requires a different cwd, the test project should still resolve its own resources deterministically or the command should set the intended project root explicitly.
7. UTF-8 normalization versus platform defaults
Text files in the Robot project that contain non-ASCII characters must be UTF-8. Custom Python code should also specify encodings for files instead of inheriting locale defaults. This matters on Windows, Linux containers, and CI runners where default encodings can differ.
# Deterministic text boundary
text = Path(path).read_text(encoding="utf-8")
Path(out).write_text(text, encoding="utf-8")
# Protocol exception example: only use another codec when the protocol/file format says so.
legacy = Path(path).read_text(encoding="iso-8859-1")
If the external system provides bytes, record the contract: HTTP charset, file specification, process encoding, database driver behavior, or binary format. Do not infer a codec from what “looks readable” on one workstation.
8. Smallest failing subset versus whole-suite rerun
Targeted execution improves signal-to-noise and reduces accidental state mutation. Start with one test or suite when the failure is independent:
python -m robot --test "Scope incident" -d results/repro tests/incident_scope.robot
Expand the selection only if evidence supports an interaction hypothesis: suite setup order, shared library scope, shared test data, Pabot collision, or CI sharding. A test that fails only after another test is valuable evidence about hidden coupling; do not “fix” it by permanently changing execution order.
9. Local reproduction versus CI-only instrumentation
Local reproduction is preferable when you can match Python/Robot/library versions, variables, fixture data, working directory, and execution command. But some failures genuinely depend on CI/container state: filesystem permissions, service DNS, environment-variable injection, ephemeral ports, locale, CPU pressure, or parallel workers.
| CI evidence | Why capture it | Privacy/safety note |
|---|---|---|
| Python/Robot/library versions | Detect runner drift or wrong interpreter. | No credentials needed. |
| Exact command + selected tests/tags | Prove what actually ran. | Redact literal secrets; prefer secret stores/environment indirection. |
cwd, ${EXECDIR}, relevant paths |
Diagnose path/layout differences. | Avoid dumping unrelated home directories. |
| Non-secret environment allowlist | Confirm intended environment mode. | Never print all environment variables blindly. |
| Worker/job ID and isolated resource names | Diagnose parallel collisions. | Use synthetic IDs; do not expose customer identifiers. |
| output.xml/log/debug artifacts | Preserve first failure. | Apply retention/access controls. |
10. Keep configuration ownership explicit
| Concern | Owner | Do not mislabel as Robot core |
|---|---|---|
| Resource and variable imports | Robot data model + project layout | Python package installation. |
| Python library discovery |
Python environment/module path + Robot
--pythonpath
|
Resource-file relative resolution. |
| Browser/API/database timeout | External library/target contract | Generic Robot parser setting. |
| Editor diagnostics/profile | RobotCode/editor tooling | Runtime execution truth. |
| CI secrets/runner/image | CI provider/container platform | Robot variable scope. |
| Pabot worker isolation | Pabot/project state design | Core Robot keyword parallelism. |
11. Worked decision: CI-only “resource not found”
Suppose local execution succeeds but CI says a resource file cannot
be found. Do not start with a retry or install everything again.
First preserve the CI error and compare repository path/case,
checked-out files, exact importing file, and command. On
case-sensitive Linux, Common.resource and
common.resource differ even if a Windows workstation
previously hid the mistake. If the resource path is correctly
project-relative, changing PYTHONPATH is not the
appropriate repair.
If the missing item is instead a Python test library, compare the actual Python interpreter, installed package, and explicit search path. Same red symptom category—different owner and evidence.
Knowledge check
When is TRACE a bad default even during an incident?
When INFO/DEBUG already distinguishes the cause, or when TRACE would capture sensitive arguments/messages and create unnecessary evidence volume. Use minimum sufficient instrumentation.
Why should project resource imports prefer paths relative to the importing file?
They remain stable across shell working directories and make project ownership explicit. Robot resolves resource paths relative to the importing file before module search paths.
A failure disappears when rerun alone. What should you investigate next?
Order dependence, shared suite/global/library state, shared external records/files/ports, or parallel interactions—not simply mark the test flaky and retry it.
When is a non-UTF-8 decoder acceptable?
When the external protocol or file format explicitly defines that codec and the boundary is documented and tested.
Summary and bridge
Troubleshooting quality comes from matching diagnostic depth to the suspected layer. Lesson 4 now applies the same model to production anti-patterns: sleeps, arbitrary search paths, globalized variables, silent decoding, broad catches/retries, evidence deletion, environment reinstallations, and parallel-state collisions.
Further reading
- Robot Framework User Guide — current syntax, imports, variable priority/scope, dry-run, log levels, debug file, search paths, output artifacts, and execution semantics.
- BuiltIn library, OperatingSystem library, and Process library — assertions, variable inspection, filesystem checks, and local process control used by the labs.
- Robot Framework releases and Robot Framework on PyPI — verify the stable/pre-release boundary before reproducing incident behavior.
- Pabot documentation and Pabot releases — parallel worker behavior and evidence when diagnosing contention or collisions.
- Python codecs documentation — text/bytes conversion and explicit encoding behavior at custom-library boundaries.
Troubleshooting commands are intentionally conservative and local. Re-check primary documentation when Robot Framework, Python, external libraries, Pabot, containers, or CI runners change.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.