Chapter 08Lesson 04170–230 min

Collections, String, DateTime, OperatingSystem, Process, and XML Libraries: Diagnostics, Failure Modes, and Production Practices

Diagnose standard-library failures from first evidence to smallest owning layer: command construction, filesystem guards, process lifecycle/output, encoding, platform paths, shared mutable state, and XML structure.

DiagnosticsCommand injectionProcess leaksEncodingFailure analysis

Learning objectives

  • Classify failures as value, path/filesystem, process, encoding, XML, Robot resolution, or environment problems before changing code.
  • Diagnose command-injection risk and replace concatenated shell text with structured Process arguments.
  • Detect and clean leaked processes and avoid output-management failure modes.
  • Repair shared collection mutation and platform-specific path assumptions without global workarounds.
  • Preserve first-failure evidence and apply the least destructive correction.

Current compatibility baseline. Verified 2026-08-31: Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. The six libraries in this chapter ship with Robot Framework core. BuiltIn is available automatically, while Collections, String, DateTime, OperatingSystem, Process, and XML must be imported explicitly. Robot Framework 7.4 added type hints across standard libraries, expanded bytes support, and Secret-aware arguments in relevant APIs. This chapter does not require pre-release syntax or third-party libraries.

1. Diagnostic sequence: isolate the owning layer

  1. Preserve the first failing output.xml, log.html, console output, process result, and any owned scratch files.
  2. Record Robot/Python/library versions and the exact command/output directory.
  3. Confirm the executed suite, input variables, current platform, and disposable target root.
  4. Validate Robot parse/import/keyword resolution before blaming the OS.
  5. Inspect variable values/scopes and collection ownership.
  6. Inspect path normalization, file existence, child process state, rc/stdout/stderr, or parsed XML model.
  7. Only then inspect CI/container/parallel resource differences.
  8. Apply one minimal correction and rerun the smallest controlled case.

2. Failure mode: command injection through concatenated shell text

Intentionally unsafe example — do not use with untrusted input. The following pattern combines data and shell syntax.

*** Settings ***
Library    Process

*** Test Cases ***
Unsafe Pattern
    ${target}=    Set Variable    harmless.txt && echo injected
    Run Process    checker ${target}    shell=True

The defect is architectural: ${target} is no longer merely data once inserted into shell text. Fix it by keeping the executable and each argument in separate cells and leaving shell=False.

Run Process    checker    ${target}

3. Failure mode: deleting outside the disposable workspace

Remove Directory ${path} recursive=True is powerful enough to remove an entire tree. Never “fix” a cleanup failure by broadening the path or suppressing the error. Normalize the candidate path, compare it to a known disposable root, require a recognizable run-specific segment, and only then remove it.

${candidate}=    Normalize Path    ${WORKSPACE}    case_normalize=True
${root}=         Normalize Path    ${TEMPDIR}    case_normalize=True
Should Start With    ${candidate}    ${root}${/}
Should Contain       ${candidate}    rf08-artifact-lab-
Remove Directory     ${candidate}    recursive=True

4. Failure mode: process leaks

A foreground Run Process normally waits for completion. A background Start Process creates a lifecycle obligation. The suite must wait for readiness, retain the handle/alias, and terminate or wait for the child during teardown. “The CI job ended” is not a cleanup strategy in persistent runners.

*** Settings ***
Library    Process
Suite Teardown    Terminate All Processes    kill=True

*** Test Cases ***
Background Fixture Skeleton
    ${handle}=    Start Process    ${PYTHON}    -c    import time; time.sleep(30)    alias=fixture
    Process Should Be Running    fixture
    Terminate Process    fixture    kill=True
    Process Should Be Stopped    fixture

5. Failure mode: stdout/stderr capacity and blocked output

For normal-sized output, Process captures streams in memory. For large or unlimited output, redirect them to files under the owned workspace. Robot Framework 7.3 improved internal handling and removed the earlier lower-limit deadlock issue, but memory and output volume are still finite resources.

${result}=    Run Process
...    ${PYTHON}
...    -c
...    print("bounded")
...    stdout=${WORKSPACE}/stdout.txt
...    stderr=${WORKSPACE}/stderr.txt
Should Be Equal As Integers    ${result.rc}    0
File Should Exist    ${result.stdout_path}

6. Failure mode: assuming every byte stream is UTF-8

A process normally decodes output using a console-oriented encoding; files have their own encodings. If a producer emits a known different encoding, configure that boundary instead of adding errors=ignore everywhere. Silent replacement can turn corruption into a green test.

Symptom Inspect first Correction
UnicodeDecodeError from file Actual file encoding/source contract Use matching Get File encoding=...
Garbled process stdout Tool/console encoding Set output_encoding deliberately
Bytes compared to text Value type and conversion boundary Decode/encode explicitly with String

7. Failure mode: platform-specific paths or accidental working-directory dependence

Use OperatingSystem path keywords and automatic variables rather than hard-coded C:\... or /tmp/... paths. Remember that automatic slash normalization applies to path arguments, not arbitrary shell strings. A suite that passes locally only because it was started from one specific directory is not reproducible.

8. Failure mode: shared temp filenames under parallel execution

${TEMPDIR}/result.txt is not worker-safe if multiple runs share the same machine. Add a run/worker identifier or use a library/API that creates unique temporary resources. Keep cleanup scoped to that unique resource. Chapter 24 will formalize Pabot worker isolation; this chapter establishes the underlying state rule.

9. Failure mode: accidental shared collection mutation

*** Settings ***
Library    Collections

*** Variables ***
@{COMPONENTS}    api    worker

*** Test Cases ***
Broken Shared Mutation
    Append To List    ${COMPONENTS}    scheduler
    # A later test now observes mutated suite data.

Repair by copying at the test/keyword ownership boundary and mutating the copy. Do not “reset” a global list after every test; that merely hides the shared-state design.

10. Failure mode: fragile XML string comparison

If the requirement is “artifact id is A-1 and name is demo,” assert those fields. Whole-file string equality makes indentation and serializer behavior part of the contract by accident. Conversely, if exact bytes are the signed artifact requirement, use byte/checksum validation deliberately rather than semantic XML equality.

11. Intentionally broken diagnostic exercise

Create a copy of Lesson 2’s workflow and introduce these three faults one at a time: use a shared ${TEMPDIR}/rf08-artifact-lab without run ID; make the child exit with code 9; change XML id to B-2. Preserve separate evidence directories for each failure. Diagnose each failure at its owning layer before editing anything else.

robot --outputdir evidence/fail-process --variable RUN_ID:diag-1 suites/broken.robot
robot --outputdir evidence/fail-xml     --variable RUN_ID:diag-2 suites/broken.robot

12. Production practices

  • Use Process with shell disabled unless shell semantics are a reviewed requirement.
  • Give each run/worker an isolated scratch root.
  • Guard recursive delete operations with normalized-path assertions.
  • Set finite process timeouts for tools that can hang; define terminate/kill behavior intentionally.
  • Redirect potentially large child output to owned files.
  • Use explicit encodings and UTC conventions where evidence crosses systems.
  • Do not log secret environment/file/process values merely for diagnosis.
  • Prefer structural XML assertions for semantic requirements.

13. Knowledge check

A process hangs and produces huge output. What should you inspect before adding a retry?

Why is deleting a fixed shared temp directory unsafe under parallelism?

Why is errors=ignore usually a poor encoding fix?

What is the correct repair for an XML formatting-only mismatch?

14. Summary and next step

You can now diagnose standard-library failures by state owner instead of applying global workarounds. Lesson 5 consolidates the chapter with a guarded local artifact-processing checkpoint, an evidence packet, prediction steps, a deliberate process failure, and cleanup proof.

Next lesson

Checkpoint Lab — Collections, String, DateTime, OperatingSystem, Process, and XML Libraries

Continue with Checkpoint Lab — Collections, String, DateTime, OperatingSystem, Process, and XML Libraries. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.