Chapter 28Lesson 04180–240 min

Performance, Large Suites, Output Management, and Execution Optimization: Diagnostics, Failure Modes, and Production Practices

A fast false-green suite is worse than a slow trustworthy one. This lesson engineers common performance failures and applies a diagnostic sequence that preserves evidence, isolates the expensive layer, and changes the smallest safe variable.

DiagnosticsFalse optimizationContentionFlaky retriesProduction practice

Learning objectives

  • Apply a repeatable diagnostic sequence to slow, large, or contention-prone executions.
  • Recognize false optimizations that improve timing by weakening correctness or evidence.
  • Diagnose giant result files, repeated setup, target saturation, and retry amplification.
  • Separate Robot/Pabot/CI/container causes from external-system latency.
  • Repair an intentionally broken shared-state parallel example without hiding the first failure.

Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.

1. Diagnostic sequence: preserve before you tune

  1. Preserve first-failure artifacts. Copy or retain raw output, worker outputs, console text, screenshots, and target-side evidence.
  2. Confirm versions. Robot, Python, Pabot, external libraries, browser/runtime, container image, and CI runner class.
  3. Confirm selection/environment/data. Same tests, tags, variables, seeds, target, and fixture state.
  4. Validate parse/import graph. Make sure the “slow” run is not importing or discovering a different tree.
  5. Inspect variable/library/external state. Look for broader sharing, connection pools, stale sessions, or data collisions.
  6. Measure phase timing. Setup, test body, external calls, teardown, result generation, Rebot, merge, upload.
  7. Inspect parallel/container/CI layers. Aggregate process count, CPU/memory limits, service capacity, network/storage bottlenecks.
  8. Change the least destructive variable. Rerun the smallest representative slice first.

2. Failure mode: “optimization” by deleting assertions

A team removes response-body checks because they are “expensive” and reports a 12% speedup. The runtime may be real, but the suite now proves less. Treat this as a coverage change, not a performance change. Before accepting it, determine which assertion is costly and why: JSON parsing, a second API call, database polling, or simply a large diagnostic message.

Possible safe fixes include asserting a smaller business invariant, reusing an already-fetched response, or moving duplicate checks to the correct layer. The acceptance criterion is equivalent defect-detection intent, not a smaller stopwatch number.

3. Failure mode: global shared session to save startup

Reusing one browser/API/database session across unrelated tests can eliminate startup but creates cross-test state. Under Pabot, one process-local “global” object is not even the same global object in another worker; if the shared resource is external—an account, record, file, port—workers can still collide.

Rule. Share only immutable/read-only infrastructure or explicitly partitioned resources. Preserve per-test mutable data ownership. A faster suite that depends on order is not production-ready.

4. Failure mode: blanket logging reduction removes first-failure evidence

Changing --loglevel WARN can shrink results, but it also removes INFO/DEBUG messages at execution time. If a failure requires a request correlation ID or synthetic fixture identifier that was logged at INFO, the compact run may no longer be diagnosable. First classify high-volume messages; do not globally lower retention to solve one noisy loop.

5. Failure mode: giant output.xml from loops

FOR/WHILE loops can create deeply repeated keyword structures and messages. Diagnose whether the growth comes from iteration count, nested keywords, large values, screenshots, library logging, or duplicated payloads. Then choose the narrowest control: redesign noisy diagnostic logging, flatten a known high-volume keyword, or generate a compact Rebot derivative.

*** Keywords ***
Process Synthetic Batch
    [Tags]    robot:flatten
    [Arguments]    ${count}=500
    FOR    ${i}    IN RANGE    ${count}
        Log    item=${i}
    END

The robot:flatten tag is strong: nested content is not written to output. Use it only for a keyword whose internal call tree is deliberately nonessential to retained evidence. Do not apply it broadly because a file is large.

6. Failure mode: overparallelizing an external service

Suppose serial tests average 150 ms API latency. At two workers they average 170 ms; at four, 400 ms; at eight, 1.2 s with throttling. The host CPU may still be idle. The bottleneck is service capacity or shared fixture contention, not Pabot scheduling. The correct process count may be two.

Workers Wall time Mean target latency Errors Interpretation
1 Baseline 150 ms 0 Reference.
2 Improved 170 ms 0 Useful overlap.
4 Flat/worse 400 ms 0–few Queueing starts.
8 Worse 1.2 s Throttle/timeouts Over capacity.

Numbers are illustrative. Build your own scaling curve with controlled data and authorization.

7. Failure mode: one warm run and mixed hardware/version changes

Caches, JIT-like warmups in external tools, browser startup, filesystem cache, container pulls, and network conditions can make one run unrepresentative. Worse, changing the runner class and Robot version at the same time as source changes makes causality unknowable. Preserve a benchmark manifest and compare multiple samples on the same class of runner.

8. Failure mode: expensive suite setup for a tiny selected slice

A suite setup that provisions a large fixture may be reasonable for fifty tests but wasteful when CI selects one test by tag. Diagnose selected-test count and setup duration. Possible fixes include moving a truly test-specific fixture down to Test Setup, lazily creating optional fixture parts, or splitting suite boundaries so selection aligns with fixture ownership. Do not simply skip cleanup or make state global.

9. Failure mode: retries dominate runtime

Retries can turn a ten-minute deterministic suite into a forty-minute flaky one while making the final report look mostly green. Measure retry count and retry time as first-class metrics. Preserve the original failure, then fix synchronization, isolation, or target instability. A blanket retry policy is both a performance problem and an evidence problem.

10. Intentionally broken example: Pabot file collision

The following test appears harmless serially but all workers mutate the same file. Under --testlevelsplit --processes 2, order and contents become nondeterministic.

*** Settings ***
Library    OperatingSystem

*** Test Cases ***
Worker A
    Append To File    ${CURDIR}/shared.txt    A\n
Worker B
    Append To File    ${CURDIR}/shared.txt    B\n

Do not “fix” this with a retry. Preserve the failing/odd output and inspect the shared path. Repair ownership by giving each execution an isolated filename derived from Pabot’s worker identity or, preferably, by designing the test so it does not share mutable file state.

*** Settings ***
Library    OperatingSystem
Suite Setup    Create Directory    ${RESULT_ROOT}

*** Variables ***
${RESULT_ROOT}    ${CURDIR}/tmp

*** Keywords ***
Worker File
    ${worker}=    Get Variable Value    \${PABOTEXECUTIONPOOLID}    serial
    ${path}=      Set Variable    ${RESULT_ROOT}/worker-${worker}.txt
    RETURN    ${path}

*** Test Cases ***
Isolated A
    ${path}=    Worker File
    Append To File    ${path}    A\n
Isolated B
    ${path}=    Worker File
    Append To File    ${path}    B\n

For a serial fallback, the default value serial keeps the example runnable without Pabot. In a richer real suite, include a test identifier as well so tests within the same worker do not collide.

11. Production anti-patterns to reject

  • Deleting assertions or checks solely to meet runtime budgets.
  • GLOBAL mutable library/session state introduced only for speed.
  • --loglevel NONE or broad removal without a failure-reconstruction test.
  • --processes all against a capacity-unknown external service.
  • Benchmarking different test selections, versions, data, or hardware and comparing only wall time.
  • Using retries, giant sleeps, or broader timeouts as “stability optimization.”
  • Deleting raw output after failure because it is large.
  • Treating Robot Framework as a load-testing tool. Use a dedicated load/performance tool when the goal is to drive target-system load.

Knowledge check

A suite is 20% faster after removing two assertions. How should the change be classified?

Why can low CPU coexist with worse runtime at higher Pabot process counts?

What is the first action when a large result file accompanies a failure?

Why is a retry a bad repair for the shared.txt collision?

Summary and bridge

Performance diagnostics now follow the same evidence discipline as functional troubleshooting: preserve, classify, measure, isolate, change one safe variable, and verify. Lesson 5 combines the chapter into a checkpoint dossier with two controlled optimizations and explicit regression thresholds.

Next lesson

Checkpoint Lab — Performance, Large Suites, Output Management, and Execution Optimization

Continue with Checkpoint Lab — Performance, Large Suites, Output Management, and Execution Optimization. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.