Performance, Large Suites, Output Management, and Execution Optimization: Diagnostics, Failure Modes, and Production Practices
A fast false-green suite is worse than a slow trustworthy one. This lesson engineers common performance failures and applies a diagnostic sequence that preserves evidence, isolates the expensive layer, and changes the smallest safe variable.
Learning objectives
- Apply a repeatable diagnostic sequence to slow, large, or contention-prone executions.
- Recognize false optimizations that improve timing by weakening correctness or evidence.
- Diagnose giant result files, repeated setup, target saturation, and retry amplification.
- Separate Robot/Pabot/CI/container causes from external-system latency.
- Repair an intentionally broken shared-state parallel example without hiding the first failure.
Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.
1. Diagnostic sequence: preserve before you tune
- Preserve first-failure artifacts. Copy or retain raw output, worker outputs, console text, screenshots, and target-side evidence.
- Confirm versions. Robot, Python, Pabot, external libraries, browser/runtime, container image, and CI runner class.
- Confirm selection/environment/data. Same tests, tags, variables, seeds, target, and fixture state.
- Validate parse/import graph. Make sure the “slow” run is not importing or discovering a different tree.
- Inspect variable/library/external state. Look for broader sharing, connection pools, stale sessions, or data collisions.
- Measure phase timing. Setup, test body, external calls, teardown, result generation, Rebot, merge, upload.
- Inspect parallel/container/CI layers. Aggregate process count, CPU/memory limits, service capacity, network/storage bottlenecks.
- Change the least destructive variable. Rerun the smallest representative slice first.
2. Failure mode: “optimization” by deleting assertions
A team removes response-body checks because they are “expensive” and reports a 12% speedup. The runtime may be real, but the suite now proves less. Treat this as a coverage change, not a performance change. Before accepting it, determine which assertion is costly and why: JSON parsing, a second API call, database polling, or simply a large diagnostic message.
Possible safe fixes include asserting a smaller business invariant, reusing an already-fetched response, or moving duplicate checks to the correct layer. The acceptance criterion is equivalent defect-detection intent, not a smaller stopwatch number.
4. Failure mode: blanket logging reduction removes first-failure evidence
Changing --loglevel WARN can shrink results, but it
also removes INFO/DEBUG messages at execution time. If a failure
requires a request correlation ID or synthetic fixture identifier
that was logged at INFO, the compact run may no longer be
diagnosable. First classify high-volume messages; do not globally
lower retention to solve one noisy loop.
5. Failure mode: giant output.xml from loops
FOR/WHILE loops can create deeply repeated keyword structures and messages. Diagnose whether the growth comes from iteration count, nested keywords, large values, screenshots, library logging, or duplicated payloads. Then choose the narrowest control: redesign noisy diagnostic logging, flatten a known high-volume keyword, or generate a compact Rebot derivative.
*** Keywords ***
Process Synthetic Batch
[Tags] robot:flatten
[Arguments] ${count}=500
FOR ${i} IN RANGE ${count}
Log item=${i}
END
The robot:flatten tag is strong: nested content is not
written to output. Use it only for a keyword whose internal call
tree is deliberately nonessential to retained evidence. Do not apply
it broadly because a file is large.
6. Failure mode: overparallelizing an external service
Suppose serial tests average 150 ms API latency. At two workers they average 170 ms; at four, 400 ms; at eight, 1.2 s with throttling. The host CPU may still be idle. The bottleneck is service capacity or shared fixture contention, not Pabot scheduling. The correct process count may be two.
| Workers | Wall time | Mean target latency | Errors | Interpretation |
|---|---|---|---|---|
| 1 | Baseline | 150 ms | 0 | Reference. |
| 2 | Improved | 170 ms | 0 | Useful overlap. |
| 4 | Flat/worse | 400 ms | 0–few | Queueing starts. |
| 8 | Worse | 1.2 s | Throttle/timeouts | Over capacity. |
Numbers are illustrative. Build your own scaling curve with controlled data and authorization.
7. Failure mode: one warm run and mixed hardware/version changes
Caches, JIT-like warmups in external tools, browser startup, filesystem cache, container pulls, and network conditions can make one run unrepresentative. Worse, changing the runner class and Robot version at the same time as source changes makes causality unknowable. Preserve a benchmark manifest and compare multiple samples on the same class of runner.
8. Failure mode: expensive suite setup for a tiny selected slice
A suite setup that provisions a large fixture may be reasonable for fifty tests but wasteful when CI selects one test by tag. Diagnose selected-test count and setup duration. Possible fixes include moving a truly test-specific fixture down to Test Setup, lazily creating optional fixture parts, or splitting suite boundaries so selection aligns with fixture ownership. Do not simply skip cleanup or make state global.
9. Failure mode: retries dominate runtime
Retries can turn a ten-minute deterministic suite into a forty-minute flaky one while making the final report look mostly green. Measure retry count and retry time as first-class metrics. Preserve the original failure, then fix synchronization, isolation, or target instability. A blanket retry policy is both a performance problem and an evidence problem.
10. Intentionally broken example: Pabot file collision
The following test appears harmless serially but all workers mutate
the same file. Under --testlevelsplit --processes 2,
order and contents become nondeterministic.
*** Settings ***
Library OperatingSystem
*** Test Cases ***
Worker A
Append To File ${CURDIR}/shared.txt A\n
Worker B
Append To File ${CURDIR}/shared.txt B\n
Do not “fix” this with a retry. Preserve the failing/odd output and inspect the shared path. Repair ownership by giving each execution an isolated filename derived from Pabot’s worker identity or, preferably, by designing the test so it does not share mutable file state.
*** Settings ***
Library OperatingSystem
Suite Setup Create Directory ${RESULT_ROOT}
*** Variables ***
${RESULT_ROOT} ${CURDIR}/tmp
*** Keywords ***
Worker File
${worker}= Get Variable Value \${PABOTEXECUTIONPOOLID} serial
${path}= Set Variable ${RESULT_ROOT}/worker-${worker}.txt
RETURN ${path}
*** Test Cases ***
Isolated A
${path}= Worker File
Append To File ${path} A\n
Isolated B
${path}= Worker File
Append To File ${path} B\n
For a serial fallback, the default value serial keeps
the example runnable without Pabot. In a richer real suite, include
a test identifier as well so tests within the same worker do not
collide.
11. Production anti-patterns to reject
- Deleting assertions or checks solely to meet runtime budgets.
- GLOBAL mutable library/session state introduced only for speed.
-
--loglevel NONEor broad removal without a failure-reconstruction test. -
--processes allagainst a capacity-unknown external service. - Benchmarking different test selections, versions, data, or hardware and comparing only wall time.
- Using retries, giant sleeps, or broader timeouts as “stability optimization.”
- Deleting raw output after failure because it is large.
- Treating Robot Framework as a load-testing tool. Use a dedicated load/performance tool when the goal is to drive target-system load.
Knowledge check
A suite is 20% faster after removing two assertions. How should the change be classified?
Primarily as a coverage/semantic change. It is not a valid performance optimization until equivalent defect-detection intent is demonstrated.
Why can low CPU coexist with worse runtime at higher Pabot process counts?
Workers may be queueing on external services, locks, databases, ports, network I/O, or memory/storage pressure rather than CPU.
What is the first action when a large result file accompanies a failure?
Preserve the raw first-failure artifact before applying removal/flattening or deletion so the failure remains reconstructable.
Why is a retry a bad repair for the shared.txt collision?
The root cause is shared mutable file ownership. Retrying changes timing but does not establish isolation, so the race remains.
Summary and bridge
Performance diagnostics now follow the same evidence discipline as functional troubleshooting: preserve, classify, measure, isolate, change one safe variable, and verify. Lesson 5 combines the chapter into a checkpoint dossier with two controlled optimizations and explicit regression thresholds.
Further reading
- Robot Framework User Guide — execution, output files, log levels, Rebot, keyword removal/flattening, and library scope.
- Robot Framework releases and Robot Framework on PyPI — verify the stable/pre-release boundary before reproducing measurements.
- Pabot documentation and Pabot releases — process count, suite/test splitting, chunking, ordering, PabotLib, and output handling.
- Robot Framework documentation portal — current ecosystem guidance and examples.
Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.