Performance, Large Suites, Output Management, and Execution Optimization: Core Concepts and Mental Model
Large Robot Framework estates become slow for many different reasons. This lesson replaces guesswork with a phase-by-phase cost model so optimization decisions preserve assertions, isolation, first-failure evidence, and reproducibility.
Learning objectives
- Model a Robot run as distinct cost phases instead of one opaque duration.
- Separate latency, throughput, CPU/memory pressure, external-system capacity, and artifact volume.
- Explain why concurrency can improve throughput while leaving per-test work unchanged or making contention worse.
- Choose evidence-preserving output controls and understand when information becomes unrecoverable.
- Define a reproducible performance baseline with versions, hardware context, selection, process count, and result sizes.
Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.
1. The practical problem: “the suite is slow” is not a diagnosis
By Chapter 27, a Robot Framework estate may contain hundreds or thousands of tests/tasks, shared resources, Python libraries, browser/API/database integrations, Pabot execution, containers, and CI artifact retention. The aggregate runtime can be expensive, but the phrase slow suite hides several distinct mechanisms. A run can spend time parsing many files, importing libraries, creating fixtures, waiting on external systems, producing log messages, serializing result trees, rendering HTML, merging Pabot outputs, or uploading artifacts.
Optimization starts by naming the phase that owns the cost. If 70%
of wall time is an external API wait, changing Robot syntax is
unlikely to help. If execution is fast but
log.html takes minutes to open, parallelism is
irrelevant. If Pabot workers saturate a database, adding processes
can make both runtime and flakiness worse.
Optimization invariant. A performance change is acceptable only if the selected tests/tasks, assertions, environment, isolation guarantees, failure exit status, and required diagnostic evidence remain equivalent—or the deliberate difference is documented and approved.
2. The end-to-end performance pipeline
flowchart TD A[Discovery and parse] --> B[Imports and library init] B --> C[Suite and test setup] C --> D[Keyword execution] D --> E[External waits and I/O] E --> F[Teardown] F --> G[Result serialization] G --> H[log/report or Rebot] H --> I[CI artifact upload] P[Pabot or CI concurrency] -. changes scheduling and contention .-> C P -. does not erase per-test work .-> E
Discovery and parse transforms selected Robot data into an executable model. Imports/library initialization resolves resources, variable files, and Python libraries and may construct library instances. Setups and teardowns prepare and clean state. Keyword execution includes Robot-level control flow and custom/external library calls. External waits often dominate browser, API, database, SSH, or process automation. Result serialization records execution evidence. Rebot/rendering creates or transforms human-facing results. CI then stores or transfers those artifacts.
Pabot inserts a scheduling layer around multiple Robot processes. It can overlap independent work, but each worker still parses, imports, executes, logs, and serializes its assigned slice. Parallelism therefore changes throughput and resource contention; it does not make an individual keyword cheaper.
3. Metrics that must not be conflated
| Metric | Meaning | Typical evidence | Common mistake |
|---|---|---|---|
| Wall-clock latency | Elapsed time from process start to completion. | Monotonic timer around the exact command. | Comparing different selections or environments. |
| CPU time / utilization | Processor work consumed by Robot, libraries, browsers, helpers, or Rebot. | OS/process metrics captured with the run. | Assuming low CPU means Robot itself is inefficient; it may be waiting on I/O. |
| Memory footprint | Resident/peak memory used during execution or result processing. | Process/runner metrics plus result-tree context. | Treating output file size as the same thing as memory. |
| Throughput | Tests/tasks completed per unit time. | Count ÷ elapsed time, with stable selection. | Calling a parallel run faster when it executed fewer tests. |
| Artifact volume | Bytes retained in output.xml/JSON, log/report, screenshots, and worker artifacts. | File sizes and counts. | Deleting evidence before diagnosing why it grew. |
| External capacity | What the target system can sustain without queueing, throttling, or corruption. | Service metrics, response latency, errors, connection pools. | Choosing Pabot process count from CPU count alone. |
A meaningful baseline records all metrics needed to explain the change. Wall time alone is insufficient when the candidate optimization moves cost from execution into Rebot or creates much larger CI artifacts.
4. Read-only inspection before benchmarking
python --version
python -m robot --version
python -m robot --help
python -m robot.rebot --help
pabot --version
pabot --help
Record the Python version, stable Robot/Pabot versions, operating system, CPU model or runner class, available memory, selected suite path/tags, process count, and whether browser/service containers are cold or already running. Do not mix dependency upgrades with a performance change; otherwise the measurement has more than one independent variable.
5. Map each cost to an owner and mutable state
| Cost area | Owner/state | Safe first question | Risky shortcut |
|---|---|---|---|
| Parse/import | Robot model + Python import system | Are many files/libraries imported unnecessarily for the selected slice? | PYTHONPATH hacks or merging unrelated resources. |
| Fixture/setup | Suite/test lifecycle + external fixture | Is expensive setup repeated more often than correctness requires? | Changing to GLOBAL state only to avoid setup. |
| Keyword/external work | Library instance + target system | Which keyword or external operation dominates measured time? | Removing assertions or replacing waits with fixed sleeps. |
| Logging/result tree | Robot result model | Which loops/keywords/messages account for result growth? | Blanket log suppression before preserving failure evidence. |
| Parallel scheduling | Pabot/CI workers + target capacity | At 1, 2, 4 workers, where does throughput stop improving? | processes=all on shared accounts/ports/data. |
| Artifact handling | Rebot/browser/CI storage | Can rendering be deferred while output remains preserved? | Deleting output.xml after a failed run. |
Library scope deserves special caution. Robot test libraries default to TEST scope; SUITE and GLOBAL scopes intentionally share instances for longer periods. Reusing a stateful instance can reduce construction cost but creates a larger sharing boundary. The speed benefit is not evidence that the new state model is safe.
6. Output management is an evidence-retention problem
Robot Framework normally creates a machine-readable output plus HTML
log and report. The User Guide explicitly supports running with
--log NONE --report NONE and generating presentations
later with Rebot. This can reduce execution-time rendering/memory
overhead while preserving the authoritative machine-readable result
for post-processing.
python -m robot -d results/raw --log NONE --report NONE tests
python -m robot.rebot --outputdir results/view results/raw/output.xml
That pattern is different from lowering the execution log level.
Messages excluded at execution time cannot be recovered later.
Likewise, --removekeywords and
--flattenkeywords change the detail retained in
generated views; when applied through Rebot they can also produce a
deliberately compact derived output. Preserve the original
first-failure result before creating a reduced derivative.
| Technique | Primary purpose | Evidence consequence |
|---|---|---|
--log NONE --report NONE during execution
|
Defer HTML generation. | Machine-readable output remains available unless output itself is disabled. |
--splitlog |
Make very large HTML logs load in smaller pieces. | Presentation is split; it is not a substitute for controlling excessive result generation. |
--loglevel |
Control which messages are retained. | Messages omitted during execution are unrecoverable. |
--removekeywords |
Discard selected keyword detail. | Can materially reduce diagnostics; warnings/errors have special preservation behavior except ALL. |
--flattenkeywords |
Collapse nested keyword structures while retaining messages. | Reduces hierarchy; Rebot flattening can reduce processing memory for deep structures. |
robot:flatten keyword tag |
Flatten during execution. | Nested content is not written to output at all; use only with an explicit evidence policy. |
7. Serial baseline first, then controlled concurrency
Pabot’s stable documentation exposes suite-level splitting by
default and test-level splitting with --testlevelsplit.
Process count is an explicit capacity choice. A controlled
experiment compares the same selection at one process, then a small
number such as two, while giving every worker isolated files, ports,
users, database rows, downloads, and artifact names.
pabot --processes 1 -d results/p1 tests
pabot --processes 2 -d results/p2 tests
If two processes cut runtime by only 5% while doubling database connections and output volume, the system is not usefully parallel at that boundary. Conversely, a mostly independent I/O-bound suite may gain substantial throughput. Measure rather than infer from CPU count.
8. A safe optimization order
- Freeze selection and environment. Record versions, command, data seed, runner, and target.
- Preserve a raw baseline. Keep first-failure output and record timing/artifact sizes.
- Find the dominant cost. Separate parse/import, setup, keyword/external work, serialization/rendering, scheduling, and upload.
- Change one variable. Example: defer HTML rendering, reduce redundant immutable setup, or test two Pabot workers.
- Re-run enough times. Compare multiple samples and include cold/warm behavior when it matters.
- Verify equivalence. Same selected test count, assertions, statuses, external state contract, and required evidence.
- Set a budget. Convert the accepted result into a regression threshold with measurement metadata.
9. Why this matters in DevOps
Automation feedback is part of delivery capacity. A twenty-minute suite on every pull request can become a queueing problem; a five-gigabyte result set can become a storage/privacy problem. But fast feedback is useful only when failures remain diagnosable and the suite still enforces the same release criteria. Production performance engineering therefore treats runtime, capacity, correctness, and evidence as a joint contract.
Knowledge check
A parallel run finishes sooner. Does that prove each test became faster?
No. Parallelism changes scheduling and throughput. Per-test work may be unchanged, and contention can even make individual tests slower.
Why preserve output.xml before aggressive Rebot removal or flattening?
Because the raw machine-readable result is the strongest reconstruction point for first-failure evidence. A compact derivative may intentionally discard hierarchy or keyword data.
A team changes Python, Robot Framework, runner size, and Pabot process count in one benchmark. What is wrong?
The experiment has several independent variables, so causality is not attributable. Freeze versions/hardware and change one factor at a time.
When can GLOBAL library scope be a dangerous “optimization”?
When it reuses mutable state across tests or suites that were previously isolated. Reduced initialization cost can introduce coupling, order dependence, or data leakage.
Summary and bridge
You now have a phase-oriented model for large-suite cost and a preservation rule for correctness and evidence. Lesson 2 builds a local benchmark harness, measures serial runtime and artifacts, changes one cost at a time, and compares a deliberately small Pabot configuration.
Further reading
- Robot Framework User Guide — execution, output files, log levels, Rebot, keyword removal/flattening, and library scope.
- Robot Framework releases and Robot Framework on PyPI — verify the stable/pre-release boundary before reproducing measurements.
- Pabot documentation and Pabot releases — process count, suite/test splitting, chunking, ordering, PabotLib, and output handling.
- Robot Framework documentation portal — current ecosystem guidance and examples.
Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.