Performance, Large Suites, Output Management, and Execution Optimization: Configuration, Design Patterns, and Trade-Offs
Performance settings are architecture decisions. This lesson connects each speed or size lever to the state it shares, the evidence it removes or preserves, the failures it can hide, and the CI capacity it consumes.
Learning objectives
- Choose suite granularity and lifecycle reuse without creating hidden coupling.
- Evaluate TEST, SUITE, and GLOBAL library scopes as state-sharing contracts, not speed switches.
- Balance log verbosity, output detail, Rebot timing, and privacy/diagnostic requirements.
- Decide when serial, Pabot, CI sharding, caching, or reuse is appropriate.
- Express performance decisions as budgets with measurement and rollback criteria.
Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.
1. Decision model: optimize the bottleneck you can prove
Each performance lever trades one resource or guarantee for another. Larger suites may reuse setup but reduce scheduling flexibility. Broader library scope may avoid construction cost but share mutable state. Lower logging may reduce output volume but remove investigative evidence. Parallelism may reduce elapsed time but consume more browser sessions, memory, database connections, ports, and CI minutes.
Therefore a production decision needs four fields: measured bottleneck, proposed change, preserved invariants, and rollback trigger. Without those fields, “optimization” is an opinion.
2. The primary trade-off matrix
| Choice | Potential benefit | Primary risk | Evidence to require |
|---|---|---|---|
| Larger suite / more shared setup | Less repeated setup and fewer process boundaries. | Coupling, longer failure domains, less Pabot flexibility. | Setup count/time plus unchanged test isolation outcomes. |
| Broader library scope | Less repeated library initialization. | Mutable state leaks and order dependence. | Instance lifecycle trace and explicit reset/cleanup proof. |
| Lower log volume | Smaller artifacts and faster rendering. | Missing first-failure detail or secret leakage policy confusion. | Failure reconstruction test using retained evidence. |
| Defer HTML to Rebot | Shorter critical execution path / lower execution memory. | Later rendering cost and artifact pipeline dependency. | Raw output retained and Rebot command/version recorded. |
| More Pabot workers | Higher throughput for independent work. | External contention, memory pressure, data/port collisions. | Scaling curve at 1/2/4 workers and target capacity metrics. |
| Cache/reuse | Avoid downloads, builds, or immutable setup. | Stale dependencies/data and non-reproducible runs. | Cache key/provenance plus clean-run comparison. |
| Deeper keyword abstraction | Potentially less source repetition. | Harder diagnosis and additional result hierarchy. | Call-depth/result-size evidence plus readability review. |
3. Suite granularity versus setup reuse
Moving expensive preparation from each test into a suite setup can be legitimate when the prepared state is immutable or each test receives a clean logical partition. It is unsafe when tests mutate shared state and depend on teardown to restore it. A suite boundary is both an organizational boundary and a lifecycle boundary.
*** Settings ***
Suite Setup Start Read-Only Fixture
Suite Teardown Stop Read-Only Fixture
Test Setup Allocate Isolated Record
Test Teardown Release Isolated Record
This pattern shares only an infrastructure fixture while preserving per-test data isolation. The performance dossier should count how often each setup executes and measure its cost rather than assuming suite-level reuse is beneficial.
4. Library scope versus isolation
Robot Framework’s documented library scopes are TEST (default), SUITE, and GLOBAL. Scope determines how long a library instance—and therefore its Python object state—survives. It does not automatically determine browser/session scope inside an external library, nor does it override external target state.
| Scope | Instance lifetime | Performance attraction | Required safety question |
|---|---|---|---|
| TEST | New instance for each test; setup/teardown have their own lifecycle context. | Maximum isolation, potentially more construction. | Is construction actually significant? |
| SUITE | One instance per suite. | Reuse within a suite. | Can every test reset/partition mutable instance state? |
| GLOBAL | One shared instance for the full execution. | Maximum reuse. | Is the library truly stateless or explicitly concurrency-safe and cleaned? |
Do not broaden scope solely because construction appears in a profiler. First determine whether the constructor is doing work that should instead be lazy, cached immutably, or moved to an external fixture. Shared mutable GLOBAL state can invalidate tests and break Pabot isolation.
5. Log verbosity versus evidence and privacy
Logging has three costs: result-tree size, rendering/browser usability, and privacy/storage exposure. The correct policy is not “INFO is too much.” Classify messages by investigative value. High-volume loop diagnostics may be flattenable; secrets and large HTTP bodies may need redaction; failed assertions and target identifiers may need stronger retention.
Robot’s --loglevel controls which messages are
retained. Information excluded at execution cannot be reconstructed
by Rebot. Consequently, a production policy might keep INFO in raw
output but create a compact human-facing derivative, or it might
retain DEBUG only for a targeted diagnostic job. The choice must be
intentional.
6. output.xml detail versus size
The machine-readable result is more than an HTML input: external tools, rerun/merge workflows, CI reporting, and post-processing may depend on it. Rebot removal/flattening can create smaller derived outputs, but a compact artifact is not a drop-in substitute if downstream tools require keyword-level detail.
# Preserve raw result
python -m robot -d results/raw --log NONE --report NONE tests
# Derive a compact view/output only after preservation
python -m robot.rebot --flattenkeywords FOR --output results/compact.xml results/raw/output.xml
Also remember that current Robot Framework supports JSON result
files when the output extension is .json. Format choice
is an ecosystem compatibility decision, not automatically a
performance win. XML remains the baseline in this chapter because it
is the academy’s established evidence model and is widely supported.
7. On-run rendering versus post-processing
Rendering log.html and report.html on
every execution is convenient, but large estates can benefit from
producing only the machine-readable output on workers and
centralizing Rebot. This makes worker critical paths smaller and
standardizes presentation generation. The trade-off is that Rebot
becomes a distinct pipeline stage whose version, failure status,
CPU/memory, and artifact inputs must be observable.
Never hide a Robot failure because report generation failed separately. Preserve both exit statuses and both logs.
8. Serial versus Pabot versus CI sharding
| Model | Best fit | Hidden cost | Do not use when |
|---|---|---|---|
| Serial Robot | Small suites, stateful targets, baseline measurement. | Long elapsed time for independent work. | Feedback budget is demonstrably missed and safe parallelism exists. |
| Pabot on one runner | Independent suites/tests with enough work per worker. | Process startup, merge, memory, target contention. | Shared accounts/files/ports/data cannot be isolated. |
| CI job sharding | Large estates with independent shards and runner capacity. | More runner startup/artifact aggregation, nested concurrency risk. | Provider limits/target capacity are lower than aggregate demand. |
| Pabot inside CI shards | Very large estates after capacity modeling. | Concurrency multiplies: jobs × processes. | No explicit aggregate resource budget exists. |
Compute aggregate concurrency before nesting layers. Four CI jobs each running Pabot with four workers can create up to sixteen concurrent Robot executors, not “four-way parallelism.”
9. Cache/reuse versus reproducibility
Immutable package/browser caches can improve startup without sharing test state, but every cache requires a key and invalidation policy. If a “fast” run succeeds only because an old dependency or stale generated artifact remains on disk, it is not reproducible. Periodically compare against a clean run.
10. Shorter keywords versus deep abstraction
Reducing keyword call count for performance is usually a weak first target unless measurement proves Robot-level dispatch/result hierarchy is material. Deep wrapper stacks can enlarge logs and obscure failures, but flattening all abstraction to procedural tests harms maintainability. Optimize domain boundaries for clarity first; then target proven hot paths.
11. Convert decisions into budgets
| Budget | Example policy | Regression trigger |
|---|---|---|
| Pull-request runtime | Median ≤ 8 min on runner class X for smoke selection. | >10% over budget for 3 comparable runs. |
| Nightly runtime | P95 ≤ 45 min at 4 workers with target capacity Y. | Scaling efficiency drops or target throttling rises. |
| Raw result size | ≤ 250 MB per shard. | >20% growth without test-count growth or approved evidence change. |
| HTML render time | ≤ 90 s in centralized Rebot stage. | Browser/report unusability or memory pressure. |
| Retry share | <5% of total runtime. | Retries mask instability or dominate elapsed time. |
Numbers are examples, not universal targets. Establish budgets from your own baseline, runner class, suite size, and diagnostic requirements.
Knowledge check
Why can a broader library scope be faster but less trustworthy?
It may reuse expensive construction, but it also extends the lifetime of mutable Python state across tests/suites and can create order dependence or data leakage.
What is the key difference between deferring HTML rendering and lowering execution log level?
Deferring log/report generation can preserve the machine-readable result for later Rebot; lowering execution log level can discard messages that cannot be recovered.
Four CI jobs each use four Pabot processes. What capacity should the target expect?
Up to sixteen concurrent executors, subject to scheduling. Capacity must be modeled at the product of concurrency layers.
Why is result format (XML vs JSON) not automatically a performance optimization?
Downstream compatibility, tooling, schemas, transformation cost, and measured size/time all matter. Choose by evidence and ecosystem requirements.
Summary and bridge
Performance architecture is a set of explicit state, evidence, and capacity trade-offs. Lesson 4 now breaks the most tempting shortcuts—deleted assertions, unsafe shared sessions, blanket log reduction, runaway loops, overparallelism, mixed benchmarks, and retry-driven runtime—and diagnoses them systematically.
Further reading
- Robot Framework User Guide — execution, output files, log levels, Rebot, keyword removal/flattening, and library scope.
- Robot Framework releases and Robot Framework on PyPI — verify the stable/pre-release boundary before reproducing measurements.
- Pabot documentation and Pabot releases — process count, suite/test splitting, chunking, ordering, PabotLib, and output handling.
- Robot Framework documentation portal — current ecosystem guidance and examples.
Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.