Chapter 28Lesson 03180–240 min

Performance, Large Suites, Output Management, and Execution Optimization: Configuration, Design Patterns, and Trade-Offs

Performance settings are architecture decisions. This lesson connects each speed or size lever to the state it shares, the evidence it removes or preserves, the failures it can hide, and the CI capacity it consumes.

Trade-offsLibrary scopeResult retentionParallelismPerformance budgets

Learning objectives

  • Choose suite granularity and lifecycle reuse without creating hidden coupling.
  • Evaluate TEST, SUITE, and GLOBAL library scopes as state-sharing contracts, not speed switches.
  • Balance log verbosity, output detail, Rebot timing, and privacy/diagnostic requirements.
  • Decide when serial, Pabot, CI sharding, caching, or reuse is appropriate.
  • Express performance decisions as budgets with measurement and rollback criteria.

Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.

1. Decision model: optimize the bottleneck you can prove

Each performance lever trades one resource or guarantee for another. Larger suites may reuse setup but reduce scheduling flexibility. Broader library scope may avoid construction cost but share mutable state. Lower logging may reduce output volume but remove investigative evidence. Parallelism may reduce elapsed time but consume more browser sessions, memory, database connections, ports, and CI minutes.

Therefore a production decision needs four fields: measured bottleneck, proposed change, preserved invariants, and rollback trigger. Without those fields, “optimization” is an opinion.

2. The primary trade-off matrix

Choice Potential benefit Primary risk Evidence to require
Larger suite / more shared setup Less repeated setup and fewer process boundaries. Coupling, longer failure domains, less Pabot flexibility. Setup count/time plus unchanged test isolation outcomes.
Broader library scope Less repeated library initialization. Mutable state leaks and order dependence. Instance lifecycle trace and explicit reset/cleanup proof.
Lower log volume Smaller artifacts and faster rendering. Missing first-failure detail or secret leakage policy confusion. Failure reconstruction test using retained evidence.
Defer HTML to Rebot Shorter critical execution path / lower execution memory. Later rendering cost and artifact pipeline dependency. Raw output retained and Rebot command/version recorded.
More Pabot workers Higher throughput for independent work. External contention, memory pressure, data/port collisions. Scaling curve at 1/2/4 workers and target capacity metrics.
Cache/reuse Avoid downloads, builds, or immutable setup. Stale dependencies/data and non-reproducible runs. Cache key/provenance plus clean-run comparison.
Deeper keyword abstraction Potentially less source repetition. Harder diagnosis and additional result hierarchy. Call-depth/result-size evidence plus readability review.

3. Suite granularity versus setup reuse

Moving expensive preparation from each test into a suite setup can be legitimate when the prepared state is immutable or each test receives a clean logical partition. It is unsafe when tests mutate shared state and depend on teardown to restore it. A suite boundary is both an organizational boundary and a lifecycle boundary.

*** Settings ***
Suite Setup       Start Read-Only Fixture
Suite Teardown    Stop Read-Only Fixture
Test Setup        Allocate Isolated Record
Test Teardown     Release Isolated Record

This pattern shares only an infrastructure fixture while preserving per-test data isolation. The performance dossier should count how often each setup executes and measure its cost rather than assuming suite-level reuse is beneficial.

4. Library scope versus isolation

Robot Framework’s documented library scopes are TEST (default), SUITE, and GLOBAL. Scope determines how long a library instance—and therefore its Python object state—survives. It does not automatically determine browser/session scope inside an external library, nor does it override external target state.

Scope Instance lifetime Performance attraction Required safety question
TEST New instance for each test; setup/teardown have their own lifecycle context. Maximum isolation, potentially more construction. Is construction actually significant?
SUITE One instance per suite. Reuse within a suite. Can every test reset/partition mutable instance state?
GLOBAL One shared instance for the full execution. Maximum reuse. Is the library truly stateless or explicitly concurrency-safe and cleaned?

Do not broaden scope solely because construction appears in a profiler. First determine whether the constructor is doing work that should instead be lazy, cached immutably, or moved to an external fixture. Shared mutable GLOBAL state can invalidate tests and break Pabot isolation.

5. Log verbosity versus evidence and privacy

Logging has three costs: result-tree size, rendering/browser usability, and privacy/storage exposure. The correct policy is not “INFO is too much.” Classify messages by investigative value. High-volume loop diagnostics may be flattenable; secrets and large HTTP bodies may need redaction; failed assertions and target identifiers may need stronger retention.

Robot’s --loglevel controls which messages are retained. Information excluded at execution cannot be reconstructed by Rebot. Consequently, a production policy might keep INFO in raw output but create a compact human-facing derivative, or it might retain DEBUG only for a targeted diagnostic job. The choice must be intentional.

6. output.xml detail versus size

The machine-readable result is more than an HTML input: external tools, rerun/merge workflows, CI reporting, and post-processing may depend on it. Rebot removal/flattening can create smaller derived outputs, but a compact artifact is not a drop-in substitute if downstream tools require keyword-level detail.

# Preserve raw result
python -m robot -d results/raw --log NONE --report NONE tests

# Derive a compact view/output only after preservation
python -m robot.rebot --flattenkeywords FOR --output results/compact.xml results/raw/output.xml

Also remember that current Robot Framework supports JSON result files when the output extension is .json. Format choice is an ecosystem compatibility decision, not automatically a performance win. XML remains the baseline in this chapter because it is the academy’s established evidence model and is widely supported.

7. On-run rendering versus post-processing

Rendering log.html and report.html on every execution is convenient, but large estates can benefit from producing only the machine-readable output on workers and centralizing Rebot. This makes worker critical paths smaller and standardizes presentation generation. The trade-off is that Rebot becomes a distinct pipeline stage whose version, failure status, CPU/memory, and artifact inputs must be observable.

Never hide a Robot failure because report generation failed separately. Preserve both exit statuses and both logs.

8. Serial versus Pabot versus CI sharding

Model Best fit Hidden cost Do not use when
Serial Robot Small suites, stateful targets, baseline measurement. Long elapsed time for independent work. Feedback budget is demonstrably missed and safe parallelism exists.
Pabot on one runner Independent suites/tests with enough work per worker. Process startup, merge, memory, target contention. Shared accounts/files/ports/data cannot be isolated.
CI job sharding Large estates with independent shards and runner capacity. More runner startup/artifact aggregation, nested concurrency risk. Provider limits/target capacity are lower than aggregate demand.
Pabot inside CI shards Very large estates after capacity modeling. Concurrency multiplies: jobs × processes. No explicit aggregate resource budget exists.

Compute aggregate concurrency before nesting layers. Four CI jobs each running Pabot with four workers can create up to sixteen concurrent Robot executors, not “four-way parallelism.”

9. Cache/reuse versus reproducibility

Immutable package/browser caches can improve startup without sharing test state, but every cache requires a key and invalidation policy. If a “fast” run succeeds only because an old dependency or stale generated artifact remains on disk, it is not reproducible. Periodically compare against a clean run.

10. Shorter keywords versus deep abstraction

Reducing keyword call count for performance is usually a weak first target unless measurement proves Robot-level dispatch/result hierarchy is material. Deep wrapper stacks can enlarge logs and obscure failures, but flattening all abstraction to procedural tests harms maintainability. Optimize domain boundaries for clarity first; then target proven hot paths.

11. Convert decisions into budgets

Budget Example policy Regression trigger
Pull-request runtime Median ≤ 8 min on runner class X for smoke selection. >10% over budget for 3 comparable runs.
Nightly runtime P95 ≤ 45 min at 4 workers with target capacity Y. Scaling efficiency drops or target throttling rises.
Raw result size ≤ 250 MB per shard. >20% growth without test-count growth or approved evidence change.
HTML render time ≤ 90 s in centralized Rebot stage. Browser/report unusability or memory pressure.
Retry share <5% of total runtime. Retries mask instability or dominate elapsed time.

Numbers are examples, not universal targets. Establish budgets from your own baseline, runner class, suite size, and diagnostic requirements.

Knowledge check

Why can a broader library scope be faster but less trustworthy?

What is the key difference between deferring HTML rendering and lowering execution log level?

Four CI jobs each use four Pabot processes. What capacity should the target expect?

Why is result format (XML vs JSON) not automatically a performance optimization?

Summary and bridge

Performance architecture is a set of explicit state, evidence, and capacity trade-offs. Lesson 4 now breaks the most tempting shortcuts—deleted assertions, unsafe shared sessions, blanket log reduction, runaway loops, overparallelism, mixed benchmarks, and retry-driven runtime—and diagnoses them systematically.

Next lesson

Performance, Large Suites, Output Management, and Execution Optimization: Diagnostics, Failure Modes, and Production Practices

Continue with Performance, Large Suites, Output Management, and Execution Optimization: Diagnostics, Failure Modes, and Production Practices. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.