Chapter 28Lesson 02240–300 min

Performance, Large Suites, Output Management, and Execution Optimization: Guided Hands-On Workflow

Turn the performance model into a reproducible laboratory: generate synthetic work, record wall time and artifact sizes, isolate one cost, change one variable, and prove that the optimized run still executes the same assertions.

Disposable benchmarkWall timeArtifact sizePabot 5.2.2Evidence ledger

Learning objectives

  • Create a disposable benchmark project with deterministic synthetic work and no production dependencies.
  • Record command, versions, wall time, return code, selected test count, and artifact sizes in a machine-readable ledger.
  • Measure a serial baseline before changing result generation or parallelism.
  • Optimize one measured cost at a time and preserve the baseline result for comparison.
  • Compare one-process and two-process Pabot runs with worker-safe state.

Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.

1. Scenario and safety boundary

You maintain a synthetic regression suite that is becoming expensive in CI. The lab contains only local Robot files and a tiny Python timing library. No browser, network, database, SSH host, secret, or production account is required. The synthetic delays represent external I/O; the repeated log messages represent result growth.

Disposable path only. Create the lab under a temporary or clearly named training directory such as rf-performance-lab. All cleanup commands below target only that directory and its results subtree.

2. Preflight and pinned environment

python -m venv .venv
# Bash / macOS / Linux
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install robotframework==7.4.2 robotframework-pabot==5.2.2
python -m robot --version
pabot --version
py -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install robotframework==7.4.2 robotframework-pabot==5.2.2
python -m robot --version
pabot --version

If package installation resolves a different version, stop and correct the environment before benchmarking. Do not compare measurements from an old environment to a new one and call the delta an optimization.

3. Build deterministic synthetic work

Create perf_lib.py. Its keyword sleeps for a requested number of milliseconds and returns a deterministic token. This models an external wait without depending on a real service.

from time import sleep
from robot.api.deco import keyword, library

@library(scope="TEST")
class PerfLib:
    @keyword
    def synthetic_io(self, milliseconds: int, token: str = "ok") -> str:
        sleep(int(milliseconds) / 1000.0)
        return token

Create tests/performance.robot. The suite deliberately contains a moderate amount of logging so result growth is visible but still safe on a laptop.

*** Settings ***
Library    ../perf_lib.py

*** Keywords ***
Assert Synthetic Transaction
    [Arguments]    ${id}
    ${value}=    Synthetic Io    80    tx-${id}
    Should Be Equal    ${value}    tx-${id}

Emit Diagnostic Batch
    [Arguments]    ${id}    ${count}=40
    FOR    ${i}    IN RANGE    ${count}
        Log    transaction=${id} iteration=${i} status=synthetic-ok
    END

*** Test Cases ***
Transaction 01
    Assert Synthetic Transaction    01
    Emit Diagnostic Batch    01

Transaction 02
    Assert Synthetic Transaction    02
    Emit Diagnostic Batch    02

Transaction 03
    Assert Synthetic Transaction    03
    Emit Diagnostic Batch    03

Transaction 04
    Assert Synthetic Transaction    04
    Emit Diagnostic Batch    04

Transaction 05
    Assert Synthetic Transaction    05
    Emit Diagnostic Batch    05

Transaction 06
    Assert Synthetic Transaction    06
    Emit Diagnostic Batch    06

The assertion remains mandatory. The log loop is intentionally verbose enough to demonstrate output behavior, but it contains no credentials or private data.

4. Add a portable benchmark wrapper

Create bench.py. It records monotonic wall time, exit status, the exact child command, Python/platform context, and sizes of result files after each run.

from __future__ import annotations
import json, platform, subprocess, sys, time
from pathlib import Path

if len(sys.argv) < 4 or sys.argv[2] != "--":
    raise SystemExit("usage: python bench.py LABEL -- COMMAND [ARGS...]")

label = sys.argv[1]
command = sys.argv[3:]
started = time.perf_counter()
completed = subprocess.run(command, check=False)
elapsed = time.perf_counter() - started

result_dir = Path("results") / label
files = {}
if result_dir.exists():
    for path in sorted(result_dir.rglob("*")):
        if path.is_file():
            files[str(path)] = path.stat().st_size

record = {
    "label": label,
    "command": command,
    "return_code": completed.returncode,
    "wall_seconds": round(elapsed, 4),
    "python": sys.version.split()[0],
    "platform": platform.platform(),
    "files_bytes": files,
}
Path("results").mkdir(exist_ok=True)
with (Path("results") / "benchmarks.jsonl").open("a", encoding="utf-8") as fh:
    fh.write(json.dumps(record, sort_keys=True) + "\n")
print(json.dumps(record, indent=2, sort_keys=True))
raise SystemExit(completed.returncode)

The wrapper propagates Robot/Pabot’s exit code. A benchmark harness that always exits zero creates false-green CI evidence.

5. Measure the serial baseline

python bench.py serial-full -- python -m robot -d results/serial-full tests
python bench.py serial-full-2 -- python -m robot -d results/serial-full-2 tests
python bench.py serial-full-3 -- python -m robot -d results/serial-full-3 tests

Do not optimize after one run. Use at least three samples for this small lab and note whether the first run is systematically colder. Record the median rather than selecting the fastest sample.

Inspect the ledger and result directory. You should observe six tests, all PASS, plus output.xml, log.html, and report.html. The exact timing and byte sizes are machine-dependent; the invariant is selection/status equivalence.

Baseline item Record
Selection 6 tests from tests/performance.robot
Expected status 6 PASS, process exit 0
Process count 1 Robot process
Synthetic wait 6 × ~80 ms plus framework/runtime overhead
Output evidence Raw output.xml + log.html + report.html
Context Python/Robot/Pabot versions and platform from preflight

6. Optimization A: defer HTML rendering, not evidence

Run the same selection while disabling log/report generation during execution. Keep output.xml.

python bench.py serial-raw -- python -m robot -d results/serial-raw --log NONE --report NONE tests
python -m robot.rebot --outputdir results/serial-view results/serial-raw/output.xml

Compare only the execution command’s wall time with the serial baseline, then separately time Rebot if your operating model defers report creation. The optimization does not eliminate work; it moves presentation generation out of the critical execution path. This is valuable when CI needs fast pass/fail feedback and can render or centralize reports afterward.

Verify that the raw run still contains the same six tests and statuses. Then open the Rebot-generated log/report and confirm the same functional evidence is accessible.

7. Optimization B: control repeated log structure with a derived result

Never overwrite the baseline first-failure result. Create a compact derivative with Rebot after preserving the raw output.

python -m robot.rebot \
  --output results/compact.xml \
  --log results/compact-log.html \
  --report results/compact-report.html \
  --flattenkeywords FOR \
  results/serial-raw/output.xml

On PowerShell, place the command on one line or use PowerShell’s backtick continuation. Measure the original and compact file sizes. Flattening collapses loop hierarchy while retaining log messages; it changes the navigational detail available to investigators. If your requirement is stronger reduction, --removekeywords can discard data, but that is a larger evidence decision and must be justified separately.

Do not claim that execution-time --flattenkeywords shrinks output.xml. The current User Guide states that command-line removal/flattening during execution affects the generated log view, not the XML result. Rebot can create a transformed output; the robot:flatten tag is the mechanism that can prevent nested content from being written during execution.

8. Controlled Pabot comparison

The suite’s six tests use TEST-scoped library instances and no shared files/ports/accounts, so test-level splitting is safe for this synthetic lab. Compare one and two workers explicitly.

python bench.py pabot-p1 -- pabot --testlevelsplit --processes 1 -d results/pabot-p1 tests
python bench.py pabot-p2 -- pabot --testlevelsplit --processes 2 -d results/pabot-p2 tests

Do not assume the two-process result must be twice as fast. Pabot adds process startup, scheduling, and result merging. On tiny workloads the overhead can dominate. On real estates, target capacity and worker memory are additional constraints.

Check p1 p2 Acceptance rule
Selected tests 6 6 Counts must match.
Statuses 6 PASS 6 PASS No false-green or lost test.
Worker capacity 1 2 No shared mutable fixture in this lab.
Wall time Measure Measure Use observed median, not expectation.
Artifacts Inspect Inspect No name collision or missing merged result.

9. Challenge: identify the dominant cost before changing code

Increase Synthetic Io from 80 ms to 250 ms or increase the diagnostic loop from 40 to 200 iterations, but not both. Predict which metric will move most: wall time, output size, Rebot time, or parallel speedup. Run the three-sample baseline and explain the observation from the phase model.

10. Verification checklist and evidence packet

  • Environment versions are pinned and recorded.
  • Every compared run selects the same six tests unless the experiment explicitly changes selection.
  • Robot/Pabot non-zero status propagates through bench.py.
  • The raw first-failure-capable output remains preserved before any compact derivative is created.
  • At least three serial samples are retained; cold/warm behavior is noted.
  • Wall time, output/log/report bytes, process count, and transformation commands are in the dossier.
  • Functional equivalence is verified by test count/status and the same assertions—not by “it looked green.”

11. Cleanup

# Run only from inside the disposable rf-performance-lab directory.
python -c "from pathlib import Path; assert Path('tests/performance.robot').exists(), 'guard failed'"
# Then remove only the lab's generated results directory with your OS file manager or guarded shell command.

Keep benchmarks.jsonl only if you want it as course evidence. Never generalize a recursive-delete command to an arbitrary repository path.

Knowledge check

Why is three-sample measurement better than one warm run?

Why does the lab keep output.xml when disabling log/report generation?

Two Pabot workers are slower than one. Is Pabot broken?

What must remain equivalent after an optimization?

Summary and bridge

You built a reproducible performance experiment instead of a stopwatch demo: raw baseline, multiple samples, artifact accounting, deferred rendering, evidence-aware compaction, and controlled Pabot comparison. Lesson 3 turns these mechanics into architecture choices and trade-offs.

Next lesson

Performance, Large Suites, Output Management, and Execution Optimization: Configuration, Design Patterns, and Trade-Offs

Continue with Performance, Large Suites, Output Management, and Execution Optimization: Configuration, Design Patterns, and Trade-Offs. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.