Performance, Large Suites, Output Management, and Execution Optimization: Guided Hands-On Workflow
Turn the performance model into a reproducible laboratory: generate synthetic work, record wall time and artifact sizes, isolate one cost, change one variable, and prove that the optimized run still executes the same assertions.
Learning objectives
- Create a disposable benchmark project with deterministic synthetic work and no production dependencies.
- Record command, versions, wall time, return code, selected test count, and artifact sizes in a machine-readable ledger.
- Measure a serial baseline before changing result generation or parallelism.
- Optimize one measured cost at a time and preserve the baseline result for comparison.
- Compare one-process and two-process Pabot runs with worker-safe state.
Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.
1. Scenario and safety boundary
You maintain a synthetic regression suite that is becoming expensive in CI. The lab contains only local Robot files and a tiny Python timing library. No browser, network, database, SSH host, secret, or production account is required. The synthetic delays represent external I/O; the repeated log messages represent result growth.
Disposable path only. Create the lab under a
temporary or clearly named training directory such as
rf-performance-lab. All cleanup commands below target
only that directory and its results subtree.
2. Preflight and pinned environment
python -m venv .venv
# Bash / macOS / Linux
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install robotframework==7.4.2 robotframework-pabot==5.2.2
python -m robot --version
pabot --version
py -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install robotframework==7.4.2 robotframework-pabot==5.2.2
python -m robot --version
pabot --version
If package installation resolves a different version, stop and correct the environment before benchmarking. Do not compare measurements from an old environment to a new one and call the delta an optimization.
3. Build deterministic synthetic work
Create perf_lib.py. Its keyword sleeps for a requested
number of milliseconds and returns a deterministic token. This
models an external wait without depending on a real service.
from time import sleep
from robot.api.deco import keyword, library
@library(scope="TEST")
class PerfLib:
@keyword
def synthetic_io(self, milliseconds: int, token: str = "ok") -> str:
sleep(int(milliseconds) / 1000.0)
return token
Create tests/performance.robot. The suite deliberately
contains a moderate amount of logging so result growth is visible
but still safe on a laptop.
*** Settings ***
Library ../perf_lib.py
*** Keywords ***
Assert Synthetic Transaction
[Arguments] ${id}
${value}= Synthetic Io 80 tx-${id}
Should Be Equal ${value} tx-${id}
Emit Diagnostic Batch
[Arguments] ${id} ${count}=40
FOR ${i} IN RANGE ${count}
Log transaction=${id} iteration=${i} status=synthetic-ok
END
*** Test Cases ***
Transaction 01
Assert Synthetic Transaction 01
Emit Diagnostic Batch 01
Transaction 02
Assert Synthetic Transaction 02
Emit Diagnostic Batch 02
Transaction 03
Assert Synthetic Transaction 03
Emit Diagnostic Batch 03
Transaction 04
Assert Synthetic Transaction 04
Emit Diagnostic Batch 04
Transaction 05
Assert Synthetic Transaction 05
Emit Diagnostic Batch 05
Transaction 06
Assert Synthetic Transaction 06
Emit Diagnostic Batch 06
The assertion remains mandatory. The log loop is intentionally verbose enough to demonstrate output behavior, but it contains no credentials or private data.
4. Add a portable benchmark wrapper
Create bench.py. It records monotonic wall time, exit
status, the exact child command, Python/platform context, and sizes
of result files after each run.
from __future__ import annotations
import json, platform, subprocess, sys, time
from pathlib import Path
if len(sys.argv) < 4 or sys.argv[2] != "--":
raise SystemExit("usage: python bench.py LABEL -- COMMAND [ARGS...]")
label = sys.argv[1]
command = sys.argv[3:]
started = time.perf_counter()
completed = subprocess.run(command, check=False)
elapsed = time.perf_counter() - started
result_dir = Path("results") / label
files = {}
if result_dir.exists():
for path in sorted(result_dir.rglob("*")):
if path.is_file():
files[str(path)] = path.stat().st_size
record = {
"label": label,
"command": command,
"return_code": completed.returncode,
"wall_seconds": round(elapsed, 4),
"python": sys.version.split()[0],
"platform": platform.platform(),
"files_bytes": files,
}
Path("results").mkdir(exist_ok=True)
with (Path("results") / "benchmarks.jsonl").open("a", encoding="utf-8") as fh:
fh.write(json.dumps(record, sort_keys=True) + "\n")
print(json.dumps(record, indent=2, sort_keys=True))
raise SystemExit(completed.returncode)
The wrapper propagates Robot/Pabot’s exit code. A benchmark harness that always exits zero creates false-green CI evidence.
5. Measure the serial baseline
python bench.py serial-full -- python -m robot -d results/serial-full tests
python bench.py serial-full-2 -- python -m robot -d results/serial-full-2 tests
python bench.py serial-full-3 -- python -m robot -d results/serial-full-3 tests
Do not optimize after one run. Use at least three samples for this small lab and note whether the first run is systematically colder. Record the median rather than selecting the fastest sample.
Inspect the ledger and result directory. You should observe six
tests, all PASS, plus output.xml,
log.html, and report.html. The exact
timing and byte sizes are machine-dependent; the invariant is
selection/status equivalence.
| Baseline item | Record |
|---|---|
| Selection | 6 tests from tests/performance.robot |
| Expected status | 6 PASS, process exit 0 |
| Process count | 1 Robot process |
| Synthetic wait | 6 × ~80 ms plus framework/runtime overhead |
| Output evidence | Raw output.xml + log.html + report.html |
| Context | Python/Robot/Pabot versions and platform from preflight |
6. Optimization A: defer HTML rendering, not evidence
Run the same selection while disabling log/report generation during
execution. Keep output.xml.
python bench.py serial-raw -- python -m robot -d results/serial-raw --log NONE --report NONE tests
python -m robot.rebot --outputdir results/serial-view results/serial-raw/output.xml
Compare only the execution command’s wall time with the serial baseline, then separately time Rebot if your operating model defers report creation. The optimization does not eliminate work; it moves presentation generation out of the critical execution path. This is valuable when CI needs fast pass/fail feedback and can render or centralize reports afterward.
Verify that the raw run still contains the same six tests and statuses. Then open the Rebot-generated log/report and confirm the same functional evidence is accessible.
7. Optimization B: control repeated log structure with a derived result
Never overwrite the baseline first-failure result. Create a compact derivative with Rebot after preserving the raw output.
python -m robot.rebot \
--output results/compact.xml \
--log results/compact-log.html \
--report results/compact-report.html \
--flattenkeywords FOR \
results/serial-raw/output.xml
On PowerShell, place the command on one line or use PowerShell’s
backtick continuation. Measure the original and compact file sizes.
Flattening collapses loop hierarchy while retaining log messages; it
changes the navigational detail available to investigators. If your
requirement is stronger reduction, --removekeywords can
discard data, but that is a larger evidence decision and must be
justified separately.
Do not claim that execution-time
--flattenkeywords shrinks output.xml.
The current User Guide states that command-line removal/flattening
during execution affects the generated log view, not the XML
result. Rebot can create a transformed output; the
robot:flatten tag is the mechanism that can prevent
nested content from being written during execution.
8. Controlled Pabot comparison
The suite’s six tests use TEST-scoped library instances and no shared files/ports/accounts, so test-level splitting is safe for this synthetic lab. Compare one and two workers explicitly.
python bench.py pabot-p1 -- pabot --testlevelsplit --processes 1 -d results/pabot-p1 tests
python bench.py pabot-p2 -- pabot --testlevelsplit --processes 2 -d results/pabot-p2 tests
Do not assume the two-process result must be twice as fast. Pabot adds process startup, scheduling, and result merging. On tiny workloads the overhead can dominate. On real estates, target capacity and worker memory are additional constraints.
| Check | p1 | p2 | Acceptance rule |
|---|---|---|---|
| Selected tests | 6 | 6 | Counts must match. |
| Statuses | 6 PASS | 6 PASS | No false-green or lost test. |
| Worker capacity | 1 | 2 | No shared mutable fixture in this lab. |
| Wall time | Measure | Measure | Use observed median, not expectation. |
| Artifacts | Inspect | Inspect | No name collision or missing merged result. |
9. Challenge: identify the dominant cost before changing code
Increase Synthetic Io from 80 ms to 250 ms
or increase the diagnostic loop from 40 to 200 iterations,
but not both. Predict which metric will move most: wall time, output
size, Rebot time, or parallel speedup. Run the three-sample baseline
and explain the observation from the phase model.
10. Verification checklist and evidence packet
- Environment versions are pinned and recorded.
- Every compared run selects the same six tests unless the experiment explicitly changes selection.
-
Robot/Pabot non-zero status propagates through
bench.py. - The raw first-failure-capable output remains preserved before any compact derivative is created.
- At least three serial samples are retained; cold/warm behavior is noted.
- Wall time, output/log/report bytes, process count, and transformation commands are in the dossier.
- Functional equivalence is verified by test count/status and the same assertions—not by “it looked green.”
11. Cleanup
# Run only from inside the disposable rf-performance-lab directory.
python -c "from pathlib import Path; assert Path('tests/performance.robot').exists(), 'guard failed'"
# Then remove only the lab's generated results directory with your OS file manager or guarded shell command.
Keep benchmarks.jsonl only if you want it as course
evidence. Never generalize a recursive-delete command to an
arbitrary repository path.
Knowledge check
Why is three-sample measurement better than one warm run?
It exposes startup/cache variance and prevents a single favorable sample from becoming the claimed baseline.
Why does the lab keep output.xml when disabling log/report generation?
It preserves the machine-readable execution result so Rebot can generate presentations later and investigators retain first-failure evidence.
Two Pabot workers are slower than one. Is Pabot broken?
Not necessarily. The workload may be too small, process/merge overhead may dominate, or the target/host may be constrained. Measure the phases before concluding.
What must remain equivalent after an optimization?
At minimum the intended selection, assertions, statuses, isolation contract, environment assumptions, exit status, and required diagnostic evidence.
Summary and bridge
You built a reproducible performance experiment instead of a stopwatch demo: raw baseline, multiple samples, artifact accounting, deferred rendering, evidence-aware compaction, and controlled Pabot comparison. Lesson 3 turns these mechanics into architecture choices and trade-offs.
Further reading
- Robot Framework User Guide — execution, output files, log levels, Rebot, keyword removal/flattening, and library scope.
- Robot Framework releases and Robot Framework on PyPI — verify the stable/pre-release boundary before reproducing measurements.
- Pabot documentation and Pabot releases — process count, suite/test splitting, chunking, ordering, PabotLib, and output handling.
- Robot Framework documentation portal — current ecosystem guidance and examples.
Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.