Chapter 28Lesson 05270–330 min

Checkpoint Lab — Performance, Large Suites, Output Management, and Execution Optimization

Checkpoint: create a defensible performance dossier for a disposable Robot Framework estate, identify the largest measured cost, optimize twice in controlled steps, and convert the evidence into enforceable runtime and artifact budgets.

Checkpoint labPerformance dossierRegression budgetFunctional equivalenceCleanup

Learning objectives

  • Produce a baseline performance dossier with reproducible command, environment, timings, and result sizes.
  • Predict two source/runtime/result changes before applying them and verify those predictions independently.
  • Apply two optimizations one at a time while preserving functional equivalence and raw evidence.
  • Set runtime, output-size, process-count, and retry budgets with regression thresholds.
  • Hand off a clean evidence packet that Chapter 29 can use for troubleshooting regressions and flakiness.

Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.

1. Checkpoint scenario

Your team owns a growing Robot Framework regression estate. CI feedback is approaching its service-level budget and result artifacts are increasing. You are asked to improve the suite, but the acceptance criteria prohibit deleting assertions, sharing mutable state across tests, hiding first failures, or adding uncontrolled concurrency.

The checkpoint uses the disposable project from Lesson 2 or an equivalent local synthetic suite. You will create a baseline dossier, identify the dominant cost, apply Optimization A, re-measure, then apply Optimization B, re-measure, and finally define a budget. Never apply A and B before measuring A alone.

2. Evidence architecture

Checkpoint evidence flow
flowchart TD
S[Fixed source + pinned environment] --> B[Baseline: 3+ samples]
B --> R[Preserve raw output]
R --> A[Optimization A: one variable]
A --> VA[Verify selection/status/evidence]
VA --> C[Optimization B: one variable]
C --> VC[Verify selection/status/evidence]
VC --> D[Dossier + budgets + rollback triggers]
D --> N[Chapter 29 troubleshooting baseline]

3. Preflight

  • Robot Framework 7.4.2 is installed in an isolated virtual environment.
  • Pabot 5.2.2 is installed only if you use the parallel option.
  • No 7.5 beta feature is required.
  • The target is local/synthetic; no production service or credential is involved.
  • Runner hardware/class and OS are recorded.
  • The exact test selection, variables, data seed, and process count are fixed.
  • results/ is empty or moved aside so old artifacts cannot contaminate size measurements.
python --version
python -m robot --version
pabot --version
python -m robot --dryrun -d results/preflight tests

4. Required predictions before execution

Write down at least two predictions before changing anything. For example:

Prediction Why How you will verify
Deferring HTML will reduce execution critical-path work but keep raw result evidence. Robot can run with log/report disabled while output remains. Compare run timing and confirm output.xml still contains the same tests/statuses.
Two Pabot workers may reduce elapsed time for independent synthetic waits. I/O-like work can overlap. Compare p1 vs p2 medians and verify identical test counts/statuses with isolated state.
Flattening loop hierarchy in a derived result will reduce hierarchy/processing cost, not raw baseline evidence. Transformation occurs after raw result preservation. Compare raw and compact bytes and inspect failure-relevant messages.

5. Build the baseline dossier

Run the unoptimized command at least three times. Preserve the exact command and raw output from each sample or at minimum the representative median sample plus the benchmark ledger.

python bench.py base-1 -- python -m robot -d results/base-1 tests
python bench.py base-2 -- python -m robot -d results/base-2 tests
python bench.py base-3 -- python -m robot -d results/base-3 tests

Record: median wall time, minimum/maximum, selected test count, PASS/FAIL/SKIP counts, output.xml bytes, log.html bytes, report.html bytes, process count, retry count, Python/Robot/Pabot versions, OS/runner class, and any notable cold-start behavior.

6. Identify the top cost with evidence

Classify the dominant cost rather than selecting an optimization you already prefer.

Observed pattern Likely dominant layer Candidate experiment
Execution time high; Rebot/rendering small Keyword/external work or setup. Time the expensive fixture/keyword; consider safe reuse or parallelism.
Execution completes, HTML generation/opening dominates Result rendering/volume. Defer HTML; measure Rebot separately; inspect loop logging.
p2 improves, p4 stalls with service latency growth External capacity/contention. Cap process count; isolate target resources.
Tiny selected slice still pays huge setup Suite boundary/lifecycle. Split or lazily initialize fixture without broadening mutable scope.
Retries are large share of wall time Flakiness/synchronization. Preserve first failure and repair root cause before performance tuning.

7. Optimization A — one variable

Choose exactly one evidence-supported change. A safe default for the synthetic lab is deferring HTML rendering:

python bench.py opt-a-1 -- python -m robot -d results/opt-a-1 --log NONE --report NONE tests
python bench.py opt-a-2 -- python -m robot -d results/opt-a-2 --log NONE --report NONE tests
python bench.py opt-a-3 -- python -m robot -d results/opt-a-3 --log NONE --report NONE tests
python -m robot.rebot --outputdir results/opt-a-view results/opt-a-2/output.xml

Verify identical selected-test count/statuses and the same assertions. Confirm the raw output remains sufficient to generate the human view. Record Rebot time separately if report generation moves to a later pipeline stage.

8. Optimization B — one additional variable

Only after Optimization A has its own evidence, choose a second change. For an independent synthetic workload, a controlled Pabot comparison is appropriate:

python bench.py opt-b-p1 -- pabot --testlevelsplit --processes 1 -d results/opt-b-p1 --log NONE --report NONE tests
python bench.py opt-b-p2 -- pabot --testlevelsplit --processes 2 -d results/opt-b-p2 --log NONE --report NONE tests

If the suite uses real browser/API/database state, do not copy this blindly. First provide unique users/records/ports/files per worker and model target capacity. If two workers offer no meaningful benefit, keep one; the purpose is to optimize, not to maximize process count.

9. Functional-equivalence gate

Invariant Baseline Optimized Pass condition
Selected test/task count Record Record Exact match unless explicitly approved.
Assertion intent Original Same No removed/weakened checks.
PASS/FAIL/SKIP semantics Record Record No failure converted to pass.
External state isolation Known Known No new shared mutable state.
Exit status Propagated Propagated Non-zero remains non-zero.
First-failure evidence Raw result retained Raw or equivalently reconstructable Investigators can diagnose the same failure class.
Versions/data/runner Fixed Fixed No hidden independent variable.

10. Optional derived-output budget

If result volume is a measured problem, create a compact derivative from the preserved raw result and state exactly what detail is lost.

python -m robot.rebot --flattenkeywords FOR \
  --output results/derived-compact.xml \
  --log results/derived-compact-log.html \
  --report results/derived-compact-report.html \
  results/opt-a-2/output.xml

Do not delete results/opt-a-2/output.xml as part of the same experiment. Retention policy can expire raw artifacts later according to governance, but performance tuning should not erase its own proof.

11. Set budgets and regression thresholds

Derive values from your measured baseline; the table below is a template, not a universal standard.

Dimension Accepted baseline Budget Alert/fail threshold Action
Serial median runtime e.g. 12.4 s ≤ 13.0 s > 14.3 s for 3 comparable samples Investigate phase regression.
Pabot p2 median e.g. 7.8 s ≤ 8.5 s > 9.5 s or target errors rise Check contention/process count.
Raw output size e.g. 4.2 MB ≤ 5 MB > 6 MB at same test count Inspect logging/result growth.
Derived log size e.g. 1.5 MB ≤ 2 MB > 2.5 MB Review flatten/remove policy.
Retry share 0% < 5% ≥ 10% Treat as reliability defect.

A CI gate should allow normal variance. Avoid failing on a single noisy sample unless the workload is extremely deterministic; use repeated or rolling measurements where feasible.

12. Required evidence packet

  • Manifest: date, source revision, Python/Robot/Pabot versions, OS/runner class, CPU/memory context if available.
  • Commands: exact baseline and optimization commands.
  • Selection: suite/test/task count, tags/variables/data seed.
  • Timing: all samples, median, range, and cold/warm note.
  • Artifacts: raw output/log/report sizes, derived sizes, Rebot timing if separated.
  • Concurrency: process count and external capacity assumptions.
  • Correctness: statuses, assertion equivalence, exit status, isolation verification.
  • Decision: accepted optimizations, rejected experiments, budget, threshold, and rollback trigger.

13. Verification and cleanup

  1. Open at least one baseline log/report and one optimized/derived view.
  2. Confirm the same tests and statuses.
  3. Confirm no output paths escaped the disposable lab.
  4. Archive the dossier if desired.
  5. Delete only generated results/ data after the evidence is reviewed.
  6. Deactivate and optionally remove the lab virtual environment.
python -c "from pathlib import Path; assert Path('tests').is_dir(); assert Path('results').is_dir(); print('cleanup guard OK')"

Do not publish logs containing real credentials, HTTP bodies, database rows, browser screenshots, or customer data. This lab uses synthetic data specifically so its evidence packet is safe to retain.

Knowledge check

Why must Optimization A be measured before Optimization B is added?

A compact derived output is 70% smaller. What else must be proven before adopting it?

Pabot p2 is faster but occasionally corrupts shared fixture data. Does it pass the checkpoint?

What does a performance budget add beyond a one-time benchmark?

What does Chapter 28 add to a production Robot Framework operating model?

Summary and bridge to Chapter 29

Chapter 28 turns speed and result size into governed engineering properties. You can now establish a reproducible baseline, find the dominant cost, optimize one variable at a time, preserve correctness/evidence, model concurrency capacity, and enforce budgets. Chapter 29 uses this disciplined baseline to troubleshoot imports, scope, timing, encoding, and flaky automation—especially the regressions that appear only under real suite scale.

Next lesson

Troubleshooting Imports, Scope, Timing, Encoding, and Flaky Automation: Core Concepts and Mental Model

Continue with Troubleshooting Imports, Scope, Timing, Encoding, and Flaky Automation: Core Concepts and Mental Model. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.