Checkpoint Lab — Performance, Large Suites, Output Management, and Execution Optimization
Checkpoint: create a defensible performance dossier for a disposable Robot Framework estate, identify the largest measured cost, optimize twice in controlled steps, and convert the evidence into enforceable runtime and artifact budgets.
Learning objectives
- Produce a baseline performance dossier with reproducible command, environment, timings, and result sizes.
- Predict two source/runtime/result changes before applying them and verify those predictions independently.
- Apply two optimizations one at a time while preserving functional equivalence and raw evidence.
- Set runtime, output-size, process-count, and retry budgets with regression thresholds.
- Hand off a clean evidence packet that Chapter 29 can use for troubleshooting regressions and flakiness.
Current compatibility baseline — verified 2026-09-01. Robot Framework 7.4.2 is the current stable release and requires Python 3.8+. Robot Framework 7.5b1 remains a pre-release and is not required by this course. Pabot 5.2.2 is the stable parallel-runner baseline; a 5.3.0 beta exists, so examples intentionally stay on 5.2.2. The mandatory examples use only local files, synthetic timing, and disposable result directories.
1. Checkpoint scenario
Your team owns a growing Robot Framework regression estate. CI feedback is approaching its service-level budget and result artifacts are increasing. You are asked to improve the suite, but the acceptance criteria prohibit deleting assertions, sharing mutable state across tests, hiding first failures, or adding uncontrolled concurrency.
The checkpoint uses the disposable project from Lesson 2 or an equivalent local synthetic suite. You will create a baseline dossier, identify the dominant cost, apply Optimization A, re-measure, then apply Optimization B, re-measure, and finally define a budget. Never apply A and B before measuring A alone.
2. Evidence architecture
flowchart TD S[Fixed source + pinned environment] --> B[Baseline: 3+ samples] B --> R[Preserve raw output] R --> A[Optimization A: one variable] A --> VA[Verify selection/status/evidence] VA --> C[Optimization B: one variable] C --> VC[Verify selection/status/evidence] VC --> D[Dossier + budgets + rollback triggers] D --> N[Chapter 29 troubleshooting baseline]
3. Preflight
- Robot Framework 7.4.2 is installed in an isolated virtual environment.
- Pabot 5.2.2 is installed only if you use the parallel option.
- No 7.5 beta feature is required.
- The target is local/synthetic; no production service or credential is involved.
- Runner hardware/class and OS are recorded.
- The exact test selection, variables, data seed, and process count are fixed.
-
results/is empty or moved aside so old artifacts cannot contaminate size measurements.
python --version
python -m robot --version
pabot --version
python -m robot --dryrun -d results/preflight tests
4. Required predictions before execution
Write down at least two predictions before changing anything. For example:
| Prediction | Why | How you will verify |
|---|---|---|
| Deferring HTML will reduce execution critical-path work but keep raw result evidence. | Robot can run with log/report disabled while output remains. | Compare run timing and confirm output.xml still contains the same tests/statuses. |
| Two Pabot workers may reduce elapsed time for independent synthetic waits. | I/O-like work can overlap. | Compare p1 vs p2 medians and verify identical test counts/statuses with isolated state. |
| Flattening loop hierarchy in a derived result will reduce hierarchy/processing cost, not raw baseline evidence. | Transformation occurs after raw result preservation. | Compare raw and compact bytes and inspect failure-relevant messages. |
5. Build the baseline dossier
Run the unoptimized command at least three times. Preserve the exact command and raw output from each sample or at minimum the representative median sample plus the benchmark ledger.
python bench.py base-1 -- python -m robot -d results/base-1 tests
python bench.py base-2 -- python -m robot -d results/base-2 tests
python bench.py base-3 -- python -m robot -d results/base-3 tests
Record: median wall time, minimum/maximum, selected test count, PASS/FAIL/SKIP counts, output.xml bytes, log.html bytes, report.html bytes, process count, retry count, Python/Robot/Pabot versions, OS/runner class, and any notable cold-start behavior.
6. Identify the top cost with evidence
Classify the dominant cost rather than selecting an optimization you already prefer.
| Observed pattern | Likely dominant layer | Candidate experiment |
|---|---|---|
| Execution time high; Rebot/rendering small | Keyword/external work or setup. | Time the expensive fixture/keyword; consider safe reuse or parallelism. |
| Execution completes, HTML generation/opening dominates | Result rendering/volume. | Defer HTML; measure Rebot separately; inspect loop logging. |
| p2 improves, p4 stalls with service latency growth | External capacity/contention. | Cap process count; isolate target resources. |
| Tiny selected slice still pays huge setup | Suite boundary/lifecycle. | Split or lazily initialize fixture without broadening mutable scope. |
| Retries are large share of wall time | Flakiness/synchronization. | Preserve first failure and repair root cause before performance tuning. |
7. Optimization A — one variable
Choose exactly one evidence-supported change. A safe default for the synthetic lab is deferring HTML rendering:
python bench.py opt-a-1 -- python -m robot -d results/opt-a-1 --log NONE --report NONE tests
python bench.py opt-a-2 -- python -m robot -d results/opt-a-2 --log NONE --report NONE tests
python bench.py opt-a-3 -- python -m robot -d results/opt-a-3 --log NONE --report NONE tests
python -m robot.rebot --outputdir results/opt-a-view results/opt-a-2/output.xml
Verify identical selected-test count/statuses and the same assertions. Confirm the raw output remains sufficient to generate the human view. Record Rebot time separately if report generation moves to a later pipeline stage.
8. Optimization B — one additional variable
Only after Optimization A has its own evidence, choose a second change. For an independent synthetic workload, a controlled Pabot comparison is appropriate:
python bench.py opt-b-p1 -- pabot --testlevelsplit --processes 1 -d results/opt-b-p1 --log NONE --report NONE tests
python bench.py opt-b-p2 -- pabot --testlevelsplit --processes 2 -d results/opt-b-p2 --log NONE --report NONE tests
If the suite uses real browser/API/database state, do not copy this blindly. First provide unique users/records/ports/files per worker and model target capacity. If two workers offer no meaningful benefit, keep one; the purpose is to optimize, not to maximize process count.
9. Functional-equivalence gate
| Invariant | Baseline | Optimized | Pass condition |
|---|---|---|---|
| Selected test/task count | Record | Record | Exact match unless explicitly approved. |
| Assertion intent | Original | Same | No removed/weakened checks. |
| PASS/FAIL/SKIP semantics | Record | Record | No failure converted to pass. |
| External state isolation | Known | Known | No new shared mutable state. |
| Exit status | Propagated | Propagated | Non-zero remains non-zero. |
| First-failure evidence | Raw result retained | Raw or equivalently reconstructable | Investigators can diagnose the same failure class. |
| Versions/data/runner | Fixed | Fixed | No hidden independent variable. |
10. Optional derived-output budget
If result volume is a measured problem, create a compact derivative from the preserved raw result and state exactly what detail is lost.
python -m robot.rebot --flattenkeywords FOR \
--output results/derived-compact.xml \
--log results/derived-compact-log.html \
--report results/derived-compact-report.html \
results/opt-a-2/output.xml
Do not delete results/opt-a-2/output.xml as part of the
same experiment. Retention policy can expire raw artifacts later
according to governance, but performance tuning should not erase its
own proof.
11. Set budgets and regression thresholds
Derive values from your measured baseline; the table below is a template, not a universal standard.
| Dimension | Accepted baseline | Budget | Alert/fail threshold | Action |
|---|---|---|---|---|
| Serial median runtime | e.g. 12.4 s | ≤ 13.0 s | > 14.3 s for 3 comparable samples | Investigate phase regression. |
| Pabot p2 median | e.g. 7.8 s | ≤ 8.5 s | > 9.5 s or target errors rise | Check contention/process count. |
| Raw output size | e.g. 4.2 MB | ≤ 5 MB | > 6 MB at same test count | Inspect logging/result growth. |
| Derived log size | e.g. 1.5 MB | ≤ 2 MB | > 2.5 MB | Review flatten/remove policy. |
| Retry share | 0% | < 5% | ≥ 10% | Treat as reliability defect. |
A CI gate should allow normal variance. Avoid failing on a single noisy sample unless the workload is extremely deterministic; use repeated or rolling measurements where feasible.
12. Required evidence packet
- Manifest: date, source revision, Python/Robot/Pabot versions, OS/runner class, CPU/memory context if available.
- Commands: exact baseline and optimization commands.
- Selection: suite/test/task count, tags/variables/data seed.
- Timing: all samples, median, range, and cold/warm note.
- Artifacts: raw output/log/report sizes, derived sizes, Rebot timing if separated.
- Concurrency: process count and external capacity assumptions.
- Correctness: statuses, assertion equivalence, exit status, isolation verification.
- Decision: accepted optimizations, rejected experiments, budget, threshold, and rollback trigger.
13. Verification and cleanup
- Open at least one baseline log/report and one optimized/derived view.
- Confirm the same tests and statuses.
- Confirm no output paths escaped the disposable lab.
- Archive the dossier if desired.
-
Delete only generated
results/data after the evidence is reviewed. - Deactivate and optionally remove the lab virtual environment.
python -c "from pathlib import Path; assert Path('tests').is_dir(); assert Path('results').is_dir(); print('cleanup guard OK')"
Do not publish logs containing real credentials, HTTP bodies, database rows, browser screenshots, or customer data. This lab uses synthetic data specifically so its evidence packet is safe to retain.
Knowledge check
Why must Optimization A be measured before Optimization B is added?
So the causal effect of A is observable. If A and B are introduced together, the resulting delta cannot be attributed reliably.
A compact derived output is 70% smaller. What else must be proven before adopting it?
That required diagnostic evidence and downstream-tool needs are still met, and that the raw first-failure result is retained according to policy.
Pabot p2 is faster but occasionally corrupts shared fixture data. Does it pass the checkpoint?
No. Isolation and correctness are invariants. Throughput improvement cannot compensate for shared-state corruption.
What does a performance budget add beyond a one-time benchmark?
It converts the accepted state into an operational regression boundary with consistent measurement context and a trigger for investigation or rollback.
What does Chapter 28 add to a production Robot Framework operating model?
A measured optimization discipline: phase ownership, reproducible baselines, evidence-aware result management, capacity-based parallelism, functional-equivalence gates, and explicit runtime/output budgets.
Summary and bridge to Chapter 29
Chapter 28 turns speed and result size into governed engineering properties. You can now establish a reproducible baseline, find the dominant cost, optimize one variable at a time, preserve correctness/evidence, model concurrency capacity, and enforce budgets. Chapter 29 uses this disciplined baseline to troubleshoot imports, scope, timing, encoding, and flaky automation—especially the regressions that appear only under real suite scale.
Further reading
- Robot Framework User Guide — execution, output files, log levels, Rebot, keyword removal/flattening, and library scope.
- Robot Framework releases and Robot Framework on PyPI — verify the stable/pre-release boundary before reproducing measurements.
- Pabot documentation and Pabot releases — process count, suite/test splitting, chunking, ordering, PabotLib, and output handling.
- Robot Framework documentation portal — current ecosystem guidance and examples.
Version-sensitive commands in this chapter were authored against the stable course baseline recorded above. Re-check current primary documentation when upgrading Robot Framework, Pabot, Python, external libraries, or CI runners.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.