Checkpoint Lab — Parallel Execution with Pabot, Sharding, and Resource Contention
Build and measure a restartable evidence-rich parallel lab: prove serial correctness, add isolated worker resources, reproduce one collision, repair it, compare Pabot configurations, and derive a safe concurrency budget that can survive CI sharding.
Learning objectives
- Predict concrete worker/resource/result changes before each run and verify them independently.
- Measure serial, suite-level Pabot, and test-level Pabot configurations using the same workload.
- Inject and preserve a shared-resource collision without turning it into a false-green retry.
- Repair the contention with explicit resource ownership or a narrow lock and prove the reason for the chosen pattern.
- Produce an evidence packet and safe-concurrency budget suitable for a CI runbook.
Current compatibility baseline — verified 2026-08-31.
Robot Framework 7.4.2 is the stable course baseline
and requires Python 3.8+. Pabot 5.2.2 is the stable
parallel-runner baseline used in commands; Pabot 5.3.0b1 is
prerelease and is not required. Pabot is an external runner, not
Robot Framework core. In current Pabot, suite-level splitting is the
default; --testlevelsplit opts into test-level
scheduling; --processes caps local executors;
--shard i/n partitions an execution for distribution;
PabotLib provides cross-process locks/resource sets; and
pabot_results/ plus
pabot_manager.log contain subprocess evidence before
final Rebot output is produced. Use Robot Framework 7.4.2 and Pabot
5.2.2. Pabot 5.3.0b1 is prerelease and excluded. The lab is
local/free and uses only temp files and Python standard library. It
does not require Pabot to be installed in this course-generation
environment; learners execute the pinned commands in their own
isolated environment and record actual timings/versions.
1. Lab setup and ownership contract
Create the project from Lesson 2: parallel.resource,
alpha.robot, beta.robot,
lease_helper.py, collision.robot, and
locked.robot. Use a unique run ID and place mutable lab
data only under
${TEMPDIR}/rf24-parallel-lab-<RUN_ID>. Result
artifacts stay under the project’s evidence/ directory.
python -m venv .venv
# Activate the environment for your shell.
python -m pip install "robotframework==7.4.2" "robotframework-pabot==5.2.2"
python --version
python -m robot --version
pabot --version
# Pick a new identifier for this checkpoint attempt.
# PowerShell: $env:RF24_RUN_ID = "checkpoint-001"
# Bash/zsh: export RF24_RUN_ID="checkpoint-001"
2. Preflight: prove source and current external state
-
Run a Robot dry run on
alpha.robotandbeta.robot. -
Verify the selected temp lab root does not point to the temp
directory itself and contains the prefix
rf24-parallel-lab-. - List the lab root before execution; it should be absent or empty for a fresh run ID.
- Record CPU count and available memory using normal OS tools. These are observations, not permission to use all capacity.
- Record expected number of tests: four in the control workload.
python -m robot --dryrun --outputdir evidence/preflight suites/alpha.robot suites/beta.robot
python -c "import os; print('cpu_count=', os.cpu_count())"
3. Required predictions before mutation
| Prediction | Write it before running | How to verify |
|---|---|---|
| Serial process identity |
All four control tests should use one OS PID and queue
fallback serial.
|
Read worker marker files + serial log. |
| Suite-level Pabot | With two suite files and two processes, work can overlap across two Robot processes. | Compare PIDs/pool/queue markers and manager log. |
| Test-level Pabot | Individual tests become scheduling units, producing finer-grained queue evidence. | Count per-queue directories/partial outputs. |
| Collision run | Two tests competing for one exclusive lock file should expose contention under overlap. | Preserve FileExistsError/keyword failure + both worker traces. |
| Locked repair | Both tests pass, but the singular critical section becomes serialized. | Manager/keyword timing plus absence of lease overlap. |
4. Phase A — serial control run
# Resolve root in your shell as shown in Lesson 2.
python -m robot --outputdir evidence/01-serial --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot
Record elapsed time, status, test count, PID set, and marker tree. If this phase fails, stop. Parallelizing a broken serial baseline makes diagnosis worse.
5. Phase B — suite-level Pabot
pabot --processes 1 --outputdir evidence/02-pabot-p1 --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot
pabot --processes 2 --outputdir evidence/03-pabot-p2 --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot
Do not compare only final seconds. Capture the command, Pabot manager log, number of Robot PIDs, test count, and any CPU/memory observation. A two-process run may approach 2× speedup for this sleep-heavy synthetic workload, but real suites will differ.
6. Phase C — test-level split
pabot --testlevelsplit --processes 2 --outputdir evidence/04-testlevel-p2 --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot
pabot --testlevelsplit --processes 4 --outputdir evidence/05-testlevel-p4 --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot
Compare throughput and overhead. Four processes cannot create more than four useful concurrent work units here because only four tests exist. On a real system, an external capacity ceiling may be lower.
7. Phase D — inject and preserve a collision
python -m robot --outputdir evidence/06-collision-serial --variable "LAB_ROOT:${root}" suites/collision.robot
# Expected diagnostic failure under overlap. Do not overwrite this directory.
pabot --testlevelsplit --processes 2 --outputdir evidence/07-collision-fail --variable "LAB_ROOT:${root}" suites/collision.robot
Save the exact failure message, pabot_manager.log, both
subprocess stdout/stderr files, partial outputs, and the state of
${LAB_ROOT}/lease. The first failure is a required
artifact, not something to clean away immediately.
8. Phase E — repair and justify
Choose one:
- Isolation repair: change the lease path to include a unique domain/test/work-item ID. Use this when the resource can be duplicated.
-
Narrow PabotLib lock: use
locked.robotwhen the fixture is deliberately singular. Keep the lock around only the exclusive operation.
pabot --testlevelsplit --processes 2 --outputdir evidence/08-repaired --variable "LAB_ROOT:${root}" suites/locked.robot
Your evidence must explain why the chosen repair matches the resource model. “It passed” is insufficient.
9. Phase F — compute a safe concurrency budget
Measure or estimate these values from your local run and synthetic target:
| Budget input | Example | Your value |
|---|---|---|
| Host CPU budget | 4 workers | _____ |
| Host memory budget | 900 MiB available / 180 MiB peak ≈ 5 | _____ |
| Synthetic resource slots | 3 independent slots | _____ |
| Target service safe clients | 4 | _____ |
| CI jobs/shards hitting same target | 2 | _____ |
If the same target will receive two CI shards, the local process budget must account for the product. For example, a target ceiling of four concurrent clients with two simultaneous shards implies at most two Pabot processes per shard unless stronger isolation/capacity evidence justifies another plan.
target_limit = 4 concurrent clients
ci_shards_running_together = 2
pabot_processes_per_shard <= floor(4 / 2) = 2
Then also apply CPU, memory, data-slot, and port-slot ceilings.
Final safe value = minimum of all ceilings.
11. Required evidence packet
| Artifact | What it proves |
|---|---|
versions.txt |
Exact Python, Robot Framework, and Pabot versions. |
| Serial timing/status | Correct baseline before parallelism. |
| Pabot p1/p2/p4 timings | Measured scheduler/process-count effect. |
| Worker marker tree | PID/queue/pool identities and unique artifact ownership. |
pabot_manager.log |
Scheduler/process lifecycle and timeout evidence. |
| Failed collision output directory | Unmodified first-failure cause and subprocess evidence. |
| Repaired run output | Ownership/lock correction restored correctness. |
| Capacity worksheet | Why chosen process/shard counts are safe. |
| Optional shard aggregate | Cross-shard result count and aggregation evidence. |
12. Cleanup / rollback
Only delete the synthetic temp root after all evidence is captured.
Before recursive deletion, normalize the path, prove it is below
${TEMPDIR}, and prove its basename begins with
rf24-parallel-lab-. Never run a recursive delete
against a variable that is empty, the temp root itself, a repository
root, or an externally supplied arbitrary path.
*** Keywords ***
Guarded Cleanup Lab Root
${root}= Normalize Path ${LAB_ROOT}
${temp}= Normalize Path ${TEMPDIR}
Should Be True $root != $temp Refusing to delete system temp root.
${root_abs}= Evaluate __import__("os").path.normcase(__import__("os").path.abspath($root))
${temp_abs}= Evaluate __import__("os").path.normcase(__import__("os").path.abspath($temp))
${common}= Evaluate __import__("os").path.commonpath([$root_abs, $temp_abs])
Should Be Equal ${common} ${temp_abs} Lab root is outside TEMPDIR.
${name}= Evaluate __import__("os").path.basename($root_abs)
Should Start With ${name} rf24-parallel-lab-
Remove Directory ${root_abs} recursive=True
Keep evidence/ until the lab has been reviewed; cleanup
of results is a separate retention decision.
13. Verification checklist
- Serial control passed with four expected tests.
- Parallel control passed with unique worker/work-item artifact ownership.
- Timings and PIDs were recorded, not inferred.
- The collision failure was preserved before repair.
- The repair changed resource ownership/serialization, not merely timeout/retry behavior.
- Final merged test counts match the selected inventory.
- No public/production system, real credential, shared customer account, or destructive data was touched.
- The safe concurrency budget includes CI shards/jobs as well as local Pabot processes.
Knowledge check
The serial suite passes and the parallel suite fails on one shared file. What is the strongest first hypothesis?
A resource-ownership race. Preserve the failed worker evidence and prove which executions touched the same path before changing timing or retries.
Why must the checkpoint include a
--processes 1 Pabot run in addition to plain
Robot?
It separates “using the Pabot runner” overhead/behavior from actual multi-process concurrency, giving a cleaner comparison against p2/p4.
Two CI shards each run four Pabot processes against a service that safely supports six clients. Is the plan safe?
No. Effective concurrency can reach eight, exceeding the service ceiling. Reduce local processes, shard concurrency, or use a different authorized environment.
If a PabotLib lock fixes the collision, why might isolation still be better?
Isolation permits concurrent progress and removes shared-lock liveness/bottleneck concerns. A lock is appropriate only when the resource is genuinely singular or has to be shared.
Which evidence should survive after a repaired run passes?
At minimum the failed run’s original outputs/logs, repaired run outputs, version/command manifest, worker markers, and the capacity/ownership decision. A passing rerun must not erase the original cause.
14. Production operating model and Chapter 25 bridge
This chapter adds a concurrency contract to the Robot Framework operating model: every parallel worker has an identity, every mutable resource has an owner or bounded allocator, total concurrency is budgeted across Pabot and CI, and every merge remains traceable to subprocess evidence. Chapter 25 will take that contract into GitHub Actions, GitLab CI, and Jenkins while preserving exit status and artifacts on both success and failure.
References and version anchors
- Robot Framework PyPI — current stable/pre-release and Python support
- Robot Framework 7.4.2 User Guide — execution, variables, result files, Rebot, and core semantics
- robotframework-pabot PyPI — stable 5.2.2 and prerelease status
- Pabot repository README — current CLI, PabotLib, ordering, global variables, sharding, and output/artifact behavior
- PabotLib keyword documentation — locks and shared resource distribution
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.