Chapter 24Lesson 05240–330 min

Checkpoint Lab — Parallel Execution with Pabot, Sharding, and Resource Contention

Build and measure a restartable evidence-rich parallel lab: prove serial correctness, add isolated worker resources, reproduce one collision, repair it, compare Pabot configurations, and derive a safe concurrency budget that can survive CI sharding.

Checkpoint labMeasured concurrencyCollision injectionSafe budgetOperator evidence

Learning objectives

  • Predict concrete worker/resource/result changes before each run and verify them independently.
  • Measure serial, suite-level Pabot, and test-level Pabot configurations using the same workload.
  • Inject and preserve a shared-resource collision without turning it into a false-green retry.
  • Repair the contention with explicit resource ownership or a narrow lock and prove the reason for the chosen pattern.
  • Produce an evidence packet and safe-concurrency budget suitable for a CI runbook.

Current compatibility baseline — verified 2026-08-31. Robot Framework 7.4.2 is the stable course baseline and requires Python 3.8+. Pabot 5.2.2 is the stable parallel-runner baseline used in commands; Pabot 5.3.0b1 is prerelease and is not required. Pabot is an external runner, not Robot Framework core. In current Pabot, suite-level splitting is the default; --testlevelsplit opts into test-level scheduling; --processes caps local executors; --shard i/n partitions an execution for distribution; PabotLib provides cross-process locks/resource sets; and pabot_results/ plus pabot_manager.log contain subprocess evidence before final Rebot output is produced. Use Robot Framework 7.4.2 and Pabot 5.2.2. Pabot 5.3.0b1 is prerelease and excluded. The lab is local/free and uses only temp files and Python standard library. It does not require Pabot to be installed in this course-generation environment; learners execute the pinned commands in their own isolated environment and record actual timings/versions.

1. Lab setup and ownership contract

Create the project from Lesson 2: parallel.resource, alpha.robot, beta.robot, lease_helper.py, collision.robot, and locked.robot. Use a unique run ID and place mutable lab data only under ${TEMPDIR}/rf24-parallel-lab-<RUN_ID>. Result artifacts stay under the project’s evidence/ directory.

python -m venv .venv
# Activate the environment for your shell.
python -m pip install "robotframework==7.4.2" "robotframework-pabot==5.2.2"
python --version
python -m robot --version
pabot --version

# Pick a new identifier for this checkpoint attempt.
# PowerShell: $env:RF24_RUN_ID = "checkpoint-001"
# Bash/zsh:   export RF24_RUN_ID="checkpoint-001"

2. Preflight: prove source and current external state

  1. Run a Robot dry run on alpha.robot and beta.robot.
  2. Verify the selected temp lab root does not point to the temp directory itself and contains the prefix rf24-parallel-lab-.
  3. List the lab root before execution; it should be absent or empty for a fresh run ID.
  4. Record CPU count and available memory using normal OS tools. These are observations, not permission to use all capacity.
  5. Record expected number of tests: four in the control workload.
python -m robot --dryrun --outputdir evidence/preflight suites/alpha.robot suites/beta.robot
python -c "import os; print('cpu_count=', os.cpu_count())"

3. Required predictions before mutation

Prediction Write it before running How to verify
Serial process identity All four control tests should use one OS PID and queue fallback serial. Read worker marker files + serial log.
Suite-level Pabot With two suite files and two processes, work can overlap across two Robot processes. Compare PIDs/pool/queue markers and manager log.
Test-level Pabot Individual tests become scheduling units, producing finer-grained queue evidence. Count per-queue directories/partial outputs.
Collision run Two tests competing for one exclusive lock file should expose contention under overlap. Preserve FileExistsError/keyword failure + both worker traces.
Locked repair Both tests pass, but the singular critical section becomes serialized. Manager/keyword timing plus absence of lease overlap.

4. Phase A — serial control run

# Resolve root in your shell as shown in Lesson 2.
python -m robot --outputdir evidence/01-serial --variable "LAB_ROOT:${root}"   suites/alpha.robot suites/beta.robot

Record elapsed time, status, test count, PID set, and marker tree. If this phase fails, stop. Parallelizing a broken serial baseline makes diagnosis worse.

5. Phase B — suite-level Pabot

pabot --processes 1 --outputdir evidence/02-pabot-p1   --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot

pabot --processes 2 --outputdir evidence/03-pabot-p2   --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot

Do not compare only final seconds. Capture the command, Pabot manager log, number of Robot PIDs, test count, and any CPU/memory observation. A two-process run may approach 2× speedup for this sleep-heavy synthetic workload, but real suites will differ.

6. Phase C — test-level split

pabot --testlevelsplit --processes 2 --outputdir evidence/04-testlevel-p2   --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot

pabot --testlevelsplit --processes 4 --outputdir evidence/05-testlevel-p4   --variable "LAB_ROOT:${root}" suites/alpha.robot suites/beta.robot

Compare throughput and overhead. Four processes cannot create more than four useful concurrent work units here because only four tests exist. On a real system, an external capacity ceiling may be lower.

7. Phase D — inject and preserve a collision

python -m robot --outputdir evidence/06-collision-serial --variable "LAB_ROOT:${root}" suites/collision.robot

# Expected diagnostic failure under overlap. Do not overwrite this directory.
pabot --testlevelsplit --processes 2 --outputdir evidence/07-collision-fail   --variable "LAB_ROOT:${root}" suites/collision.robot

Save the exact failure message, pabot_manager.log, both subprocess stdout/stderr files, partial outputs, and the state of ${LAB_ROOT}/lease. The first failure is a required artifact, not something to clean away immediately.

8. Phase E — repair and justify

Choose one:

  • Isolation repair: change the lease path to include a unique domain/test/work-item ID. Use this when the resource can be duplicated.
  • Narrow PabotLib lock: use locked.robot when the fixture is deliberately singular. Keep the lock around only the exclusive operation.
pabot --testlevelsplit --processes 2 --outputdir evidence/08-repaired   --variable "LAB_ROOT:${root}" suites/locked.robot

Your evidence must explain why the chosen repair matches the resource model. “It passed” is insufficient.

9. Phase F — compute a safe concurrency budget

Measure or estimate these values from your local run and synthetic target:

Budget input Example Your value
Host CPU budget 4 workers _____
Host memory budget 900 MiB available / 180 MiB peak ≈ 5 _____
Synthetic resource slots 3 independent slots _____
Target service safe clients 4 _____
CI jobs/shards hitting same target 2 _____

If the same target will receive two CI shards, the local process budget must account for the product. For example, a target ceiling of four concurrent clients with two simultaneous shards implies at most two Pabot processes per shard unless stronger isolation/capacity evidence justifies another plan.

target_limit = 4 concurrent clients
ci_shards_running_together = 2
pabot_processes_per_shard <= floor(4 / 2) = 2

Then also apply CPU, memory, data-slot, and port-slot ceilings.
Final safe value = minimum of all ceilings.

10. Optional local shard simulation

pabot --shard 1/2 --processes 2 --outputdir evidence/09-shard-1 suites/alpha.robot suites/beta.robot
pabot --shard 2/2 --processes 2 --outputdir evidence/10-shard-2 suites/alpha.robot suites/beta.robot

rebot --name "RF24 Local Shard Simulation" --output evidence/11-aggregate/output.xml   --log evidence/11-aggregate/log.html --report evidence/11-aggregate/report.html   evidence/09-shard-1/output.xml evidence/10-shard-2/output.xml

Verify that the two shard inputs together represent the intended test inventory and that the aggregate count matches the plan. In real CI these commands run on separate workers and the outputs must be collected before aggregation.

11. Required evidence packet

Artifact What it proves
versions.txt Exact Python, Robot Framework, and Pabot versions.
Serial timing/status Correct baseline before parallelism.
Pabot p1/p2/p4 timings Measured scheduler/process-count effect.
Worker marker tree PID/queue/pool identities and unique artifact ownership.
pabot_manager.log Scheduler/process lifecycle and timeout evidence.
Failed collision output directory Unmodified first-failure cause and subprocess evidence.
Repaired run output Ownership/lock correction restored correctness.
Capacity worksheet Why chosen process/shard counts are safe.
Optional shard aggregate Cross-shard result count and aggregation evidence.

12. Cleanup / rollback

Only delete the synthetic temp root after all evidence is captured. Before recursive deletion, normalize the path, prove it is below ${TEMPDIR}, and prove its basename begins with rf24-parallel-lab-. Never run a recursive delete against a variable that is empty, the temp root itself, a repository root, or an externally supplied arbitrary path.

*** Keywords ***
Guarded Cleanup Lab Root
    ${root}=    Normalize Path    ${LAB_ROOT}
    ${temp}=    Normalize Path    ${TEMPDIR}
    Should Be True    $root != $temp    Refusing to delete system temp root.
    ${root_abs}=    Evaluate    __import__("os").path.normcase(__import__("os").path.abspath($root))
    ${temp_abs}=    Evaluate    __import__("os").path.normcase(__import__("os").path.abspath($temp))
    ${common}=    Evaluate    __import__("os").path.commonpath([$root_abs, $temp_abs])
    Should Be Equal    ${common}    ${temp_abs}    Lab root is outside TEMPDIR.
    ${name}=    Evaluate    __import__("os").path.basename($root_abs)
    Should Start With    ${name}    rf24-parallel-lab-
    Remove Directory    ${root_abs}    recursive=True

Keep evidence/ until the lab has been reviewed; cleanup of results is a separate retention decision.

13. Verification checklist

  • Serial control passed with four expected tests.
  • Parallel control passed with unique worker/work-item artifact ownership.
  • Timings and PIDs were recorded, not inferred.
  • The collision failure was preserved before repair.
  • The repair changed resource ownership/serialization, not merely timeout/retry behavior.
  • Final merged test counts match the selected inventory.
  • No public/production system, real credential, shared customer account, or destructive data was touched.
  • The safe concurrency budget includes CI shards/jobs as well as local Pabot processes.

Knowledge check

The serial suite passes and the parallel suite fails on one shared file. What is the strongest first hypothesis?

Why must the checkpoint include a --processes 1 Pabot run in addition to plain Robot?

Two CI shards each run four Pabot processes against a service that safely supports six clients. Is the plan safe?

If a PabotLib lock fixes the collision, why might isolation still be better?

Which evidence should survive after a repaired run passes?

14. Production operating model and Chapter 25 bridge

This chapter adds a concurrency contract to the Robot Framework operating model: every parallel worker has an identity, every mutable resource has an owner or bounded allocator, total concurrency is budgeted across Pabot and CI, and every merge remains traceable to subprocess evidence. Chapter 25 will take that contract into GitHub Actions, GitLab CI, and Jenkins while preserving exit status and artifacts on both success and failure.

Next lesson

CI/CD Integration with GitHub Actions, GitLab CI, and Jenkins: Core Concepts and Mental Model

Continue with CI/CD Integration with GitHub Actions, GitLab CI, and Jenkins: Core Concepts and Mental Model. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

References and version anchors

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.