Chapter 24Lesson 01180–240 min

Parallel Execution with Pabot, Sharding, and Resource Contention: Core Concepts and Mental Model

Build a correct concurrency mental model before adding workers: Pabot schedules independent Robot processes, while test data, ports, accounts, files, external services, CI shards, and result aggregation each retain their own capacity and ownership rules.

Robot Framework 7.4.2Pabot 5.2.2Isolation firstShardingEvidence integrity

Learning objectives

  • Explain the execution path from suite inventory through Pabot scheduling, independent Robot processes, per-process outputs, and final merged evidence.
  • Separate Robot core execution, Pabot local process parallelism, PabotLib synchronization, and CI-level sharding/concurrency.
  • Identify every mutable resource that must be uniquely owned or explicitly serialized before parallel execution.
  • Use Pabot-provided execution variables as diagnostic identities without treating them as durable business identifiers.
  • Derive a safe concurrency limit from CPU, memory, external-system capacity, and available isolated resource slots rather than CPU count alone.

Current compatibility baseline — verified 2026-08-31. Robot Framework 7.4.2 is the stable course baseline and requires Python 3.8+. Pabot 5.2.2 is the stable parallel-runner baseline used in commands; Pabot 5.3.0b1 is prerelease and is not required. Pabot is an external runner, not Robot Framework core. In current Pabot, suite-level splitting is the default; --testlevelsplit opts into test-level scheduling; --processes caps local executors; --shard i/n partitions an execution for distribution; PabotLib provides cross-process locks/resource sets; and pabot_results/ plus pabot_manager.log contain subprocess evidence before final Rebot output is produced. The mandatory examples use only local files, short sleeps, synthetic identifiers, and Python standard-library helpers; they never require a shared production account, database, port, browser cloud, or hosted CI runner.

1. The practical problem: serial correctness does not imply parallel correctness

A Robot suite can be perfectly deterministic when one process owns all resources and still become flaky the moment two processes run simultaneously. The usual cause is not “Pabot is unstable.” It is that serial execution concealed an ownership assumption: two tests use the same filename, row, user account, TCP port, browser profile, approval queue, or cleanup directory.

Parallelism therefore starts as a state-design problem, not a command-line optimization. Pabot can schedule work faster, but it cannot infer which external resources are safe to share.

Safety boundary. Do not use this chapter to discover the concurrency limits of production systems. Capacity experiments belong to disposable local fixtures or explicitly authorized performance/test environments. A faster test command is not authorization to multiply requests against real services.

2. Read-only inspection before adding concurrency

python --version
python -m robot --version
pabot --version

# Read the suite inventory without mutating external state.
python -m robot --dryrun --outputdir evidence/dryrun suites

# Inspect available CPU only as one input to a capacity decision.
python -c "import os; print('cpu_count=', os.cpu_count())"

These checks establish interpreter ownership, Robot/Pabot versions, parse/import health, suite names, and a rough CPU ceiling. The dry run does not prove that runtime resources are isolated; it only proves that Robot can build and resolve the selected model.

3. Mental model: three concurrency layers, several state stores

Local Pabot and CI sharding are different schedulers
flowchart TD
A[Suite / test inventory] --> B[CI shard selection]
B --> C[Pabot scheduler in shard]
C --> D1[Robot process 1]
C --> D2[Robot process 2]
D1 --> E1[Owned files / data / ports]
D2 --> E2[Owned files / data / ports]
D1 --> F1[Partial output.xml + stdout/stderr]
D2 --> F2[Partial output.xml + stdout/stderr]
F1 --> G[Local Pabot / Rebot aggregation]
F2 --> G
G --> H[Shard output.xml / log / report]
H --> I[Cross-shard aggregation / CI evidence]
J[PabotLib] -. narrow locks or value sets .-> D1
J -. narrow locks or value sets .-> D2

The first scheduler may be CI: two jobs receive disjoint shards. Inside each job, Pabot may launch multiple Robot processes. Each Robot process has its own Python interpreter state, Robot variables, library instances, and output directory. External systems do not automatically become independent; you must allocate them.

PabotLib is a coordination service. Its locks can serialize a truly singular fixture, and its value sets can distribute a finite pool of credentials or endpoints. Neither changes the fact that independent resources are preferable because locking reduces throughput and can create deadlock/liveness concerns.

4. State inventory: what must have an owner?

Layer Example Correct ownership question Failure if ignored
Robot process Python/library instance state Can one process observe another process's in-memory state? Usually no. Assuming GLOBAL library scope crosses processes. It does not.
Filesystem artifact or scratch path Does every scheduled execution write a unique path? Overwrite, cleanup races, missing screenshots.
Network local port Who allocates and releases this port? Bind collisions or test connects to wrong server.
Data row/order/work-item/account Is each test given a unique identity? Duplicate mutation, stale reads, order dependence.
External service rate/concurrency limit How many simultaneous operations are allowed? 429s, throttling, queue saturation, cascading retries.
PabotLib lock/value set Is sharing unavoidable and bounded? Lock contention, starvation, deadlock if lock order differs.
CI jobs × Pabot processes What is total effective concurrency? 4 CI jobs × 8 processes becomes 32 clients.
Results pabot_results and final output.xml Where is first-failure evidence preserved? Rerun overwrites diagnostic data or artifacts collide.

5. Pabot execution identity is diagnostic metadata

Pabot injects variables such as ${PABOTQUEUEINDEX}, ${PABOTEXECUTIONPOOLID}, ${PABOTNUMBEROFPROCESSES}, ${PABOTLIBURI}, and ${CALLER_ID}. Queue indexes start from zero and identify scheduled execution items; pool IDs identify executor slots. These are useful for worker-specific artifact paths and traces.

*** Keywords ***
Resolve Execution Identity
    ${queue}=    Get Variable Value    \${PABOTQUEUEINDEX}    serial
    ${pool}=     Get Variable Value    \${PABOTEXECUTIONPOOLID}    serial
    ${pid}=      Evaluate    __import__("os").getpid()
    Log    queue=${queue} pool=${pool} pid=${pid}
    RETURN    ${queue}    ${pool}    ${pid}

The escaped variable names are deliberate: a normal serial robot run does not define Pabot variables, so Get Variable Value must receive the variable name as data and provide a fallback. Do not store business records under a Pabot queue index and expect that identity to survive a different scheduling plan or rerun.

6. Suite-level versus test-level splitting

Mode Scheduling unit Benefits Costs / caveats
Default Suite Preserves tests of one suite in one Robot process; suite setup/teardown amortized. A single slow suite can dominate the critical path.
--testlevelsplit Test Finer-grained balancing when tests inside suites are independent. Suite-level assumptions and expensive setups may be repeated/handled differently across executions; isolation pressure rises.
CI shard + Pabot Shard, then local suite/test units Scales across machines and cores. Two schedulers multiply concurrency; shard outputs require explicit aggregation and artifact ownership.

Choose the coarsest split that meets the feedback-time goal. More granular splitting is not automatically better; it increases scheduling overhead and makes hidden suite-level coupling more visible.

7. Result topology: preserve subprocess evidence before judging the merge

A current Pabot run writes normal final Robot artifacts and a pabot_results/ tree containing per-queue-index partial output.xml, robot_stdout.out, robot_stderr.out, robot_argfile.txt, and pabot_manager.log. The manager log is especially useful for reconstructing process start/finish and timeouts.

Evidence rule. Pabot documents pabot_results/ as temporary and it can be overwritten by a later run using the same output directory. Use a unique output directory per diagnostic attempt, or archive/copy the failed run before rerunning. Never delete the first-failure evidence merely to get a clean report.

8. Capacity model: the smallest ceiling wins

A safe process count is constrained by more than cores. Suppose one worker needs about 180 MiB of available memory, the test host has 900 MiB reserved for this job, the local service permits four concurrent clients, and only three isolated synthetic accounts exist. The memory ceiling is 5, service ceiling 4, account ceiling 3; even on an eight-core host, the safe budget is at most 3.

safe_processes = min(
    CPU budget,
    memory budget / memory per worker,
    external-service concurrent capacity,
    number of isolated accounts/data slots,
    number of safe port/browser/profile slots
)

Measure, then increase gradually. The objective is throughput with deterministic evidence—not maximum process count.

9. DevOps connection: parallelism is an operating contract

Once parallel execution enters CI, the run contract must name process count, shard count, data-allocation strategy, per-worker artifact path, external capacity, lock use, merge strategy, and failure-artifact retention. This makes a parallel failure explainable instead of “rerun it; it is probably timing.”

Knowledge check

If a Python test library has GLOBAL scope, is the object automatically shared by all Pabot workers?

Why is --processes all risky even when the machine has plenty of CPU?

What does --testlevelsplit change?

A failed parallel run is immediately rerun into the same output directory. What evidence can be lost?

When is a PabotLib lock justified?

10. Summary and bridge

Pabot gives Robot Framework a scheduler and multiple worker processes; it does not grant extra capacity to the system under test. Correct parallelism comes from explicit resource ownership, bounded concurrency, per-process evidence, and disciplined aggregation. Lesson 2 turns this model into a local experiment with measurable serial and parallel runs.

Next lesson

Parallel Execution with Pabot, Sharding, and Resource Contention: Guided Hands-On Workflow

Continue with Parallel Execution with Pabot, Sharding, and Resource Contention: Guided Hands-On Workflow. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

References and version anchors

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.