Parallel Execution with Pabot, Sharding, and Resource Contention: Core Concepts and Mental Model
Build a correct concurrency mental model before adding workers: Pabot schedules independent Robot processes, while test data, ports, accounts, files, external services, CI shards, and result aggregation each retain their own capacity and ownership rules.
Learning objectives
- Explain the execution path from suite inventory through Pabot scheduling, independent Robot processes, per-process outputs, and final merged evidence.
- Separate Robot core execution, Pabot local process parallelism, PabotLib synchronization, and CI-level sharding/concurrency.
- Identify every mutable resource that must be uniquely owned or explicitly serialized before parallel execution.
- Use Pabot-provided execution variables as diagnostic identities without treating them as durable business identifiers.
- Derive a safe concurrency limit from CPU, memory, external-system capacity, and available isolated resource slots rather than CPU count alone.
Current compatibility baseline — verified 2026-08-31.
Robot Framework 7.4.2 is the stable course baseline
and requires Python 3.8+. Pabot 5.2.2 is the stable
parallel-runner baseline used in commands; Pabot 5.3.0b1 is
prerelease and is not required. Pabot is an external runner, not
Robot Framework core. In current Pabot, suite-level splitting is the
default; --testlevelsplit opts into test-level
scheduling; --processes caps local executors;
--shard i/n partitions an execution for distribution;
PabotLib provides cross-process locks/resource sets; and
pabot_results/ plus
pabot_manager.log contain subprocess evidence before
final Rebot output is produced. The mandatory examples use only
local files, short sleeps, synthetic identifiers, and Python
standard-library helpers; they never require a shared production
account, database, port, browser cloud, or hosted CI runner.
1. The practical problem: serial correctness does not imply parallel correctness
A Robot suite can be perfectly deterministic when one process owns all resources and still become flaky the moment two processes run simultaneously. The usual cause is not “Pabot is unstable.” It is that serial execution concealed an ownership assumption: two tests use the same filename, row, user account, TCP port, browser profile, approval queue, or cleanup directory.
Parallelism therefore starts as a state-design problem, not a command-line optimization. Pabot can schedule work faster, but it cannot infer which external resources are safe to share.
Safety boundary. Do not use this chapter to discover the concurrency limits of production systems. Capacity experiments belong to disposable local fixtures or explicitly authorized performance/test environments. A faster test command is not authorization to multiply requests against real services.
2. Read-only inspection before adding concurrency
python --version
python -m robot --version
pabot --version
# Read the suite inventory without mutating external state.
python -m robot --dryrun --outputdir evidence/dryrun suites
# Inspect available CPU only as one input to a capacity decision.
python -c "import os; print('cpu_count=', os.cpu_count())"
These checks establish interpreter ownership, Robot/Pabot versions, parse/import health, suite names, and a rough CPU ceiling. The dry run does not prove that runtime resources are isolated; it only proves that Robot can build and resolve the selected model.
3. Mental model: three concurrency layers, several state stores
flowchart TD A[Suite / test inventory] --> B[CI shard selection] B --> C[Pabot scheduler in shard] C --> D1[Robot process 1] C --> D2[Robot process 2] D1 --> E1[Owned files / data / ports] D2 --> E2[Owned files / data / ports] D1 --> F1[Partial output.xml + stdout/stderr] D2 --> F2[Partial output.xml + stdout/stderr] F1 --> G[Local Pabot / Rebot aggregation] F2 --> G G --> H[Shard output.xml / log / report] H --> I[Cross-shard aggregation / CI evidence] J[PabotLib] -. narrow locks or value sets .-> D1 J -. narrow locks or value sets .-> D2
The first scheduler may be CI: two jobs receive disjoint shards. Inside each job, Pabot may launch multiple Robot processes. Each Robot process has its own Python interpreter state, Robot variables, library instances, and output directory. External systems do not automatically become independent; you must allocate them.
PabotLib is a coordination service. Its locks can serialize a truly singular fixture, and its value sets can distribute a finite pool of credentials or endpoints. Neither changes the fact that independent resources are preferable because locking reduces throughput and can create deadlock/liveness concerns.
4. State inventory: what must have an owner?
| Layer | Example | Correct ownership question | Failure if ignored |
|---|---|---|---|
| Robot process | Python/library instance state | Can one process observe another process's in-memory state? Usually no. | Assuming GLOBAL library scope crosses processes. It does not. |
| Filesystem | artifact or scratch path | Does every scheduled execution write a unique path? | Overwrite, cleanup races, missing screenshots. |
| Network | local port | Who allocates and releases this port? | Bind collisions or test connects to wrong server. |
| Data | row/order/work-item/account | Is each test given a unique identity? | Duplicate mutation, stale reads, order dependence. |
| External service | rate/concurrency limit | How many simultaneous operations are allowed? | 429s, throttling, queue saturation, cascading retries. |
| PabotLib | lock/value set | Is sharing unavoidable and bounded? | Lock contention, starvation, deadlock if lock order differs. |
| CI | jobs × Pabot processes | What is total effective concurrency? | 4 CI jobs × 8 processes becomes 32 clients. |
| Results | pabot_results and final output.xml | Where is first-failure evidence preserved? | Rerun overwrites diagnostic data or artifacts collide. |
5. Pabot execution identity is diagnostic metadata
Pabot injects variables such as ${PABOTQUEUEINDEX},
${PABOTEXECUTIONPOOLID},
${PABOTNUMBEROFPROCESSES}, ${PABOTLIBURI},
and ${CALLER_ID}. Queue indexes start from zero and
identify scheduled execution items; pool IDs identify executor
slots. These are useful for worker-specific artifact paths and
traces.
*** Keywords ***
Resolve Execution Identity
${queue}= Get Variable Value \${PABOTQUEUEINDEX} serial
${pool}= Get Variable Value \${PABOTEXECUTIONPOOLID} serial
${pid}= Evaluate __import__("os").getpid()
Log queue=${queue} pool=${pool} pid=${pid}
RETURN ${queue} ${pool} ${pid}
The escaped variable names are deliberate: a normal serial
robot run does not define Pabot variables, so
Get Variable Value must receive the variable name as
data and provide a fallback. Do not store business records under a
Pabot queue index and expect that identity to survive a different
scheduling plan or rerun.
6. Suite-level versus test-level splitting
| Mode | Scheduling unit | Benefits | Costs / caveats |
|---|---|---|---|
| Default | Suite | Preserves tests of one suite in one Robot process; suite setup/teardown amortized. | A single slow suite can dominate the critical path. |
--testlevelsplit |
Test | Finer-grained balancing when tests inside suites are independent. | Suite-level assumptions and expensive setups may be repeated/handled differently across executions; isolation pressure rises. |
| CI shard + Pabot | Shard, then local suite/test units | Scales across machines and cores. | Two schedulers multiply concurrency; shard outputs require explicit aggregation and artifact ownership. |
Choose the coarsest split that meets the feedback-time goal. More granular splitting is not automatically better; it increases scheduling overhead and makes hidden suite-level coupling more visible.
7. Result topology: preserve subprocess evidence before judging the merge
A current Pabot run writes normal final Robot artifacts and a
pabot_results/ tree containing per-queue-index partial
output.xml, robot_stdout.out,
robot_stderr.out, robot_argfile.txt, and
pabot_manager.log. The manager log is especially useful
for reconstructing process start/finish and timeouts.
Evidence rule. Pabot documents
pabot_results/ as temporary and it can be overwritten
by a later run using the same output directory. Use a unique
output directory per diagnostic attempt, or archive/copy the
failed run before rerunning. Never delete the first-failure
evidence merely to get a clean report.
8. Capacity model: the smallest ceiling wins
A safe process count is constrained by more than cores. Suppose one worker needs about 180 MiB of available memory, the test host has 900 MiB reserved for this job, the local service permits four concurrent clients, and only three isolated synthetic accounts exist. The memory ceiling is 5, service ceiling 4, account ceiling 3; even on an eight-core host, the safe budget is at most 3.
safe_processes = min(
CPU budget,
memory budget / memory per worker,
external-service concurrent capacity,
number of isolated accounts/data slots,
number of safe port/browser/profile slots
)
Measure, then increase gradually. The objective is throughput with deterministic evidence—not maximum process count.
9. DevOps connection: parallelism is an operating contract
Once parallel execution enters CI, the run contract must name process count, shard count, data-allocation strategy, per-worker artifact path, external capacity, lock use, merge strategy, and failure-artifact retention. This makes a parallel failure explainable instead of “rerun it; it is probably timing.”
Knowledge check
If a Python test library has GLOBAL scope, is the object automatically shared by all Pabot workers?
No. GLOBAL scope is global inside one Robot process. Pabot launches separate Robot processes, so Python memory is not automatically shared across them.
Why is --processes all risky even when the machine
has plenty of CPU?
Because external systems, test accounts, ports, memory, browser profiles, or other finite resources may have a lower capacity than the number of executable suites/tests.
What does --testlevelsplit change?
It changes the Pabot scheduling unit from the default suite-level split to test-level split. It does not make shared external state safe.
A failed parallel run is immediately rerun into the same output directory. What evidence can be lost?
Current Pabot treats pabot_results as temporary and it may be overwritten. Preserve or use a unique output directory for the first failed run before rerunning.
When is a PabotLib lock justified?
When a resource is genuinely singular or has a deliberately limited pool and redesigning for independent ownership is not feasible. Isolation is preferable when possible.
10. Summary and bridge
Pabot gives Robot Framework a scheduler and multiple worker processes; it does not grant extra capacity to the system under test. Correct parallelism comes from explicit resource ownership, bounded concurrency, per-process evidence, and disciplined aggregation. Lesson 2 turns this model into a local experiment with measurable serial and parallel runs.
References and version anchors
- Robot Framework PyPI — current stable/pre-release and Python support
- Robot Framework 7.4.2 User Guide — execution, variables, result files, Rebot, and core semantics
- robotframework-pabot PyPI — stable 5.2.2 and prerelease status
- Pabot repository README — current CLI, PabotLib, ordering, global variables, sharding, and output/artifact behavior
- PabotLib keyword documentation — locks and shared resource distribution
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.