Parallel Execution with Pabot, Sharding, and Resource Contention: Diagnostics, Failure Modes, and Production Practices
Diagnose contention from evidence instead of adding sleeps and retries: reconstruct the worker schedule, identify the shared resource, bound total concurrency, preserve partial outputs, and apply the smallest ownership correction.
Learning objectives
- Apply a repeatable diagnostic sequence to parallel-only failures.
- Recognize shared account/row/file/port collisions and distinguish them from ordinary Robot failures.
- Detect concurrency multiplication when Pabot and CI sharding are combined.
- Use Pabot manager/subprocess evidence to investigate stale workers, timeouts, and merge discrepancies.
- Repair a shared-artifact race by changing ownership rather than hiding it with retries.
Current compatibility baseline — verified 2026-08-31.
Robot Framework 7.4.2 is the stable course baseline
and requires Python 3.8+. Pabot 5.2.2 is the stable
parallel-runner baseline used in commands; Pabot 5.3.0b1 is
prerelease and is not required. Pabot is an external runner, not
Robot Framework core. In current Pabot, suite-level splitting is the
default; --testlevelsplit opts into test-level
scheduling; --processes caps local executors;
--shard i/n partitions an execution for distribution;
PabotLib provides cross-process locks/resource sets; and
pabot_results/ plus
pabot_manager.log contain subprocess evidence before
final Rebot output is produced. Pabot 5.2.2 writes
pabot_manager.log and per-queue subprocess evidence.
Its --processtimeout can terminate overlong worker
processes; current 5.2 logging records these termination events. The
lab does not recommend using process timeout as a substitute for
domain-level bounded waits.
1. Diagnostic sequence for a parallel failure
- Preserve the failed output directory before any rerun.
-
Record
python --version,python -m robot --version, andpabot --version. - Record exact selection, split mode, process count, shard index/total, argument files, and output directory.
- Validate parse/import state with a dry run if the failure could be model-related.
-
Inspect the first failing keyword/test and its worker’s
robot_stdout.out,robot_stderr.out, partialoutput.xml, and marker identities. - Identify the external state touched by that worker: file, port, account, row, queue, browser profile, service rate limit.
-
Inspect
pabot_manager.logfor overlap, process termination, or scheduling clues. - If CI is involved, multiply job/shard concurrency by local Pabot processes and compare with capacity.
- Apply the least destructive correction: unique ownership, smaller process count, explicit pool, or smallest possible lock.
- Rerun the smallest controlled slice into a new output directory.
2. Failure taxonomy
| Symptom | Likely layer | Evidence to inspect | Correction direction |
|---|---|---|---|
| Random wrong record/user | Shared test data/account | IDs in logs + external fixture state | Unique data allocation/value sets |
| Address already in use | Port ownership | Worker PID/queue + listener/service log | Dynamic/allocated port per worker |
| File content changes between write/read | Shared path | Per-worker output + file timestamps | Queue/test-specific directories |
| 429/throttle/timeouts only in CI | External capacity | Effective concurrency, service metrics | Reduce workers/shards; authorized capacity plan |
| All workers waiting | Lock deadlock | Pabot manager log + lock order | Single lock order; smaller critical sections |
| Final report missing expected items | Selection/shard/merge boundary | Shard commands + partial output counts | Reconcile inventories before aggregation |
| Machine swaps heavily | CPU/memory overcommit | OS metrics + process count | Reduce processes; measure footprint |
| Rerun “fixes” failure | Race/flakiness likely | Compare first-failure worker evidence | Fix isolation; do not institutionalize rerun |
3. Hidden multiplication: Pabot × CI
A common production mistake is independently tuning each layer. A team decides that eight Pabot processes are safe on a laptop; later a CI matrix launches six jobs against the same shared test environment. The effective concurrency can become 48. If the environment supports ten clients, failures may look like random API or database instability.
Do not solve capacity errors with retries. Retries can increase the number of requests precisely when the service is saturated. Reduce concurrency or allocate more authorized test capacity first.
4. Artifact collisions and missing evidence
Pabot separates its subprocess output directories and can collect configured artifact extensions. Problems reappear when test code overrides that design with one hard-coded shared path. The most reliable evidence layout includes run ID, shard ID if any, and execution/test identity.
evidence/
└── run-20260831-001/
├── shard-1/
│ ├── output.xml
│ └── pabot_results/...
├── shard-2/
│ ├── output.xml
│ └── pabot_results/...
└── aggregate/
├── output.xml
├── log.html
└── report.html
5. Lock deadlocks: correctness without liveness is still failure
Consider worker A acquiring account then
database, while worker B acquires
database then account. Both can wait
forever. Prevent this by defining a single global lock order,
avoiding nested locks where possible, and keeping critical sections
short. A lock is not a reason to use giant process timeouts.
Bad:
Worker A: lock(account) -> lock(database)
Worker B: lock(database) -> lock(account)
Safer:
All workers: lock(account) -> lock(database)
Better when possible: allocate independent account + database namespace and use no lock.
6. Stale workers and process timeout
--processtimeout is a Pabot guardrail for a worker
process that fails to finish. It is not the same as Robot
test/keyword timeout and not the same as an external HTTP/database
timeout. If it fires, inspect why the process is stuck before
increasing it. Current Pabot manager logging records process
kills/timeouts, which should be part of the evidence packet.
7. Intentionally broken example: a shared filename race
*** Settings ***
Library OperatingSystem
*** Test Cases ***
Unsafe Parallel Artifact
# BROKEN ON PURPOSE: all workers share one filename.
Create Directory ${LAB_ROOT}${/}shared
Create File ${LAB_ROOT}${/}shared${/}result.txt ${TEST_NAME}
Sleep 0.5s
${actual}= Get File ${LAB_ROOT}${/}shared${/}result.txt
Should Be Equal ${actual} ${TEST_NAME}
Run two copies/tests at test-level parallelism. Both write the same path; whichever worker writes last can cause the other worker’s assertion to observe the wrong test name. The parser is fine, the Robot assertion is doing exactly what it should, and the operating system is not “random.” The ownership model is wrong.
Repair: derive an owned path
*** Settings ***
Library OperatingSystem
*** Test Cases ***
Owned Parallel Artifact
${queue}= Get Variable Value \${PABOTQUEUEINDEX} serial
${owned}= Set Variable ${LAB_ROOT}${/}workers${/}${queue}
Create Directory ${owned}
Create File ${owned}${/}${TEST_NAME}.txt ${TEST_NAME}
${actual}= Get File ${owned}${/}${TEST_NAME}.txt
Should Be Equal ${actual} ${TEST_NAME}
The repair changes the resource identity rather than timing. Because queue index is diagnostic scheduler identity, production code should prefer a domain/test-specific unique ID when the artifact must remain meaningful across reruns; for this disposable worker-evidence lab, queue ownership is appropriate.
8. Merged report mismatch
If the final test count differs from expectation, do not start by
rerunning. Compare the selected inventory,
.pabotsuitenames/ordering inputs, per-worker partial
outputs, shard commands, and final Rebot inputs. A merge cannot
create a test that was never selected or whose shard output was
never collected.
9. Performance diagnosis by layer
| Time component | Typical symptom | Action |
|---|---|---|
| Parse/import/startup | Parallel run not faster for tiny tests | Use coarser units/chunking only after measuring |
| Robot keyword time | Workers busy doing local logic | Profile keyword/domain code |
| External latency | Workers mostly waiting | Check service capacity and request timing |
| Logging/output | Large output.xml and disk I/O | Reduce unnecessary logging; preserve useful evidence |
| Pabot scheduling | Many tiny execution units | Use suite-level/coarser split |
| CI/container startup | Shards spend time bootstrapping | Cache/pin responsibly; avoid excess shards |
| Retry/rerun | Total request/load grows after failures | Fix isolation/capacity before retry policy |
Knowledge check
A test fails with “address already in use” only under four workers. Which layer should be inspected first?
The external port allocation/ownership layer. Verify which worker/PID attempted which port; do not add a Robot retry until ownership is fixed.
Why can a lock make a test correct but the suite still operationally bad?
A broad lock can serialize the critical path, create long waits, and introduce deadlock/starvation. Correctness and liveness/throughput are separate concerns.
What should you do before rerunning a failed Pabot run into the same output directory?
Preserve/archive the first failed output directory or choose a new output directory, because pabot_results is temporary and can be overwritten.
A final aggregate has fewer tests than planned. Why is “Rebot is broken” not the first conclusion?
Selection, sharding, missing shard artifacts, or scheduling inventory may have omitted tests before aggregation. Reconcile those inputs first.
10. Summary and bridge
Parallel failures become tractable when every worker, resource, and result artifact has an identity. The repair hierarchy is: preserve evidence, classify the layer, isolate state, bound capacity, then use locks only when sharing is unavoidable. Lesson 5 combines these principles into a checkpoint lab and produces a reviewable concurrency budget.
References and version anchors
- Robot Framework PyPI — current stable/pre-release and Python support
- Robot Framework 7.4.2 User Guide — execution, variables, result files, Rebot, and core semantics
- robotframework-pabot PyPI — stable 5.2.2 and prerelease status
- Pabot repository README — current CLI, PabotLib, ordering, global variables, sharding, and output/artifact behavior
- PabotLib keyword documentation — locks and shared resource distribution
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.