Chapter 24Lesson 04180–240 min

Parallel Execution with Pabot, Sharding, and Resource Contention: Diagnostics, Failure Modes, and Production Practices

Diagnose contention from evidence instead of adding sleeps and retries: reconstruct the worker schedule, identify the shared resource, bound total concurrency, preserve partial outputs, and apply the smallest ownership correction.

Race diagnosisDeadlocksArtifact collisionsCapacityFirst-failure evidence

Learning objectives

  • Apply a repeatable diagnostic sequence to parallel-only failures.
  • Recognize shared account/row/file/port collisions and distinguish them from ordinary Robot failures.
  • Detect concurrency multiplication when Pabot and CI sharding are combined.
  • Use Pabot manager/subprocess evidence to investigate stale workers, timeouts, and merge discrepancies.
  • Repair a shared-artifact race by changing ownership rather than hiding it with retries.

Current compatibility baseline — verified 2026-08-31. Robot Framework 7.4.2 is the stable course baseline and requires Python 3.8+. Pabot 5.2.2 is the stable parallel-runner baseline used in commands; Pabot 5.3.0b1 is prerelease and is not required. Pabot is an external runner, not Robot Framework core. In current Pabot, suite-level splitting is the default; --testlevelsplit opts into test-level scheduling; --processes caps local executors; --shard i/n partitions an execution for distribution; PabotLib provides cross-process locks/resource sets; and pabot_results/ plus pabot_manager.log contain subprocess evidence before final Rebot output is produced. Pabot 5.2.2 writes pabot_manager.log and per-queue subprocess evidence. Its --processtimeout can terminate overlong worker processes; current 5.2 logging records these termination events. The lab does not recommend using process timeout as a substitute for domain-level bounded waits.

1. Diagnostic sequence for a parallel failure

  1. Preserve the failed output directory before any rerun.
  2. Record python --version, python -m robot --version, and pabot --version.
  3. Record exact selection, split mode, process count, shard index/total, argument files, and output directory.
  4. Validate parse/import state with a dry run if the failure could be model-related.
  5. Inspect the first failing keyword/test and its worker’s robot_stdout.out, robot_stderr.out, partial output.xml, and marker identities.
  6. Identify the external state touched by that worker: file, port, account, row, queue, browser profile, service rate limit.
  7. Inspect pabot_manager.log for overlap, process termination, or scheduling clues.
  8. If CI is involved, multiply job/shard concurrency by local Pabot processes and compare with capacity.
  9. Apply the least destructive correction: unique ownership, smaller process count, explicit pool, or smallest possible lock.
  10. Rerun the smallest controlled slice into a new output directory.

2. Failure taxonomy

Symptom Likely layer Evidence to inspect Correction direction
Random wrong record/user Shared test data/account IDs in logs + external fixture state Unique data allocation/value sets
Address already in use Port ownership Worker PID/queue + listener/service log Dynamic/allocated port per worker
File content changes between write/read Shared path Per-worker output + file timestamps Queue/test-specific directories
429/throttle/timeouts only in CI External capacity Effective concurrency, service metrics Reduce workers/shards; authorized capacity plan
All workers waiting Lock deadlock Pabot manager log + lock order Single lock order; smaller critical sections
Final report missing expected items Selection/shard/merge boundary Shard commands + partial output counts Reconcile inventories before aggregation
Machine swaps heavily CPU/memory overcommit OS metrics + process count Reduce processes; measure footprint
Rerun “fixes” failure Race/flakiness likely Compare first-failure worker evidence Fix isolation; do not institutionalize rerun

3. Hidden multiplication: Pabot × CI

A common production mistake is independently tuning each layer. A team decides that eight Pabot processes are safe on a laptop; later a CI matrix launches six jobs against the same shared test environment. The effective concurrency can become 48. If the environment supports ten clients, failures may look like random API or database instability.

Do not solve capacity errors with retries. Retries can increase the number of requests precisely when the service is saturated. Reduce concurrency or allocate more authorized test capacity first.

4. Artifact collisions and missing evidence

Pabot separates its subprocess output directories and can collect configured artifact extensions. Problems reappear when test code overrides that design with one hard-coded shared path. The most reliable evidence layout includes run ID, shard ID if any, and execution/test identity.

evidence/
└── run-20260831-001/
    ├── shard-1/
    │   ├── output.xml
    │   └── pabot_results/...
    ├── shard-2/
    │   ├── output.xml
    │   └── pabot_results/...
    └── aggregate/
        ├── output.xml
        ├── log.html
        └── report.html

5. Lock deadlocks: correctness without liveness is still failure

Consider worker A acquiring account then database, while worker B acquires database then account. Both can wait forever. Prevent this by defining a single global lock order, avoiding nested locks where possible, and keeping critical sections short. A lock is not a reason to use giant process timeouts.

Bad:
Worker A: lock(account) -> lock(database)
Worker B: lock(database) -> lock(account)

Safer:
All workers: lock(account) -> lock(database)
Better when possible: allocate independent account + database namespace and use no lock.

6. Stale workers and process timeout

--processtimeout is a Pabot guardrail for a worker process that fails to finish. It is not the same as Robot test/keyword timeout and not the same as an external HTTP/database timeout. If it fires, inspect why the process is stuck before increasing it. Current Pabot manager logging records process kills/timeouts, which should be part of the evidence packet.

7. Intentionally broken example: a shared filename race

*** Settings ***
Library    OperatingSystem

*** Test Cases ***
Unsafe Parallel Artifact
    # BROKEN ON PURPOSE: all workers share one filename.
    Create Directory    ${LAB_ROOT}${/}shared
    Create File    ${LAB_ROOT}${/}shared${/}result.txt    ${TEST_NAME}
    Sleep    0.5s
    ${actual}=    Get File    ${LAB_ROOT}${/}shared${/}result.txt
    Should Be Equal    ${actual}    ${TEST_NAME}

Run two copies/tests at test-level parallelism. Both write the same path; whichever worker writes last can cause the other worker’s assertion to observe the wrong test name. The parser is fine, the Robot assertion is doing exactly what it should, and the operating system is not “random.” The ownership model is wrong.

Repair: derive an owned path

*** Settings ***
Library    OperatingSystem

*** Test Cases ***
Owned Parallel Artifact
    ${queue}=    Get Variable Value    \${PABOTQUEUEINDEX}    serial
    ${owned}=    Set Variable    ${LAB_ROOT}${/}workers${/}${queue}
    Create Directory    ${owned}
    Create File    ${owned}${/}${TEST_NAME}.txt    ${TEST_NAME}
    ${actual}=    Get File    ${owned}${/}${TEST_NAME}.txt
    Should Be Equal    ${actual}    ${TEST_NAME}

The repair changes the resource identity rather than timing. Because queue index is diagnostic scheduler identity, production code should prefer a domain/test-specific unique ID when the artifact must remain meaningful across reruns; for this disposable worker-evidence lab, queue ownership is appropriate.

8. Merged report mismatch

If the final test count differs from expectation, do not start by rerunning. Compare the selected inventory, .pabotsuitenames/ordering inputs, per-worker partial outputs, shard commands, and final Rebot inputs. A merge cannot create a test that was never selected or whose shard output was never collected.

9. Performance diagnosis by layer

Time component Typical symptom Action
Parse/import/startup Parallel run not faster for tiny tests Use coarser units/chunking only after measuring
Robot keyword time Workers busy doing local logic Profile keyword/domain code
External latency Workers mostly waiting Check service capacity and request timing
Logging/output Large output.xml and disk I/O Reduce unnecessary logging; preserve useful evidence
Pabot scheduling Many tiny execution units Use suite-level/coarser split
CI/container startup Shards spend time bootstrapping Cache/pin responsibly; avoid excess shards
Retry/rerun Total request/load grows after failures Fix isolation/capacity before retry policy

Knowledge check

A test fails with “address already in use” only under four workers. Which layer should be inspected first?

Why can a lock make a test correct but the suite still operationally bad?

What should you do before rerunning a failed Pabot run into the same output directory?

A final aggregate has fewer tests than planned. Why is “Rebot is broken” not the first conclusion?

10. Summary and bridge

Parallel failures become tractable when every worker, resource, and result artifact has an identity. The repair hierarchy is: preserve evidence, classify the layer, isolate state, bound capacity, then use locks only when sharing is unavoidable. Lesson 5 combines these principles into a checkpoint lab and produces a reviewable concurrency budget.

Next lesson

Checkpoint Lab — Parallel Execution with Pabot, Sharding, and Resource Contention

Continue with Checkpoint Lab — Parallel Execution with Pabot, Sharding, and Resource Contention. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

References and version anchors

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.