Chapter 24Lesson 03180–240 min

Parallel Execution with Pabot, Sharding, and Resource Contention: Configuration, Design Patterns, and Trade-Offs

Turn the lab into production-oriented decisions: choose split granularity, process and shard budgets, isolation versus locking, setup strategy, ordering, and artifact handling based on the state that each choice actually changes.

Design trade-offsCapacity budgetPabot vs CIPabotLibArtifacts

Learning objectives

  • Choose suite-level or test-level splitting from fixture/lifecycle boundaries rather than preference.
  • Separate local Pabot process count from CI shard/job concurrency and compute effective concurrency.
  • Prefer independent resource allocation over locking and recognize cases where a lock or value set is justified.
  • Use current ordering/sharding/artifact options without confusing selection, scheduling, and result aggregation.
  • Document a measurable safe-concurrency budget that can be reviewed in CI.

Current compatibility baseline — verified 2026-08-31. Robot Framework 7.4.2 is the stable course baseline and requires Python 3.8+. Pabot 5.2.2 is the stable parallel-runner baseline used in commands; Pabot 5.3.0b1 is prerelease and is not required. Pabot is an external runner, not Robot Framework core. In current Pabot, suite-level splitting is the default; --testlevelsplit opts into test-level scheduling; --processes caps local executors; --shard i/n partitions an execution for distribution; PabotLib provides cross-process locks/resource sets; and pabot_results/ plus pabot_manager.log contain subprocess evidence before final Rebot output is produced. Pabot 5.2.2 includes current --ordering static/dynamic modes, --shard, --chunk, artifact collection, and PabotLib. Dynamic ordering was introduced in the 5.2 line. The lesson treats these as Pabot configuration, separate from Robot Framework core selection and from CI-provider matrix/job syntax.

1. Suite-level or test-level split?

Question Prefer suite-level default Consider --testlevelsplit
Fixture ownership Suite setup owns one isolated fixture used by tests in that suite. Each test can independently allocate and release its own fixture.
Setup cost Expensive setup should be amortized within a suite. Setup is cheap or test-owned.
Load balance Suites have similar duration. One suite contains many uneven long-running tests.
Diagnostics Suite grouping is meaningful operational evidence. Per-test work units give better scheduling without losing interpretability.
Risk A long suite may create a tail. Hidden order/shared-state coupling becomes visible.

Test-level splitting is not a correctness feature. It is a scheduling choice that should be enabled only when tests are independently runnable.

2. Process count: CPU is a ceiling, not a target

Pabot’s documented default process count is the smaller of two and CPU count. You can set an explicit value or all, but a production run should normally use a reviewed number tied to capacity measurements.

Constraint How to measure Example ceiling
CPU Observe utilization during representative run 4
Memory Available job memory ÷ peak worker footprint 6
Synthetic account pool Count independently resettable accounts 3
Local service Known safe concurrent clients 5
Database/API test quota Authorized test-environment limit 2

For the example above, safe local Pabot concurrency is 2. Increasing processes beyond that merely creates contention or throttling.

3. Pabot versus CI sharding: multiply deliberately

Pabot uses multiple Robot processes inside one job/host. CI sharding uses multiple jobs/machines. If three CI shards each run --processes 4, the system under test can see up to twelve concurrent Robot executors. That number—not “4”—must fit the external capacity budget.

effective_concurrency ≈ ci_jobs_running_same_target × pabot_processes_per_job

Example:
3 shards × 4 Pabot processes = up to 12 Robot executors
If the test API supports 6 clients safely, reduce either layer.

Use Pabot’s --shard i/n when the same suite inventory must be deterministically partitioned across machines. Give every shard a distinct output directory. Aggregate disjoint shard outputs only after all shards finish; the CI system owns that cross-machine collection boundary.

# Two independent jobs/machines receive the same sources and options.
pabot --shard 1/2 --processes 2 --outputdir evidence/shard-1 suites
pabot --shard 2/2 --processes 2 --outputdir evidence/shard-2 suites

# After collecting both output.xml files in one aggregation job:
rebot --name "RF24 Sharded Run" --output evidence/sharded-output.xml   --log evidence/sharded-log.html --report evidence/sharded-report.html   evidence/shard-1/output.xml evidence/shard-2/output.xml

4. Isolation versus PabotLib locking

Approach Use when Throughput Main operational risk
Unique resource per execution Files, rows, accounts, ports can be provisioned independently Highest Leaked resources if cleanup/TTL is weak
PabotLib value set/resource pool Finite interchangeable resources exist Bounded by pool size Exhaustion if sets are not released
PabotLib lock One singular non-replicable fixture Serial within critical section Deadlock/starvation and hidden bottleneck
Disable parallelism for tagged group Semantics are inherently exclusive Lowest for that group Long tail if overused

Locks should be small and named after the resource, not the test. Acquire them as late as possible and release them in teardown/FINALLY. If tests need several locks, define one lock order to prevent circular waits.

5. Ordering and dependencies: scheduling is not selection

Pabot can use a separate --ordering file. Current 5.2.x supports static and dynamic modes; dynamic mode can honor dependency relationships and choose whether dependents are skipped or still run after dependency failure. This controls scheduling/order, but it does not replace Robot selection options such as --include, --exclude, --suite, or --test.

Maintain an explicit ordering file when ordering is intentional. Pabot’s .pabotsuitenames is an internal/cache file and may be overwritten. Do not treat hand-edits to it as source-controlled policy.

6. One large setup versus duplicated setup

Parallelism can duplicate suite startup work. Pabot’s --chunk option groups execution into a limited number of Robot runs, which can reduce repeated setup/teardown cost, but it also changes execution grouping. Measure before using it and do not use chunking to conceal state-sharing defects.

Pattern Benefit Trade-off
Many small suites Fine scheduling Repeated imports/setup can dominate
Chunked Robot runs Amortizes setup across chunks Coarser isolation and different lifecycle grouping
External reusable fixture Workers attach to pre-created test fixture Requires explicit ownership/capacity/cleanup outside Robot
Run Setup Only Once via PabotLib Avoids duplicate shared setup Creates cross-worker dependency; failure propagates and resource must be safe to share

7. Artifact strategy under concurrency

Pabot copies configured artifact extensions and uses queue information in renamed artifacts. It also keeps per-execution subprocess output under pabot_results. Prefer one of two patterns: emit evidence into each subprocess output directory, or embed artifacts when the library supports safe embedding. Avoid a manually shared global screenshot/video directory.

When debugging, preserve pabot_manager.log, each failing subprocess’s stdout/stderr and partial output, and the final output. A final merged report tells you the logical result; subprocess evidence tells you how it was scheduled and where a worker failed.

8. Worked design decision

Scenario Decision Reason
20 API suites, each owns unique tenant; API permits 8 concurrent tenants Suite-level Pabot, --processes 6 Isolation matches suite boundary; capacity margin remains.
One suite has 40 independent 30-second tests Test-level split Finer scheduling reduces tail if each test creates unique data.
Four physical devices shared by 12 tests PabotLib value sets or explicit device allocator Capacity is four devices; process count alone cannot allocate them.
One destructive migration fixture Keep exclusive/serial Locking may make it technically possible, but semantic risk and rollback cost dominate.
Large fleet across 4 CI machines Shard, then modest Pabot in each shard Separate machine distribution from local CPU concurrency; aggregate outputs afterward.

Knowledge check

Four CI jobs each run Pabot with six processes. What concurrency should the target capacity model consider?

Why is an ordering file not a test-selection mechanism?

When is a value set better than a lock?

Why can --chunk improve time but also change risk?

9. Summary and bridge

Parallel configuration is a resource-allocation design. Split level, process count, sharding, ordering, locking, setup strategy, and artifact layout must all map to explicit state ownership and measurable capacity. Lesson 4 focuses on diagnosing the failures that appear when those contracts are wrong.

Next lesson

Parallel Execution with Pabot, Sharding, and Resource Contention: Diagnostics, Failure Modes, and Production Practices

Continue with Parallel Execution with Pabot, Sharding, and Resource Contention: Diagnostics, Failure Modes, and Production Practices. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

References and version anchors

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.