Listeners, PreRunModifiers, Visitors, and Programmatic Execution: Diagnostics, Failure Modes, and Production Practices
Diagnose callback-version mismatches, silent suite removal, extension errors, ignored return codes, private-API coupling, global state, and destructive result mutation without losing first-failure evidence.
Current compatibility baseline — verified 2026-08-31.
Robot Framework 7.4.2 is the stable course baseline
and requires Python 3.8+. Listener API version 3 is the default from
Robot Framework 7.0 and is generally recommended; version 2 remains
supported for compatibility. Optional ListenerV2/ListenerV3
base classes are public APIs from 6.1. Pre-run modifiers normally
extend robot.api.SuiteVisitor and are applied before
ordinary selection such as
--include/--exclude. Result processing
uses robot.api.ExecutionResult and
ResultVisitor. robot.run() returns an
integer status code, not an ExecutionResult;
run_cli(..., exit=False) returns the code instead of
exiting. The mandatory labs require no third-party library beyond
Robot Framework itself.
1. Evidence-first diagnostic sequence
- Preserve the first failing output.xml, log/report, console output, listener files, and modifier manifests.
- Confirm Python/Robot versions and executable paths.
- Confirm executed source, selection, output directory, extension import names and constructor arguments.
- Validate parse/import graph.
- Confirm listener API version and callback data model.
- Inspect modifier before/after evidence.
- Inspect output/result state separately from live callback state.
- Repair the smallest layer and rerun into a new evidence directory.
2. Broken listener: v3 object treated as v2 dict
# INTENTIONALLY BROKEN: v3 data is treated like a v2 dictionary.
ROBOT_LISTENER_API_VERSION = 3
def end_test(name, attrs):
print(attrs["status"])
With listener v3, the second argument is a result model object.
Treating it as attrs["status"] is a contract mismatch.
Repair the data model rather than blindly changing the API version:
ROBOT_LISTENER_API_VERSION = 3
def end_test(data, result):
print(f"{data.id} {result.status}")
3. Broken modifier: silent test removal
from robot.api import SuiteVisitor
class SilentDrop(SuiteVisitor):
# INTENTIONALLY BROKEN: removes everything and records no reason/evidence.
def start_suite(self, suite):
suite.tests = []
A run with no failed tests is meaningless if a modifier silently removed the intended tests. Repair with explicit criteria, before/after counts or names, policy version, and an expected minimum selected set when appropriate.
4. Broken programmatic gate: ignored return code
from robot import run
run("suites/extension_demo.robot", outputdir="evidence/bad")
print("release gate: allowed") # Wrong: Robot status was discarded.
from robot import run
rc = run("suites/extension_demo.robot", outputdir="evidence/good")
if rc != 0:
raise SystemExit(rc)
print("release gate: allowed")
The outermost CI/process status must reflect Robot's status unless the workflow explicitly models an expected diagnostic failure.
5. Private API coupling
Internal model classes may be importable and still be inappropriate
external contracts. Prefer robot.api, listener
interfaces, and robot.run/run_cli. Private
dependencies should be isolated behind a compatibility adapter and
version pin.
6. Listener logs recursively or leaks sensitive data
Do not serialize every keyword argument, variable, environment entry, request body, screenshot, or message merely because callbacks expose rich objects. Low-level message callbacks can multiply output and may create recursion/noise if the listener itself logs through observable channels. Collect the minimum telemetry needed.
7. Global mutable listener state breaks isolation
Module-level “current test” lists may appear harmless in serial runs but fail with multiple listener instances, nested suites, programmatic reuse, or future Pabot workers. Keep state instance-local and use run/worker-specific evidence paths. Chapter 24 formalizes parallel ownership.
8. Result mutation destroys audit evidence
evidence/failure-original/output.xml # immutable first result
evidence/policy-derivative/output.xml # transformed derivative
evidence/policy-derivative/change.json # what/why/version
Never save a status-changing post-processor over the only original result. The derivative may be useful; it is not a replacement for the source evidence.
9. Measure the causal layer
| Symptom | Likely layer | Measure first |
|---|---|---|
| Slow startup | Parsing/import/modifier | suite build + visitor traversal time |
| Slow runtime | Keywords/external systems | keyword elapsed time |
| Huge log/output.xml | Logging/output | message volume and listener duplication |
| Slow post-processing | Result visitor | result node count and serialization |
| Slow CI | Runner/container setup | provisioning vs Robot time |
| Slow parallel run | Pabot scheduling/contention | worker ownership; Chapter 24 |
10. Controlled broken example
Run the broken listener only against the synthetic candidate tests
and write to evidence/broken. Preserve console/Test
Execution Errors and any generated output. Repair the callback and
rerun to evidence/repaired. Do not overwrite the
original failure directory.
11. Production anti-patterns
- Listeners that implement hidden business workflows.
- Modifiers that silently drop tests.
-
Broad
except Exception: passaround extension failures. - Blanket retries or giant timeouts for import/callback errors.
- Arbitrary PYTHONPATH changes instead of fixing packaging/imports.
- Ignoring Robot return codes.
- Global mutable listener state.
- Overwriting first-failure output.xml.
- Depending on private model attributes as a public contract.
12. Knowledge check
A v3 listener fails because code indexes attrs["status"]. What layer is wrong?
The listener interface/data-model contract. V3 passes model objects, not the v2 attribute dictionary.
Why is zero failed tests insufficient after adding a modifier?
The modifier may have removed all intended tests. Verify before/after model counts and executed result evidence.
Why is broad exception swallowing dangerous in governance listeners?
It can silently remove observability or policy evidence while the run appears healthy.
How should a status-changing result transform be stored?
As a separate derivative with a change manifest, preserving the original output.xml.
13. Summary and bridge
Extension failures are diagnosable when each lifecycle phase has separate evidence and ownership. Lesson 5 packages those rules into a checkpoint.
References and version anchors
- Listener interface — callback models and versions
- Errors and warnings during execution — extension error evidence
- run.py source — status propagation
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.