RPA Tasks, Long-Running Automation, and Human-in-the-Loop Workflows: Diagnostics, Failure Modes, and Production Practices
Diagnose long-running Robot Framework automation without false-green recovery, duplicate irreversible actions, infinite polling, unsafe Dialogs in CI, secret-bearing work items, or destructive cleanup.
Learning objectives
- Diagnose duplicate effects, ambiguous retries, polling hangs, approval bypasses, and false-green external outcomes.
- Preserve first-failure Robot artifacts and durable workflow state before attempting repair.
- Recognize security failures involving real targets, secrets in work items, and interactive approval shortcuts.
- Repair cleanup so it never deletes unprocessed inputs or an unowned workflow root.
- Separate Robot execution cost from external latency, scheduler delay, parallel-worker contention, and retry/rerun cost.
Current compatibility baseline — verified 2026-08-31.
Robot Framework 7.4.2 is the stable course baseline
and requires Python 3.8+. Task files use *** Tasks ***;
task execution uses the same engine as tests, while generic
automation/RPA mode changes terminology and enables task-oriented
settings and selectors such as Task Setup,
Task Teardown, Task Timeout,
--task, and --rpa. The standard
Dialogs 7.4.2 library is interactive and blocks
while waiting for a desktop user, so it is optional only. The
mandatory learning path is fully local and uses Robot Framework core
plus OperatingSystem, String, DateTime, Collections, and BuiltIn.
The external RPA Framework 33.0.1 release
(2026-08-23; Python 3.10–3.13) is mentioned only as an optional
ecosystem layer; no commercial Control Room, cloud orchestrator,
email account, finance system, browser, database, SSH service,
Pabot, container runtime, CI provider, or production target is
required.
1. Diagnostic sequence: preserve → identify state owner → correct minimally
-
Preserve first-failure artifacts: copy or retain
output.xml, log/report, checkpoint/ledger, approval record, and safe external-state snapshot. - Confirm versions: Python, Robot Framework, and any external RPA library/tool.
- Confirm trigger/selection/data: run ID, task name, work-item ID, approval identity, execution directory, environment.
- Validate parse/import graph: resources/libraries/variables resolved from the intended project.
- Inspect Robot scope: current variables and keyword resolution.
- Inspect durable state: input, output, checkpoint, idempotency record, approval, queue/DB state.
- Inspect timing/orchestration: bounded wait, scheduler, parallel workers, container/CI namespace if relevant.
- Apply least destructive repair and rerun only the smallest controlled work item.
Do not start by deleting state and “trying again.” In RPA, the missing state may be the only evidence that an external action already happened.
2. Failure: retry duplicates an irreversible action
*** Tasks ***
Broken Submission
Wait Until Keyword Succeeds 3x 1 second Submit Irreversible Action
This wrapper cannot know whether a timeout means “nothing happened” or “the action happened but the response was lost.” Repair the design by giving the target operation a stable idempotency key and querying its state, or stop for reconciliation when the outcome is ambiguous. Never teach generic retry as a transaction protocol.
3. Failure: an approval loop has no upper bound
# BROKEN: this can occupy a worker forever.
WHILE not $approved
${approved}= Approval Exists ${item_id}
Sleep 5 seconds
END
Add a time/iteration limit and a defined timeout result, or redesign the workflow so an external event triggers a fresh task when approval arrives. A “long-running” process is not justification for an unbounded loop.
4. Failure: Dialogs appears in unattended CI
Symptoms include a task that never completes, no useful log after the dialog keyword, and a runner waiting until the job-level timeout. The root cause is architectural: Dialogs assumes an interactive user interface. Repair by moving the decision to an external approval record/queue/API and making Robot poll only briefly or exit until retriggered.
Do not solve this by auto-selecting “Approve” from an environment default. That converts a human control into an unattended bypass.
5. Failure: task PASS without verifying the external outcome
# BROKEN: keyword returned without exception, but external outcome was never checked.
Process Work Item WI-001
Log completed
A task status reflects Robot keyword execution, not hidden business truth. Repair by querying or reading the externally observable result, asserting the expected state, and only then writing the completion checkpoint. If verification cannot be performed, the task must not claim durable completion.
6. Failure: durable workflow state exists only in suite variables
VAR ${LAST_COMPLETED} WI-001 scope=SUITE can coordinate
keywords in one process. It cannot help a later process determine
whether WI-001 completed. After a crash, a fresh run may start from
the beginning and repeat work. Repair by using a durable completion
record keyed by the work-item identity.
7. Failure: a training task points at real email, finance, or production systems
Reject this before execution. Verify an allowlisted target in preflight; use synthetic accounts/data or a local simulator; use least-privilege credentials; and never make production mutation the way to “prove” an RPA concept. If the target cannot be proven safe, do not run the mutation.
8. Failure: work items and evidence contain secrets
Input files, approval records, screenshots, logs, and result XML can all leak credentials or personal data. Synthetic labs should use fake values. Production designs need explicit redaction, access control, retention, and secret injection. Robot Framework's masking features or external vault integrations do not make arbitrary work-item files safe by themselves.
9. Failure: cleanup deletes unprocessed inputs
# BROKEN: recursive deletion runs merely because the task ended.
[Teardown] Remove Directory ${WORK_ROOT} recursive=${True}
The repair is the Chapter 19 ownership contract: durable state is
not ordinary teardown state. A separate cleanup action validates the
owner/run ID, confirms every expected item has a completion marker,
confirms no pending/approved markers remain, preserves the ledger,
normalizes the path below ${TEMPDIR}, and only then
removes it.
10. Failure: “human approval” is implemented as an automatic bypass
# BROKEN security boundary.
${approval}= Get Environment Variable APPROVAL APPROVE
Should Be Equal ${approval} APPROVE
The default silently approves when no human decision exists. If environment input is used in a controlled local simulation, require the value to be explicitly present and record who/what supplied it. Production approval should come from an authenticated, auditable external decision system.
11. Performance: measure the layer that actually waits
Separate parsing/import time, Robot keyword execution, external-system latency, approval wait, log/output cost, scheduler delay, Pabot/worker contention, container startup, and rerun/retry cost. A task that spends 99% of its time waiting for human approval will not be fixed by optimizing Robot keyword dispatch.
Long loops and retries can also inflate output.xml and
log.html. Preserve enough detail for diagnostics but
avoid generating repetitive wait noise indefinitely.
12. Intentionally broken checkpoint order: interpret before repair
# BROKEN: checkpoint is written before the external output is verified.
Create File ${COMPLETED_DIR}${/}${item_id}.done completed
Create File ${OUTBOX_DIR}${/}${item_id}.txt ${payload}
Should Contain ${payload} required-field
If the output write or assertion fails, the durable marker lies: the next run sees “done” and skips the item. Preserve the failed run and marker as evidence, then repair the order to output → independent verification → completion marker. Do not simply delete the failure log and rerun.
13. Troubleshooting shortcuts to reject
- blanket retries around unknown side effects;
- giant sleeps or disabled timeouts;
- catch-all EXCEPT that turns unknown failures into PASS;
- global variables used as a pseudo-workflow database;
- automatic “approve” defaults;
- production-target experiments;
- deleting original output/checkpoint files before reconciliation;
- recursive cleanup without normalized-path and ownership checks.
Knowledge check
A rerun passes after a payment keyword timed out. Why can the run still be unsafe?
The first request may already have produced the payment. A later PASS does not prove there was only one effect; reconcile by idempotency key/transaction identity.
What is the first artifact to preserve when a task fails after an ambiguous external action?
The original Robot output/log/report plus the durable checkpoint/ledger and a safe snapshot/query of the external state. Preserve before mutation.
Why is a default environment value of APPROVE a security bug in a human-in-the-loop workflow?
It converts absence of human evidence into approval, bypassing the control boundary.
How do you distinguish a Robot performance problem from orchestration latency?
Measure parse/import/keyword time separately from external I/O, approval waiting, scheduler/worker delay, container startup, and retry time.
What is wrong with writing the checkpoint before verifying output?
A failure after the marker creates a false durable completion record. Resume then skips unfinished or invalid work.
14. Summary and checkpoint bridge
Production RPA diagnostics start with evidence and state ownership, not retries. Duplicate prevention, approval integrity, bounded waiting, truthful completion, secret hygiene, and guarded cleanup are all observable contracts. Lesson 5 combines them in a restart-and-resume checkpoint exercise.
References and version anchors
- Robot Framework 7.4.2 — Stopping execution gracefully — stop/fatal/teardown semantics
- Robot Framework BuiltIn — Fatal Error and bounded retry behavior
- Robot Framework Dialogs — interactive blocking behavior
- Robot Framework OperatingSystem — guarded local state
- RPA Framework release notes — optional ecosystem status/security changes
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.