Chapter 19Lesson 04180–240 min

RPA Tasks, Long-Running Automation, and Human-in-the-Loop Workflows: Diagnostics, Failure Modes, and Production Practices

Diagnose long-running Robot Framework automation without false-green recovery, duplicate irreversible actions, infinite polling, unsafe Dialogs in CI, secret-bearing work items, or destructive cleanup.

False greenDuplicate effectsInfinite pollingDialogs in CIEvidence preservation

Learning objectives

  • Diagnose duplicate effects, ambiguous retries, polling hangs, approval bypasses, and false-green external outcomes.
  • Preserve first-failure Robot artifacts and durable workflow state before attempting repair.
  • Recognize security failures involving real targets, secrets in work items, and interactive approval shortcuts.
  • Repair cleanup so it never deletes unprocessed inputs or an unowned workflow root.
  • Separate Robot execution cost from external latency, scheduler delay, parallel-worker contention, and retry/rerun cost.

Current compatibility baseline — verified 2026-08-31. Robot Framework 7.4.2 is the stable course baseline and requires Python 3.8+. Task files use *** Tasks ***; task execution uses the same engine as tests, while generic automation/RPA mode changes terminology and enables task-oriented settings and selectors such as Task Setup, Task Teardown, Task Timeout, --task, and --rpa. The standard Dialogs 7.4.2 library is interactive and blocks while waiting for a desktop user, so it is optional only. The mandatory learning path is fully local and uses Robot Framework core plus OperatingSystem, String, DateTime, Collections, and BuiltIn. The external RPA Framework 33.0.1 release (2026-08-23; Python 3.10–3.13) is mentioned only as an optional ecosystem layer; no commercial Control Room, cloud orchestrator, email account, finance system, browser, database, SSH service, Pabot, container runtime, CI provider, or production target is required.

1. Diagnostic sequence: preserve → identify state owner → correct minimally

  1. Preserve first-failure artifacts: copy or retain output.xml, log/report, checkpoint/ledger, approval record, and safe external-state snapshot.
  2. Confirm versions: Python, Robot Framework, and any external RPA library/tool.
  3. Confirm trigger/selection/data: run ID, task name, work-item ID, approval identity, execution directory, environment.
  4. Validate parse/import graph: resources/libraries/variables resolved from the intended project.
  5. Inspect Robot scope: current variables and keyword resolution.
  6. Inspect durable state: input, output, checkpoint, idempotency record, approval, queue/DB state.
  7. Inspect timing/orchestration: bounded wait, scheduler, parallel workers, container/CI namespace if relevant.
  8. Apply least destructive repair and rerun only the smallest controlled work item.

Do not start by deleting state and “trying again.” In RPA, the missing state may be the only evidence that an external action already happened.

2. Failure: retry duplicates an irreversible action

*** Tasks ***
Broken Submission
    Wait Until Keyword Succeeds    3x    1 second    Submit Irreversible Action

This wrapper cannot know whether a timeout means “nothing happened” or “the action happened but the response was lost.” Repair the design by giving the target operation a stable idempotency key and querying its state, or stop for reconciliation when the outcome is ambiguous. Never teach generic retry as a transaction protocol.

3. Failure: an approval loop has no upper bound

# BROKEN: this can occupy a worker forever.
WHILE    not $approved
    ${approved}=    Approval Exists    ${item_id}
    Sleep    5 seconds
END

Add a time/iteration limit and a defined timeout result, or redesign the workflow so an external event triggers a fresh task when approval arrives. A “long-running” process is not justification for an unbounded loop.

4. Failure: Dialogs appears in unattended CI

Symptoms include a task that never completes, no useful log after the dialog keyword, and a runner waiting until the job-level timeout. The root cause is architectural: Dialogs assumes an interactive user interface. Repair by moving the decision to an external approval record/queue/API and making Robot poll only briefly or exit until retriggered.

Do not solve this by auto-selecting “Approve” from an environment default. That converts a human control into an unattended bypass.

5. Failure: task PASS without verifying the external outcome

# BROKEN: keyword returned without exception, but external outcome was never checked.
Process Work Item    WI-001
Log    completed

A task status reflects Robot keyword execution, not hidden business truth. Repair by querying or reading the externally observable result, asserting the expected state, and only then writing the completion checkpoint. If verification cannot be performed, the task must not claim durable completion.

6. Failure: durable workflow state exists only in suite variables

VAR ${LAST_COMPLETED} WI-001 scope=SUITE can coordinate keywords in one process. It cannot help a later process determine whether WI-001 completed. After a crash, a fresh run may start from the beginning and repeat work. Repair by using a durable completion record keyed by the work-item identity.

7. Failure: a training task points at real email, finance, or production systems

Reject this before execution. Verify an allowlisted target in preflight; use synthetic accounts/data or a local simulator; use least-privilege credentials; and never make production mutation the way to “prove” an RPA concept. If the target cannot be proven safe, do not run the mutation.

8. Failure: work items and evidence contain secrets

Input files, approval records, screenshots, logs, and result XML can all leak credentials or personal data. Synthetic labs should use fake values. Production designs need explicit redaction, access control, retention, and secret injection. Robot Framework's masking features or external vault integrations do not make arbitrary work-item files safe by themselves.

9. Failure: cleanup deletes unprocessed inputs

# BROKEN: recursive deletion runs merely because the task ended.
[Teardown]    Remove Directory    ${WORK_ROOT}    recursive=${True}

The repair is the Chapter 19 ownership contract: durable state is not ordinary teardown state. A separate cleanup action validates the owner/run ID, confirms every expected item has a completion marker, confirms no pending/approved markers remain, preserves the ledger, normalizes the path below ${TEMPDIR}, and only then removes it.

10. Failure: “human approval” is implemented as an automatic bypass

# BROKEN security boundary.
${approval}=    Get Environment Variable    APPROVAL    APPROVE
Should Be Equal    ${approval}    APPROVE

The default silently approves when no human decision exists. If environment input is used in a controlled local simulation, require the value to be explicitly present and record who/what supplied it. Production approval should come from an authenticated, auditable external decision system.

11. Performance: measure the layer that actually waits

Separate parsing/import time, Robot keyword execution, external-system latency, approval wait, log/output cost, scheduler delay, Pabot/worker contention, container startup, and rerun/retry cost. A task that spends 99% of its time waiting for human approval will not be fixed by optimizing Robot keyword dispatch.

Long loops and retries can also inflate output.xml and log.html. Preserve enough detail for diagnostics but avoid generating repetitive wait noise indefinitely.

12. Intentionally broken checkpoint order: interpret before repair

# BROKEN: checkpoint is written before the external output is verified.
Create File    ${COMPLETED_DIR}${/}${item_id}.done    completed
Create File    ${OUTBOX_DIR}${/}${item_id}.txt    ${payload}
Should Contain    ${payload}    required-field

If the output write or assertion fails, the durable marker lies: the next run sees “done” and skips the item. Preserve the failed run and marker as evidence, then repair the order to output → independent verification → completion marker. Do not simply delete the failure log and rerun.

13. Troubleshooting shortcuts to reject

  • blanket retries around unknown side effects;
  • giant sleeps or disabled timeouts;
  • catch-all EXCEPT that turns unknown failures into PASS;
  • global variables used as a pseudo-workflow database;
  • automatic “approve” defaults;
  • production-target experiments;
  • deleting original output/checkpoint files before reconciliation;
  • recursive cleanup without normalized-path and ownership checks.

Knowledge check

A rerun passes after a payment keyword timed out. Why can the run still be unsafe?

What is the first artifact to preserve when a task fails after an ambiguous external action?

Why is a default environment value of APPROVE a security bug in a human-in-the-loop workflow?

How do you distinguish a Robot performance problem from orchestration latency?

What is wrong with writing the checkpoint before verifying output?

14. Summary and checkpoint bridge

Production RPA diagnostics start with evidence and state ownership, not retries. Duplicate prevention, approval integrity, bounded waiting, truthful completion, secret hygiene, and guarded cleanup are all observable contracts. Lesson 5 combines them in a restart-and-resume checkpoint exercise.

Next lesson

Checkpoint Lab — RPA Tasks, Long-Running Automation, and Human-in-the-Loop Workflows

Continue with Checkpoint Lab — RPA Tasks, Long-Running Automation, and Human-in-the-Loop Workflows. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

References and version anchors

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.