Checkpoint Lab — RPA Tasks, Long-Running Automation, and Human-in-the-Loop Workflows
Prove a restartable Robot Framework RPA operating model by interrupting and resuming a local work-item workflow without duplicate effects, then preserve an evidence ledger and operator runbook.
Checkpoint outcomes
- Run a restartable local task through pending, approved, and completed states.
- Predict durable mutations before execution and independently verify them after each run.
- Interrupt after one item, preserve the failed run, resume without duplication, and prove a third run is idempotent.
- Exercise a real human-decision boundary through an explicit synthetic approval record.
- Produce an operator runbook and evidence ledger, then clean up only after ownership/completion guards pass.
Current compatibility baseline — verified 2026-08-31.
Robot Framework 7.4.2 is the stable course baseline
and requires Python 3.8+. Task files use *** Tasks ***;
task execution uses the same engine as tests, while generic
automation/RPA mode changes terminology and enables task-oriented
settings and selectors such as Task Setup,
Task Teardown, Task Timeout,
--task, and --rpa. The standard
Dialogs 7.4.2 library is interactive and blocks
while waiting for a desktop user, so it is optional only. The
mandatory learning path is fully local and uses Robot Framework core
plus OperatingSystem, String, DateTime, Collections, and BuiltIn.
The external RPA Framework 33.0.1 release
(2026-08-23; Python 3.10–3.13) is mentioned only as an optional
ecosystem layer; no commercial Control Room, cloud orchestrator,
email account, finance system, browser, database, SSH service,
Pabot, container runtime, CI provider, or production target is
required.
1. Checkpoint charter: finite runs, durable state, zero production effects
Your final Chapter 19 checkpoint is an operational rehearsal. Two synthetic work items start pending. WI-001 is preapproved. The first Robot process completes WI-001 and then intentionally fails. A human/operator creates the approval record for WI-002. The second Robot process reopens the same durable root, no-ops WI-001, completes WI-002, and exits PASS. A third run proves repeat invocation creates no additional completion rows. Cleanup is separate and guarded.
Do not substitute real email, finance, ticketing, browser, database, SSH, or cloud targets. The point is workflow correctness, not integration breadth.
2. Exact preflight and assumptions
python --version
python -m robot --version
python -c "import tempfile; print(tempfile.gettempdir())"
# Required runtime:
# Python 3.8+ compatible with Robot Framework 7.4.2
# robotframework==7.4.2
# No external library is mandatory.
# Optional ecosystem observation only:
# rpaframework==33.0.1 requires Python 3.10+ and is not used by this lab.
Record these versions in your evidence notes. If your current environment differs materially, re-check the current Robot Framework/library compatibility before running.
3. Project files and dependency diagram
rf19-checkpoint/
├── resources/workflow.resource # use the exact resource from Lesson 2
├── tasks/workflow.robot # normal/resume task
├── tasks/cleanup.robot # explicit cleanup task
├── OPERATOR_RUNBOOK.md # commands, expected state, recovery
└── evidence/ # Robot result directories only
flowchart TD
A["Operator / scheduler trigger"] --> B["workflow.robot"]
B --> C["workflow.resource"]
C --> D["${TEMPDIR}/rf19-rpa-RUN_ID"]
D --> E["inbox + state markers"]
D --> F["approval records"]
D --> G["deterministic outbox"]
B --> H["Robot output.xml / log.html / report.html"]
I["Second Robot process"] --> C
J["cleanup.robot"] --> C
C --> K["guarded cleanup after completion"]
K --> L["ledger/tree copied to cleanup output dir"]
The diagram has two independent evidence streams: Robot result
artifacts under evidence/, and durable work state under
the system temp directory. Cleanup copies the durable ledger/tree
into its own Robot output directory before deleting the temporary
root.
4. Predict before you run
| Prediction | Expected change | Independent verification |
|---|---|---|
| First run after WI-001 | WI-001 outbox + completed marker; WI-002 pending | Inspect temp tree + first log |
| Injected interruption | Task status FAIL; state root remains | first/output.xml + owner/checkpoint files |
| Manual approval | WI-002 approval file appears; no Robot status changes yet | Read approval file directly |
| Resume | WI-001 no-op; WI-002 completes | resume log + ledger has exactly two completion rows |
| Third run | No new business-state mutation | ledger entry files unchanged; both items take no-op branch |
| Cleanup | State root removed only after guards; ledger/tree retained in cleanup evidence | cleanup output files + Directory Should Not Exist |
5. Operator runbook — exact sequence
# 0) Start from a clean disposable project directory and install RF 7.4.2.
python -m pip install "robotframework==7.4.2"
# 1) First run — expected FAIL after WI-001 is durably checkpointed.
python -m robot --rpa --variable RUN_ID:demo-001 --variable INTERRUPT_AFTER:WI-001 --outputdir evidence/first tasks/workflow.robot
# 2) PRESERVE evidence/first. Do not overwrite or delete it.
# 3) Human approval simulation for WI-002.
python -c "from pathlib import Path; import tempfile; p=Path(tempfile.gettempdir())/'rf19-rpa-demo-001'/'approvals'/'WI-002.approved'; p.write_text('decision=APPROVE\noperator=training-user\n', encoding='utf-8'); print(p)"
# 4) Resume — expected PASS.
python -m robot --rpa --variable RUN_ID:demo-001 --variable INTERRUPT_AFTER:NONE --outputdir evidence/resume tasks/workflow.robot
# 5) Idempotency proof — expected PASS and no new ledger rows.
python -m robot --rpa --variable RUN_ID:demo-001 --outputdir evidence/idempotent tasks/workflow.robot
# 6) Cleanup — only after all verification below succeeds.
python -m robot --rpa --variable RUN_ID:demo-001 --outputdir evidence/cleanup tasks/cleanup.robot
Windows PowerShell can use the same one-line Robot commands without the Bash continuation backslashes. The Python approval command is cross-platform and deliberately explicit.
6. Interruption evidence: what must be true after the first FAIL
-
evidence/first/output.xml,log.html, andreport.htmlexist and show the injected Fatal Error; - the durable root still exists;
-
outbox/WI-001.txtandstate/completed/WI-001.doneexist; - the durable ledger directory contains exactly one WI-001 entry file;
- WI-002 remains pending and has no output/completion marker;
- the owner file still contains
demo-001.
If any of these predictions is false, stop and diagnose before adding the second approval. Do not turn an unexpected state into a resume experiment.
7. Approval boundary: verify the decision independently
After the manual approval command, read the approval file directly
(for example with your editor or python -c) before
running Robot again. Confirm both decision=APPROVE and
the synthetic operator identity. This is deliberately outside the
Robot process: the second execution must consume existing approval
evidence, not invent it.
8. Resume proof: no duplicate WI-001 effect
In evidence/resume/log.html, locate the WI-001 branch
and confirm it reports the durable completion marker and returns
before output/checkpoint/ledger mutation. Then confirm WI-002
advances through approval and completion. The durable ledger
directory should now contain exactly two deterministic entry
files—one per work item; cleanup later assembles them into a single
retained ledger.
A passing resume without this state inspection is insufficient. The core checkpoint claim is no duplicate externally visible effect.
9. Controlled failure variant: reject or remove WI-002 approval
For one additional diagnostic run on a fresh RUN_ID, do
not create WI-002 approval (or create a file with
decision=REJECT). Predict the final status before
running. The task must not create WI-002 outbox/completion state.
Preserve that failure evidence, then restore the intended approval
record and resume. This tests the approval boundary rather than
bypassing it.
10. Required evidence ledger
| Artifact | What it proves | Retention note |
|---|---|---|
| first/output.xml + log/report | Initial work item completed, then process intentionally failed | Keep immutable for comparison |
| resume/output.xml + log/report | Second process reopened durable state and completed remaining item | Keep with first run |
| idempotent/output.xml + log/report | Repeated invocation produced safe no-ops | Keep as idempotency proof |
| cleanup/workflow-ledger.txt | Exactly one deterministic durable ledger entry per item | Retained after temp cleanup |
| cleanup/workflow-tree.txt | Final durable-state structure before deletion | Retained after temp cleanup |
| Approval record observation | Human boundary existed before WI-002 mutation | Record synthetic operator/decision only |
| OPERATOR_RUNBOOK.md | Reproducible trigger, resume, verification, cleanup contract | Version with automation source |
11. Verification checklist
-
Task files use
*** Tasks ***and execute in RPA/generic automation mode. - No real target, real credential, public service, or paid platform is required.
-
Durable root path is derived from
${TEMPDIR}and stableRUN_ID. - Owner marker is validated on resume and before cleanup.
- WI-001 completion survives the first process failure.
- WI-001 does not receive a second ledger row on resume or third run.
- WI-002 cannot complete without an explicit approval record.
- Every wait is bounded; no infinite polling or giant timeout is introduced.
- Completion marker is written only after deterministic output verification.
- Cleanup refuses unfinished state, preserves ledger/tree, validates normalized path ownership, then deletes.
12. Cleanup and rollback
The cleanup task is the rollback for this disposable lab. It is intentionally impossible to run successfully while items remain pending/approved. If cleanup fails, do not broaden the delete command. Inspect the durable state and reconcile the unfinished item first.
If the lab is abandoned before completion, manually inspect the
owner file and exact normalized path before deleting the specific
rf19-rpa-<RUN_ID> directory. Never use a wildcard
such as rf19-rpa-* in a shared temporary directory.
13. What Chapter 19 adds to the production Robot operating model
Chapters 01–18 established readable Robot syntax, scopes, lifecycle, control flow, imports, failure semantics, result engineering, selection/reruns, browser/API automation, and guarded infrastructure integration. Chapter 19 adds the missing operational layer for RPA: finite task execution around durable work-item state, idempotent commit points, explicit approval evidence, bounded waits, restartability, and operator-controlled retention/cleanup.
Chapter 20 now moves down one abstraction layer to writing Python test libraries, where domain-specific stateful capabilities can be implemented behind clean Robot keywords without leaking Python complexity into business-facing tasks.
Knowledge check
What exact observation proves WI-001 was resumed safely?
The second/third run detects its durable completion marker and returns before creating output/checkpoint/ledger state; the ledger still has exactly one WI-001 entry.
Why must the first failed output directory remain untouched?
It proves where execution stopped and what Robot observed. Overwriting it would destroy the evidence needed to distinguish a safe resume from a hidden duplicate.
What should happen if WI-002 approval is missing?
The bounded approval wait fails; WI-002 must remain without outbox/completion state. The workflow preserves durable state for later approval/resume.
A cleanup task sees one pending marker. What is the correct action?
Fail cleanup and reconcile/complete or explicitly disposition that work item. Do not delete the root or broaden the cleanup guard.
What is the main boundary Chapter 19 establishes before Chapter 20?
Robot user-facing tasks orchestrate a finite run, while durable workflow state and domain-specific low-level behavior need explicit external stores/libraries. Chapter 20 teaches how to implement reusable Python libraries behind that boundary.
References and version anchors
- Robot Framework 7.4.2 — Task execution — finite task execution / RPA terminology
- Robot Framework 7.4.2 — Timeouts — bounded task/keyword execution
- Robot Framework 7.4.2 — Stopping execution gracefully — failure/interruption behavior
- Robot Framework Dialogs — optional interactive boundary only
- RPA Framework release notes — optional external RPA ecosystem; 33.0.1 current at generation time
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.