Checkpoint Lab — Capstone: Build a Governed End-to-End Robot Framework Automation Platform
Deliver the complete automation product, prove it in serial and one alternate mode, inject four failures across different layers, recover from preserved evidence, and finish with an operating/governance runbook rather than a demo-only success screenshot.
Learning objectives
- Run the full capstone from documented setup/preflight commands and preserve a complete evidence packet.
- Exercise the serial baseline plus one alternate mode—Pabot or the private container path—without changing test meaning.
- Inject at least four controlled failures across source/import, secret/evidence, parallel/runtime, and container/configuration layers.
- Write root-cause statements, apply least-invasive repairs, and prove recovery with the smallest slice before full regression.
- Set measurable runtime/output budgets and complete a production-readiness review covering ownership, upgrades, privacy, capacity, and incident response.
Current compatibility baseline — verified 2026-09-01. The mandatory capstone path uses Python 3.12, Robot Framework 7.4.2, RequestsLibrary 0.9.7, Requests 2.34.2, and Robocop 8.5.0. Pabot 5.2.2 is optional for the alternate execution slice. Robot Framework 7.5b1 and RequestsLibrary 1.0a14 are pre-releases and are not required. Browser automation is an optional architecture branch: Browser 20.4.0 and SeleniumLibrary 6.9.0 are discussed in Lesson 3, not installed by the mandatory lab. Record the versions actually used before adopting the platform outside this dated course lab.
1. Checkpoint scenario and acceptance contract
You are the maintainer receiving a Robot Framework estate for production adoption. The code is useful, but adoption is blocked until the team can prove safe target selection, reproducible dependencies, secret-safe integration, deterministic serial behavior, one alternate execution mode, retained evidence on failure, style/static quality, performance budgets, and a repeatable incident process.
Lab guardrails. Use only
127.0.0.1 or the private Compose service
fixture, the exact fake token marker, disposable
result/evidence directories, and harmless child processes. Stop if
any command resolves to a public/production target or a real
credential appears.
2. Preflight and predictions
Before execution, write two predictions in
evidence/predictions.txt. Example: (1) setting
RF_CAPSTONE_TOKEN will create a Secret object in Robot
data, while the raw value should not appear in normal Robot logs;
(2) switching the container base URL from loopback to
fixture will change only the external target resolution
layer, not suite/test names or expected statuses.
python --version
python -m robot --version
python -m pip freeze
robocop --version
pabot --version
python -m robot --dryrun --outputdir results/checkpoint-dryrun tests
Record all command output or version files. The checkpoint fails if the environment is reconstructed from memory after an incident.
3. Establish the clean serial baseline
Start the local fixture with the fake token in Terminal A. In Terminal B, set the same fake token and run the quality + execution + post-processing gates. Time the serial Robot command using a shell-appropriate timing mechanism and record the wall-clock value, suite/test counts, and output size.
robocop check tests resources
robocop format --check tests resources
python -m robot --outputdir results/final-serial --log NONE --report NONE tests
python -m robot.rebot --outputdir results/final-published --name "Governed Capstone Final" results/final-serial/output.xml
Do not delete a failing output.xml before the
investigation. If the clean run fails unexpectedly, that real
failure becomes the first incident and must be explained before
seeded failures continue.
4. Exercise one alternate execution mode
Option A — Pabot
pabot --processes 2 --outputdir results/final-pabot tests/parallel
Compare suite/test counts, pass/fail semantics, and mutable-resource ownership to serial execution. Preserve worker/manager evidence.
Option B — private container path
docker compose -f containers/compose.yaml up --build --abort-on-container-exit --exit-code-from robot
Confirm the Robot service uses http://fixture:8765, the
result volume persists outside the ephemeral container, and no
fixture port is published publicly. The mandatory course outcome is
architecture understanding; if Docker is unavailable, Pabot provides
the free local alternate path.
5. Inject four failures and recover
| Incident | Seed | Expected first bad layer | Required repair proof |
|---|---|---|---|
| I1 Import drift | Rename resource without updating suite | Source/import graph | Dry-run fails before external action; source path repaired; dry-run passes. |
| I2 Secret disclosure |
Log ${API_TOKEN.value} using fake token only
|
Secret/logging boundary | Incident artifact proves leak; source unwrapping removed; clean artifact search finds no fake marker. |
| I3 Shared-state parallel collision | Two Pabot shards use same mutable file | Worker/external state ownership | Unique worker/test path; two-process run passes repeatedly without lock/sleep folklore. |
| I4 Container target mismatch | Use loopback base URL inside Robot container | Container networking/config | Switch to private service DNS; smallest API suite passes; no public port workaround. |
For every incident create
evidence/incidents/Ix.md with: exact command; versions;
expected versus observed state; first bad layer; preserved artifact
paths; root-cause statement; repair; smallest rerun;
regression-prevention control; and cleanup. Do not write “flaky” as
the root cause.
6. Optional advanced incidents
If the four required incidents finish cleanly, add one of these:
make the CI wrapper ignore Robot’s exit code; change the custom
library scope to GLOBAL and introduce mutable per-test
state; skip artifact retention on failure; add a Robocop bypass; or
inject a measurable 30% wait regression. The repair must preserve
correctness and evidence, not merely restore green status.
7. Audit retained artifacts for the fake marker
Search only the disposable capstone tree. The expected source locations containing the literal fake marker are documentation/lab configuration where deliberately declared; generated Robot evidence from a clean run should not reveal the runtime Secret value. Classify findings instead of deleting them blindly.
# Bash example: scoped to disposable results/evidence only
grep -R --line-number --fixed-strings 'capstone-FAKE_DO_NOT_USE' results evidence || true
# PowerShell equivalent
# Get-ChildItem results,evidence -Recurse -File | Select-String -SimpleMatch 'capstone-FAKE_DO_NOT_USE'
If a generated artifact contains the marker, identify which keyword/library/source boundary emitted it, preserve a restricted incident copy, repair the source, and rerun. Do not claim Robot Secret values are encrypted.
8. Set runtime and output budgets from the measured baseline
Budgets should be relative to a known clean baseline and stable hardware class, not universal numbers copied from this course. Example policy:
| Metric | Baseline | Regression threshold | Action |
|---|---|---|---|
| Serial wall time | Median of 3 clean runs | >20% slower under same environment | Block adoption until dominant phase is explained or budget intentionally revised. |
| Pabot wall time | Measured two-process slice | Must improve or justify overhead; no correctness change | Remove/retune parallelism if it adds contention without useful throughput. |
output.xml size |
Clean serial artifact | >25% growth without planned evidence change | Inspect loops/logging/result model before suppressing evidence. |
| Test status/count | Clean serial run | Any unexplained difference | Hard stop; investigate selection/merge/semantic drift. |
| Robocop | 0 unexplained gate failures | Any new unexplained issue/bypass | Repair or document narrow reviewed suppression. |
| Secret audit | No fake runtime token in generated evidence | Any generated-artifact hit | Security stop; repair logging/transport boundary. |
9. Production-readiness review
| Area | Readiness question | Evidence |
|---|---|---|
| Architecture | Are suite/resource/library/system boundaries documented? | diagram + project tree + owners |
| Targets | Can automation refuse unauthorized environments? | guard test + environment policy |
| Secrets | Are values injected externally and kept out of retained evidence? | Secret design + artifact audit |
| Reproducibility | Can another runner recreate the environment? | pins + freeze + Python/runtime/container manifest |
| Failure semantics | Does CI preserve Robot failure status? | wrapper/job evidence |
| Artifacts | Are first-failure outputs retained even on failure? | raw output + retention rule |
| Parallelism | Are worker resources isolated and capacity bounded? | serial/Pabot comparison + isolation map |
| Containers | Are DNS/ports/users/volumes explicit? | compose/Docker evidence |
| Style | Are current Robocop checks enforced without blanket bypass? | check/format gate |
| Performance | Are budgets measured and regression thresholds owned? | budget table + baseline data |
| Upgrades | Is there a compatibility lane and rollback policy? | matrix/cadence/runbook |
| Incidents | Can maintainers reproduce evidence-first diagnosis? | four completed incident records |
10. Deliver the operating/governance runbook
RUNBOOK — Governed Robot Framework Platform
Owners
- Platform owner: dependency/runtime/evidence/CI contracts
- Domain owner: suite intent and system-under-automation contracts
- Security owner: target/secret/evidence policy
Normal operation
1. Activate the approved environment and record versions.
2. Resolve approved target + Secret injection.
3. Run dry-run and Robocop gates.
4. Run serial Robot command; preserve raw output.xml.
5. Post-process with Rebot.
6. Run alternate mode only when required by the gate.
7. Upload/retain evidence even if Robot failed.
Incident response
1. Freeze first-failure artifacts and exact command/environment.
2. Identify first bad layer using the Chapter 29/30 diagnostic sequence.
3. Apply least destructive repair.
4. Rerun smallest affected slice.
5. Run full regression + secret/style/performance checks.
6. Add a prevention control and update the runbook.
Upgrade policy
- Evaluate candidate Python/Robot/external libraries/tools in a compatibility lane.
- Change one dependency family at a time where practical.
- Compare status counts, warnings/deprecations, output schema/size, runtime, Pabot/container behavior.
- Roll back unexplained semantic/evidence changes.
- Update pins + matrix + runbook together.
11. Cleanup and rollback
Stop the local fixture with Ctrl+C. Remove only the disposable capstone virtual environment/results if desired; retain the evidence packet required for grading/review. For containers, stop/remove the lab stack and verify no host port or persistent secret remains. Never generalize these cleanup commands to unrelated directories or running processes.
docker compose -f containers/compose.yaml down --remove-orphans
# Review paths before deleting anything:
python -c "from pathlib import Path; print(Path('results').resolve()); print(Path('evidence').resolve())"
12. What this chapter adds to the production operating model
The completed course now ends with an automation platform that can explain itself: what it is allowed to target, where state lives, how secrets cross boundaries, how source becomes results, which execution modes are supported, which artifacts prove the outcome, how quality/performance are governed, and how a maintainer responds when any layer fails. That operating model—not the demo fixture—is the transferable capstone.
Knowledge check
Why must the checkpoint establish a clean serial baseline before testing Pabot or containers?
Serial execution is the correctness reference. If the baseline is not deterministic, alternate modes add scheduling/network/runtime layers that make cause attribution harder.
What makes a performance budget valid for this platform?
It is measured against a documented environment/workload, preserves functional/evidence equivalence, and has an owned regression threshold rather than an arbitrary universal time limit.
During the secret audit, the fake token appears in a generated
log because a test accessed ${API_TOKEN.value}.
What is the correct sequence?
Preserve the incident artifact under restricted lab evidence, remove the source unwrapping/logging, rerun the smallest slice, and verify clean generated evidence no longer contains the marker. Do not merely delete the log.
What is the final deliverable of the capstone beyond passing tests?
A governed operating model and evidence packet: architecture, version manifest, target/secret controls, clean and alternate execution evidence, incident records, style/performance budgets, upgrade policy, and runbook.
Further reading
- Robot Framework User Guide 7.4.2 — authoritative core behavior for this dated capstone.
- Robot Framework releases — upgrade lane input.
- RequestsLibrary, Pabot, and Robocop.
- Browser Library and SeleniumLibrary — optional UI architecture path.
- Robot Framework course curriculum — use the earlier chapters as the detailed reference for each layer summarized in this capstone.
This final lesson intentionally keeps cloud/CI-provider administration, browser-stack administration, Docker/Kubernetes operation, and secret-manager administration at their Robot integration boundaries. Use the dedicated DevOps Academy courses for those platforms.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.