Chapter 30Lesson 04240–330 min

Capstone: Build a Governed End-to-End Robot Framework Automation Platform: Diagnostics, Failure Modes, and Production Practices

Production incidents cross abstraction boundaries. This lesson applies the Chapter 29 diagnostic ladder to the whole platform and rejects repairs that create green results by destroying evidence, isolation, or trust.

Incident responseFalse greenSecret leakContainer networkingRegression

Learning objectives

  • Apply one evidence-first diagnostic sequence across source/import, variable/secret, library, external-system, timing, parallel, container, CI, and result-processing layers.
  • Distinguish a symptom-reducing workaround from a root-cause repair that preserves test meaning.
  • Diagnose secret disclosure, shared-state parallel collision, container networking mismatch, output merge problems, and missing CI artifacts.
  • Recognize unsafe target selection, custom-library scope leaks, style-gate bypass, and performance regressions before they become platform folklore.
  • Write incident records that identify first bad layer, evidence, corrective action, regression prevention, and ownership.

Current compatibility baseline — verified 2026-09-01. The mandatory capstone path uses Python 3.12, Robot Framework 7.4.2, RequestsLibrary 0.9.7, Requests 2.34.2, and Robocop 8.5.0. Pabot 5.2.2 is optional for the alternate execution slice. Robot Framework 7.5b1 and RequestsLibrary 1.0a14 are pre-releases and are not required. Browser automation is an optional architecture branch: Browser 20.4.0 and SeleniumLibrary 6.9.0 are discussed in Lesson 3, not installed by the mandatory lab. Record the versions actually used before adopting the platform outside this dated course lab.

1. One diagnostic sequence for the entire platform

Cross-layer incident sequence
flowchart TD
A[Preserve first-failure artifacts] --> B[Record versions]
B --> C[Confirm command / path / selection / env]
C --> D[Parse + import graph]
D --> E[Variable / Secret / keyword resolution]
E --> F[Library + external state]
F --> G[Timing / Pabot / container / CI]
G --> H[Least destructive repair]
H --> I[Smallest controlled rerun]
I --> J[Full regression + governance update]

Do not start by reinstalling everything, increasing timeouts, adding PYTHONPATH, globalizing variables, disabling TLS/SSH verification, or rerunning until green. Those actions modify multiple layers before the cause is known. The first investigation step is to preserve the failing run and write down the exact command/environment that produced it.

2. Failure-mode map

Failure Likely first layer Evidence Wrong shortcut
Import/config drift Source/import/environment dry-run error, path, interpreter, manifest append arbitrary PYTHONPATH
False-green recovery Failure semantics/CI wrapper original exit code + first output catch/retry until pass
Secret leak Data/logging/library boundary artifact search + source location delete artifacts without fixing source
Parallel collision Worker/external mutable state worker IDs, paths/records, serial comparison add sleeps
Container networking mismatch Runtime/network boundary resolved URL, service DNS, compose config assume container localhost means host
Dependency drift Environment/library boundary freeze/pins/library versions reinstall latest everything
Output merge/post-process error Result layer raw outputs + Rebot/Pabot command edit generated report
Style gate bypass Governance/CI Robocop command/config + job status blanket disable rules
Performance regression Measured phase/runtime layer baseline vs current versions/timings remove assertions/logs blindly
Artifacts missing on failure CI control flow job log + artifact step conditions rerun and hope to reproduce
Unsafe target selection Configuration/trust boundary resolved base URL + guard result disable guard for convenience
Custom library scope leak Library instance state scope, instance mutations, test order reset global state in random teardowns

3. Incident A — import/config drift

Seed: rename resources/platform.resource to platform-v2.resource without updating the suite. The run fails before domain execution. Preserve the console/output, run the smallest --dryrun slice, and inspect the importing suite path.

python -m robot --dryrun --outputdir results/incidents/import tests/api.robot

Root cause statement: “tests/api.robot imports a resource path that no longer exists relative to the suite; no API request occurred.” The repair is to update the source import (or revert the unintended rename) and add the structural gate to review/CI. Adding unrelated paths to PYTHONPATH would not repair a file-resource contract.

4. Incident B — fake-secret disclosure

Use only the fake token in this destructive-to-evidence exercise. Seed an unsafe line that explicitly unwraps the Secret in Robot data:

*** Test Cases ***
Unsafe Debug Example — LAB ONLY
    Log    ${API_TOKEN.value}

Why this is intentionally broken. Robot’s Secret protection does not help after you explicitly access .value in data. The fake token may appear in logs. Never perform this incident with real credentials.

Preserve one failing/unsafe artifact copy for the restricted incident record, then repair the source by removing the unwrapping/logging behavior. Search the retained evidence for the exact fake marker capstone-FAKE_DO_NOT_USE. The acceptance condition is not “we deleted the log”; it is “a clean rerun does not emit the marker, while required status evidence remains available.”

5. Incident C — Pabot shared-state collision

Seed two parallel tests that write to the same filename and then assert ownership. Serial execution may pass because writes do not overlap; Pabot can expose the architectural defect.

*** Settings ***
Library    OperatingSystem

*** Test Cases ***
Shard A Owns File
    Create File    ${OUTPUT DIR}${/}shared.txt    A
    ${value}=    Get File    ${OUTPUT DIR}${/}shared.txt
    Should Be Equal    ${value}    A

A second shard writes B to the same path. The correct repair is per-worker/per-test isolation—for example a unique directory derived from suite/test identity—not a sleep or lock by default. A lock can serialize the collision, but it also removes the throughput benefit and preserves shared-state coupling. Prefer redesign when practical.

6. Incident D — container networking mismatch

Seed: run the Robot service in Compose with RF_CAPSTONE_BASE_URL=http://127.0.0.1:8765. Inside the Robot container, loopback points back to the Robot container, so the fixture is unreachable.

Observed evidence
- target guard accepts loopback syntactically
- HTTP connection fails
- fixture container is healthy
- Robot container resolves service name `fixture`
- compose.yaml places both services on the same private network

The repair is configuration: use http://fixture:8765 for the container execution profile while retaining loopback locally. Do not publish the fixture port publicly merely to make the wrong localhost assumption work.

7. Result, CI, and governance incidents

False-green wrapper

A wrapper that runs Robot, catches a non-zero status, runs Rebot, then exits 0 creates a false-green pipeline. Preserve the original Robot return code and exit with it after post-processing/artifact work.

Output merge error

If Rebot/Pabot reports wrong counts, preserve every raw input and the exact merge command. Diagnose whether inputs represent reruns versus independent suites before choosing merge/combine semantics. Never edit report.html to “fix” source results.

CI artifacts missing on failure

Artifact retention must run under an “always”/finally-style condition appropriate to the CI provider. The Robot command’s failure must not skip upload. Provider syntax belongs to the dedicated CI courses; the Robot operating requirement is portable: raw result + diagnostics are retained regardless of pass/fail.

Style gate bypass

A blanket robocop disable or ignored exit status is not governance. Suppress only a specific rule with a documented reason when the source cannot reasonably conform, and keep formatter/linter status visible in CI.

8. Performance regression without sacrificing correctness

Suppose the serial median increases 35% after a dependency update. Compare the same suite, data, hardware class, Python/Robot/library versions, output settings, and external fixture. Attribute cost to parse/import/setup/keyword/external wait/output/post-processing before changing behavior.

Do not remove assertions, share a global browser/session, reduce evidence below the incident requirement, or add concurrency before proving the dominant phase. A performance repair is accepted only when the functional status/counts and evidence contract remain equivalent.

9. Custom-library scope leak

If a stateful library is changed from TEST to GLOBAL to “save initialization,” mutable attributes now survive across all suites in the process. Failures can become order-dependent. Diagnose by recording scope, constructor calls, state mutations, and execution order. If reuse is needed, split immutable expensive configuration from mutable per-test state rather than broadening scope blindly.

10. Unsafe target selection is a stop condition

If RF_CAPSTONE_BASE_URL resolves to a public or production host, the target guard should fail before domain activity. Do not disable the guard to make a demonstration run. In real organizations, approved environment IDs/allowlists and separate credentials should provide defense in depth; this lab’s loopback/private-service check is intentionally narrow.

11. Write the incident record

Incident ID: RF-CAP-004
First bad layer: container networking/configuration
Observed: Robot container cannot reach 127.0.0.1:8765; fixture service healthy
Preserved evidence: raw output.xml, Robot console, compose config, service logs, version manifest
Root cause: localhost referred to Robot container, not fixture service
Repair: RF_CAPSTONE_BASE_URL=http://fixture:8765 in container mode
Smallest rerun: tests/api.robot in compose
Regression prevention: container-mode smoke in compatibility lane
Owner: automation platform
Secrets: fake token marker only; no real credential retained

Knowledge check

A Pabot-only file collision disappears after adding a one-second sleep. Is the incident fixed?

Why is deleting a log containing a leaked fake secret insufficient as the repair?

A containerized test cannot reach a service on host-style localhost. What layer should you inspect before changing Robot keywords?

Why should the original Robot exit status survive Rebot/post-processing?

Summary and bridge

The production practice is consistent across every failure: preserve evidence, confirm versions and exact execution context, isolate the first bad layer, make the smallest repair, rerun the smallest slice, then run the full governed gate and update prevention. Lesson 5 requires you to operate the complete platform under this discipline and deliver a final readiness dossier.

Next lesson

Checkpoint Lab — Capstone: Build a Governed End-to-End Robot Framework Automation Platform

Continue with Checkpoint Lab — Capstone: Build a Governed End-to-End Robot Framework Automation Platform. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.