Chapter 14Lesson 04180–240 min

Logs, Reports, output.xml, Rebot, and Result Post-Processing: Diagnostics, Failure Modes, and Production Practices

Diagnose stale, incompatible, oversized, privacy-sensitive, or incorrectly merged result artifacts without deleting first-failure evidence or hiding the original cause.

DiagnosticsStale evidenceMerge failuresPrivacyProduction

Learning objectives

  • Apply a repeatable evidence-first diagnostic sequence before rerunning or rewriting results.
  • Recognize wrong-root merges, stale reports, missing rerun replacement, and schema/version incompatibility.
  • Diagnose oversized outputs without blind deletion or blanket evidence suppression.
  • Respond to secret/PII discovery without falsely claiming post-processing erased the disclosure.
  • Separate execution defects, Rebot transformation defects, and CI artifact-publishing defects.

Current compatibility baseline. Verified 2026-08-31: Robot Framework 7.4.2 is the current stable release and requires Python 3.8+; 7.5b1 is a pre-release and is not required here. The 7.4.2 result.xsd is schema version 5; default XML outputs carry a schemaversion attribute and Robot Framework 7.x can also read legacy output. --legacyoutput exists for Robot Framework 6.x-compatible consumers. Rebot combines independent outputs by creating a new parent suite; --merge is for the same logical top-level suite, including reruns, where later matching test results replace earlier results and later SKIP results do not replace originals. The built-in Testdoc tool is deprecated in 7.4.2 in favor of external Testdoc and is not a result-reporting substitute.

1. Diagnostic sequence: preserve before transforming

  1. Preserve first-failure artifacts. Copy or checksum the original output/log/report before any Rebot operation.
  2. Confirm Robot/Python/Rebot versions. Record exact versions.
  3. Confirm executed path, selection, environment, and test data. A report cannot repair the wrong run.
  4. Validate parse/import graph and source revision. Separate source defects from result processing.
  5. Inspect variable scope, keyword resolution, library/external state. Result files describe what Robot observed, not necessarily why external state changed.
  6. Inspect timing/parallel/CI/container state if relevant. Identify worker and workspace ownership.
  7. Inspect artifact generator/schema/root suite/test identities.
  8. Apply the least destructive correction. Generate a new derivative instead of overwriting the original.
  9. Rerun the smallest controlled slice only when execution truly needs to change.

2. Failure mode: merging unrelated roots

# Intentionally wrong diagnostic example — do not use as a fix.
python -m robot.rebot --merge results/run-a/output.xml results/run-b/output.xml

If the roots represent independent suites, --merge is the wrong semantic operation. Depending on their hierarchy Robot can reject the merge or produce a structure that does not represent the intended reporting model. The repair is not to rename things until the command succeeds. Confirm the logical identity: independent roots should be combined; same-root reruns/pieces should be merged.

3. Failure mode: a fresh-looking report from stale XML

HTML timestamps can mislead if a report was regenerated from an old input. Diagnose by recording the input path, file modification time, checksum, root generated value, generator/version, suite/test counts, and source/CI run ID. A correct production pipeline should make the relationship between raw input and derived report explicit in an artifact manifest.

from pathlib import Path
import hashlib
import xml.etree.ElementTree as ET

p = Path("results/run-a/output.xml")
root = ET.parse(p).getroot()
print("sha256", hashlib.sha256(p.read_bytes()).hexdigest())
print("generated", root.attrib.get("generated"))
print("generator", root.attrib.get("generator"))
print("schema", root.attrib.get("schemaversion"))
print("tests", len(root.findall(".//test")))

4. Failure mode: consumer assumes the wrong XML generation

Robot Framework 7.0 introduced incompatible XML-format changes. A tool written around older output structure may fail or silently misread fields. First inspect generator and schemaversion, then test the consumer against the documented XSD. Use --legacyoutput only when the consumer contract explicitly requires it. Do not rewrite XML by hand to “make it look old.”

5. Failure mode: rerun exists but final report still shows the old failure

Common causes include normal combine instead of --merge, different top-level suite identity, a rerun that selected no matching failures, stale input paths, or processing the original XML after the merged file was generated. Diagnose all three files: original, rerun, and merged. Count tests and statuses in each, verify suite names, and check the merged test message for replacement evidence. Never delete the original until the merged derivative is verified.

6. Failure mode: huge logs from loops or library internals

Start by measuring which artifacts are large and why. A large HTML log may be driven by many repeated keyword/message nodes in raw output. Options such as --removekeywords PASSED, loop-specific removal, or --flattenkeywords can reduce derived artifacts, but each removes diagnostic structure. Prefer narrow name/tag/loop policies. If one keyword is inherently noisy, consider a documented robot:flatten keyword tag only after confirming the messages retained are sufficient.

Performance here is result serialization/Rebot/browser rendering/artifact upload. It is not a justification to remove meaningful assertions or hide failing evidence.

7. Failure mode: secret or PII appears in result artifacts

Stop broad publication, preserve a restricted copy according to incident policy, identify where the value entered Robot/library logs, rotate a real credential if necessary, and repair the source. Do not “fix” the incident by deleting the only evidence or by generating a reduced public report and claiming the original leak no longer matters. Secret masking is not encryption and external libraries can still disclose underlying values.

Never practice with real secrets. Use only markers such as token-FAKE_DO_NOT_USE in a disposable private workspace. Do not paste Authorization headers, private keys, production cookies, personal data, or customer payloads into lessons or CI logs.

8. Failure mode: deleting originals before post-processing succeeds

A fragile pipeline runs Rebot, uploads HTML, then deletes raw output regardless of post-processing status. If Rebot fails, the investigation source is gone. Production ordering should be: capture raw → validate/checksum → post-process → validate derivatives → upload/retain according to policy → cleanup only after all required evidence is safely stored. Cleanup should target the run-owned workspace, not a shared parent directory.

9. Separate failure layers

Symptom Likely layer First check
Test failed in output.xml Execution/test/system Keyword message, variables, external state
Rebot rejects input Result schema/file integrity Generator/schema, truncation, path
Merged report has duplicates Wrong combine/merge identity Top-level suite and command
CI dashboard missing tests but Robot report is correct xUnit/publisher integration Generated xUnit and CI parser
Artifact not found after run Workspace/upload retention Outputdir, worker path, upload conditions
HTML contains sensitive value Logging/privacy boundary Source/library messages and access policy

10. Intentionally broken diagnostic exercise

Take the Chapter 14 Lesson 2 evidence directory. Copy results/run-a/output.xml to results/stale/input.xml. Regenerate a report from it after running newer tests elsewhere. Without opening the HTML first, use the inspector/checksum tool to prove the report is based on the stale XML. Then regenerate from the intended current input into a new directory. Keep both derivatives until you can explain the provenance difference.

The lesson is not “rerun everything.” It is “prove which artifact is wrong, then change only that layer.”

11. Knowledge check

A Rebot merge produces duplicate-looking hierarchy. What should you check before changing test names?

Why should you checksum raw output before post-processing?

Is --legacyoutput a generic fix for any Rebot problem?

What is the correct reaction if a real token appears in output.xml?

Why can blanket keyword removal be dangerous after a failure?

12. Summary and next step

Production result diagnostics preserve originals, verify provenance/schema/identity, separate execution from post-processing and publishing, and make the least destructive correction. Lesson 5 combines these practices into an evidence-pipeline checkpoint.

Next lesson

Checkpoint Lab — Logs, Reports, output.xml, Rebot, and Result Post-Processing

Continue with Checkpoint Lab — Logs, Reports, output.xml, Rebot, and Result Post-Processing. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.