Chapter 14Lesson 03180–240 min

Logs, Reports, output.xml, Rebot, and Result Post-Processing: Configuration, Design Patterns, and Trade-Offs

Choose a result-retention and post-processing strategy that balances diagnosability, size, privacy, portability, shard identity, and CI publishing without weakening status truth.

Design trade-offsRetentionFlatten/removeCI publishingPrivacy

Learning objectives

  • Choose between raw output retention and presentation-only strategies based on evidence needs.
  • Use combine versus merge according to suite identity, not convenience.
  • Explain remove/flatten trade-offs and when they affect log only versus a generated output file.
  • Separate local viewing, CI publishing, and long-term retention responsibilities.
  • Define a policy for shard outputs that preserves first-failure diagnosis and privacy.

Current compatibility baseline. Verified 2026-08-31: Robot Framework 7.4.2 is the current stable release and requires Python 3.8+; 7.5b1 is a pre-release and is not required here. The 7.4.2 result.xsd is schema version 5; default XML outputs carry a schemaversion attribute and Robot Framework 7.x can also read legacy output. --legacyoutput exists for Robot Framework 6.x-compatible consumers. Rebot combines independent outputs by creating a new parent suite; --merge is for the same logical top-level suite, including reruns, where later matching test results replace earlier results and later SKIP results do not replace originals. The built-in Testdoc tool is deprecated in 7.4.2 in favor of external Testdoc and is not a result-reporting substitute.

1. Configuration begins with an evidence question, not a CLI flag

The right output strategy depends on who needs to answer what later. A developer debugging a failed keyword needs more detail than a release manager looking at test counts. A compliance archive may require immutable raw evidence and checksums. A public CI job may need only sanitized summaries. Starting with --removekeywords ALL because logs are large is backwards: first define required evidence, then reduce only what is safe to lose.

2. Decision matrix

Decision Prefer A when… Prefer B when… Evidence consequence
Raw output vs HTML only You need Rebot/tooling/rerun diagnosis Only disposable local viewing is required Raw output preserves richer machine-readable evidence
Combine vs merge Inputs are independent roots Inputs are same logical top-level suite Wrong choice distorts hierarchy or replacement semantics
Detailed log vs reduced log First-failure diagnosis is important High-volume passed details have low value Reduction can save size but may remove context
Remove vs flatten You can discard matching keyword/message data You want parent messages but not deep child structure Flatten retains messages while collapsing structure
Separate shard outputs vs immediate merge You need worker-level provenance A validated aggregate view is needed after capture Keep shard originals even after aggregate generation
Local HTML vs CI publish Sensitive detail is for engineers only Approved summaries can be shared to broader users Publishing is an access-control decision

3. Raw output retention versus HTML-only workflows

Robot Framework can generate report and xUnit directly during execution, and Rebot can later generate them from raw output. Keeping output.xml gives you the option to regenerate with different filters, titles, or newer Robot tooling. Disabling raw output may be reasonable for truly disposable runs, but it prevents later Robot-aware post-processing. Conversely, keeping every raw output forever can create privacy and storage risk. Treat retention duration and access level as policy, not defaults.

4. Remove and flatten are evidence transformations

--removekeywords discards matching keyword data. In current Robot Framework it can remove passed keywords, loop iterations, WUKS internals, or name/tag matches. --flattenkeywords keeps messages but collapses selected child keyword structure into the matching parent. During normal execution these options affect the log presentation while the XML output retains the original execution data. With Rebot, they affect the processed model and therefore any new output XML explicitly written with --output.

# Derive a smaller report while preserving the original raw output.
python -m robot.rebot   --removekeywords PASSED   --log reduced-log.html   --report reduced-report.html   --outputdir results/reduced   results/raw/output.xml

# If you also request --output, the new XML is a transformed derivative.
python -m robot.rebot   --flattenkeywords FOR   --output flattened.xml   --outputdir results/flattened   results/raw/output.xml

Label transformed XML explicitly. Do not overwrite the original with a reduced artifact and later call it raw evidence.

5. Size and memory: measure the right stage

Large loops and deeply nested libraries can make output and log artifacts large. Measure execution-time serialization, file size, Rebot processing time/memory, and CI artifact-upload time separately. Current documentation notes that flattening can save more memory during post-processing because it happens while parsing, whereas removal occurs after the result model has already been built. That is a causal performance distinction—not a reason to flatten everything.

6. Privacy is not equivalent to smaller files

A smaller log can still contain one critical secret, and a large log can contain none. Prevent sensitive data from being logged at the source; use Robot 7.4 Secret semantics only where appropriate; remember that Secret is masking, not encryption; treat external-library output as a separate disclosure surface. Apply artifact access controls and retention limits. If broad sharing requires a sanitized derivative, preserve the restricted original according to incident/compliance policy and record the transformation.

7. Schema/version compatibility policy

The XML format changed incompatibly in Robot Framework 7.0. Robot Framework 7.4.2 can process both current and legacy formats, and --legacyoutput can generate a 6.x-compatible output for external consumers that have not upgraded. That flag is an interoperability fallback—not the default design. A robust consumer inspects generator/schemaversion and tests against representative files rather than assuming a filename named output.xml means one timeless schema.

8. Shards: aggregate late, preserve originals first

Evidence pipeline for parallel/sharded execution
flowchart TD
S1[Worker/shard 1 output] --> V[Validate provenance/schema]
S2[Worker/shard 2 output] --> V
S3[Worker/shard 3 output] --> V
V --> O[Retain original shard artifacts]
O --> A{Same logical root?}
A -->|Yes| M[Rebot --merge]
A -->|No| C[Rebot combine]
M --> P[Published aggregate]
C --> P
P --> R[Retention / access / checksums]

Parallelism does not change the identity rule. Pabot and CI sharding are covered later, but Chapter 14 establishes the artifact boundary: every worker must have collision-free output paths; originals should be captured before aggregation; and merge/combine choice depends on logical suite identity, not on the fact that files arrived from separate workers.

9. Testdoc is a different product concern

Built-in Testdoc is deprecated in 7.4.2 in favor of external Testdoc. Testdoc documents test data; it does not replace output.xml, log/report, or Rebot. Avoid operational documentation that says “generate Testdoc to archive execution results.” Source documentation and execution evidence are different artifacts.

10. Worked scenario: choose a policy

A regulated CI job runs 12 shards. Failed runs must be diagnosable for 30 days; passing runs need trend summaries for one year; raw artifacts contain synthetic IDs but no production secrets. A defensible policy is: capture each shard raw output with worker/run IDs; checksum and restrict access; aggregate only after all originals exist; publish report/xUnit to normal CI viewers; retain failed raw/log artifacts for 30 days; retain long-term aggregate summaries without detailed keyword payloads; and validate schema/version on ingestion. This design keeps the evidence needed for failures without making detailed data universally visible forever.

11. Knowledge check

Why can --removekeywords PASSED be acceptable for a derivative but risky for the only retained artifact?

Why is flattening not a general privacy control?

What should happen before shard outputs are merged or combined?

When should --legacyoutput be used?

12. Summary and next step

You can now design evidence intentionally: retain what diagnosis and governance require, reduce only derived artifacts with explicit policy, preserve shard provenance, use merge/combine according to identity, and keep schema/privacy boundaries visible. Lesson 4 focuses on diagnosing what goes wrong when those controls are ignored.

Next lesson

Logs, Reports, output.xml, Rebot, and Result Post-Processing: Diagnostics, Failure Modes, and Production Practices

Continue with Logs, Reports, output.xml, Rebot, and Result Post-Processing: Diagnostics, Failure Modes, and Production Practices. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.