Chapter 04Lesson 04130–175 min

Test Cases, Tasks, Suites, Names, Documentation, and Metadata: Diagnostics, Failure Modes, and Production Practices

Diagnose suite-identity and hierarchy failures systematically by preserving first evidence, checking the executed root, validating initialization boundaries, inspecting full names, and repairing the smallest causal layer.

DiagnosticsExecution rootScope failureSecret safetyFirst evidence

Learning objectives

  • Diagnose wrong-root, wrong-full-name, initialization-boundary, duplicate-name, and mixed-responsibility failures without changing external systems.
  • Preserve output.xml, console output, commands, and source snapshots before repair.
  • Recognize metadata/documentation as potentially sensitive result evidence and remove secret material safely.
  • Separate hierarchy/selection failures from parser/import/keyword/external-system failures using a fixed diagnostic sequence.
  • Apply minimal repairs and verify them by rerunning the smallest controlled slice from the intended production root.

Diagnostic safety boundary. All failures in this lesson are engineered inside a disposable local suite tree using synthetic strings. Do not “test” hierarchy fixes by changing production targets, real credentials, CI secrets, database state, browser profiles, or Remote services.

1. Use the diagnostic sequence before editing

Hierarchy failures should be proven before external-system troubleshooting begins
flowchart TD
A[Preserve command + first artifacts] --> B[Confirm Robot/Python versions]
B --> C[Confirm executed root and selection]
C --> D[Validate parse/import graph]
D --> E[Inspect suite/test full names]
E --> F[Inspect __init__.robot participation]
F --> G[Inspect tags/metadata/lifecycle]
G --> H{External integration involved?}
H -- No --> I[Repair smallest hierarchy/config cause]
H -- Yes --> J[Only then inspect library/external state]
I --> K[Rerun controlled slice]
J --> K

This chapter sits before external integration chapters for a reason. If the wrong root was executed, restarting a browser driver or changing an API timeout is irrelevant. Preserve the exact command and first result tree first.

2. Failure mode: executing the wrong directory changes identity

Suppose CI expects Release Acceptance.Platform Acceptance.Api Health.Service Reports Healthy State, but a developer runs the file directly and sees Api Health.Service Reports Healthy State. Both can PASS, yet they are not the same execution context.

# Intended production-style root
python -m robot --dryrun --outputdir evidence/intended acceptance

# Accidental direct-file root
python -m robot --dryrun --outputdir evidence/direct acceptance/platform/api_health.robot

Repair: do not rename the test to force a string match. Restore the intended root, or explicitly decide that direct-file execution is the desired contract and update consumers. The cause is execution selection, not test syntax.

3. Failure mode: treating a filesystem path as a full test name

A path such as acceptance/platform/api_health.robot and a full name such as Release Acceptance.Platform Acceptance.Api Health.Service Reports Healthy State encode different information. CLI selection by --suite/--test uses suite/test names, not repository path syntax.

When a selector fails, print the actual result/model full names and compare them literally. Do not guess that every underscore or directory segment maps one-for-one to a name; custom Name settings and naming normalization can change that mapping.

from robot.api import ExecutionResult

result = ExecutionResult("evidence/intended/output.xml")
for child in result.suite.suites:
    print(child.source, "=>", child.full_name)

4. Failure mode: assuming parent initialization applies outside the executed tree

Create acceptance/__init__.robot with root metadata and a tag, then execute acceptance/platform directly. If you expect the root metadata/tag to appear, your model is wrong: the parent initialization is above the root and has no effect.

The correct fix is to run the intended higher root and select the child, or move the policy to the suite boundary that truly owns it. Do not duplicate root metadata into every file merely to hide inconsistent execution roots.

5. Failure mode: assuming variables/keywords in __init__.robot are inherited

This deliberately broken example defines a variable in the directory initialization file and tries to use it in a child suite:

# acceptance/platform/__init__.robot
*** Variables ***
${INIT ONLY}    not inherited by child suite files

# acceptance/platform/api_health.robot
*** Test Cases ***
Child Assumes Parent Variable Scope
    Should Be Equal    ${INIT ONLY}    not inherited by child suite files

The child can fail variable resolution even though the directory initialization was processed. Repair the architecture by putting shared data in an explicit resource/variable file and importing it into the child, or by passing state through the supported lifecycle mechanism. Do not solve this with global variables or arbitrary PYTHONPATH hacks.

6. Failure mode: duplicate or vague identities

Two sibling suite files both named smoke.robot under differently named directories can still be distinguishable by full name, but repeated generic test names such as Works degrade reports and incident communication. More dangerous is a custom Name strategy that intentionally gives different suite nodes the same human identity inside one hierarchy.

Repair by expressing meaningful domain intent in names and relying on the parent chain for context. Avoid embedding environment timestamps or random IDs in names; those destroy trend continuity. Runtime-specific details belong in variables/evidence, not stable suite identity.

7. Failure mode: metadata contains a secret

Metadata is attractive because it appears in results, which is exactly why it is the wrong place for secrets.

*** Settings ***
# WRONG — fake value shown only to teach the leak pattern.
Metadata    Api Token    token-FAKE_DO_NOT_USE

Even with a fake value, treat this as a leak pattern. In a real project, stop further artifact publication, preserve access-controlled incident evidence, rotate/revoke the real credential through the owning secret system, remove it from source/history according to organizational procedure, and replace it with a non-sensitive identifier such as Credential Profile=release-test-user. Do not delete all result files blindly and pretend the exposure never happened.

8. Failure mode: a giant suite mixes responsibilities

A 150-case file spanning identity, checkout, catalog, and RPA maintenance may technically run, but failures lack a clear owner and setup/teardown scope becomes risky. The diagnostic symptom is not a parser error; it is operational ambiguity: long reports, broad reruns, frequent merge conflicts, and one suite identity representing unrelated responsibilities.

Repair incrementally. First establish semantic child boundaries and preserve names/selectors in a migration map. Then split files and update CI selectors in the same reviewed change. Do not reformat and reorganize hundreds of cases at once with no before/after result-tree evidence.

9. Failure mode: mixing test and RPA semantics without an explicit mode

If a team calls a nightly cleanup “a test” solely to place it in the release report, a failure can block releases even though no product assertion failed. Conversely, treating a release assertion as “just a task” can weaken quality semantics.

Repair by classifying intent first. Keep a test root and task root separate, with independent CI meaning and artifact retention where appropriate. The syntax similarity is not evidence that the organizational meaning is interchangeable.

10. Performance: diagnose hierarchy cost only when it is causal

Suite organization can affect discovery/parse time and initialization cost, but this chapter is not a concurrency-tuning chapter. If a large tree is slow, measure discovery/parsing/import/setup versus keyword/external time before restructuring. A deeply nested report may be inconvenient without being the runtime bottleneck.

Do not “optimize” by skipping parent roots if that removes required suite lifecycle, tags, or metadata. A faster run with a different identity/lifecycle contract is not the same execution.

11. Intentionally broken diagnostic lab

Use the Chapter 02/03/04 disposable environment. Create this tree:

broken-hierarchy/
├── acceptance/
│   ├── __init__.robot
│   └── component/
│       ├── __init__.robot
│       └── smoke.robot
└── evidence/

Root initialization:

*** Settings ***
Name         Release Gate
Metadata     Owner    quality-platform
Test Tags    synthetic

Child initialization:

*** Variables ***
${PARENT VALUE}    hidden-scope-assumption

Child suite:

*** Test Cases ***
Scope Assumption Fails
    Should Be Equal    ${PARENT VALUE}    hidden-scope-assumption

Run from the root and preserve the failure:

python -m robot --outputdir evidence/first broken-hierarchy/acceptance

Expected cause: variable resolution fails in the child suite. Preserve evidence/first. Repair by creating acceptance/component/shared.resource containing the variable and importing it explicitly in smoke.robot. Rerun into evidence/repaired. The minimal change proves that the problem was scope architecture, not the suite name or external state.

12. Production diagnostic checklist

  • Record exact Python/Robot versions and the exact command.
  • Record the execution input root and all --suite/--test/--include/--exclude selectors.
  • Preserve first output.xml, console output, log/report, and source snapshot.
  • Print actual suite/test full names before editing selectors.
  • Identify which __init__.robot files are inside the executed tree.
  • Check metadata/tags for unsafe values.
  • Confirm tests versus tasks and ownership semantics.
  • Only after hierarchy/import/scope checks pass, investigate external systems.
  • Apply one minimal correction and rerun a controlled slice into a new evidence directory.

13. Knowledge check

A direct-file run passes, but CI running the root fails in suite setup. Is the direct-file PASS proof that CI is wrong?

Why is deleting output.xml a poor first response to a secret leak?

A child suite cannot resolve a variable defined in its parent __init__.robot. What is the architectural fix?

Why can a rename be the cause of a CI selection failure even when all test bodies are unchanged?

When should external browser/API/database troubleshooting begin in this diagnostic sequence?

14. Summary and next step

Production diagnosis starts with identity: exact command, root, full names, initialization participation, scope, and first result evidence. Wrong-root and hidden-initialization assumptions can produce misleading green or red runs without any external-system defect.

Lesson 5 combines the chapter into a three-level checkpoint: build a hierarchy, run it from two deliberate roots, explain the resulting identities, and publish a naming/boundary convention with evidence.

Next lesson

Checkpoint Lab — Test Cases, Tasks, Suites, Names, Documentation, and Metadata

Continue with Checkpoint Lab — Test Cases, Tasks, Suites, Names, Documentation, and Metadata. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.