Chapter 25Lesson 04180–240 min

CI/CD Integration with GitHub Actions, GitLab CI, and Jenkins: Diagnostics, Failure Modes, and Production Practices

Diagnose CI-only Robot failures by preserving the first failed run, separating wrapper errors from Robot failures and SUT failures, and resisting shortcuts that hide status, secrets, version drift, or shared-state contention.

DiagnosticsFirst-failure evidenceSecretsCache driftConcurrencyProduction practices

Learning objectives

  • Follow a deterministic diagnostic sequence from first-failure preservation through versions, command/selection, imports, external state, concurrency, and minimal rerun.
  • Recognize false-green pipelines caused by ignored return codes or CI retries.
  • Diagnose missing artifacts, working-directory/import mismatches, dependency drift, and stale caches without deleting evidence.
  • Prevent credential leakage and shared-environment collisions in CI logs and parallel jobs.
  • Repair an intentionally broken workflow while preserving the original failed output.
Production rule

Do not troubleshoot CI by deleting output.xml, widening timeouts, adding blanket retries, disabling TLS/SSH verification, editing global PYTHONPATH, or pointing the job at a production system. Preserve the failed run and isolate the layer that diverged.

1. Diagnostic sequence

  1. Preserve first-failure artifacts. Copy or retain console output, output.xml, log/report/xUnit, runtime manifest, and job identity before rerunning.
  2. Confirm versions. Python, Robot, external libraries, Pabot if used, relevant action/plugin/runner versions.
  3. Confirm execution contract. Commit/SHA, working directory, exact command, selected tests/tags/profile, environment name, synthetic input IDs.
  4. Validate parse/import graph. Resource/library paths, ${CURDIR}/${EXECDIR}, package installation, current directory.
  5. Inspect variable and secret boundaries. Names/presence, not values; check accidental provider-to-Robot mapping.
  6. Inspect external state. Target, fixture ownership, cleanup, correlation IDs.
  7. Inspect timing/concurrency/container state only if relevant. Pabot workers, CI matrix jobs, ports, shared accounts, service DNS, runner capacity.
  8. Apply the least destructive correction. Rerun the smallest controlled slice, keeping the original evidence immutable.

2. Common failure modes and what they actually mean

Symptom Likely broken layer Evidence to inspect Do not do
Robot log is red but job is green Wrapper/status propagation command, shell operators, exit-code.txt Add more assertions—the gate is already miswired
No log/report on failed job Provider finalization/artifact rule workflow/Jenkins post config, artifact paths Force Robot to pass so upload runs
Passes locally, import fails in CI cwd/package/PYTHONPATH mismatch manifest, pwd, pip list, import path Append arbitrary directories to PYTHONPATH
Different failure after cache hit Dependency/cache drift cache key, lock hash, pip freeze Delete all evidence and rerun repeatedly
Secret appears in console logging/command/env misuse redacted log lines, command construction Assume provider masking catches every derived value
Parallel jobs fail intermittently shared mutable SUT/workspace resource resource IDs, ports, files, Pabot/CI job IDs Increase retries
Later stage overwrites earlier results artifact naming/working directory run/job/stage artifact names and paths Reuse one mutable artifact folder across unrelated jobs

3. Intentionally broken example: false green + missing evidence

# DO NOT COPY — intentionally broken GitHub Actions fragment
- name: Run Robot
  run: python tools/ci_run.py || true

- name: Upload results
  uses: actions/upload-artifact@v7
  with:
    name: robot-results
    path: artifacts/robot/

There are two independent defects. First, || true discards the Robot failure code, so the delivery gate can become green. Second, the upload step has no if: always(); because the first defect masks failure it happens to run, but if installation or another earlier step fails, the artifact step can still be skipped. The apparently “working” pipeline is therefore doubly misleading.

Repair

- name: Run Robot
  run: python tools/ci_run.py

- name: Upload results
  if: always()
  uses: actions/upload-artifact@v7
  with:
    name: robot-results-${{ github.run_id }}
    path: artifacts/robot/
    if-no-files-found: warn

The Robot step can fail naturally. GitHub still evaluates the final artifact step because of always(). The job remains failed; evidence remains available. If no files exist because failure occurred before Robot started, the upload warning does not replace the original cause.

4. Working-directory and import mismatches

CI often starts from repository root, but a job-level working-directory, Jenkins dir(...), container workdir, or wrapper cwd can change path resolution. Do not “fix” this with a global PYTHONPATH export until you know which import is intended to be package-based and which resource path should be relative to the importing file.

# Minimal evidence, safe to print
python -c "import os,sys; print('cwd=', os.getcwd()); print('python=', sys.executable)"
python -m pip freeze
python -m robot --version

# Then run the smallest suite/file that reproduces the import problem.
python -m robot --dryrun --outputdir artifacts/dryrun suites/ci_contract.robot

5. Secret failures: masking is not a data-flow policy

Provider secret stores reduce accidental exposure, but they cannot guarantee safety after a secret is transformed, concatenated, written into a test fixture, included in an HTTP body, or returned by a library. Keep credentials out of command-line arguments where process listings or logs may expose them; inject through approved environment/credential mechanisms and pass them only to the component that needs them.

For diagnostic evidence, log credential_present=true, credential alias/name, or a non-secret correlation identifier—not the token. Use fake values for this chapter's labs.

6. CI retry is not a flaky-test strategy

A provider-level “retry job” button can be useful after infrastructure incidents, but an automatic retry can hide nondeterminism by replacing the first failure with a later pass. Preserve the first run's Robot output and job identity. If retry is part of an approved design, make it bounded and report both attempts.

Likewise, Robot's Wait Until Keyword Succeeds and rerun/merge workflows solve specific problems; they should not be added globally to make a gate greener.

7. Performance diagnosis by layer

Layer Measurement Typical action
Checkout/runtime setup job step duration shallow checkout where appropriate; prebuilt approved image only if justified
Dependency install pip timing + cache hit/miss lock dependencies; cache with correct key
Robot parse/import/setup console/output timing reduce import side effects; keep setup scoped
External system latency keyword timings/correlation IDs fix service/test data; avoid giant sleeps
Robot logging/output output size + report generation time target excessive logging; do not delete failure evidence
Pabot scheduling worker utilization + contention adjust processes only after isolation
CI/container startup runner/image/service start time measure separately from Robot execution
Retry/rerun first + subsequent attempt time fix cause before normalizing retry cost

8. Production practices

  • Pin Robot and external test dependencies in a reviewed manifest or lock file.
  • Record runtime versions and the exact repository commit with every run.
  • Keep target allowlists and environment promotion explicit; never use production as a troubleshooting sandbox.
  • Name artifacts with immutable run/job/shard identity so later stages cannot overwrite first-failure evidence.
  • Use least-privilege provider permissions and credentials; fork/untrusted contribution workflows need stricter secret and checkout boundaries.
  • Make cleanup idempotent and scoped to resources owned by the current test run.
  • Measure before adding caches, Pabot, matrices, containers, or retries.

Knowledge check

A failed GitLab job has xunit.xml in the Tests UI but no log.html artifact. Is Robot evidence missing?

A Jenkins pipeline uses sh 'python tools/ci_run.py || true' so that junit can execute. What is the safer structure?

Why is deleting a stale cache not the first diagnostic step?

Two Pabot workers and two CI matrix jobs all use the same account and output folder. What kind of failure is this?

Summary and bridge

A reliable CI investigation begins with the failed run, not with a rerun. Preserve evidence, prove versions and command identity, isolate provider/workspace/dependency/SUT/concurrency layers, and repair the smallest broken contract. Lesson 5 turns those practices into a checkpoint where a forced failure must close the gate and still leave a complete evidence packet.

Next lesson

Checkpoint Lab — CI/CD Integration with GitHub Actions, GitLab CI, and Jenkins

Continue with Checkpoint Lab — CI/CD Integration with GitHub Actions, GitLab CI, and Jenkins. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Current primary references

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.