Chapter 37Lesson 03~125 minutes

Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Configuration, Design Choices, and Tradeoffs

Choose retry, new pipeline, debug tracing, cleanup, mitigation, and remediation strategies by their effect on GitLab and external state.

Design tradeoffsRetry semanticsDebug safetyTargeted repair

Learning objectives

  • Distinguish a job retry from a new pipeline when configuration/includes can differ.
  • Use automatic retry only for understood transient classes with bounded side effects.
  • Treat CI_DEBUG_TRACE and service debugging as security-sensitive diagnostics.
  • Prefer targeted evidence-preserving repair over broad cleanup or delete-and-recreate actions.
  • Separate temporary mitigation from permanent remediation with owners and verification.

Design rule: a recovery action is part of the incident, not an administrative convenience. Before retrying, cleaning, enabling trace, or rebuilding, state exactly which GitLab/external state will change and what evidence could be lost.

1. Retry a job, retry failed jobs, or create a new pipeline?

Current GitLab behavior makes this distinction operationally important. A job retry stays in the same pipeline. GitLab’s YAML reference states that included configuration is not fetched again when a job is rerun; jobs in that pipeline use the configuration resolved when the pipeline was created. A new/rerun pipeline resolves includes again, so changed remote/project includes can change the experiment.

Choice Best fit State preserved/changed Evidence requirement
Retry one job Transient, side-effect-safe failure with the original pipeline/configuration still desired Same pipeline identity; new job attempt; original source/config snapshot remains the basis Keep original job ID/trace; compare retried job ID/attempt
Retry failed/canceled jobs in pipeline Several independent transient failures under same pipeline intent Same pipeline; retries failed/canceled jobs; successful jobs normally stay untouched Record which jobs were retried and side effects
Create a new pipeline Configuration/include/source/policy changed, or you need a clean experiment New pipeline ID; includes/rules may resolve differently Record old/new SHA, config, includes, variables and pipeline IDs
Do not retry yet Destructive deploy, unknown external side effects, secret exposure, or missing first-failure evidence No execution mutation Preserve state and determine idempotency/recovery path first

2. Retry APIs do not mean “safe to retry”

GitLab exposes a job retry endpoint and a pipeline retry endpoint. The pipeline endpoint retries failed or canceled jobs; if none exist, it has no effect. Those APIs provide mechanics, not idempotency analysis. A failed deployment job may already have partially changed an external system.

# Examples only — use an authorized disposable project.
# Retry one job:
# POST /projects/:id/jobs/:job_id/retry
# Retry failed/canceled jobs in one pipeline:
# POST /projects/:id/pipelines/:pipeline_id/retry

# Before either action, record:
# project/ref/SHA, pipeline ID, job ID/status/failure_reason,
# environment/deployment record, artifact digest, and external target state.

3. Automatic retry should target known transient classes

The retry keyword defaults to zero when omitted and supports a maximum of two retries. Use failure-specific retry only when the failure is understood and side effects are safe. Current documentation is especially version-sensitive: GitLab 19.1 refines several runner failure reasons and deprecates some older aggregate names. If your estate spans versions, build the retry policy against the oldest supported instance/Runner semantics and test it.

test:
  script: ./run-tests.sh
  retry:
    max: 1
    when:
      - runner_external_dependency_failure
      - runner_interrupted

memory_sensitive_test:
  script: ./run-memory-test.sh
  retry:
    max: 1
    exit_codes: 137

Do not use retry: 2 on every script failure to hide flaky tests. A retry policy should state the expected transient cause and preserve first-attempt evidence.

4. Debug tracing versus secret exposure

CI_DEBUG_TRACE=true makes Runner output much more verbose, but GitLab explicitly warns that it exposes all variables and secrets available to the job. CI_DEBUG_SERVICES is also dangerous: service logs can interleave with the job trace in ways that defeat masking. Treat both as security-sensitive incident actions.

Need Safer first step Escalation only if necessary
See failing shell command Add narrow, non-secret command/result logging Short-lived CI_DEBUG_TRACE on disposable/controlled job
Inspect service container Raise service app log level with secrets redacted CI_DEBUG_SERVICES only with access/retention controls
Inspect runner manager Normal runner logs, list/verify/status, metrics Temporary debug log level on authorized runner manager
Inspect variable selection Print names/presence/length/hash where safe, never values Review protected/masked/file variable configuration via authorized UI/API

Never paste debug traces into public issues before secret review. If a secret may have been exposed, rotate/revoke it and treat the trace as sensitive evidence.

5. Broad cleanup versus targeted repair

Deleting caches, artifacts, logs, workspaces, or whole pipelines can make symptoms disappear while destroying the evidence needed to explain them. GitLab documents artifact/job-log deletion as destructive and irreversible. Preserve first, then clean the smallest state that is demonstrably corrupt.

Symptom Bad shortcut Targeted response
Dependency cache suspected Delete all project caches and rerun everything Record cache key/object evidence; bypass or invalidate the single affected key
Runner workspace dirty Recreate all runners Prove workspace/executor issue; clean/replace one disposable worker or job path
Artifact corrupt Rebuild release from newer source Preserve producer SHA/job/artifact digest; re-transfer or reproduce only under controlled release rules
Debug trace leaked secrets Erase everything immediately Restrict access, preserve authorized incident evidence, rotate secret, then follow retention/erase policy
Deploy partially failed Delete environment and redeploy Inspect external target and deployment record; roll forward/back with verified artifact identity

6. Temporary mitigation is not permanent remediation

A mitigation restores flow while containing risk: route a job to a known-good runner, disable one failing optional integration, or pin a known-good immutable image. Remediation removes the root cause: repair DNS trust, fix runner fleet configuration, eliminate flake, or correct the deployment contract. Every mitigation needs an owner, expiry/review trigger, and evidence that it does not silently become permanent.

7. Worked decision table

Use this table before applying a recovery action during the checkpoint.

Case Choice Tier/trust prerequisite Predicted state change Proof
YAML invalid before pipeline Fix config, then lint/simulate Free; project source access New commit/config; no runner change Lint valid + merged config/jobs
Pending: missing gpu tag Repair routing or requirement; no blind retry Runner/project admin appropriate to chosen repair Runner match changes; source may not Job status + runner tag/scope evidence
Transient registry DNS outage, no side effect yet Retry once after dependency health restored Job retriable; dependency independently healthy New job attempt in same pipeline Old/new job traces + resolver/registry evidence
Remote include changed after original pipeline Create new pipeline if new config is intended Authorized source/config change New pipeline can resolve different include Old/new merged config + pipeline IDs
Deploy may have partially mutated target Do not retry until target inspected External target access + rollback authority None until safe plan selected Provider/target state + deployment record
Need verbose debugging with secrets present Prefer narrow logging; trace only in controlled scope Restricted log viewers + rotation plan Potential sensitive log creation Trace access/retention + secret rotation evidence

8. Keep the state boundaries explicit

  • Repository YAML and merged pipeline configuration are not runner-host config.toml.
  • GitLab pipeline/job/report/artifact/cache state is not the same as an external registry, Kubernetes cluster, cloud account, or IaC state backend.
  • CI_JOB_TOKEN, personal/deploy/trigger tokens, and OIDC identities have different trust and lifetime models.
  • A build artifact digest is evidence of what was produced; deployment authorization is evidence of who/what could deploy it; target verification is evidence of what is actually running.
  • Copied YAML is a local fork; a versioned include/component is an external dependency whose exact revision must be part of evidence.
Next lesson

Next: Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Diagnostics, Failure Modes, Security, and Performance

Apply the decision model to deliberately broken examples, including blind retries, runner tags, DNS/TLS, secret-bearing debug logs, evidence-destroying cleanup, and unverified rollback.

Knowledge check

What evidence should you preserve before retrying a failed GitLab job?

A job remains pending with no trace. Should you start by debugging the application script?

Why is blind retry a poor default for a flaky failure?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Further reading — current official GitLab sources

Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.