Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Configuration, Design Choices, and Tradeoffs
Choose retry, new pipeline, debug tracing, cleanup, mitigation, and remediation strategies by their effect on GitLab and external state.
Learning objectives
- Distinguish a job retry from a new pipeline when configuration/includes can differ.
- Use automatic retry only for understood transient classes with bounded side effects.
- Treat CI_DEBUG_TRACE and service debugging as security-sensitive diagnostics.
- Prefer targeted evidence-preserving repair over broad cleanup or delete-and-recreate actions.
- Separate temporary mitigation from permanent remediation with owners and verification.
Design rule: a recovery action is part of the incident, not an administrative convenience. Before retrying, cleaning, enabling trace, or rebuilding, state exactly which GitLab/external state will change and what evidence could be lost.
1. Retry a job, retry failed jobs, or create a new pipeline?
Current GitLab behavior makes this distinction operationally important. A job retry stays in the same pipeline. GitLab’s YAML reference states that included configuration is not fetched again when a job is rerun; jobs in that pipeline use the configuration resolved when the pipeline was created. A new/rerun pipeline resolves includes again, so changed remote/project includes can change the experiment.
| Choice | Best fit | State preserved/changed | Evidence requirement |
|---|---|---|---|
| Retry one job | Transient, side-effect-safe failure with the original pipeline/configuration still desired | Same pipeline identity; new job attempt; original source/config snapshot remains the basis | Keep original job ID/trace; compare retried job ID/attempt |
| Retry failed/canceled jobs in pipeline | Several independent transient failures under same pipeline intent | Same pipeline; retries failed/canceled jobs; successful jobs normally stay untouched | Record which jobs were retried and side effects |
| Create a new pipeline | Configuration/include/source/policy changed, or you need a clean experiment | New pipeline ID; includes/rules may resolve differently | Record old/new SHA, config, includes, variables and pipeline IDs |
| Do not retry yet | Destructive deploy, unknown external side effects, secret exposure, or missing first-failure evidence | No execution mutation | Preserve state and determine idempotency/recovery path first |
2. Retry APIs do not mean “safe to retry”
GitLab exposes a job retry endpoint and a pipeline retry endpoint. The pipeline endpoint retries failed or canceled jobs; if none exist, it has no effect. Those APIs provide mechanics, not idempotency analysis. A failed deployment job may already have partially changed an external system.
# Examples only — use an authorized disposable project.
# Retry one job:
# POST /projects/:id/jobs/:job_id/retry
# Retry failed/canceled jobs in one pipeline:
# POST /projects/:id/pipelines/:pipeline_id/retry
# Before either action, record:
# project/ref/SHA, pipeline ID, job ID/status/failure_reason,
# environment/deployment record, artifact digest, and external target state.
3. Automatic retry should target known transient classes
The retry keyword defaults to zero when omitted and
supports a maximum of two retries. Use failure-specific retry only
when the failure is understood and side effects are safe. Current
documentation is especially version-sensitive: GitLab 19.1 refines
several runner failure reasons and deprecates some older aggregate
names. If your estate spans versions, build the retry policy against
the oldest supported instance/Runner semantics and test it.
test:
script: ./run-tests.sh
retry:
max: 1
when:
- runner_external_dependency_failure
- runner_interrupted
memory_sensitive_test:
script: ./run-memory-test.sh
retry:
max: 1
exit_codes: 137
Do not use retry: 2 on every script failure to hide
flaky tests. A retry policy should state the expected transient
cause and preserve first-attempt evidence.
4. Debug tracing versus secret exposure
CI_DEBUG_TRACE=true makes Runner output much more
verbose, but GitLab explicitly warns that it exposes all variables
and secrets available to the job. CI_DEBUG_SERVICES is
also dangerous: service logs can interleave with the job trace in
ways that defeat masking. Treat both as security-sensitive incident
actions.
| Need | Safer first step | Escalation only if necessary |
|---|---|---|
| See failing shell command | Add narrow, non-secret command/result logging |
Short-lived CI_DEBUG_TRACE on
disposable/controlled job
|
| Inspect service container | Raise service app log level with secrets redacted |
CI_DEBUG_SERVICES only with access/retention
controls
|
| Inspect runner manager |
Normal runner logs,
list/verify/status,
metrics
|
Temporary debug log level on authorized runner manager |
| Inspect variable selection | Print names/presence/length/hash where safe, never values | Review protected/masked/file variable configuration via authorized UI/API |
Never paste debug traces into public issues before secret review. If a secret may have been exposed, rotate/revoke it and treat the trace as sensitive evidence.
5. Broad cleanup versus targeted repair
Deleting caches, artifacts, logs, workspaces, or whole pipelines can make symptoms disappear while destroying the evidence needed to explain them. GitLab documents artifact/job-log deletion as destructive and irreversible. Preserve first, then clean the smallest state that is demonstrably corrupt.
| Symptom | Bad shortcut | Targeted response |
|---|---|---|
| Dependency cache suspected | Delete all project caches and rerun everything | Record cache key/object evidence; bypass or invalidate the single affected key |
| Runner workspace dirty | Recreate all runners | Prove workspace/executor issue; clean/replace one disposable worker or job path |
| Artifact corrupt | Rebuild release from newer source | Preserve producer SHA/job/artifact digest; re-transfer or reproduce only under controlled release rules |
| Debug trace leaked secrets | Erase everything immediately | Restrict access, preserve authorized incident evidence, rotate secret, then follow retention/erase policy |
| Deploy partially failed | Delete environment and redeploy | Inspect external target and deployment record; roll forward/back with verified artifact identity |
6. Temporary mitigation is not permanent remediation
A mitigation restores flow while containing risk: route a job to a known-good runner, disable one failing optional integration, or pin a known-good immutable image. Remediation removes the root cause: repair DNS trust, fix runner fleet configuration, eliminate flake, or correct the deployment contract. Every mitigation needs an owner, expiry/review trigger, and evidence that it does not silently become permanent.
7. Worked decision table
Use this table before applying a recovery action during the checkpoint.
| Case | Choice | Tier/trust prerequisite | Predicted state change | Proof |
|---|---|---|---|---|
| YAML invalid before pipeline | Fix config, then lint/simulate | Free; project source access | New commit/config; no runner change | Lint valid + merged config/jobs |
Pending: missing gpu tag |
Repair routing or requirement; no blind retry | Runner/project admin appropriate to chosen repair | Runner match changes; source may not | Job status + runner tag/scope evidence |
| Transient registry DNS outage, no side effect yet | Retry once after dependency health restored | Job retriable; dependency independently healthy | New job attempt in same pipeline | Old/new job traces + resolver/registry evidence |
| Remote include changed after original pipeline | Create new pipeline if new config is intended | Authorized source/config change | New pipeline can resolve different include | Old/new merged config + pipeline IDs |
| Deploy may have partially mutated target | Do not retry until target inspected | External target access + rollback authority | None until safe plan selected | Provider/target state + deployment record |
| Need verbose debugging with secrets present | Prefer narrow logging; trace only in controlled scope | Restricted log viewers + rotation plan | Potential sensitive log creation | Trace access/retention + secret rotation evidence |
8. Keep the state boundaries explicit
-
Repository YAML and merged pipeline configuration are not
runner-host
config.toml. - GitLab pipeline/job/report/artifact/cache state is not the same as an external registry, Kubernetes cluster, cloud account, or IaC state backend.
-
CI_JOB_TOKEN, personal/deploy/trigger tokens, and OIDC identities have different trust and lifetime models. - A build artifact digest is evidence of what was produced; deployment authorization is evidence of who/what could deploy it; target verification is evidence of what is actually running.
- Copied YAML is a local fork; a versioned include/component is an external dependency whose exact revision must be part of evidence.
Knowledge check
What evidence should you preserve before retrying a failed GitLab job?
Preserve pipeline/job IDs, source ref and SHA, compiled configuration/rule outcome, first-failure trace, runner/executor/image identity, and relevant artifact/deployment/external evidence.
A job remains pending with no trace. Should you start by debugging the application script?
No. A pending job has not executed the script. Inspect job eligibility, requested tags, runner scope/protection/status, manager connectivity, and executor capacity first.
Why is blind retry a poor default for a flaky failure?
It can erase the timing and first-failure evidence needed to distinguish a real flake from runner, network, configuration, or external-system failure.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Further reading — current official GitLab sources
Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.
- GitLab Docs — Validate GitLab CI/CD configuration
- GitLab Docs — CI Lint API
- GitLab Docs — Pipeline editor / full configuration
- GitLab Docs — CI/CD YAML syntax and retry
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — CI/CD jobs and statuses
- GitLab Docs — Troubleshooting CI/CD variables / debug trace
- GitLab Docs — Services and CI_DEBUG_SERVICES
- GitLab Docs — GitLab Runner commands
- GitLab Docs — Troubleshooting GitLab Runner
- GitLab Docs — Runner advanced configuration
- GitLab Docs — Docker executor image pull errors
- GitLab Docs — Docker build troubleshooting (DNS/TLS)
- GitLab Docs — Job artifacts and destructive deletion
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.