Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Diagnostics, Failure Modes, Security, and Performance
Diagnose realistic GitLab CI/CD failure modes without blind retries, secret leakage, TLS bypass, evidence destruction, or unverified rollback.
Learning objectives
- Apply the full evidence-first sequence from source/config through external target verification.
- Diagnose flake, wrong runner tags, DNS/TLS failures, debug-log leakage, destructive cleanup, and rollback verification gaps.
- Interpret an intentionally broken pending-job example and reject irrelevant fixes.
- Use Runner diagnostics and network evidence without broadening privileges or disabling trust checks.
- Write a minimal recovery playbook another engineer can follow under pressure.
Evidence-first rule: preserve the original pipeline/job IDs, source SHA, merged configuration, first failing trace/report, runner/executor identity, and external target state before any retry, erase, cleanup, or rollback. A repaired run is new evidence; it does not replace the original failure.
1. The full diagnostic sequence
- Preserve pipeline/job IDs and first-failure evidence.
- Confirm pipeline source/ref/SHA and policy/config revision.
- Validate/inspect merged configuration and workflow/job rule decisions.
- Inspect job graph and current status: created, pending, preparing, running, waiting, failed.
- Confirm runner scope/tags/protection, Runner version, executor, image/helper/toolchain.
- Inspect the first failing script/tool/network operation and non-secret effective inputs.
- Inspect reports, artifacts, caches, registry state and digests.
- Inspect environment/deployment and independently verify the external target.
- Apply the least destructive causal correction.
- Retry only the smallest safe scope, or create a new pipeline when source/config changed.
- Verify recovery and retain both before/after evidence.
2. Failure mode: blind retry hides a flake
Suppose attempt 1 fails an integration test and attempt 2 passes without any source/configuration change. The correct conclusion is nondeterminism exists, not “the pipeline is fixed.” Preserve attempt 1’s job ID, test report and timing; compare dependencies and environment; then decide whether a narrow transient retry policy is justified.
# Bad: hides every kind of script failure.
test:
script: ./integration-tests.sh
retry: 2
# Better only after evidence shows a specific transient class.
test:
script: ./integration-tests.sh
retry:
max: 1
when: runner_external_dependency_failure
4. Failure mode: DNS/TLS error misdiagnosed as application failure
Image pulls, package downloads, Git fetches, secret providers, and
deployment APIs all depend on name resolution and TLS trust. An x509
failure or no such host occurs before application
tests. Record the hostname, resolver, proxy path, certificate
issuer/chain and which container/process makes the request.
# Examples for an authorized disposable environment; do not disable TLS verification.
python - <<'PY'
import socket
for host in ['gitlab.com','registry.gitlab.com']:
try: print(host, sorted({x[4][0] for x in socket.getaddrinfo(host,443)}))
except Exception as e: print(host, type(e).__name__, e)
PY
# If available, inspect certificate metadata rather than using -k/--insecure:
# openssl s_client -connect registry.gitlab.com:443 -servername registry.gitlab.com </dev/null
Unsafe shortcut: curl -k,
GIT_SSL_NO_VERIFY=true, or disabling registry TLS
turns a trust failure into an interception risk. Repair CA
distribution/proxy/hostname configuration instead.
5. Failure mode: debug logging leaks secrets
A job fails before a deployment API call, so someone sets
CI_DEBUG_TRACE=true. The trace now contains exported
variables. GitLab warns that debug trace exposes all
variables/secrets available to the job. The incident has become both
a pipeline failure and a possible credential-exposure event.
Repair sequence: restrict log visibility, preserve evidence according to your incident process, identify exposed credentials, rotate/revoke them, replace broad trace with narrow redacted diagnostics, and then erase sensitive logs only under authorized retention policy. Erasing first can destroy the timeline.
6. Failure mode: cleanup destroys evidence
Deleting an artifact, erasing a job, or deleting a pipeline can be irreversible. If the artifact/report is the only proof of what failed, broad cleanup sabotages root-cause analysis. Hash/export the necessary evidence first, record access controls, then remove only what retention/security policy requires.
| Evidence | Preserve | Destructive action to delay |
|---|---|---|
| Job trace | Job ID, source SHA, safe copy/hash, timestamps | Erase job log |
| Test/security report | Artifact/report ID, type, digest, producer job | Delete artifacts |
| Merged config | Lint/full config, include refs/digests | Create new pipeline and forget original config |
| Runner evidence | Runner ID/version/executor/tags/log excerpt | Unregister/rebuild entire runner fleet |
| External target | Current version/digest/health/provider event | Delete/recreate environment/resource |
7. Failure mode: rollback declared successful but never verified
A rollback job returning exit code zero proves only that its script believed it completed. Verify the external system: deployed digest/version, health/readiness, traffic destination, schema compatibility, and any stateful side effects. If the rollback rebuilt an artifact from source rather than using the previously verified release artifact, you also changed artifact identity and must treat it as a new release experiment.
8. Intentionally broken example: classify before repair
This local record represents a real-looking incident: the job is pending, the requested runner tags cannot be satisfied, and someone incorrectly proposes enabling debug trace.
{
"pipeline_id": 3701,
"job_id": 3714,
"source":"push",
"sha":"abc123-demo",
"status":"pending",
"job_tags":["linux","docker","gpu"],
"available_runner":{"id":88,"tags":["linux","docker"],"executor":"docker"},
"first_failure_evidence":"job never entered preparing/running",
"proposed_fix":"CI_DEBUG_TRACE=true"
}
Interpretation: debug trace cannot help because the
script never started. The missing gpu match is
sufficient scheduling evidence. Preserve the original job, repair
the routing requirement/capacity, then allow the existing job to be
retried only if no source/configuration change is required. If you
change YAML tags, create a new pipeline for the new configuration.
9. Troubleshooting performance without causing a second outage
Slow pipelines can be queue, long-polling, cache/artifact I/O, image
pull, network, or critical-path problems. Runner documentation now
surfaces long-polling bottlenecks involving concurrent,
limit and request_concurrency. Change one
capacity/routing parameter at a time, compare queue/execution
metrics on equivalent workloads, and retain rollback values. Do not
respond with unbounded autoscaling.
10. Minimal recovery playbook template
Write the playbook so another engineer can use it under pressure.
INCIDENT: <short symptom>
1. Freeze destructive actions; record time/owner.
2. Capture project/ref/SHA, pipeline ID/source, policy/config revision.
3. Capture CI Lint/full config + jobs/rules evidence.
4. Capture job ID/status/failure_reason and first trace/report.
5. Capture runner ID/version/tags/executor/image/tool versions.
6. Capture DNS/TLS/API/artifact/cache/deployment/external target evidence.
7. State causal layer + rejected alternative hypotheses.
8. State proposed minimal change + side effects + rollback.
9. Retry one safe job OR create a new pipeline if source/config changed.
10. Independently verify final GitLab + external target state.
11. Preserve before/after packet; create permanent remediation owner/date.
Knowledge check
What evidence should you preserve before retrying a failed GitLab job?
Preserve pipeline/job IDs, source ref and SHA, compiled configuration/rule outcome, first-failure trace, runner/executor/image identity, and relevant artifact/deployment/external evidence.
A job remains pending with no trace. Should you start by debugging the application script?
No. A pending job has not executed the script. Inspect job eligibility, requested tags, runner scope/protection/status, manager connectivity, and executor capacity first.
Why is blind retry a poor default for a flaky failure?
It can erase the timing and first-failure evidence needed to distinguish a real flake from runner, network, configuration, or external-system failure.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Further reading — current official GitLab sources
Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.
- GitLab Docs — Validate GitLab CI/CD configuration
- GitLab Docs — CI Lint API
- GitLab Docs — Pipeline editor / full configuration
- GitLab Docs — CI/CD YAML syntax and retry
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — CI/CD jobs and statuses
- GitLab Docs — Troubleshooting CI/CD variables / debug trace
- GitLab Docs — Services and CI_DEBUG_SERVICES
- GitLab Docs — GitLab Runner commands
- GitLab Docs — Troubleshooting GitLab Runner
- GitLab Docs — Runner advanced configuration
- GitLab Docs — Docker executor image pull errors
- GitLab Docs — Docker build troubleshooting (DNS/TLS)
- GitLab Docs — Job artifacts and destructive deletion
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.