Chapter 37Lesson 04~145 minutes

Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Diagnostics, Failure Modes, Security, and Performance

Diagnose realistic GitLab CI/CD failure modes without blind retries, secret leakage, TLS bypass, evidence destruction, or unverified rollback.

Failure modesRunner diagnosticsDNS/TLSPlaybooks

Learning objectives

  • Apply the full evidence-first sequence from source/config through external target verification.
  • Diagnose flake, wrong runner tags, DNS/TLS failures, debug-log leakage, destructive cleanup, and rollback verification gaps.
  • Interpret an intentionally broken pending-job example and reject irrelevant fixes.
  • Use Runner diagnostics and network evidence without broadening privileges or disabling trust checks.
  • Write a minimal recovery playbook another engineer can follow under pressure.

Evidence-first rule: preserve the original pipeline/job IDs, source SHA, merged configuration, first failing trace/report, runner/executor identity, and external target state before any retry, erase, cleanup, or rollback. A repaired run is new evidence; it does not replace the original failure.

1. The full diagnostic sequence

  1. Preserve pipeline/job IDs and first-failure evidence.
  2. Confirm pipeline source/ref/SHA and policy/config revision.
  3. Validate/inspect merged configuration and workflow/job rule decisions.
  4. Inspect job graph and current status: created, pending, preparing, running, waiting, failed.
  5. Confirm runner scope/tags/protection, Runner version, executor, image/helper/toolchain.
  6. Inspect the first failing script/tool/network operation and non-secret effective inputs.
  7. Inspect reports, artifacts, caches, registry state and digests.
  8. Inspect environment/deployment and independently verify the external target.
  9. Apply the least destructive causal correction.
  10. Retry only the smallest safe scope, or create a new pipeline when source/config changed.
  11. Verify recovery and retain both before/after evidence.

2. Failure mode: blind retry hides a flake

Suppose attempt 1 fails an integration test and attempt 2 passes without any source/configuration change. The correct conclusion is nondeterminism exists, not “the pipeline is fixed.” Preserve attempt 1’s job ID, test report and timing; compare dependencies and environment; then decide whether a narrow transient retry policy is justified.

# Bad: hides every kind of script failure.
test:
  script: ./integration-tests.sh
  retry: 2

# Better only after evidence shows a specific transient class.
test:
  script: ./integration-tests.sh
  retry:
    max: 1
    when: runner_external_dependency_failure

3. Failure mode: wrong runner tags

A job requests [linux, docker, gpu]. The only authorized runner has [linux, docker]. The job remains pending. Restarting the runner cannot manufacture the missing tag/capability. First prove the required tags and runner inventory. Then either provision/restore the correct trusted runner or remove the unjustified job requirement through review.

# Authorized runner-manager host, read-only first:
gitlab-runner --version
gitlab-runner list
gitlab-runner verify
# Remember: verify tests GitLab connectivity; it does not prove service usage.
gitlab-runner status

4. Failure mode: DNS/TLS error misdiagnosed as application failure

Image pulls, package downloads, Git fetches, secret providers, and deployment APIs all depend on name resolution and TLS trust. An x509 failure or no such host occurs before application tests. Record the hostname, resolver, proxy path, certificate issuer/chain and which container/process makes the request.

# Examples for an authorized disposable environment; do not disable TLS verification.
python - <<'PY'
import socket
for host in ['gitlab.com','registry.gitlab.com']:
    try: print(host, sorted({x[4][0] for x in socket.getaddrinfo(host,443)}))
    except Exception as e: print(host, type(e).__name__, e)
PY

# If available, inspect certificate metadata rather than using -k/--insecure:
# openssl s_client -connect registry.gitlab.com:443 -servername registry.gitlab.com </dev/null

Unsafe shortcut: curl -k, GIT_SSL_NO_VERIFY=true, or disabling registry TLS turns a trust failure into an interception risk. Repair CA distribution/proxy/hostname configuration instead.

5. Failure mode: debug logging leaks secrets

A job fails before a deployment API call, so someone sets CI_DEBUG_TRACE=true. The trace now contains exported variables. GitLab warns that debug trace exposes all variables/secrets available to the job. The incident has become both a pipeline failure and a possible credential-exposure event.

Repair sequence: restrict log visibility, preserve evidence according to your incident process, identify exposed credentials, rotate/revoke them, replace broad trace with narrow redacted diagnostics, and then erase sensitive logs only under authorized retention policy. Erasing first can destroy the timeline.

6. Failure mode: cleanup destroys evidence

Deleting an artifact, erasing a job, or deleting a pipeline can be irreversible. If the artifact/report is the only proof of what failed, broad cleanup sabotages root-cause analysis. Hash/export the necessary evidence first, record access controls, then remove only what retention/security policy requires.

Evidence Preserve Destructive action to delay
Job trace Job ID, source SHA, safe copy/hash, timestamps Erase job log
Test/security report Artifact/report ID, type, digest, producer job Delete artifacts
Merged config Lint/full config, include refs/digests Create new pipeline and forget original config
Runner evidence Runner ID/version/executor/tags/log excerpt Unregister/rebuild entire runner fleet
External target Current version/digest/health/provider event Delete/recreate environment/resource

7. Failure mode: rollback declared successful but never verified

A rollback job returning exit code zero proves only that its script believed it completed. Verify the external system: deployed digest/version, health/readiness, traffic destination, schema compatibility, and any stateful side effects. If the rollback rebuilt an artifact from source rather than using the previously verified release artifact, you also changed artifact identity and must treat it as a new release experiment.

8. Intentionally broken example: classify before repair

This local record represents a real-looking incident: the job is pending, the requested runner tags cannot be satisfied, and someone incorrectly proposes enabling debug trace.

{
  "pipeline_id": 3701,
  "job_id": 3714,
  "source":"push",
  "sha":"abc123-demo",
  "status":"pending",
  "job_tags":["linux","docker","gpu"],
  "available_runner":{"id":88,"tags":["linux","docker"],"executor":"docker"},
  "first_failure_evidence":"job never entered preparing/running",
  "proposed_fix":"CI_DEBUG_TRACE=true"
}

Interpretation: debug trace cannot help because the script never started. The missing gpu match is sufficient scheduling evidence. Preserve the original job, repair the routing requirement/capacity, then allow the existing job to be retried only if no source/configuration change is required. If you change YAML tags, create a new pipeline for the new configuration.

9. Troubleshooting performance without causing a second outage

Slow pipelines can be queue, long-polling, cache/artifact I/O, image pull, network, or critical-path problems. Runner documentation now surfaces long-polling bottlenecks involving concurrent, limit and request_concurrency. Change one capacity/routing parameter at a time, compare queue/execution metrics on equivalent workloads, and retain rollback values. Do not respond with unbounded autoscaling.

10. Minimal recovery playbook template

Write the playbook so another engineer can use it under pressure.

INCIDENT: <short symptom>
1. Freeze destructive actions; record time/owner.
2. Capture project/ref/SHA, pipeline ID/source, policy/config revision.
3. Capture CI Lint/full config + jobs/rules evidence.
4. Capture job ID/status/failure_reason and first trace/report.
5. Capture runner ID/version/tags/executor/image/tool versions.
6. Capture DNS/TLS/API/artifact/cache/deployment/external target evidence.
7. State causal layer + rejected alternative hypotheses.
8. State proposed minimal change + side effects + rollback.
9. Retry one safe job OR create a new pipeline if source/config changed.
10. Independently verify final GitLab + external target state.
11. Preserve before/after packet; create permanent remediation owner/date.
Next lesson

Next: Checkpoint Lab — Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks

Run a three-layer incident drill, recover with minimal changes, and produce a reusable evidence-backed troubleshooting playbook.

Knowledge check

What evidence should you preserve before retrying a failed GitLab job?

A job remains pending with no trace. Should you start by debugging the application script?

Why is blind retry a poor default for a flaky failure?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Further reading — current official GitLab sources

Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.