Chapter 34Lesson 04~240 minutes

CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Diagnostics, Failure Modes, Security, and Performance

Diagnose queue bottlenecks, lost first-failure context, unsafe debug traces, misleading failure labels, log truncation, and telemetry gaps using an evidence-first sequence.

DiagnosticsJob logsSecurityRoot causeFirst failure

Learning objectives

  • Apply the evidence-first diagnostic sequence without destroying first-failure context.
  • Diagnose queue-dominated latency, log truncation and telemetry gaps causally.
  • Recognize unsafe debug logging and secret exposure.
  • Separate system-layer taxonomy from human blame.
  • Retry only the smallest safe scope after preserving evidence.

1. Evidence-first diagnostic sequence

Record source identity first: preserve CI_PIPELINE_SOURCE, the ref, and CI_COMMIT_SHA alongside the pipeline/job IDs before rerunning or changing configuration. Those fields bind the failure evidence to the exact pipeline creation context and revision.

Use the same order regardless of symptom:

  1. Preserve pipeline/job IDs and first-failure trace/artifacts.
  2. Confirm pipeline source, ref, exact SHA and compiled configuration.
  3. Confirm workflow/job rules and effective non-secret inputs.
  4. Inspect graph and queue timing.
  5. Confirm runner/executor/image/tool versions.
  6. Inspect the failing script/tool/network interval.
  7. Inspect reports/artifacts/cache/registry evidence.
  8. Inspect deployment/provider/external target state.
  9. Apply the least destructive correction.
  10. Retry only the smallest safe scope and compare new IDs/metrics with the preserved original.

2. Failure mode: job duration looks healthy while queue dominates

Suppose integration runs in 90 seconds so a dashboard calls it healthy. Jobs API shows queued_duration=520 seconds. The developer still waits more than ten minutes. The causal layer is runner eligibility/capacity/request flow, not test runtime.

python - <<'PY'
import json
for j in json.load(open("jobs.json")):
    print(j["id"],j["name"],"queue",j.get("queued_duration"),"run",j.get("duration"))
PY

Preserve IDs, inspect tags/runner state/fleet metrics, then change capacity/settings only if evidence proves that layer.

3. Failure mode: rerun hides first-failure context

A retry is a new observation. External services can recover, runners can differ, mutable images can move, and cache state can change.

Before retry: capture exact job trace, failure reason, runner identity, source SHA, compiled-config/rule context, artifact/report metadata and external operation ID. A successful rerun does not retroactively explain the first failure.

4. Failure mode: debug logging leaks sensitive values

CI_DEBUG_TRACE is a troubleshooting feature, not a telemetry collector. Current GitLab docs warn that it exposes variables/secrets in verbose execution details. CI_DEBUG_SERVICES can also interleave service output in ways that defeat masking.

If a secret appears, treat it as compromised: restrict evidence access, rotate/revoke it, preserve an incident-safe record, and erase sensitive trace only after authorized preservation. Do not assume masking fixes an already exposed credential.

5. Failure mode: taxonomy blames people instead of layers

“Developer failure” does not tell an operator what to inspect. Use configuration, queue/capacity, runner/executor, script/tool/network, identity/authorization, artifact/report/cache, deployment/provider, or API/policy. Ownership comes after causal classification.

6. Failure mode: missing log tail is mistaken for missing execution

Broken interpretation: “The log ends at 4 MB, therefore the process died.” Runner output_limit defaults to 4096 KB and can truncate output while execution continues. GitLab has a separate default 100 MB trace limit that can fail/drop a job.

Evidence Interpretation
Trace stops near Runner output limit; job later has terminal status possible Runner truncation; inspect status/artifacts
GitLab reports trace-size limit failure server-side trace limit affected job
Process exit/error before limit actual script/tool failure evidence

Reduce noisy output and place structured diagnostics in bounded artifacts instead of simply raising every limit.

7. Intentionally broken diagnostic example

Do not run this pattern:

# INTENTIONALLY BROKEN — do not copy.
diagnose_everything:
  variables:
    CI_DEBUG_TRACE: "true"
    CI_DEBUG_SERVICES: "true"
  script:
    - echo "rerunning everything without preserving the failed job ID"
    - ./retry-all.sh

It broadens sensitive trace output, creates new state before preserving the original, and retries unbounded side effects. Safer: download exact failed trace/metadata first, add only narrowly scoped non-secret component logging if needed, then retry one idempotent job or operation.

8. API/rate-limit failures belong to their own layer

For 401/403, preserve HTTP status and request identity but never print the token. For 429, honor backoff and avoid making the collector a load amplifier. Observability that overloads the control plane creates a second incident.

9. CI success versus deployment/provider health

A deployment script can submit successfully while asynchronous rollout fails later. Preserve environment/deployment record and target operation/rollout/health evidence. “Job success” and “external target healthy” are different claims.

10. Least-destructive corrections

Cause Bad shortcut Bounded correction
Queue saturation rerun repeatedly adjust proven capacity/request bottleneck; compare queue p95
Noisy trace increase every limit reduce output; artifact structured diagnostics
Secret leaked leave log public because masking exists restrict, rotate/revoke, preserve incident evidence, authorized erase
External 503 blind retry whole pipeline bounded retry only idempotent affected call/job
Wrong taxonomy rename blame category classify failing state/evidence first

11. Observability overhead is real

Timestamped logs add storage overhead, service logs increase trace volume, high-cardinality metrics increase time-series cost, and large evidence artifacts lengthen transfers. Preserve high-value identifiers and bounded structured evidence rather than disabling evidence entirely.

12. Diagnostic exercise

Pipeline 990 failed; job 9912 waited 410 seconds, ran 12 seconds and got HTTP 403 uploading a report. Runner metrics show no CPU saturation. A rerun succeeds after someone grants a broad token.

Diagnosis: queue/capacity is a latency symptom; identity/authorization is the terminal failure. The broad token changes the trust model and is not acceptable. Preserve job 9912 evidence, determine the minimum required report-upload permission, restore least privilege, and rerun only that safe scope.

Knowledge check

Why is a successful rerun insufficient incident evidence?

What if a live secret appears in a trace?

A log ends near Runner output limit. Next check?

Why is “developer error” weak taxonomy?

A job waits 410 s, runs 12 s, then gets 403. Which two layers?

13. Summary

Capture the first failure, confirm source/config/rules, separate queue from execution, confirm runner/tool identity, inspect the exact failing interval, then follow reports and external state. Avoid secret-heavy debug tracing, blind reruns and blame categories. Correct only the proven layer.

Next lesson

Checkpoint observability worksheet

Lesson 5 correlates synthetic pipeline data, Runner metrics and first-failure evidence into a bounded corrective action.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs and compute-usage concepts are available across Free/Premium/Ultimate unless a narrower capability is explicitly identified. The newer GLQL Pipeline Analytics data source is Premium/Ultimate. Runner metrics are exposed without built-in authorization when enabled, so labs bind or simulate them locally rather than publishing the endpoint. Current job logs support line timestamps; Runner 18.7+ is required to control them with FF_TIMESTAMPS. Runner output_limit defaults to 4096 KB, while GitLab server-side job trace size defaults to 100 MB; those are distinct limits with different truncation/failure behavior. Debug trace and service debug logging are security-sensitive because secret material can appear in logs. The mandatory labs use only synthetic local evidence and Python standard-library tooling; no live token, paid analytics feature, cloud account or public Runner endpoint is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.