CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Diagnostics, Failure Modes, Security, and Performance
Diagnose queue bottlenecks, lost first-failure context, unsafe debug traces, misleading failure labels, log truncation, and telemetry gaps using an evidence-first sequence.
Learning objectives
- Apply the evidence-first diagnostic sequence without destroying first-failure context.
- Diagnose queue-dominated latency, log truncation and telemetry gaps causally.
- Recognize unsafe debug logging and secret exposure.
- Separate system-layer taxonomy from human blame.
- Retry only the smallest safe scope after preserving evidence.
1. Evidence-first diagnostic sequence
Record source identity first: preserve
CI_PIPELINE_SOURCE, the ref, and
CI_COMMIT_SHA alongside the pipeline/job IDs before
rerunning or changing configuration. Those fields bind the failure
evidence to the exact pipeline creation context and revision.
Use the same order regardless of symptom:
- Preserve pipeline/job IDs and first-failure trace/artifacts.
- Confirm pipeline source, ref, exact SHA and compiled configuration.
- Confirm workflow/job rules and effective non-secret inputs.
- Inspect graph and queue timing.
- Confirm runner/executor/image/tool versions.
- Inspect the failing script/tool/network interval.
- Inspect reports/artifacts/cache/registry evidence.
- Inspect deployment/provider/external target state.
- Apply the least destructive correction.
- Retry only the smallest safe scope and compare new IDs/metrics with the preserved original.
2. Failure mode: job duration looks healthy while queue dominates
Suppose integration runs in 90 seconds so a dashboard
calls it healthy. Jobs API shows
queued_duration=520 seconds. The developer still waits
more than ten minutes. The causal layer is runner
eligibility/capacity/request flow, not test runtime.
python - <<'PY'
import json
for j in json.load(open("jobs.json")):
print(j["id"],j["name"],"queue",j.get("queued_duration"),"run",j.get("duration"))
PY
Preserve IDs, inspect tags/runner state/fleet metrics, then change capacity/settings only if evidence proves that layer.
3. Failure mode: rerun hides first-failure context
A retry is a new observation. External services can recover, runners can differ, mutable images can move, and cache state can change.
Before retry: capture exact job trace, failure reason, runner identity, source SHA, compiled-config/rule context, artifact/report metadata and external operation ID. A successful rerun does not retroactively explain the first failure.
4. Failure mode: debug logging leaks sensitive values
CI_DEBUG_TRACE is a troubleshooting feature, not a
telemetry collector. Current GitLab docs warn that it exposes
variables/secrets in verbose execution details.
CI_DEBUG_SERVICES can also interleave service output in
ways that defeat masking.
If a secret appears, treat it as compromised: restrict evidence access, rotate/revoke it, preserve an incident-safe record, and erase sensitive trace only after authorized preservation. Do not assume masking fixes an already exposed credential.
5. Failure mode: taxonomy blames people instead of layers
“Developer failure” does not tell an operator what to inspect. Use configuration, queue/capacity, runner/executor, script/tool/network, identity/authorization, artifact/report/cache, deployment/provider, or API/policy. Ownership comes after causal classification.
6. Failure mode: missing log tail is mistaken for missing execution
Broken interpretation: “The log ends at 4 MB, therefore the process
died.” Runner output_limit defaults to 4096 KB and can
truncate output while execution continues. GitLab has a separate
default 100 MB trace limit that can fail/drop a job.
| Evidence | Interpretation |
|---|---|
| Trace stops near Runner output limit; job later has terminal status | possible Runner truncation; inspect status/artifacts |
| GitLab reports trace-size limit failure | server-side trace limit affected job |
| Process exit/error before limit | actual script/tool failure evidence |
Reduce noisy output and place structured diagnostics in bounded artifacts instead of simply raising every limit.
7. Intentionally broken diagnostic example
Do not run this pattern:
# INTENTIONALLY BROKEN — do not copy.
diagnose_everything:
variables:
CI_DEBUG_TRACE: "true"
CI_DEBUG_SERVICES: "true"
script:
- echo "rerunning everything without preserving the failed job ID"
- ./retry-all.sh
It broadens sensitive trace output, creates new state before preserving the original, and retries unbounded side effects. Safer: download exact failed trace/metadata first, add only narrowly scoped non-secret component logging if needed, then retry one idempotent job or operation.
8. API/rate-limit failures belong to their own layer
For 401/403, preserve HTTP status and request identity but never print the token. For 429, honor backoff and avoid making the collector a load amplifier. Observability that overloads the control plane creates a second incident.
9. CI success versus deployment/provider health
A deployment script can submit successfully while asynchronous rollout fails later. Preserve environment/deployment record and target operation/rollout/health evidence. “Job success” and “external target healthy” are different claims.
10. Least-destructive corrections
| Cause | Bad shortcut | Bounded correction |
|---|---|---|
| Queue saturation | rerun repeatedly | adjust proven capacity/request bottleneck; compare queue p95 |
| Noisy trace | increase every limit | reduce output; artifact structured diagnostics |
| Secret leaked | leave log public because masking exists | restrict, rotate/revoke, preserve incident evidence, authorized erase |
| External 503 | blind retry whole pipeline | bounded retry only idempotent affected call/job |
| Wrong taxonomy | rename blame category | classify failing state/evidence first |
11. Observability overhead is real
Timestamped logs add storage overhead, service logs increase trace volume, high-cardinality metrics increase time-series cost, and large evidence artifacts lengthen transfers. Preserve high-value identifiers and bounded structured evidence rather than disabling evidence entirely.
12. Diagnostic exercise
Pipeline 990 failed; job 9912 waited 410 seconds, ran 12 seconds and got HTTP 403 uploading a report. Runner metrics show no CPU saturation. A rerun succeeds after someone grants a broad token.
Diagnosis: queue/capacity is a latency symptom; identity/authorization is the terminal failure. The broad token changes the trust model and is not acceptable. Preserve job 9912 evidence, determine the minimum required report-upload permission, restore least privilege, and rerun only that safe scope.
Knowledge check
Why is a successful rerun insufficient incident evidence?
It is a new job with potentially different runner/cache/external state; the first failure still needs explanation.
What if a live secret appears in a trace?
Treat it as compromised: restrict access, rotate/revoke, preserve authorized incident evidence, then follow policy for erasure.
A log ends near Runner output limit. Next check?
Terminal job status, artifacts and Runner/server limit context; truncated trace alone does not prove execution stopped.
Why is “developer error” weak taxonomy?
It names a person/group rather than failing state, so it does not direct reproducible diagnosis.
A job waits 410 s, runs 12 s, then gets 403. Which two layers?
Queue/capacity for latency and identity/authorization for terminal failure.
13. Summary
Capture the first failure, confirm source/config/rules, separate queue from execution, confirm runner/tool identity, inspect the exact failing interval, then follow reports and external state. Avoid secret-heavy debug tracing, blind reruns and blame categories. Correct only the proven layer.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-13.
Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the
timestamped baseline where a concrete version matters. Project CI/CD
analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs
and compute-usage concepts are available across
Free/Premium/Ultimate unless a narrower capability is explicitly
identified. The newer GLQL Pipeline Analytics data source is
Premium/Ultimate. Runner metrics are exposed without built-in
authorization when enabled, so labs bind or simulate them locally
rather than publishing the endpoint. Current job logs support line
timestamps; Runner 18.7+ is required to control them with
FF_TIMESTAMPS. Runner
output_limit defaults to 4096 KB, while GitLab
server-side job trace size defaults to 100 MB; those are distinct
limits with different truncation/failure behavior. Debug trace and
service debug logging are security-sensitive because secret material
can appear in logs. The mandatory labs use only synthetic local
evidence and Python standard-library tooling; no live token, paid
analytics feature, cloud account or public Runner endpoint is
required.
- CI/CD analytics — official reference.
- Pipeline analytics (GLQL) — official reference.
- Jobs API — official reference.
- Pipelines API — official reference.
- Runner monitoring and Prometheus metrics — official reference.
- Runner advanced configuration — official reference.
- CI/CD job logs — official reference.
- Self-Managed job-log storage and limits — official reference.
- CI/CD limits — official reference.
- CI/CD variable security and masking — official reference.
- Troubleshooting CI/CD variables / debug trace — official reference.
- Service-container debug logs — official reference.
- Job artifacts — official reference.
- Job Artifacts API — official reference.
- Compute minutes — official reference.
- Compute usage for instance runners — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.