CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Concepts, Architecture, and Mental Model
Observe GitLab pipelines as distributed delivery systems by correlating pipeline and job identity, queue time, runner metrics, logs, artifacts, deployment telemetry, retention, and causal failure classes.
Learning objectives
- Explain why queue time, execution time, pipeline wall time and compute consumption answer different questions.
- Correlate pipeline/job IDs, source SHA, runner identity, traces, artifacts and external telemetry without collapsing them into one success state.
- Interpret GitLab CI/CD analytics and Runner Prometheus metrics while respecting tier and security boundaries.
- Design a causal failure taxonomy that classifies system layers rather than people.
- Preserve enough evidence to diagnose the first failure before rerunning or erasing anything.
1. The practical problem: a green or slow pipeline does not explain itself
After Chapter 33, an external system can create a pipeline and reconcile its final status. The next operational question is harder: why did delivery take twelve minutes, why did a job fail, and which layer owns the correction? A pipeline can be slow because no runner was free, because a runner started slowly, because a test took a long time, because an artifact upload saturated the network, or because a deployment waited on an external service. Looking only at final pipeline duration hides those causes.
Observability means preserving and correlating enough evidence to distinguish those states. In GitLab CI/CD, pipeline records, job timestamps, Runner metrics, traces, artifacts/reports and external deployment telemetry are related but separate evidence streams.
Operating rule: preserve identity before
interpretation. Record pipeline ID, job ID,
CI_PIPELINE_SOURCE, exact source SHA and
runner/executor identity before computing a metric or proposing a
fix.
2. Mental model: event → queue → runner → trace → evidence → telemetry → improvement
A source event creates a pipeline record. The compiled configuration determines the job graph. Each runnable job enters a queue until an eligible runner accepts it. Execution produces a trace and possibly reports/artifacts. Deployment jobs can create GitLab environment/deployment records, while the target has its own health telemetry. Analytics aggregates selected facts; an SLO or failure taxonomy converts those facts into an operating decision.
Each arrow is causal: an overloaded fleet increases queue time before execution; a slow script increases running duration; an unavailable artifact backend affects transfer; an unhealthy target can fail after the CI script succeeds. If these states are merged into one “pipeline time,” optimization becomes guesswork.
flowchart TD
A[Pipeline event + SHA] --> B[Compiled graph]
B --> C[Queued job]
C --> D[Runner / executor]
D --> E[Timestamped trace]
E --> F[Artifacts + reports]
F --> G[Deployment / external telemetry]
C -. queue metrics .-> H[Analytics + taxonomy]
D -. runner metrics .-> H
E -. failure evidence .-> H
G -. health evidence .-> H
H --> I[Bounded improvement]
3. Define state before changing it
| State layer | Evidence to preserve | Question it answers |
|---|---|---|
| Source/revision |
CI_PIPELINE_SOURCE, ref, exact
CI_COMMIT_SHA
|
Which code and event created this pipeline? |
| Compiled configuration | merged YAML, workflow/job rule outcome, graph | Which jobs existed before any runner touched them? |
| Pipeline/job record | pipeline ID, job ID, status and timestamps | What object are we measuring? |
| Runner/executor | runner identity, manager version, executor/image | Was waiting caused by capacity or execution? |
| Trace/log | timestamped trace, sections, first error | What did the process actually emit, and when? |
| Artifacts/reports | names, sizes, digests, expiry and ingestion state | What structured evidence survived the job? |
| Deployment/external | environment/deployment ID, rollout/health telemetry | Did the outside system reach the intended state? |
| Governance/retention | access, masking, trace/artifact retention, erase actions | Who can inspect evidence and for how long? |
4. Start with read-only inspection
Before rerunning, canceling, erasing, scaling or changing YAML,
collect the current record. The Jobs API exposes useful timing
fields including created_at, started_at,
finished_at, duration,
queued_duration, failure_reason, pipeline
identity and runner metadata. The trace endpoint returns the log for
an exact job ID.
# Read-only examples; use only an authorized disposable project.
curl --silent --header "PRIVATE-TOKEN: $LAB_TOKEN" "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs?scope[]=failed&per_page=20"
curl --silent --location --header "PRIVATE-TOKEN: $LAB_TOKEN" "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs/$LAB_JOB_ID/trace" --output "evidence/job-$LAB_JOB_ID.log"
Do not select “latest failed job” when you already know the exact ID. Reruns create new job IDs and can obscure which attempt contained the original symptom.
5. Queue time, execution time and wall-clock latency are different
Queued duration measures waiting before a job executes. Job duration measures running time. Pipeline wall-clock duration spans the pipeline graph and is not the sum of every job because jobs can run concurrently. Compute usage can sum runner execution across jobs and therefore exceed end-to-end wall time.
| Observation | Likely layer | Do not conclude yet |
|---|---|---|
| Queue 420 s, run 25 s | runner eligibility/capacity/request flow | “tests are slow” |
| Queue 3 s, run 420 s | script/tool/network inside executor | “buy more runners” |
| Jobs each 120 s but pipeline 125 s | parallel graph functioning | “pipeline used only 125 s of compute” |
| CI job succeeds; target health later degrades | deployment/external system | “green job proves healthy service” |
A useful SLI set starts small: median/p95 queue duration, median/p95 job duration for a stable job family, pipeline success rate, first-failure class, and post-deployment target health.
6. Runner metrics explain capacity behavior
GitLab Runner exposes native Prometheus metrics when its embedded metrics server is enabled. Signals include running-job counts, API request metrics, process statistics, build-version information and executor/autoscaler metrics. These are operator-side signals: they describe Runner infrastructure, not repository semantics.
# Safe local-only example in config.toml
listen_address = "127.0.0.1:9252"
# Read-only inspection on the runner host
curl --silent http://127.0.0.1:9252/metrics | grep -E '^gitlab_runner_|^process_' | head -n 40
Security boundary: the embedded metrics endpoint has no built-in authorization. Keep it on loopback/private monitoring networks or place an authenticated proxy/firewall boundary in front of it.
7. Logs are event evidence, not a metric store
Current GitLab job logs can include an ISO 8601 timestamp on each line. Timestamps are valuable for step-level sequencing, but logs remain human-oriented evidence. A line saying “download took 15s” is not a stable time-series metric unless format, units and collection are controlled.
Runner and GitLab also impose separate trace limits. Runner
output_limit defaults to 4096 KB; reaching it truncates
further Runner output while execution can continue. GitLab’s
server-side job trace size defaults to 100 MB; exceeding that server
limit can fail/drop the job. A missing tail can therefore be an
observability failure rather than proof that later steps never ran.
Never enable verbose debug tracing as routine
observability.
CI_DEBUG_TRACE can expose all variables and secrets
available to a job. Service debug logging can also interfere with
masking.
8. Artifacts and reports preserve structured evidence
Logs tell the story of execution; artifacts/reports preserve machine-readable outputs. Useful bounded observability artifacts include timing JSON, test summaries, failure-classification records, digests and a small external-health snapshot. Give them explicit expiry and restricted access when needed.
observe:
script:
- python tools/write_observation.py --out evidence/observation.json
artifacts:
when: always
expire_in: 7 days
paths:
- evidence/observation.json
when: always can preserve bounded evidence from a
failed job; it does not change the script exit code or make a failed
job successful.
9. Deployment records and external telemetry close the loop
A deployment job can succeed after submitting a change while the actual service is unhealthy. Preserve GitLab environment/deployment identity where applicable, then query the target through its own safe health signal: rollout status, an HTTP health endpoint, a synthetic check or provider operation ID.
Store the deployment/environment identity beside pipeline/job ID and source SHA. If the target uses its own trace/operation ID, record that too. Correlation does not require centralizing every log; it requires stable identifiers across systems.
10. Built-in analytics versus detailed records
Project CI/CD analytics is available across current GitLab tiers and summarizes pipeline success/failure and duration trends. It is useful for orientation but does not replace job-level evidence. The newer GLQL Pipeline Analytics data source adds aggregated query controls and is Premium/Ultimate. A Free-compatible operating model can still compute a worksheet from Jobs/Pipelines API exports or local evidence.
Use aggregates to find where to investigate, then exact pipeline/job IDs to prove a cause. “p95 pipeline duration rose” is a trend; “p95 queue duration rose while execution stayed flat for one runner pool” is a causal clue.
11. Failure taxonomy: classify the layer before assigning ownership
| Class | Typical evidence | Primary correction layer |
|---|---|---|
| Configuration/compilation | lint/compile error, missing job, unexpected rule result | repository CI configuration |
| Queue/capacity |
high queued_duration, no eligible runner,
saturation
|
runner/fleet capacity |
| Runner/executor | prepare failure, image pull/executor setup error | runner platform |
| Script/tool/network | non-zero command, test/compiler/network trace | job/tool/service dependency |
| Identity/authorization | 401/403, protected-data denial, scope mismatch | trust/permissions boundary |
| Artifact/report/cache | upload/download/parse/retention failure | evidence/data-flow layer |
| Deployment/provider | provider timeout, rollout failure, unhealthy target | deployment/external system |
| API/policy/governance | rate limit, approval/policy denial | control-plane/governance layer |
This taxonomy intentionally avoids “developer error.” People may introduce defects, but operations needs the system state that failed and evidence that proves it.
12. Retention, privacy and erasure are observability choices
More telemetry is not automatically better. Logs can contain source paths, user names, internal URLs and accidental secrets. Artifacts consume storage. High-cardinality metric labels can make observability expensive or unusable. Retain enough for incident analysis/trends, then expire or erase under policy.
On GitLab.com, CI/CD job logs are retained by default and do not have configurable automatic expiry; they can be erased through the Jobs API or by deleting the containing pipeline. Self-Managed operators likewise need explicit storage-management decisions. Erasing a job is destructive because it removes both log and artifacts. Preserve authorized evidence first.
13. Compute usage is a cost signal, not a latency signal
For instance runners, compute usage is based on job running duration and a cost factor; created/pending time is not counted as running compute. Parallel jobs can consume more total compute minutes even when they reduce wall-clock latency. A cost dashboard and a latency dashboard answer different questions.
14. Common misconceptions
| Misconception | Why it fails | Safer model |
|---|---|---|
| “Job duration is pipeline latency.” | It ignores queue and graph dependencies. | Separate queue, run and pipeline wall-clock time. |
| “Rerun first; evidence will still be there.” | New attempts get new IDs and may alter the symptom. | Capture exact first-failure IDs/logs/artifacts first. |
| “More debug logs always improve diagnosis.” | Verbose modes can expose secrets and add noise. | Use bounded, non-secret evidence. |
| “Runner metrics explain test failures.” | They describe infrastructure, not test semantics. | Correlate infrastructure and job evidence. |
| “A failure category should identify who made the mistake.” | Blame labels do not identify failing state. | Classify causal layer and observable evidence. |
15. Mini lab: identify the hidden bottleneck
Consider three jobs:
| Job | Queued | Running | Result |
|---|---|---|---|
| unit | 8 s | 95 s | success |
| integration | 470 s | 105 s | success |
| deploy | 6 s | 20 s | CI success; later health failure |
Before changing test code, identify the integration job as queue-dominated. Before declaring deployment healthy, retain external-health failure separately. Capacity diagnosis should inspect runner eligibility/tags/concurrency; deployment diagnosis should correlate exact deployment/pipeline/job IDs with target health.
Knowledge check
A job shows queued_duration=360 and
duration=20. Which layer first?
Runner eligibility/capacity/request flow, because most latency occurred before execution.
Why can total compute exceed pipeline wall time?
Parallel jobs can consume runner time simultaneously; compute sums job execution while wall time measures elapsed graph time.
What does hitting Runner
output_limit prove?
That further Runner log output can be truncated at that configured threshold; it does not prove execution stopped.
Why is CI_DEBUG_TRACE unsafe as routine
observability?
It can expose job variables and secrets in trace output.
What should a failure taxonomy classify first?
The causal system layer and evidence, not a person or team to blame.
16. Summary
Pipeline observability is correlation across state boundaries. Preserve source and pipeline identity, separate queue from execution, correlate Runner metrics with exact job records, use traces for sequence evidence, retain bounded artifacts/reports, and verify deployment health independently. Aggregates reveal trends; exact IDs prove causes. Retention, masking and erasure are part of the design.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-13.
Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the
timestamped baseline where a concrete version matters. Project CI/CD
analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs
and compute-usage concepts are available across
Free/Premium/Ultimate unless a narrower capability is explicitly
identified. The newer GLQL Pipeline Analytics data source is
Premium/Ultimate. Runner metrics are exposed without built-in
authorization when enabled, so labs bind or simulate them locally
rather than publishing the endpoint. Current job logs support line
timestamps; Runner 18.7+ is required to control them with
FF_TIMESTAMPS. Runner
output_limit defaults to 4096 KB, while GitLab
server-side job trace size defaults to 100 MB; those are distinct
limits with different truncation/failure behavior. Debug trace and
service debug logging are security-sensitive because secret material
can appear in logs. The mandatory labs use only synthetic local
evidence and Python standard-library tooling; no live token, paid
analytics feature, cloud account or public Runner endpoint is
required.
- CI/CD analytics — official reference.
- Pipeline analytics (GLQL) — official reference.
- Jobs API — official reference.
- Pipelines API — official reference.
- Runner monitoring and Prometheus metrics — official reference.
- Runner advanced configuration — official reference.
- CI/CD job logs — official reference.
- Self-Managed job-log storage and limits — official reference.
- CI/CD limits — official reference.
- CI/CD variable security and masking — official reference.
- Troubleshooting CI/CD variables / debug trace — official reference.
- Service-container debug logs — official reference.
- Job artifacts — official reference.
- Job Artifacts API — official reference.
- Compute minutes — official reference.
- Compute usage for instance runners — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.