Chapter 34Lesson 01~215 minutes

CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Concepts, Architecture, and Mental Model

Observe GitLab pipelines as distributed delivery systems by correlating pipeline and job identity, queue time, runner metrics, logs, artifacts, deployment telemetry, retention, and causal failure classes.

ObservabilityAnalyticsQueue timeRunner metricsEvidence

Learning objectives

  • Explain why queue time, execution time, pipeline wall time and compute consumption answer different questions.
  • Correlate pipeline/job IDs, source SHA, runner identity, traces, artifacts and external telemetry without collapsing them into one success state.
  • Interpret GitLab CI/CD analytics and Runner Prometheus metrics while respecting tier and security boundaries.
  • Design a causal failure taxonomy that classifies system layers rather than people.
  • Preserve enough evidence to diagnose the first failure before rerunning or erasing anything.

1. The practical problem: a green or slow pipeline does not explain itself

After Chapter 33, an external system can create a pipeline and reconcile its final status. The next operational question is harder: why did delivery take twelve minutes, why did a job fail, and which layer owns the correction? A pipeline can be slow because no runner was free, because a runner started slowly, because a test took a long time, because an artifact upload saturated the network, or because a deployment waited on an external service. Looking only at final pipeline duration hides those causes.

Observability means preserving and correlating enough evidence to distinguish those states. In GitLab CI/CD, pipeline records, job timestamps, Runner metrics, traces, artifacts/reports and external deployment telemetry are related but separate evidence streams.

Operating rule: preserve identity before interpretation. Record pipeline ID, job ID, CI_PIPELINE_SOURCE, exact source SHA and runner/executor identity before computing a metric or proposing a fix.

2. Mental model: event → queue → runner → trace → evidence → telemetry → improvement

A source event creates a pipeline record. The compiled configuration determines the job graph. Each runnable job enters a queue until an eligible runner accepts it. Execution produces a trace and possibly reports/artifacts. Deployment jobs can create GitLab environment/deployment records, while the target has its own health telemetry. Analytics aggregates selected facts; an SLO or failure taxonomy converts those facts into an operating decision.

Each arrow is causal: an overloaded fleet increases queue time before execution; a slow script increases running duration; an unavailable artifact backend affects transfer; an unhealthy target can fail after the CI script succeeds. If these states are merged into one “pipeline time,” optimization becomes guesswork.

Delivery observability chain
            flowchart TD
              A[Pipeline event + SHA] --> B[Compiled graph]
              B --> C[Queued job]
              C --> D[Runner / executor]
              D --> E[Timestamped trace]
              E --> F[Artifacts + reports]
              F --> G[Deployment / external telemetry]
              C -. queue metrics .-> H[Analytics + taxonomy]
              D -. runner metrics .-> H
              E -. failure evidence .-> H
              G -. health evidence .-> H
              H --> I[Bounded improvement]
          

3. Define state before changing it

State layer Evidence to preserve Question it answers
Source/revision CI_PIPELINE_SOURCE, ref, exact CI_COMMIT_SHA Which code and event created this pipeline?
Compiled configuration merged YAML, workflow/job rule outcome, graph Which jobs existed before any runner touched them?
Pipeline/job record pipeline ID, job ID, status and timestamps What object are we measuring?
Runner/executor runner identity, manager version, executor/image Was waiting caused by capacity or execution?
Trace/log timestamped trace, sections, first error What did the process actually emit, and when?
Artifacts/reports names, sizes, digests, expiry and ingestion state What structured evidence survived the job?
Deployment/external environment/deployment ID, rollout/health telemetry Did the outside system reach the intended state?
Governance/retention access, masking, trace/artifact retention, erase actions Who can inspect evidence and for how long?

4. Start with read-only inspection

Before rerunning, canceling, erasing, scaling or changing YAML, collect the current record. The Jobs API exposes useful timing fields including created_at, started_at, finished_at, duration, queued_duration, failure_reason, pipeline identity and runner metadata. The trace endpoint returns the log for an exact job ID.

# Read-only examples; use only an authorized disposable project.
curl --silent --header "PRIVATE-TOKEN: $LAB_TOKEN"   "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs?scope[]=failed&per_page=20"

curl --silent --location --header "PRIVATE-TOKEN: $LAB_TOKEN"   "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs/$LAB_JOB_ID/trace"   --output "evidence/job-$LAB_JOB_ID.log"

Do not select “latest failed job” when you already know the exact ID. Reruns create new job IDs and can obscure which attempt contained the original symptom.

5. Queue time, execution time and wall-clock latency are different

Queued duration measures waiting before a job executes. Job duration measures running time. Pipeline wall-clock duration spans the pipeline graph and is not the sum of every job because jobs can run concurrently. Compute usage can sum runner execution across jobs and therefore exceed end-to-end wall time.

Observation Likely layer Do not conclude yet
Queue 420 s, run 25 s runner eligibility/capacity/request flow “tests are slow”
Queue 3 s, run 420 s script/tool/network inside executor “buy more runners”
Jobs each 120 s but pipeline 125 s parallel graph functioning “pipeline used only 125 s of compute”
CI job succeeds; target health later degrades deployment/external system “green job proves healthy service”

A useful SLI set starts small: median/p95 queue duration, median/p95 job duration for a stable job family, pipeline success rate, first-failure class, and post-deployment target health.

6. Runner metrics explain capacity behavior

GitLab Runner exposes native Prometheus metrics when its embedded metrics server is enabled. Signals include running-job counts, API request metrics, process statistics, build-version information and executor/autoscaler metrics. These are operator-side signals: they describe Runner infrastructure, not repository semantics.

# Safe local-only example in config.toml
listen_address = "127.0.0.1:9252"

# Read-only inspection on the runner host
curl --silent http://127.0.0.1:9252/metrics | grep -E '^gitlab_runner_|^process_' | head -n 40

Security boundary: the embedded metrics endpoint has no built-in authorization. Keep it on loopback/private monitoring networks or place an authenticated proxy/firewall boundary in front of it.

7. Logs are event evidence, not a metric store

Current GitLab job logs can include an ISO 8601 timestamp on each line. Timestamps are valuable for step-level sequencing, but logs remain human-oriented evidence. A line saying “download took 15s” is not a stable time-series metric unless format, units and collection are controlled.

Runner and GitLab also impose separate trace limits. Runner output_limit defaults to 4096 KB; reaching it truncates further Runner output while execution can continue. GitLab’s server-side job trace size defaults to 100 MB; exceeding that server limit can fail/drop the job. A missing tail can therefore be an observability failure rather than proof that later steps never ran.

Never enable verbose debug tracing as routine observability. CI_DEBUG_TRACE can expose all variables and secrets available to a job. Service debug logging can also interfere with masking.

8. Artifacts and reports preserve structured evidence

Logs tell the story of execution; artifacts/reports preserve machine-readable outputs. Useful bounded observability artifacts include timing JSON, test summaries, failure-classification records, digests and a small external-health snapshot. Give them explicit expiry and restricted access when needed.

observe:
  script:
    - python tools/write_observation.py --out evidence/observation.json
  artifacts:
    when: always
    expire_in: 7 days
    paths:
      - evidence/observation.json

when: always can preserve bounded evidence from a failed job; it does not change the script exit code or make a failed job successful.

9. Deployment records and external telemetry close the loop

A deployment job can succeed after submitting a change while the actual service is unhealthy. Preserve GitLab environment/deployment identity where applicable, then query the target through its own safe health signal: rollout status, an HTTP health endpoint, a synthetic check or provider operation ID.

Store the deployment/environment identity beside pipeline/job ID and source SHA. If the target uses its own trace/operation ID, record that too. Correlation does not require centralizing every log; it requires stable identifiers across systems.

10. Built-in analytics versus detailed records

Project CI/CD analytics is available across current GitLab tiers and summarizes pipeline success/failure and duration trends. It is useful for orientation but does not replace job-level evidence. The newer GLQL Pipeline Analytics data source adds aggregated query controls and is Premium/Ultimate. A Free-compatible operating model can still compute a worksheet from Jobs/Pipelines API exports or local evidence.

Use aggregates to find where to investigate, then exact pipeline/job IDs to prove a cause. “p95 pipeline duration rose” is a trend; “p95 queue duration rose while execution stayed flat for one runner pool” is a causal clue.

11. Failure taxonomy: classify the layer before assigning ownership

Class Typical evidence Primary correction layer
Configuration/compilation lint/compile error, missing job, unexpected rule result repository CI configuration
Queue/capacity high queued_duration, no eligible runner, saturation runner/fleet capacity
Runner/executor prepare failure, image pull/executor setup error runner platform
Script/tool/network non-zero command, test/compiler/network trace job/tool/service dependency
Identity/authorization 401/403, protected-data denial, scope mismatch trust/permissions boundary
Artifact/report/cache upload/download/parse/retention failure evidence/data-flow layer
Deployment/provider provider timeout, rollout failure, unhealthy target deployment/external system
API/policy/governance rate limit, approval/policy denial control-plane/governance layer

This taxonomy intentionally avoids “developer error.” People may introduce defects, but operations needs the system state that failed and evidence that proves it.

12. Retention, privacy and erasure are observability choices

More telemetry is not automatically better. Logs can contain source paths, user names, internal URLs and accidental secrets. Artifacts consume storage. High-cardinality metric labels can make observability expensive or unusable. Retain enough for incident analysis/trends, then expire or erase under policy.

On GitLab.com, CI/CD job logs are retained by default and do not have configurable automatic expiry; they can be erased through the Jobs API or by deleting the containing pipeline. Self-Managed operators likewise need explicit storage-management decisions. Erasing a job is destructive because it removes both log and artifacts. Preserve authorized evidence first.

13. Compute usage is a cost signal, not a latency signal

For instance runners, compute usage is based on job running duration and a cost factor; created/pending time is not counted as running compute. Parallel jobs can consume more total compute minutes even when they reduce wall-clock latency. A cost dashboard and a latency dashboard answer different questions.

14. Common misconceptions

Misconception Why it fails Safer model
“Job duration is pipeline latency.” It ignores queue and graph dependencies. Separate queue, run and pipeline wall-clock time.
“Rerun first; evidence will still be there.” New attempts get new IDs and may alter the symptom. Capture exact first-failure IDs/logs/artifacts first.
“More debug logs always improve diagnosis.” Verbose modes can expose secrets and add noise. Use bounded, non-secret evidence.
“Runner metrics explain test failures.” They describe infrastructure, not test semantics. Correlate infrastructure and job evidence.
“A failure category should identify who made the mistake.” Blame labels do not identify failing state. Classify causal layer and observable evidence.

15. Mini lab: identify the hidden bottleneck

Consider three jobs:

Job Queued Running Result
unit 8 s 95 s success
integration 470 s 105 s success
deploy 6 s 20 s CI success; later health failure

Before changing test code, identify the integration job as queue-dominated. Before declaring deployment healthy, retain external-health failure separately. Capacity diagnosis should inspect runner eligibility/tags/concurrency; deployment diagnosis should correlate exact deployment/pipeline/job IDs with target health.

Knowledge check

A job shows queued_duration=360 and duration=20. Which layer first?

Why can total compute exceed pipeline wall time?

What does hitting Runner output_limit prove?

Why is CI_DEBUG_TRACE unsafe as routine observability?

What should a failure taxonomy classify first?

16. Summary

Pipeline observability is correlation across state boundaries. Preserve source and pipeline identity, separate queue from execution, correlate Runner metrics with exact job records, use traces for sequence evidence, retain bounded artifacts/reports, and verify deployment health independently. Aggregates reveal trends; exact IDs prove causes. Retention, masking and erasure are part of the design.

Next lesson

Guided timing, Runner-metric and failure-taxonomy workflow

Lesson 2 turns the model into a disposable Jobs-API-style dataset, local Prometheus snapshot and first-failure exercise.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs and compute-usage concepts are available across Free/Premium/Ultimate unless a narrower capability is explicitly identified. The newer GLQL Pipeline Analytics data source is Premium/Ultimate. Runner metrics are exposed without built-in authorization when enabled, so labs bind or simulate them locally rather than publishing the endpoint. Current job logs support line timestamps; Runner 18.7+ is required to control them with FF_TIMESTAMPS. Runner output_limit defaults to 4096 KB, while GitLab server-side job trace size defaults to 100 MB; those are distinct limits with different truncation/failure behavior. Debug trace and service debug logging are security-sensitive because secret material can appear in logs. The mandatory labs use only synthetic local evidence and Python standard-library tooling; no live token, paid analytics feature, cloud account or public Runner endpoint is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.