Chapter 34Lesson 02~255 minutes

CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Guided Hands-On Workflow and Core Operations

Build a disposable observability workflow that computes queue and execution time, interprets simulated Runner Prometheus metrics, classifies failures, and preserves evidence without production access.

Hands-onJobs APIPrometheusFailure taxonomyPython

Learning objectives

  • Create a disposable local dataset that mirrors timing and identity fields from GitLab Jobs/Pipelines APIs.
  • Compute queue and execution time without mixing their meanings.
  • Interpret a safe local Runner Prometheus snapshot and correlate it with job timing.
  • Classify several failures by causal layer and preserve first-failure evidence.
  • Build a small evidence artifact for before/after comparison.

1. Lab scenario and safety boundary

You operate a disposable project model named glci-ch34-observe-lab. Instead of requiring a live GitLab token or exposing a Runner metrics port, the mandatory path uses synthetic API exports and a local Prometheus text file. An optional mapping shows the read-only requests you would use on an authorized lab instance.

Preflight: use only a temporary directory. No credentials, runner registration, pipeline deletion, log erasure, deployment or network listener is required.

2. Create the disposable evidence set

mkdir -p ch34-lab/evidence
cd ch34-lab
cat > evidence/jobs.json <<'JSON'
[
  {"id":3401,"name":"unit","pipeline":{"id":340},"source":"push","sha":"1111111111111111111111111111111111111111","queued_duration":8.0,"duration":95.0,"status":"success","failure_reason":null,"runner":{"description":"linux-small"}},
  {"id":3402,"name":"integration","pipeline":{"id":340},"source":"push","sha":"1111111111111111111111111111111111111111","queued_duration":470.0,"duration":105.0,"status":"success","failure_reason":null,"runner":{"description":"linux-small"}},
  {"id":3403,"name":"package","pipeline":{"id":340},"source":"push","sha":"1111111111111111111111111111111111111111","queued_duration":6.0,"duration":30.0,"status":"failed","failure_reason":"script_failure","runner":{"description":"linux-small"}}
]
JSON
cat > evidence/runner_metrics.prom <<'PROM'
# synthetic lab snapshot, not a live endpoint
gitlab_runner_concurrent 1
gitlab_runner_errors_total{level="error"} 0
process_resident_memory_bytes 94371840
PROM
cat > evidence/job-3403.log <<'LOG'
2026-09-13T08:09:41Z package: start
2026-09-13T08:09:48Z artifact upload: HTTP 503 from disposable mock store
2026-09-13T08:10:11Z ERROR package evidence upload failed
LOG

3. Compute queue/execution timing and classify the bottleneck

The analyzer reads IDs from evidence instead of inferring order from filenames.

cat > analyze.py <<'PY'
import json, statistics
from pathlib import Path
jobs=json.loads(Path("evidence/jobs.json").read_text())
for j in jobs:
    ratio=j["queued_duration"] / max(j["duration"],0.001)
    if j["status"] != "success": layer="script/tool/network" if j["failure_reason"]=="script_failure" else "unknown"
    elif ratio > 1: layer="queue/capacity"
    else: layer="execution"
    print(f'job={j["id"]} name={j["name"]} queue={j["queued_duration"]:.0f}s run={j["duration"]:.0f}s class={layer}')
print("median_queue=",statistics.median(j["queued_duration"] for j in jobs))
print("median_run=",statistics.median(j["duration"] for j in jobs))
PY
python analyze.py | tee evidence/analysis.txt

Expected interpretation: integration is queue/capacity dominated. package is not automatically a runner failure; its trace names HTTP 503 during artifact upload, so inspect script/network/artifact flow.

4. Preserve pipeline source, SHA and exact IDs

Write the identity tuple before any rerun:

python - <<'PY'
import json
j=json.load(open("evidence/jobs.json"))[0]
print({"pipeline_id":j["pipeline"]["id"],"job_id":j["id"],"pipeline_source":j["source"],"source_sha":j["sha"],"runner":j["runner"]["description"]})
PY

In a real pipeline also preserve CI_PIPELINE_SOURCE and CI_COMMIT_SHA. A branch can move and rules can compile different graphs for different sources.

5. Optional mapping to an authorized disposable GitLab project

Use exact IDs and a narrow identity. These requests are read-only:

curl --silent --header "PRIVATE-TOKEN: $LAB_TOKEN"   "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs?per_page=100&page=1"   --output evidence/jobs-page1.json

curl --silent --location --header "PRIVATE-TOKEN: $LAB_TOKEN"   "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs/$LAB_JOB_ID/trace"   --output "evidence/job-$LAB_JOB_ID.log"

Do not paste token-bearing commands into shared logs and do not use a broad PAT when narrower authorized access exists.

6. Interpret the Runner Prometheus snapshot

The local .prom file teaches that Runner metrics are infrastructure evidence. A real Runner can expose /metrics after listen_address is configured; that endpoint has no built-in authorization.

python - <<'PY'
from pathlib import Path
for line in Path("evidence/runner_metrics.prom").read_text().splitlines():
    if line and not line.startswith("#"): print(line)
PY

A single gitlab_runner_concurrent 1 sample does not prove sustained saturation. Correlate repeated metrics with high queued_duration, eligible tags and manager settings.

7. Classify failures by the failing state

Evidence Class Next safe inspection
YAML does not compile configuration/compilation CI lint / merged configuration
Job waits 470 s queue/capacity runner tags/state/concurrency
prepare executor fails runner/executor manager/executor logs
HTTP 503 in script script/tool/network dependency endpoint and bounded retry
403 reading artifact identity/authorization token/role/allowlist
report absent after failed job artifact/report artifact when/path/upload trace
CI deploy green, target unhealthy deployment/external environment + target health

8. Preserve the first-failure packet before retry

A retry is a new attempt. If an external dependency recovers, it can succeed and hide the original symptom.

sha256sum evidence/jobs.json evidence/job-3403.log evidence/analysis.txt > evidence/SHA256SUMS
cat evidence/SHA256SUMS

The checksum detects later change; it does not prove the original evidence was truthful. Collection identity and authorization are separate.

9. Optional YAML for bounded observation artifacts

This keeps the observation artifact separate from the test result.

stages: [test]
observed_test:
  stage: test
  image: python:3.13-alpine
  script:
    - mkdir -p evidence
    - printf 'pipeline=%s job=%s source=%s sha=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" > evidence/identity.txt
    - python -c 'import time; time.sleep(2); print("synthetic test complete")'
  artifacts:
    when: always
    expire_in: 7 days
    paths: [evidence/]

Do not add CI_DEBUG_TRACE=true; bounded identity evidence is sufficient.

10. Challenge: choose the layer, not the command

A code optimization cuts execution from 110 to 70 seconds, but queued duration rises from 5 to 190 seconds. Did end-to-end wait improve? A strong answer computes wait + run, identifies queue/capacity as the new dominant layer, and asks for repeated job timing plus Runner-capacity evidence.

Knowledge check

Why use a synthetic Prometheus file in the mandatory lab?

Does failure_reason=script_failure prove application code is wrong?

What must be saved before retrying job 3403?

Why hash the evidence packet?

If queue is 470 s and run is 105 s, which layer is likely to improve developer wait most?

11. Summary

You built the workflow without privileged infrastructure: preserve identity, separate queue/run time, interpret Runner metrics as infrastructure evidence, classify causal layer, retain first-failure traces and hash a bounded packet. The same model maps to authorized API/Runner endpoints.

Next lesson

Observability design choices and tradeoffs

Lesson 3 decides when logs, metrics, GitLab analytics, external observability and different retention policies are appropriate.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs and compute-usage concepts are available across Free/Premium/Ultimate unless a narrower capability is explicitly identified. The newer GLQL Pipeline Analytics data source is Premium/Ultimate. Runner metrics are exposed without built-in authorization when enabled, so labs bind or simulate them locally rather than publishing the endpoint. Current job logs support line timestamps; Runner 18.7+ is required to control them with FF_TIMESTAMPS. Runner output_limit defaults to 4096 KB, while GitLab server-side job trace size defaults to 100 MB; those are distinct limits with different truncation/failure behavior. Debug trace and service debug logging are security-sensitive because secret material can appear in logs. The mandatory labs use only synthetic local evidence and Python standard-library tooling; no live token, paid analytics feature, cloud account or public Runner endpoint is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.