CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Guided Hands-On Workflow and Core Operations
Build a disposable observability workflow that computes queue and execution time, interprets simulated Runner Prometheus metrics, classifies failures, and preserves evidence without production access.
Learning objectives
- Create a disposable local dataset that mirrors timing and identity fields from GitLab Jobs/Pipelines APIs.
- Compute queue and execution time without mixing their meanings.
- Interpret a safe local Runner Prometheus snapshot and correlate it with job timing.
- Classify several failures by causal layer and preserve first-failure evidence.
- Build a small evidence artifact for before/after comparison.
1. Lab scenario and safety boundary
You operate a disposable project model named
glci-ch34-observe-lab. Instead of requiring a live
GitLab token or exposing a Runner metrics port, the mandatory path
uses synthetic API exports and a local Prometheus text file. An
optional mapping shows the read-only requests you would use on an
authorized lab instance.
Preflight: use only a temporary directory. No credentials, runner registration, pipeline deletion, log erasure, deployment or network listener is required.
2. Create the disposable evidence set
mkdir -p ch34-lab/evidence
cd ch34-lab
cat > evidence/jobs.json <<'JSON'
[
{"id":3401,"name":"unit","pipeline":{"id":340},"source":"push","sha":"1111111111111111111111111111111111111111","queued_duration":8.0,"duration":95.0,"status":"success","failure_reason":null,"runner":{"description":"linux-small"}},
{"id":3402,"name":"integration","pipeline":{"id":340},"source":"push","sha":"1111111111111111111111111111111111111111","queued_duration":470.0,"duration":105.0,"status":"success","failure_reason":null,"runner":{"description":"linux-small"}},
{"id":3403,"name":"package","pipeline":{"id":340},"source":"push","sha":"1111111111111111111111111111111111111111","queued_duration":6.0,"duration":30.0,"status":"failed","failure_reason":"script_failure","runner":{"description":"linux-small"}}
]
JSON
cat > evidence/runner_metrics.prom <<'PROM'
# synthetic lab snapshot, not a live endpoint
gitlab_runner_concurrent 1
gitlab_runner_errors_total{level="error"} 0
process_resident_memory_bytes 94371840
PROM
cat > evidence/job-3403.log <<'LOG'
2026-09-13T08:09:41Z package: start
2026-09-13T08:09:48Z artifact upload: HTTP 503 from disposable mock store
2026-09-13T08:10:11Z ERROR package evidence upload failed
LOG
3. Compute queue/execution timing and classify the bottleneck
The analyzer reads IDs from evidence instead of inferring order from filenames.
cat > analyze.py <<'PY'
import json, statistics
from pathlib import Path
jobs=json.loads(Path("evidence/jobs.json").read_text())
for j in jobs:
ratio=j["queued_duration"] / max(j["duration"],0.001)
if j["status"] != "success": layer="script/tool/network" if j["failure_reason"]=="script_failure" else "unknown"
elif ratio > 1: layer="queue/capacity"
else: layer="execution"
print(f'job={j["id"]} name={j["name"]} queue={j["queued_duration"]:.0f}s run={j["duration"]:.0f}s class={layer}')
print("median_queue=",statistics.median(j["queued_duration"] for j in jobs))
print("median_run=",statistics.median(j["duration"] for j in jobs))
PY
python analyze.py | tee evidence/analysis.txt
Expected interpretation: integration is queue/capacity
dominated. package is not automatically a runner
failure; its trace names HTTP 503 during artifact upload, so inspect
script/network/artifact flow.
4. Preserve pipeline source, SHA and exact IDs
Write the identity tuple before any rerun:
python - <<'PY'
import json
j=json.load(open("evidence/jobs.json"))[0]
print({"pipeline_id":j["pipeline"]["id"],"job_id":j["id"],"pipeline_source":j["source"],"source_sha":j["sha"],"runner":j["runner"]["description"]})
PY
In a real pipeline also preserve CI_PIPELINE_SOURCE and
CI_COMMIT_SHA. A branch can move and rules can compile
different graphs for different sources.
5. Optional mapping to an authorized disposable GitLab project
Use exact IDs and a narrow identity. These requests are read-only:
curl --silent --header "PRIVATE-TOKEN: $LAB_TOKEN" "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs?per_page=100&page=1" --output evidence/jobs-page1.json
curl --silent --location --header "PRIVATE-TOKEN: $LAB_TOKEN" "$CI_API_V4_URL/projects/$LAB_PROJECT_ID/jobs/$LAB_JOB_ID/trace" --output "evidence/job-$LAB_JOB_ID.log"
Do not paste token-bearing commands into shared logs and do not use a broad PAT when narrower authorized access exists.
6. Interpret the Runner Prometheus snapshot
The local .prom file teaches that Runner metrics are
infrastructure evidence. A real Runner can expose
/metrics after listen_address is
configured; that endpoint has no built-in authorization.
python - <<'PY'
from pathlib import Path
for line in Path("evidence/runner_metrics.prom").read_text().splitlines():
if line and not line.startswith("#"): print(line)
PY
A single gitlab_runner_concurrent 1 sample does not
prove sustained saturation. Correlate repeated metrics with high
queued_duration, eligible tags and manager settings.
7. Classify failures by the failing state
| Evidence | Class | Next safe inspection |
|---|---|---|
| YAML does not compile | configuration/compilation | CI lint / merged configuration |
| Job waits 470 s | queue/capacity | runner tags/state/concurrency |
| prepare executor fails | runner/executor | manager/executor logs |
| HTTP 503 in script | script/tool/network | dependency endpoint and bounded retry |
| 403 reading artifact | identity/authorization | token/role/allowlist |
| report absent after failed job | artifact/report | artifact when/path/upload trace |
| CI deploy green, target unhealthy | deployment/external | environment + target health |
8. Preserve the first-failure packet before retry
A retry is a new attempt. If an external dependency recovers, it can succeed and hide the original symptom.
sha256sum evidence/jobs.json evidence/job-3403.log evidence/analysis.txt > evidence/SHA256SUMS
cat evidence/SHA256SUMS
The checksum detects later change; it does not prove the original evidence was truthful. Collection identity and authorization are separate.
9. Optional YAML for bounded observation artifacts
This keeps the observation artifact separate from the test result.
stages: [test]
observed_test:
stage: test
image: python:3.13-alpine
script:
- mkdir -p evidence
- printf 'pipeline=%s job=%s source=%s sha=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" > evidence/identity.txt
- python -c 'import time; time.sleep(2); print("synthetic test complete")'
artifacts:
when: always
expire_in: 7 days
paths: [evidence/]
Do not add CI_DEBUG_TRACE=true; bounded identity
evidence is sufficient.
10. Challenge: choose the layer, not the command
A code optimization cuts execution from 110 to 70 seconds, but queued duration rises from 5 to 190 seconds. Did end-to-end wait improve? A strong answer computes wait + run, identifies queue/capacity as the new dominant layer, and asks for repeated job timing plus Runner-capacity evidence.
Knowledge check
Why use a synthetic Prometheus file in the mandatory lab?
It keeps the path free/local and avoids exposing an unauthenticated operator endpoint.
Does failure_reason=script_failure prove
application code is wrong?
No; inspect the trace to separate test/tool/network/artifact causes inside the script layer.
What must be saved before retrying job 3403?
Pipeline/source/SHA, exact job ID, timing, runner identity, first-failure trace and relevant artifact/external evidence.
Why hash the evidence packet?
To detect later modification; the hash alone does not prove collection truth.
If queue is 470 s and run is 105 s, which layer is likely to improve developer wait most?
Runner eligibility/capacity/queue behavior, verified with repeated timing and Runner metrics.
11. Summary
You built the workflow without privileged infrastructure: preserve identity, separate queue/run time, interpret Runner metrics as infrastructure evidence, classify causal layer, retain first-failure traces and hash a bounded packet. The same model maps to authorized API/Runner endpoints.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-13.
Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the
timestamped baseline where a concrete version matters. Project CI/CD
analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs
and compute-usage concepts are available across
Free/Premium/Ultimate unless a narrower capability is explicitly
identified. The newer GLQL Pipeline Analytics data source is
Premium/Ultimate. Runner metrics are exposed without built-in
authorization when enabled, so labs bind or simulate them locally
rather than publishing the endpoint. Current job logs support line
timestamps; Runner 18.7+ is required to control them with
FF_TIMESTAMPS. Runner
output_limit defaults to 4096 KB, while GitLab
server-side job trace size defaults to 100 MB; those are distinct
limits with different truncation/failure behavior. Debug trace and
service debug logging are security-sensitive because secret material
can appear in logs. The mandatory labs use only synthetic local
evidence and Python standard-library tooling; no live token, paid
analytics feature, cloud account or public Runner endpoint is
required.
- CI/CD analytics — official reference.
- Pipeline analytics (GLQL) — official reference.
- Jobs API — official reference.
- Pipelines API — official reference.
- Runner monitoring and Prometheus metrics — official reference.
- Runner advanced configuration — official reference.
- CI/CD job logs — official reference.
- Self-Managed job-log storage and limits — official reference.
- CI/CD limits — official reference.
- CI/CD variable security and masking — official reference.
- Troubleshooting CI/CD variables / debug trace — official reference.
- Service-container debug logs — official reference.
- Job artifacts — official reference.
- Job Artifacts API — official reference.
- Compute minutes — official reference.
- Compute usage for instance runners — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.