Checkpoint Lab — CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability
Build a reproducible observability worksheet from disposable pipeline data, identify the true bottleneck and failure layer, and justify a bounded corrective action with retained evidence.
Learning objectives
- Build an observability worksheet from disposable pipeline/job data.
- Predict state changes before analysis and verify them independently.
- Identify the real bottleneck and first-failure layer using queue/run/log/metric evidence.
- Produce an evidence packet with assumptions, limitations and digests.
- Propose one corrective action and one rejected action with evidence-based reasoning.
1. Checkpoint mission
You are given two disposable pipeline attempts from the same source SHA. Identify the true bottleneck and failure layer without a live GitLab account, produce a small worksheet, and recommend a correction. The exercise includes a misleading “slow test” narrative so evidence must win over labels.
Safety: the checkpoint runs entirely on local synthetic JSON/Prometheus/log files. No token, runner, cloud account, production project, destructive API call or external monitoring system is required.
2. Timestamped assumptions and preflight
For a real GitLab run, the evidence packet must record
CI_PIPELINE_SOURCE, the ref, and
CI_COMMIT_SHA before analysis begins. The synthetic
dataset below carries equivalent pipeline_source and
sha fields so the local checkpoint preserves the same
source-to-evidence correlation.
- Documentation verified: 2026-09-13.
- GitLab/GitLab Runner behavior baseline: 19.3.2.
- Python 3.11+ standard library is sufficient.
- All IDs, SHAs, runners, logs and metrics are synthetic.
- No paid feature is required; Premium/Ultimate Pipeline Analytics is only an optional equivalent surface.
python --version
mkdir -p ch34-checkpoint/input ch34-checkpoint/evidence
cd ch34-checkpoint
3. Predict state changes before running anything
Write at least two predictions:
cat > evidence/predictions.txt <<'TXT'
1. Analysis will not mutate pipeline/job state; it will create only local evidence files.
2. Dominant latency will be queue/capacity, not test execution.
3. The failed attempt will classify to script/tool/network rather than runner CPU.
4. Hash comparison will detect later evidence-file modification.
TXT
Verify these predictions after analysis rather than rewriting them silently.
4. Create the disposable pipeline data
cat > input/jobs.json <<'JSON'
[
{"pipeline_id":3400,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34001,"name":"unit","queued":7,"run":85,"status":"success","failure_reason":null,"runner":"linux-small"},
{"pipeline_id":3400,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34002,"name":"integration","queued":440,"run":100,"status":"success","failure_reason":null,"runner":"linux-small"},
{"pipeline_id":3400,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34003,"name":"report","queued":4,"run":14,"status":"failed","failure_reason":"script_failure","runner":"linux-small"},
{"pipeline_id":3401,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34011,"name":"report","queued":5,"run":13,"status":"success","failure_reason":null,"runner":"linux-medium"}
]
JSON
cat > input/runner.prom <<'PROM'
# synthetic point snapshot
lab_runner_capacity{pool="linux-small"} 1
lab_runner_pending_jobs{pool="linux-small"} 6
lab_runner_capacity{pool="linux-medium"} 2
lab_runner_pending_jobs{pool="linux-medium"} 0
PROM
cat > input/job-34003.log <<'LOG'
2026-09-13T10:09:04Z report upload begin
2026-09-13T10:09:05Z POST https://mock-artifacts.invalid/report -> 503
2026-09-13T10:09:18Z ERROR report upload failed after bounded retry
LOG
cat > input/external-health.json <<'JSON'
{"deployment":null,"note":"No deployment occurred; external target state is not applicable."}
JSON
5. Build the worksheet/dashboard
cat > analyze.py <<'PY'
import csv,json,statistics
from pathlib import Path
jobs=json.loads(Path("input/jobs.json").read_text())
rows=[]
for j in jobs:
if j["status"] != "success": layer="script/tool/network"
elif j["queued"] > j["run"]: layer="queue/capacity"
else: layer="execution"
rows.append({"pipeline_id":j["pipeline_id"],"job_id":j["job_id"],"name":j["name"],"source":j["pipeline_source"],"sha":j["sha"],"runner":j["runner"],"queue_s":j["queued"],"run_s":j["run"],"status":j["status"],"class":layer})
with open("evidence/worksheet.csv","w",newline="") as f:
w=csv.DictWriter(f,fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)
summary={"pipeline_ids":sorted({r["pipeline_id"] for r in rows}),"sha":rows[0]["sha"],"median_queue_s":statistics.median(r["queue_s"] for r in rows),"max_queue_s":max(r["queue_s"] for r in rows),"max_queue_job":max(rows,key=lambda r:r["queue_s"])["job_id"],"failed_jobs":[r["job_id"] for r in rows if r["status"]!="success"],"dominant_bottleneck":"queue/capacity","first_failure_layer":"script/tool/network"}
Path("evidence/summary.json").write_text(json.dumps(summary,indent=2)+"\n")
print(json.dumps(summary,indent=2))
PY
python analyze.py
Expected: job 34002 has 440 seconds queue versus 100 seconds run. Failed job 34003 is a separate script/tool/network failure supported by its HTTP 503 trace. Do not merge both into “tests are slow.”
6. Correlate Runner snapshot and first-failure trace
The snapshot shows one small-pool capacity unit with six pending jobs while the medium pool has spare capacity. That supports—but does not alone prove—the queue explanation. Job timing supplies execution-side evidence; the preserved trace supports the upload/network failure.
grep -E 'capacity|pending' input/runner.prom
cat input/job-34003.log
7. Propose one correction and reject one shortcut
Evidence-backed correction: investigate why
integration is restricted to the saturated small pool.
If requirements permit, adjust eligible capacity/routing, then
compare queue p50/p95 over a bounded sample. For job 34003, keep
report upload retry bounded/idempotent; do not scale runners to
solve HTTP 503.
Rejected shortcut: global
CI_DEBUG_TRACE adds secret-exposure risk and does not
address queue saturation. Blindly rerunning the whole pipeline can
also hide the original 503 without proving a fix.
8. Produce the evidence packet
cp input/job-34003.log evidence/
cp input/runner.prom evidence/
cp input/external-health.json evidence/
cat > evidence/assumptions.txt <<'TXT'
GitLab/Runner semantics baseline: 19.3.2; docs verified 2026-09-13.
Mandatory path uses synthetic data; no live API/metrics endpoint queried.
Runner snapshot is illustrative and not a time series; queue conclusion also relies on Jobs-style timing data.
No deployment occurred, so external health is not applicable.
TXT
sha256sum evidence/* > evidence/SHA256SUMS
cat evidence/SHA256SUMS
The packet contains predictions, pipeline/job/source/SHA worksheet, summary, first-failure trace, Runner snapshot, external-state note, assumptions/limitations and digests.
9. Verification checklist
- Worksheet preserves pipeline IDs 3400/3401 and exact job IDs.
- Every row preserves pipeline source and exact SHA.
- Queue and run remain separate columns.
- Job 34002 is the queue maximum; job 34003 remains original failure evidence.
- Runner snapshot is labeled synthetic and not treated as a time series.
- No token/secret appears in evidence.
- Hashes can be recomputed to detect modification.
10. Intentional integrity failure
Alter a copy, not retained evidence:
cp evidence/summary.json /tmp/ch34-summary-tampered.json
printf '\n' >> /tmp/ch34-summary-tampered.json
sha256sum evidence/summary.json /tmp/ch34-summary-tampered.json
The changed digest demonstrates integrity checking. It does not prove the source data was truthful; collection provenance remains separate.
11. Cleanup and rollback
No remote state exists. After reviewing evidence, remove exactly the disposable directory:
cd ..
test -d ch34-checkpoint && rm -rf ch34-checkpoint
On a real project, do not erase failed jobs merely to “clean up” unless retention policy authorizes it and required incident evidence has already been preserved.
12. Bridge to Chapter 35: optimize only after measurement
Chapter 34 gives a trustworthy baseline: queue, execution, failure
layer, Runner evidence, retained artifacts and external-state notes.
Chapter 35—Pipeline Performance, Cost Control, Caching Strategy, Queue
Reduction, Selective Execution, and Optimization—will change needs, cache/artifact scope, rules and
runner usage one variable at a time, then prove whether end-to-end
latency improves without sacrificing correctness, coverage or
isolation.
Knowledge check
What is the real latency bottleneck?
Queue/capacity: job 34002 waits 440 seconds and runs for only 100.
What layer caused the first failed job?
The trace supports script/tool/network—HTTP 503 during report upload—not Runner CPU or queue capacity.
Why keep job 34003 after pipeline 3401 succeeds?
The later success is a separate attempt; 34003 is first-failure evidence.
Why is Runner snapshot supporting evidence only?
It is a synthetic point snapshot, not a historical series; queue conclusion also needs job timing/eligibility evidence.
Which “optimization” should be rejected immediately?
Global debug tracing, because it adds secret exposure and does not address the measured queue bottleneck.
13. Checkpoint summary
You turned raw pipeline-style data into an evidence-backed operating decision. You preserved exact IDs/source/SHA, separated queue/execution, correlated Runner and trace evidence, retained first failure, documented limitations, hashed the packet and rejected an unsafe non-solution. That is the baseline required before Chapter 35 changes performance or cost.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-13.
Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the
timestamped baseline where a concrete version matters. Project CI/CD
analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs
and compute-usage concepts are available across
Free/Premium/Ultimate unless a narrower capability is explicitly
identified. The newer GLQL Pipeline Analytics data source is
Premium/Ultimate. Runner metrics are exposed without built-in
authorization when enabled, so labs bind or simulate them locally
rather than publishing the endpoint. Current job logs support line
timestamps; Runner 18.7+ is required to control them with
FF_TIMESTAMPS. Runner
output_limit defaults to 4096 KB, while GitLab
server-side job trace size defaults to 100 MB; those are distinct
limits with different truncation/failure behavior. Debug trace and
service debug logging are security-sensitive because secret material
can appear in logs. The mandatory labs use only synthetic local
evidence and Python standard-library tooling; no live token, paid
analytics feature, cloud account or public Runner endpoint is
required.
- CI/CD analytics — official reference.
- Pipeline analytics (GLQL) — official reference.
- Jobs API — official reference.
- Pipelines API — official reference.
- Runner monitoring and Prometheus metrics — official reference.
- Runner advanced configuration — official reference.
- CI/CD job logs — official reference.
- Self-Managed job-log storage and limits — official reference.
- CI/CD limits — official reference.
- CI/CD variable security and masking — official reference.
- Troubleshooting CI/CD variables / debug trace — official reference.
- Service-container debug logs — official reference.
- Job artifacts — official reference.
- Job Artifacts API — official reference.
- Compute minutes — official reference.
- Compute usage for instance runners — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.