Chapter 34Lesson 05~285 minutes

Checkpoint Lab — CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability

Build a reproducible observability worksheet from disposable pipeline data, identify the true bottleneck and failure layer, and justify a bounded corrective action with retained evidence.

CheckpointWorksheetSLOEvidence packetChapter 35 bridge

Learning objectives

  • Build an observability worksheet from disposable pipeline/job data.
  • Predict state changes before analysis and verify them independently.
  • Identify the real bottleneck and first-failure layer using queue/run/log/metric evidence.
  • Produce an evidence packet with assumptions, limitations and digests.
  • Propose one corrective action and one rejected action with evidence-based reasoning.

1. Checkpoint mission

You are given two disposable pipeline attempts from the same source SHA. Identify the true bottleneck and failure layer without a live GitLab account, produce a small worksheet, and recommend a correction. The exercise includes a misleading “slow test” narrative so evidence must win over labels.

Safety: the checkpoint runs entirely on local synthetic JSON/Prometheus/log files. No token, runner, cloud account, production project, destructive API call or external monitoring system is required.

2. Timestamped assumptions and preflight

For a real GitLab run, the evidence packet must record CI_PIPELINE_SOURCE, the ref, and CI_COMMIT_SHA before analysis begins. The synthetic dataset below carries equivalent pipeline_source and sha fields so the local checkpoint preserves the same source-to-evidence correlation.

  • Documentation verified: 2026-09-13.
  • GitLab/GitLab Runner behavior baseline: 19.3.2.
  • Python 3.11+ standard library is sufficient.
  • All IDs, SHAs, runners, logs and metrics are synthetic.
  • No paid feature is required; Premium/Ultimate Pipeline Analytics is only an optional equivalent surface.
python --version
mkdir -p ch34-checkpoint/input ch34-checkpoint/evidence
cd ch34-checkpoint

3. Predict state changes before running anything

Write at least two predictions:

cat > evidence/predictions.txt <<'TXT'
1. Analysis will not mutate pipeline/job state; it will create only local evidence files.
2. Dominant latency will be queue/capacity, not test execution.
3. The failed attempt will classify to script/tool/network rather than runner CPU.
4. Hash comparison will detect later evidence-file modification.
TXT

Verify these predictions after analysis rather than rewriting them silently.

4. Create the disposable pipeline data

cat > input/jobs.json <<'JSON'
[
 {"pipeline_id":3400,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34001,"name":"unit","queued":7,"run":85,"status":"success","failure_reason":null,"runner":"linux-small"},
 {"pipeline_id":3400,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34002,"name":"integration","queued":440,"run":100,"status":"success","failure_reason":null,"runner":"linux-small"},
 {"pipeline_id":3400,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34003,"name":"report","queued":4,"run":14,"status":"failed","failure_reason":"script_failure","runner":"linux-small"},
 {"pipeline_id":3401,"pipeline_source":"merge_request_event","sha":"2222222222222222222222222222222222222222","job_id":34011,"name":"report","queued":5,"run":13,"status":"success","failure_reason":null,"runner":"linux-medium"}
]
JSON
cat > input/runner.prom <<'PROM'
# synthetic point snapshot
lab_runner_capacity{pool="linux-small"} 1
lab_runner_pending_jobs{pool="linux-small"} 6
lab_runner_capacity{pool="linux-medium"} 2
lab_runner_pending_jobs{pool="linux-medium"} 0
PROM
cat > input/job-34003.log <<'LOG'
2026-09-13T10:09:04Z report upload begin
2026-09-13T10:09:05Z POST https://mock-artifacts.invalid/report -> 503
2026-09-13T10:09:18Z ERROR report upload failed after bounded retry
LOG
cat > input/external-health.json <<'JSON'
{"deployment":null,"note":"No deployment occurred; external target state is not applicable."}
JSON

5. Build the worksheet/dashboard

cat > analyze.py <<'PY'
import csv,json,statistics
from pathlib import Path
jobs=json.loads(Path("input/jobs.json").read_text())
rows=[]
for j in jobs:
    if j["status"] != "success": layer="script/tool/network"
    elif j["queued"] > j["run"]: layer="queue/capacity"
    else: layer="execution"
    rows.append({"pipeline_id":j["pipeline_id"],"job_id":j["job_id"],"name":j["name"],"source":j["pipeline_source"],"sha":j["sha"],"runner":j["runner"],"queue_s":j["queued"],"run_s":j["run"],"status":j["status"],"class":layer})
with open("evidence/worksheet.csv","w",newline="") as f:
    w=csv.DictWriter(f,fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)
summary={"pipeline_ids":sorted({r["pipeline_id"] for r in rows}),"sha":rows[0]["sha"],"median_queue_s":statistics.median(r["queue_s"] for r in rows),"max_queue_s":max(r["queue_s"] for r in rows),"max_queue_job":max(rows,key=lambda r:r["queue_s"])["job_id"],"failed_jobs":[r["job_id"] for r in rows if r["status"]!="success"],"dominant_bottleneck":"queue/capacity","first_failure_layer":"script/tool/network"}
Path("evidence/summary.json").write_text(json.dumps(summary,indent=2)+"\n")
print(json.dumps(summary,indent=2))
PY
python analyze.py

Expected: job 34002 has 440 seconds queue versus 100 seconds run. Failed job 34003 is a separate script/tool/network failure supported by its HTTP 503 trace. Do not merge both into “tests are slow.”

6. Correlate Runner snapshot and first-failure trace

The snapshot shows one small-pool capacity unit with six pending jobs while the medium pool has spare capacity. That supports—but does not alone prove—the queue explanation. Job timing supplies execution-side evidence; the preserved trace supports the upload/network failure.

grep -E 'capacity|pending' input/runner.prom
cat input/job-34003.log

7. Propose one correction and reject one shortcut

Evidence-backed correction: investigate why integration is restricted to the saturated small pool. If requirements permit, adjust eligible capacity/routing, then compare queue p50/p95 over a bounded sample. For job 34003, keep report upload retry bounded/idempotent; do not scale runners to solve HTTP 503.

Rejected shortcut: global CI_DEBUG_TRACE adds secret-exposure risk and does not address queue saturation. Blindly rerunning the whole pipeline can also hide the original 503 without proving a fix.

8. Produce the evidence packet

cp input/job-34003.log evidence/
cp input/runner.prom evidence/
cp input/external-health.json evidence/
cat > evidence/assumptions.txt <<'TXT'
GitLab/Runner semantics baseline: 19.3.2; docs verified 2026-09-13.
Mandatory path uses synthetic data; no live API/metrics endpoint queried.
Runner snapshot is illustrative and not a time series; queue conclusion also relies on Jobs-style timing data.
No deployment occurred, so external health is not applicable.
TXT
sha256sum evidence/* > evidence/SHA256SUMS
cat evidence/SHA256SUMS

The packet contains predictions, pipeline/job/source/SHA worksheet, summary, first-failure trace, Runner snapshot, external-state note, assumptions/limitations and digests.

9. Verification checklist

  • Worksheet preserves pipeline IDs 3400/3401 and exact job IDs.
  • Every row preserves pipeline source and exact SHA.
  • Queue and run remain separate columns.
  • Job 34002 is the queue maximum; job 34003 remains original failure evidence.
  • Runner snapshot is labeled synthetic and not treated as a time series.
  • No token/secret appears in evidence.
  • Hashes can be recomputed to detect modification.

10. Intentional integrity failure

Alter a copy, not retained evidence:

cp evidence/summary.json /tmp/ch34-summary-tampered.json
printf '\n' >> /tmp/ch34-summary-tampered.json
sha256sum evidence/summary.json /tmp/ch34-summary-tampered.json

The changed digest demonstrates integrity checking. It does not prove the source data was truthful; collection provenance remains separate.

11. Cleanup and rollback

No remote state exists. After reviewing evidence, remove exactly the disposable directory:

cd ..
test -d ch34-checkpoint && rm -rf ch34-checkpoint

On a real project, do not erase failed jobs merely to “clean up” unless retention policy authorizes it and required incident evidence has already been preserved.

12. Bridge to Chapter 35: optimize only after measurement

Chapter 34 gives a trustworthy baseline: queue, execution, failure layer, Runner evidence, retained artifacts and external-state notes. Chapter 35—Pipeline Performance, Cost Control, Caching Strategy, Queue Reduction, Selective Execution, and Optimization—will change needs, cache/artifact scope, rules and runner usage one variable at a time, then prove whether end-to-end latency improves without sacrificing correctness, coverage or isolation.

Knowledge check

What is the real latency bottleneck?

What layer caused the first failed job?

Why keep job 34003 after pipeline 3401 succeeds?

Why is Runner snapshot supporting evidence only?

Which “optimization” should be rejected immediately?

13. Checkpoint summary

You turned raw pipeline-style data into an evidence-backed operating decision. You preserved exact IDs/source/SHA, separated queue/execution, correlated Runner and trace evidence, retained first failure, documented limitations, hashed the packet and rejected an unsafe non-solution. That is the baseline required before Chapter 35 changes performance or cost.

Next lesson

Chapter 35 — measured performance and cost optimization

Use the Chapter 34 baseline to optimize critical paths, cache/artifact flow, selective execution and runner usage without sacrificing correctness.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs and compute-usage concepts are available across Free/Premium/Ultimate unless a narrower capability is explicitly identified. The newer GLQL Pipeline Analytics data source is Premium/Ultimate. Runner metrics are exposed without built-in authorization when enabled, so labs bind or simulate them locally rather than publishing the endpoint. Current job logs support line timestamps; Runner 18.7+ is required to control them with FF_TIMESTAMPS. Runner output_limit defaults to 4096 KB, while GitLab server-side job trace size defaults to 100 MB; those are distinct limits with different truncation/failure behavior. Debug trace and service debug logging are security-sensitive because secret material can appear in logs. The mandatory labs use only synthetic local evidence and Python standard-library tooling; no live token, paid analytics feature, cloud account or public Runner endpoint is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.