Chapter 34Lesson 03~225 minutes

CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Configuration, Design Choices, and Tradeoffs

Choose between logs and metrics, GitLab analytics and external observability, useful correlation and harmful cardinality, and retention versus privacy/cost with explicit state ownership.

TradeoffsRetentionCardinalityPrivacyCost

Learning objectives

  • Choose the right mix of raw traces, structured metrics and artifacts for an operating question.
  • Separate Free project analytics from Premium/Ultimate aggregated Pipeline Analytics.
  • Control metric cardinality while preserving correlation identifiers.
  • Design retention around incident evidence, privacy and storage cost.
  • Explain why cost, latency and reliability metrics can move in opposite directions.

1. Design problem: every signal has a cost and a blind spot

Logs preserve detail but are expensive to search and easy to contaminate with secrets. Metrics aggregate well but lose context. GitLab analytics is convenient but cannot know external deployment health. External observability can join systems but introduces identity, retention and access controls. The goal is not “collect everything”; it is “collect enough to answer defined questions.”

2. Raw logs versus structured metrics

Choice Strength Risk Good use
Timestamped job trace step sequence / first error secret leakage, noise, storage exact job incident diagnosis
Structured observation artifact portable schema tied to job/pipeline schema ownership and retention evidence packet
Prometheus Runner metric cheap time-series aggregation infrastructure-only view; endpoint security capacity/process health
GitLab analytics aggregate low setup; project context less causal detail trend/success-rate orientation

3. Project analytics versus external observability

Current project CI/CD analytics is Free/Premium/Ultimate and shows pipeline success/failure and duration trends. The newer GLQL Pipeline Analytics data source is Premium/Ultimate. Neither automatically includes application runtime or cloud/Kubernetes health.

Use external observability for cross-project, runner-fleet and deployment telemetry when needed. Preserve bounded dimensions such as project, runner pool and environment. Keep unique pipeline/job IDs mainly in logs/artifacts/exemplars/links rather than Prometheus labels.

4. High-cardinality labels versus useful correlation

A metric label like job_name="unit" may have bounded values; job_id="9283746501" creates a new series for every job. Keep highly unique IDs in evidence records and use bounded dimensions in metrics.

Field Metric label? Better home
runner pool / executor usually yes metric label
job family/name often yes if bounded metric label
pipeline source yes, small enum metric label
pipeline ID usually no log/artifact/exemplar/link
commit SHA no evidence record/link
full URL / error text no trace/log

5. Retention versus cost and privacy

Define retention per evidence class. Artifacts can use explicit expire_in; job logs have separate retention/erasure behavior. On GitLab.com, job logs are retained by default without configurable automatic expiry. Erasing an exact job removes both trace and artifacts and is irreversible. Guard destructive cleanup with exact IDs and policy authorization.

6. Trace limits change what “complete evidence” means

Two current defaults are easy to confuse:

Limit Default Effect
Runner output_limit 4096 KB Runner truncates additional job output; job can continue.
GitLab server job trace size 100 MB Exceeding server limit can mark job failed/drop it.

If diagnostics are large, write bounded artifacts instead of flooding stdout.

7. Timestamped logs: useful but not sufficient

Current job-log timestamps are generally available, and Runner 18.7+ can control them through FF_TIMESTAMPS. Timestamps can isolate a five-minute pause but cannot tell whether it was CPU, network, lock contention or an external API. Correlate the interval with infrastructure/tool/network/target evidence.

8. Latency versus compute cost

Instance-runner compute usage is based on job running duration and cost factor; pending/created time is excluded. More parallelism can reduce wall time while increasing total compute. Conversely, queue delay can hurt developer latency without consuming running compute. Track end-to-end latency, runner execution/compute and correctness evidence separately.

9. Decision table: pick the observation design

Scenario Recommended evidence Tier/trust prerequisite Why
Small Free project wants failure trend Free CI/CD analytics + exact Jobs API IDs project access minimal setup; drill down with IDs
Runner fleet queue spikes private Runner /metrics + Jobs API queue duration runner operator; private endpoint correlates capacity with wait
Cross-project portfolio SLO external metrics/warehouse or Premium/Ultimate Pipeline Analytics organization data access aggregates bounded fields
Possible secret in log restricted exact trace + credential rotation + incident record security authorization do not copy secret into dashboards
CI deploy green; service unhealthy deployment ID + external health target read access CI success and target health differ

10. Keep state and identity boundaries explicit

Repository YAML plus CI_PIPELINE_SOURCE/CI_COMMIT_SHA describe source/configuration. Runner metrics describe manager/executor infrastructure. Jobs API describes GitLab-owned execution state. Artifacts/reports describe retained outputs. Cloud/Kubernetes/application telemetry describes external state. A dashboard can join them; correlation never makes them the same state.

11. Worked choice: p95 queue regression

p95 pipeline duration rises from 6 to 11 minutes while script duration stays flat; p95 queue for linux-small rises from 12 to 280 seconds. The next useful signal is Runner capacity/request metrics, not more application debug logging. If metrics confirm saturation, correct fleet capacity/configuration. If eligible capacity is idle, inspect tags and request concurrency.

Knowledge check

Why should pipeline ID usually not be a Prometheus label?

Which built-in analytics is available on Free?

Does artifact expire_in configure job-log expiry?

Why can a faster pipeline consume more compute?

What does a timestamped five-minute pause prove?

12. Summary

Choose signals by question and owner. Logs provide detail, metrics bounded aggregation, artifacts structured evidence, GitLab analytics platform trends, and external telemetry target behavior. Control cardinality, retention, privacy and cost explicitly.

Next lesson

Evidence-first diagnostics and security

Lesson 4 diagnoses queue bottlenecks, rerun evidence loss, trace truncation and unsafe debug logging without hiding the original cause.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs and compute-usage concepts are available across Free/Premium/Ultimate unless a narrower capability is explicitly identified. The newer GLQL Pipeline Analytics data source is Premium/Ultimate. Runner metrics are exposed without built-in authorization when enabled, so labs bind or simulate them locally rather than publishing the endpoint. Current job logs support line timestamps; Runner 18.7+ is required to control them with FF_TIMESTAMPS. Runner output_limit defaults to 4096 KB, while GitLab server-side job trace size defaults to 100 MB; those are distinct limits with different truncation/failure behavior. Debug trace and service debug logging are security-sensitive because secret material can appear in logs. The mandatory labs use only synthetic local evidence and Python standard-library tooling; no live token, paid analytics feature, cloud account or public Runner endpoint is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.