CI/CD Analytics, Job Logs, Runner Metrics, Queue Time, Failure Taxonomy, and Observability: Configuration, Design Choices, and Tradeoffs
Choose between logs and metrics, GitLab analytics and external observability, useful correlation and harmful cardinality, and retention versus privacy/cost with explicit state ownership.
Learning objectives
- Choose the right mix of raw traces, structured metrics and artifacts for an operating question.
- Separate Free project analytics from Premium/Ultimate aggregated Pipeline Analytics.
- Control metric cardinality while preserving correlation identifiers.
- Design retention around incident evidence, privacy and storage cost.
- Explain why cost, latency and reliability metrics can move in opposite directions.
1. Design problem: every signal has a cost and a blind spot
Logs preserve detail but are expensive to search and easy to contaminate with secrets. Metrics aggregate well but lose context. GitLab analytics is convenient but cannot know external deployment health. External observability can join systems but introduces identity, retention and access controls. The goal is not “collect everything”; it is “collect enough to answer defined questions.”
2. Raw logs versus structured metrics
| Choice | Strength | Risk | Good use |
|---|---|---|---|
| Timestamped job trace | step sequence / first error | secret leakage, noise, storage | exact job incident diagnosis |
| Structured observation artifact | portable schema tied to job/pipeline | schema ownership and retention | evidence packet |
| Prometheus Runner metric | cheap time-series aggregation | infrastructure-only view; endpoint security | capacity/process health |
| GitLab analytics aggregate | low setup; project context | less causal detail | trend/success-rate orientation |
3. Project analytics versus external observability
Current project CI/CD analytics is Free/Premium/Ultimate and shows pipeline success/failure and duration trends. The newer GLQL Pipeline Analytics data source is Premium/Ultimate. Neither automatically includes application runtime or cloud/Kubernetes health.
Use external observability for cross-project, runner-fleet and deployment telemetry when needed. Preserve bounded dimensions such as project, runner pool and environment. Keep unique pipeline/job IDs mainly in logs/artifacts/exemplars/links rather than Prometheus labels.
4. High-cardinality labels versus useful correlation
A metric label like job_name="unit" may have bounded
values; job_id="9283746501" creates a new series for
every job. Keep highly unique IDs in evidence records and use
bounded dimensions in metrics.
| Field | Metric label? | Better home |
|---|---|---|
| runner pool / executor | usually yes | metric label |
| job family/name | often yes if bounded | metric label |
| pipeline source | yes, small enum | metric label |
| pipeline ID | usually no | log/artifact/exemplar/link |
| commit SHA | no | evidence record/link |
| full URL / error text | no | trace/log |
5. Retention versus cost and privacy
Define retention per evidence class. Artifacts can use explicit
expire_in; job logs have separate retention/erasure
behavior. On GitLab.com, job logs are retained by default without
configurable automatic expiry. Erasing an exact job removes both
trace and artifacts and is irreversible. Guard destructive cleanup
with exact IDs and policy authorization.
6. Trace limits change what “complete evidence” means
Two current defaults are easy to confuse:
| Limit | Default | Effect |
|---|---|---|
Runner output_limit |
4096 KB | Runner truncates additional job output; job can continue. |
| GitLab server job trace size | 100 MB | Exceeding server limit can mark job failed/drop it. |
If diagnostics are large, write bounded artifacts instead of flooding stdout.
7. Timestamped logs: useful but not sufficient
Current job-log timestamps are generally available, and Runner 18.7+
can control them through FF_TIMESTAMPS. Timestamps can
isolate a five-minute pause but cannot tell whether it was CPU,
network, lock contention or an external API. Correlate the interval
with infrastructure/tool/network/target evidence.
8. Latency versus compute cost
Instance-runner compute usage is based on job running duration and cost factor; pending/created time is excluded. More parallelism can reduce wall time while increasing total compute. Conversely, queue delay can hurt developer latency without consuming running compute. Track end-to-end latency, runner execution/compute and correctness evidence separately.
9. Decision table: pick the observation design
| Scenario | Recommended evidence | Tier/trust prerequisite | Why |
|---|---|---|---|
| Small Free project wants failure trend | Free CI/CD analytics + exact Jobs API IDs | project access | minimal setup; drill down with IDs |
| Runner fleet queue spikes |
private Runner /metrics + Jobs API queue
duration
|
runner operator; private endpoint | correlates capacity with wait |
| Cross-project portfolio SLO | external metrics/warehouse or Premium/Ultimate Pipeline Analytics | organization data access | aggregates bounded fields |
| Possible secret in log | restricted exact trace + credential rotation + incident record | security authorization | do not copy secret into dashboards |
| CI deploy green; service unhealthy | deployment ID + external health | target read access | CI success and target health differ |
10. Keep state and identity boundaries explicit
Repository YAML plus CI_PIPELINE_SOURCE/CI_COMMIT_SHA
describe source/configuration. Runner metrics describe
manager/executor infrastructure. Jobs API describes GitLab-owned
execution state. Artifacts/reports describe retained outputs.
Cloud/Kubernetes/application telemetry describes external state. A
dashboard can join them; correlation never makes them the same
state.
11. Worked choice: p95 queue regression
p95 pipeline duration rises from 6 to 11 minutes while script
duration stays flat; p95 queue for linux-small rises
from 12 to 280 seconds. The next useful signal is Runner
capacity/request metrics, not more application debug logging. If
metrics confirm saturation, correct fleet capacity/configuration. If
eligible capacity is idle, inspect tags and request concurrency.
Knowledge check
Why should pipeline ID usually not be a Prometheus label?
It is extremely high-cardinality and creates a new series for nearly every pipeline.
Which built-in analytics is available on Free?
Project CI/CD analytics; the newer GLQL Pipeline Analytics data source is Premium/Ultimate.
Does artifact expire_in configure job-log
expiry?
No. Artifact and job-log retention are separate lifecycles.
Why can a faster pipeline consume more compute?
Parallel execution can shorten critical-path wall time while increasing summed runner execution.
What does a timestamped five-minute pause prove?
Only the interval between emitted log events; further telemetry is needed to identify cause.
12. Summary
Choose signals by question and owner. Logs provide detail, metrics bounded aggregation, artifacts structured evidence, GitLab analytics platform trends, and external telemetry target behavior. Control cardinality, retention, privacy and cost explicitly.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-13.
Executable examples use GitLab/GitLab Runner 19.3.2 semantics as the
timestamped baseline where a concrete version matters. Project CI/CD
analytics, Jobs/Pipelines APIs, Runner monitoring, ordinary job logs
and compute-usage concepts are available across
Free/Premium/Ultimate unless a narrower capability is explicitly
identified. The newer GLQL Pipeline Analytics data source is
Premium/Ultimate. Runner metrics are exposed without built-in
authorization when enabled, so labs bind or simulate them locally
rather than publishing the endpoint. Current job logs support line
timestamps; Runner 18.7+ is required to control them with
FF_TIMESTAMPS. Runner
output_limit defaults to 4096 KB, while GitLab
server-side job trace size defaults to 100 MB; those are distinct
limits with different truncation/failure behavior. Debug trace and
service debug logging are security-sensitive because secret material
can appear in logs. The mandatory labs use only synthetic local
evidence and Python standard-library tooling; no live token, paid
analytics feature, cloud account or public Runner endpoint is
required.
- CI/CD analytics — official reference.
- Pipeline analytics (GLQL) — official reference.
- Jobs API — official reference.
- Pipelines API — official reference.
- Runner monitoring and Prometheus metrics — official reference.
- Runner advanced configuration — official reference.
- CI/CD job logs — official reference.
- Self-Managed job-log storage and limits — official reference.
- CI/CD limits — official reference.
- CI/CD variable security and masking — official reference.
- Troubleshooting CI/CD variables / debug trace — official reference.
- Service-container debug logs — official reference.
- Job artifacts — official reference.
- Job Artifacts API — official reference.
- Compute minutes — official reference.
- Compute usage for instance runners — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.