Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics: Configuration, Design Choices, and Tradeoffs
Design an observability and audit strategy that balances evidence depth, cardinality, privacy, storage cost, external monitoring, SLOs, and operational ownership.
Learning objectives
- Design evidence collection around operational questions instead of collecting every available field.
- Choose when audit records, debug logs, metrics, traces/correlation IDs, or synthetic/user-journey checks are the primary source.
- Control metric label cardinality and log detail to balance diagnosability, privacy, storage, and performance.
- Separate GitLab application metrics from infrastructure, storage, network, and Runner metrics and assign clear ownership.
- Derive an SLO/alert plan from user journeys and baselines rather than arbitrary single thresholds.
1. Start with the operational decision, then choose evidence
Observability architecture fails when teams begin with “we have Prometheus, so scrape everything” or “ship every log forever.” Start with decisions: Can users clone? Can merge requests load? Are jobs starting within the expected time? Is Gitaly saturating? Who changed a protected resource? Can the instance accept traffic after maintenance?
Each question has a lowest-cost primary signal and supporting evidence. User-journey availability should not be inferred solely from process health. Security accountability should not be reconstructed from verbose debug logs when an audit event exists. High-frequency latency trends belong in metrics rather than millions of nearly identical log lines.
| Question | Primary evidence | Supporting evidence |
|---|---|---|
| Who changed governed project/group state? | Audit event | Related API/UI state; change ticket |
| Why did one HTTP request fail? | Correlation-linked structured logs | Request header, dependency metrics |
| Is the failure widespread or rising? | Metrics / SLI | Sampled logs, synthetic checks |
| Can the app and dependencies accept traffic? | Readiness / user journey | Component metrics and logs |
| Why are jobs pending? | Pipeline/job/runner state | Runner metrics/logs; tags/protection/config |
2. Audit evidence versus application debug logs
Audit events are curated governance records. Debug logs are implementation diagnostics. Increasing log verbosity can expose more fields and impose storage/CPU/IO costs; it also makes search harder when every request emits high-volume details. Audit events are more stable for accountability but do not contain every technical detail required for root-cause analysis.
Therefore, use audit events to establish who/what/when for supported GitLab-governed changes, and use logs/metrics to explain how the platform behaved. If an audit feature is tier-gated, a Free-tier organization can still maintain external change control/OS configuration audit records; the course does not pretend application debug logs are a perfect substitute.
| Design choice | Benefit | Cost / risk | Use when |
|---|---|---|---|
| Normal structured logs | Good operational detail, predictable volume | May omit deep internals | Default troubleshooting baseline |
| Temporary debug logging | More diagnostic detail | Sensitive data and volume risk; performance overhead | Time-boxed, scoped incident with approval |
| Audit event database/API | Governance-focused, queryable history | Paid scope for many events; query constraints | Access review, change accountability, compliance |
| External audit streaming | Near-real-time external retention/search | Ultimate; destination security and duplicate handling | Central SIEM/compliance architecture |
3. High detail and high cardinality can destroy a monitoring system
A metric label creates a new time series for every unique label combination. Labels such as bounded HTTP method or status are often manageable. Labels such as raw request ID, user ID, arbitrary project path, full URL, commit SHA, or error text can create unbounded series growth. Those identifiers belong in logs or tracing-style correlation, not as default Prometheus labels.
GOOD BOUNDED LABEL IDEA
http_requests_total{method="GET",status="500"}
DANGEROUS HIGH-CARDINALITY IDEA
http_requests_total{request_id="01J...",user="alice",project="every/project/path",sha="..."}
Rule of thumb: metrics answer aggregate questions; logs carry per-event identity.
Cardinality is not merely storage cost. It increases ingestion, memory, query, and dashboard complexity. A “more detailed metric” can make the monitoring platform less reliable during the incident when you need it most.
4. Platform metrics and external infrastructure metrics are complementary
GitLab application/exporter metrics show queues, request behavior, background jobs, and service-specific behavior. They cannot by themselves explain packet loss, disk firmware errors, cloud load-balancer health, filesystem saturation, host steal time, or object-storage service degradation. Conversely, a host CPU graph cannot tell you which Gitaly RPC is queueing or whether pipeline jobs are waiting for eligible runners.
flowchart TB UX[User journey / SLI] --> APP[GitLab application metrics] UX --> CI[Runner / pipeline metrics] APP --> DEP[DB / Redis / Gitaly / Registry evidence] CI --> HOST[Runner host / executor infrastructure] DEP --> INFRA[CPU / memory / disk / network / object storage] AUDIT[Audit / change evidence] --> APP CHANGE[OS / cloud change records] --> INFRA APP --> LOGS[Structured logs + correlation IDs] INFRA --> LOGS
Build dashboards that cross these boundaries explicitly. A repository latency panel should pair user-facing latency/error rate with Gitaly queue/drop metrics and host/storage saturation. That makes the dashboard diagnostic rather than decorative.
5. Reactive troubleshooting versus SLO- and alert-driven operations
A service-level indicator (SLI) is a measured behavior, such as successful clone requests divided by total clone requests or the fraction of web requests below a latency threshold. An SLO is the target for that indicator over a window. Alerts should represent meaningful risk to the objective, not every transient metric wiggle.
GitLab’s built-in metrics can feed such calculations, but your organization must define the user journeys and objective. Do not copy a threshold from a tutorial without a baseline. A 500 ms Gitaly operation can be normal for one workload and severe for another.
| Journey | Possible SLI | Supporting alert signal |
|---|---|---|
| Web/API requests | Success ratio + latency distribution | 5xx rate, dependency readiness, saturation |
| Git clone/fetch | Success ratio + completion latency | Gitaly queued/dropped requests, storage/network |
| CI scheduling | Time from job creation to runner pickup | Eligible runner count, request concurrency/errors |
| Background processing | Queue latency / success ratio | Sidekiq queue size, failures, Redis health |
6. Retention, aggregation, and privacy are one design problem
Retention should follow operational and governance needs. High-volume application logs may need shorter hot retention plus cheaper archive; audit evidence can require long retention; metrics may be downsampled; raw debug captures may have the shortest lifetime because of sensitivity. The product’s own retention statement does not define your external logging/SIEM policy.
For audit events, current GitLab documentation says stored events are retained indefinitely, but API/UI date-range/query limitations and export caps still matter. Instance CSV export on Self-Managed is currently capped at 100,000 events per export. For long-term external analysis, stream or export according to current tier/feature availability and design deduplication/retention in the destination.
id to deduplicate and
design the receiver to tolerate retries.
7. Prometheus topology: scrape narrowly and protect endpoints
For Linux-package Self-Managed GitLab, Prometheus and exporters are
integrated. GitLab application metrics can be served from the
Rails/Puma /-/metrics endpoint, and larger
installations can use a dedicated web exporter process for
fault/performance isolation. Runner exposes its own metrics HTTP
server when configured. These are separate targets and separate
network trust decisions.
scrape_configs:
- job_name: gitlab-app
static_configs:
- targets: ['gitlab-metrics.example.invalid:PORT'] # illustrative only
- job_name: gitlab-runner
static_configs:
- targets: ['runner-internal.example.invalid:9252']
# Keep targets on a restricted monitoring network.
# Do not copy endpoint/port choices blindly; use the documented exporter for your install.
The exact targets vary by installation method and scale. The important design rule is that monitoring reachability is intentional: localhost-only by default is safer than Internet exposure, and any wider reach must be controlled by firewall/network policy or an authenticated proxy.
8. Worked scenario: choose evidence for a 30-minute repository-latency SLO breach
Assume the repository browse/clone success SLI falls below target for 30 minutes. The Rails 5xx rate rises, all-dependency readiness intermittently reports Gitaly failure, Gitaly queue/dropped metrics rise, and host disk latency rises at the same time. Sidekiq and Runner metrics remain normal.
A good response is not “increase all retention and scrape every per-project metric.” Preserve a sample of correlation-linked failed requests, retain aggregate Gitaly/drop/disk metrics at useful resolution, and document the external storage change history. The evidence now spans user impact, application dependency, service saturation, infrastructure, and change context.
| Option | Maintainability | Security/privacy | Diagnostic value | Decision |
|---|---|---|---|---|
| Enable debug globally for a week | Poor | High exposure risk | High but noisy | Reject |
| Add request ID as Prometheus label | Poor | Identity/cardinality risk | Misuses metrics | Reject |
| Keep normal structured logs + correlation search | Good | Manageable with controls | High per-request value | Adopt |
| Alert on Gitaly queue/drop + user SLI | Good | Low sensitivity | High aggregate value | Adopt |
| Add storage latency panel from infra monitoring | Good | Low/moderate | Explains underlying pressure | Adopt |
9. Architecture decision table
| Decision | Free-compatible baseline | Optional advanced path | Guardrail |
|---|---|---|---|
| Governance evidence | Authentication log + local change records + fixtures | Premium/Ultimate audit dashboards/API; Ultimate streaming | Normalize time; protect PII; deduplicate streams |
| Application logs | Self-Managed structured logs | Central log platform | Minimize debug; redact stable identifiers carefully |
| Metrics | Self-Managed bundled Prometheus/exporters or existing monitoring | Federated/managed monitoring | Restrict endpoints; control cardinality |
| Runner monitoring | Runner embedded metrics on private interface | Central fleet dashboard | No public unauthenticated endpoint |
| Alerting | Baseline + user-journey SLI | Multi-window burn-rate/SLO tooling | Page on actionable impact, not arbitrary noise |
| Retention | Document purpose and minimum necessary data | Tiered hot/archive storage | Access control + deletion policy + legal/compliance review |
Knowledge check
Why is a request ID a poor default Prometheus label?
It is effectively unbounded and creates a new time series per request. Keep per-request identity in logs/correlation evidence, while metrics aggregate behavior.
When is an audit event preferable to a debug log for an access change?
When you need governance/accountability about who changed a supported GitLab resource. Debug logs are implementation detail and may be incomplete or too sensitive for that purpose.
Why should a repository SLO dashboard include infrastructure evidence?
GitLab application metrics can show dependency symptoms such as Gitaly queueing but host/storage/network metrics may explain the underlying resource pressure.
What does current GitLab documentation say about duplicate audit stream delivery?
The same event can be streamed more than once to a destination, so receivers should deduplicate using the event ID and tolerate retries.
What is wrong with choosing an alert threshold before establishing a baseline?
The threshold may page on normal workload or miss meaningful degradation. Alerts should be connected to known behavior, capacity, and user-impact objectives.
10. Lesson summary and bridge
- Collect evidence to support decisions, not because a telemetry source exists.
- Metrics aggregate behavior; logs carry per-event detail; audit events carry governance meaning; infrastructure monitoring explains resources outside GitLab application scope.
- Cardinality, retention, and debug verbosity have reliability, privacy, and cost consequences.
- SLOs and alerts should represent user journeys and actionable risk, with baselines and supporting component signals.
Next, you will use this design discipline to diagnose failure modes where operators commonly correlate the wrong events, trust the wrong health signal, or expose sensitive data while troubleshooting.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. External log/SIEM products are intentionally vendor-neutral in the required path; the lesson focuses on GitLab evidence semantics and secure integration boundaries.
- GitLab 19.3 release
- Audit events
- Audit events API
- Audit events administration and CSV export
- Audit event streaming for top-level groups
- Audit event streaming for instances
- GitLab log system
- Trace logs with a correlation ID
- Health checks
- Monitoring GitLab with Prometheus
- GitLab Prometheus metrics
- GitLab performance monitoring
- Monitor GitLab Runner usage
- Monitoring Gitaly
- GitLab Admin area monitoring
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.