Chapter 32Lesson 03~320 minutes

Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics: Configuration, Design Choices, and Tradeoffs

Design an observability and audit strategy that balances evidence depth, cardinality, privacy, storage cost, external monitoring, SLOs, and operational ownership.

Observability DesignSLOPrivacyCardinalityRetention

Learning objectives

  • Design evidence collection around operational questions instead of collecting every available field.
  • Choose when audit records, debug logs, metrics, traces/correlation IDs, or synthetic/user-journey checks are the primary source.
  • Control metric label cardinality and log detail to balance diagnosability, privacy, storage, and performance.
  • Separate GitLab application metrics from infrastructure, storage, network, and Runner metrics and assign clear ownership.
  • Derive an SLO/alert plan from user journeys and baselines rather than arbitrary single thresholds.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. The design exercises are free-compatible and can be completed from fixtures and official documentation. Premium/Ultimate audit features and external streaming/SIEM are optional architecture branches. Self-Managed metrics/log collection is available across tiers but requires administrator/operator access and secure infrastructure.

1. Start with the operational decision, then choose evidence

Observability architecture fails when teams begin with “we have Prometheus, so scrape everything” or “ship every log forever.” Start with decisions: Can users clone? Can merge requests load? Are jobs starting within the expected time? Is Gitaly saturating? Who changed a protected resource? Can the instance accept traffic after maintenance?

Each question has a lowest-cost primary signal and supporting evidence. User-journey availability should not be inferred solely from process health. Security accountability should not be reconstructed from verbose debug logs when an audit event exists. High-frequency latency trends belong in metrics rather than millions of nearly identical log lines.

Question Primary evidence Supporting evidence
Who changed governed project/group state? Audit event Related API/UI state; change ticket
Why did one HTTP request fail? Correlation-linked structured logs Request header, dependency metrics
Is the failure widespread or rising? Metrics / SLI Sampled logs, synthetic checks
Can the app and dependencies accept traffic? Readiness / user journey Component metrics and logs
Why are jobs pending? Pipeline/job/runner state Runner metrics/logs; tags/protection/config

2. Audit evidence versus application debug logs

Audit events are curated governance records. Debug logs are implementation diagnostics. Increasing log verbosity can expose more fields and impose storage/CPU/IO costs; it also makes search harder when every request emits high-volume details. Audit events are more stable for accountability but do not contain every technical detail required for root-cause analysis.

Therefore, use audit events to establish who/what/when for supported GitLab-governed changes, and use logs/metrics to explain how the platform behaved. If an audit feature is tier-gated, a Free-tier organization can still maintain external change control/OS configuration audit records; the course does not pretend application debug logs are a perfect substitute.

Design choice Benefit Cost / risk Use when
Normal structured logs Good operational detail, predictable volume May omit deep internals Default troubleshooting baseline
Temporary debug logging More diagnostic detail Sensitive data and volume risk; performance overhead Time-boxed, scoped incident with approval
Audit event database/API Governance-focused, queryable history Paid scope for many events; query constraints Access review, change accountability, compliance
External audit streaming Near-real-time external retention/search Ultimate; destination security and duplicate handling Central SIEM/compliance architecture

3. High detail and high cardinality can destroy a monitoring system

A metric label creates a new time series for every unique label combination. Labels such as bounded HTTP method or status are often manageable. Labels such as raw request ID, user ID, arbitrary project path, full URL, commit SHA, or error text can create unbounded series growth. Those identifiers belong in logs or tracing-style correlation, not as default Prometheus labels.

GOOD BOUNDED LABEL IDEA
http_requests_total{method="GET",status="500"}

DANGEROUS HIGH-CARDINALITY IDEA
http_requests_total{request_id="01J...",user="alice",project="every/project/path",sha="..."}

Rule of thumb: metrics answer aggregate questions; logs carry per-event identity.

Cardinality is not merely storage cost. It increases ingestion, memory, query, and dashboard complexity. A “more detailed metric” can make the monitoring platform less reliable during the incident when you need it most.

4. Platform metrics and external infrastructure metrics are complementary

GitLab application/exporter metrics show queues, request behavior, background jobs, and service-specific behavior. They cannot by themselves explain packet loss, disk firmware errors, cloud load-balancer health, filesystem saturation, host steal time, or object-storage service degradation. Conversely, a host CPU graph cannot tell you which Gitaly RPC is queueing or whether pipeline jobs are waiting for eligible runners.

Observability layers and ownership
flowchart TB
  UX[User journey / SLI] --> APP[GitLab application metrics]
  UX --> CI[Runner / pipeline metrics]
  APP --> DEP[DB / Redis / Gitaly / Registry evidence]
  CI --> HOST[Runner host / executor infrastructure]
  DEP --> INFRA[CPU / memory / disk / network / object storage]
  AUDIT[Audit / change evidence] --> APP
  CHANGE[OS / cloud change records] --> INFRA
  APP --> LOGS[Structured logs + correlation IDs]
  INFRA --> LOGS

Build dashboards that cross these boundaries explicitly. A repository latency panel should pair user-facing latency/error rate with Gitaly queue/drop metrics and host/storage saturation. That makes the dashboard diagnostic rather than decorative.

5. Reactive troubleshooting versus SLO- and alert-driven operations

A service-level indicator (SLI) is a measured behavior, such as successful clone requests divided by total clone requests or the fraction of web requests below a latency threshold. An SLO is the target for that indicator over a window. Alerts should represent meaningful risk to the objective, not every transient metric wiggle.

GitLab’s built-in metrics can feed such calculations, but your organization must define the user journeys and objective. Do not copy a threshold from a tutorial without a baseline. A 500 ms Gitaly operation can be normal for one workload and severe for another.

Journey Possible SLI Supporting alert signal
Web/API requests Success ratio + latency distribution 5xx rate, dependency readiness, saturation
Git clone/fetch Success ratio + completion latency Gitaly queued/dropped requests, storage/network
CI scheduling Time from job creation to runner pickup Eligible runner count, request concurrency/errors
Background processing Queue latency / success ratio Sidekiq queue size, failures, Redis health

6. Retention, aggregation, and privacy are one design problem

Retention should follow operational and governance needs. High-volume application logs may need shorter hot retention plus cheaper archive; audit evidence can require long retention; metrics may be downsampled; raw debug captures may have the shortest lifetime because of sensitivity. The product’s own retention statement does not define your external logging/SIEM policy.

For audit events, current GitLab documentation says stored events are retained indefinitely, but API/UI date-range/query limitations and export caps still matter. Instance CSV export on Self-Managed is currently capped at 100,000 events per export. For long-term external analysis, stream or export according to current tier/feature availability and design deduplication/retention in the destination.

Streaming is not exactly-once delivery. GitLab documents that an audit event can be streamed more than once to the same destination; use the event id to deduplicate and design the receiver to tolerate retries.

7. Prometheus topology: scrape narrowly and protect endpoints

For Linux-package Self-Managed GitLab, Prometheus and exporters are integrated. GitLab application metrics can be served from the Rails/Puma /-/metrics endpoint, and larger installations can use a dedicated web exporter process for fault/performance isolation. Runner exposes its own metrics HTTP server when configured. These are separate targets and separate network trust decisions.

scrape_configs:
  - job_name: gitlab-app
    static_configs:
      - targets: ['gitlab-metrics.example.invalid:PORT']  # illustrative only

  - job_name: gitlab-runner
    static_configs:
      - targets: ['runner-internal.example.invalid:9252']
    # Keep targets on a restricted monitoring network.

# Do not copy endpoint/port choices blindly; use the documented exporter for your install.

The exact targets vary by installation method and scale. The important design rule is that monitoring reachability is intentional: localhost-only by default is safer than Internet exposure, and any wider reach must be controlled by firewall/network policy or an authenticated proxy.

8. Worked scenario: choose evidence for a 30-minute repository-latency SLO breach

Assume the repository browse/clone success SLI falls below target for 30 minutes. The Rails 5xx rate rises, all-dependency readiness intermittently reports Gitaly failure, Gitaly queue/dropped metrics rise, and host disk latency rises at the same time. Sidekiq and Runner metrics remain normal.

A good response is not “increase all retention and scrape every per-project metric.” Preserve a sample of correlation-linked failed requests, retain aggregate Gitaly/drop/disk metrics at useful resolution, and document the external storage change history. The evidence now spans user impact, application dependency, service saturation, infrastructure, and change context.

Option Maintainability Security/privacy Diagnostic value Decision
Enable debug globally for a week Poor High exposure risk High but noisy Reject
Add request ID as Prometheus label Poor Identity/cardinality risk Misuses metrics Reject
Keep normal structured logs + correlation search Good Manageable with controls High per-request value Adopt
Alert on Gitaly queue/drop + user SLI Good Low sensitivity High aggregate value Adopt
Add storage latency panel from infra monitoring Good Low/moderate Explains underlying pressure Adopt

9. Architecture decision table

Decision Free-compatible baseline Optional advanced path Guardrail
Governance evidence Authentication log + local change records + fixtures Premium/Ultimate audit dashboards/API; Ultimate streaming Normalize time; protect PII; deduplicate streams
Application logs Self-Managed structured logs Central log platform Minimize debug; redact stable identifiers carefully
Metrics Self-Managed bundled Prometheus/exporters or existing monitoring Federated/managed monitoring Restrict endpoints; control cardinality
Runner monitoring Runner embedded metrics on private interface Central fleet dashboard No public unauthenticated endpoint
Alerting Baseline + user-journey SLI Multi-window burn-rate/SLO tooling Page on actionable impact, not arbitrary noise
Retention Document purpose and minimum necessary data Tiered hot/archive storage Access control + deletion policy + legal/compliance review

Knowledge check

Why is a request ID a poor default Prometheus label?

When is an audit event preferable to a debug log for an access change?

Why should a repository SLO dashboard include infrastructure evidence?

What does current GitLab documentation say about duplicate audit stream delivery?

What is wrong with choosing an alert threshold before establishing a baseline?

10. Lesson summary and bridge

  • Collect evidence to support decisions, not because a telemetry source exists.
  • Metrics aggregate behavior; logs carry per-event detail; audit events carry governance meaning; infrastructure monitoring explains resources outside GitLab application scope.
  • Cardinality, retention, and debug verbosity have reliability, privacy, and cost consequences.
  • SLOs and alerts should represent user journeys and actionable risk, with baselines and supporting component signals.

Next, you will use this design discipline to diagnose failure modes where operators commonly correlate the wrong events, trust the wrong health signal, or expose sensitive data while troubleshooting.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. External log/SIEM products are intentionally vendor-neutral in the required path; the lesson focuses on GitLab evidence semantics and secure integration boundaries.

Next lesson

Diagnostics, Failure Modes, Security, and Performance

Diagnose false correlation, misleading health signals, unsafe debug logging, and poorly tuned alerts.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.