Chapter 32Lesson 01~330 minutes

Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics: Concepts, Architecture, and Mental Model

Build an evidence-first mental model that separates audit history, logs, metrics, health probes, and user-visible symptoms, then correlates them without confusing coincidence with causality.

Audit EventsLogsMetricsHealthEvidence

Learning objectives

  • Distinguish audit events, application/service logs, metrics, health probes, and user-visible symptoms by the question each evidence type can answer.
  • Map Rails/Puma, Sidekiq, Workhorse, Gitaly, NGINX or proxy, Runner, database, Redis, and storage evidence to their service boundaries.
  • Explain how GitLab correlation IDs connect one request across multiple Self-Managed component logs.
  • Interpret /-/health, /health_check, /-/readiness, and /-/liveness without treating one successful probe as whole-platform health.
  • Build a safe evidence-first troubleshooting loop that preserves originals, normalizes time, tests hypotheses, and verifies state after a change.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. Successful sign-in audit events are available at all tiers, while project/group/instance audit-event dashboards and audit-event APIs are currently Premium/Ultimate. External audit-event streaming is currently Ultimate. Self-Managed logs, health probes, GitLab Prometheus metrics, and Runner Prometheus metrics have Free-compatible paths. The required exercises use synthetic fixtures, so no paid tier, production instance, external SIEM, or runner fleet is required.

1. The practical problem: incidents produce many facts, but facts are not yet a diagnosis

A user reports, “GitLab is down.” That sentence is a symptom, not a component diagnosis. The browser might be unable to load a repository page while the Rails process is alive. A pipeline can remain pending while the web application is healthy. A push can time out because Gitaly is saturated even though PostgreSQL and Redis are fine. An audit event can show that a member role changed at the same minute without proving that the role change caused a storage timeout.

Operational discipline begins by asking what kind of evidence would falsify or support a hypothesis. Audit events answer governance questions such as who changed a GitLab-governed resource and when. Logs explain individual requests, jobs, RPCs, and service errors. Metrics summarize behavior over time. Health probes answer narrow reachability/readiness questions. The user-visible symptom defines impact. None of these should silently replace the others.

Evidence Best question Typical scope Common misuse
Audit event Who changed governed state, what was targeted, and when? User, project, group, instance; availability varies Treating an event near the incident as proof of causation
Structured log What happened in this request/job/RPC? Rails, Sidekiq, Workhorse, Gitaly, proxy, Runner, etc. Searching every log without first defining time/scope
Metric Is a rate, latency, queue, error, or resource signal abnormal over time? Service/process/fleet/system Alerting on an unbaselined threshold or high-cardinality labels
Health probe Can this component/process accept a narrow class of work now? Application and dependencies depending on endpoint Calling a green probe “the whole platform is healthy”
User symptom What capability is degraded and for whom? Workflow, project, endpoint, pipeline, Git operation Assuming the complaint identifies the failing service

2. Mental model: symptom → scope → evidence → hypothesis → controlled change → verification

A strong runbook is a loop rather than a pile of commands. First define the affected offering, instance, namespace/project, Git ref or pipeline/job, time window, and user-visible operation. Preserve evidence before restarts or configuration edits rotate it away. Then gather the smallest set of sources that can distinguish competing hypotheses.

Evidence-driven operational loop
flowchart TD
  S[User-visible symptom] --> C[Define scope + time window]
  C --> P[Preserve original evidence]
  P --> A[Audit events]
  P --> L[Logs + correlation IDs]
  P --> M[Metrics / trends]
  P --> H[Health / readiness]
  A --> X[Hypothesis]
  L --> X
  M --> X
  H --> X
  X --> T[Least-destructive test or correction]
  T --> V[Independent verification]
  V -->|not fixed| C
  V -->|fixed| R[Timeline + root cause + prevention]

The arrows matter. You do not jump from “502” directly to “restart Gitaly.” You narrow scope, preserve the evidence that can disappear, and then change state only after the current evidence points to a component or dependency.

3. Audit events: governance evidence, not debug traces

GitLab audit events are designed to record security- and governance-relevant actions. A typical event contains an identity, entity/target, timestamp, IP or contextual details, and an event type. The current documentation states that audit events are retained indefinitely in GitLab, but the UI and API constrain how you query them: project/group/instance audit queries use bounded date windows, and external exports or streams have their own limits and retention policies. “Retained” and “easy to query in one request” are different statements.

Availability is equally important. Successful sign-in audit information is the Free-tier path. Project and group audit-event views and the audit-event API are currently Premium/Ultimate, and instance audit events are administrative surfaces on Self-Managed/Dedicated. Audit-event streaming is currently Ultimate. A missing audit page can therefore mean tier, role, offering, or policy—not necessarily a broken installation.

Time-zone trap: current GitLab documentation notes that the UI presents audit times in the viewer’s local time, the API returns UTC by default (or the configured Self-Managed time zone), and CSV exports use UTC. Normalize all incident evidence to one timeline before correlating events.
{
  "id": 4102,
  "author_id": 42,
  "entity_type": "Project",
  "entity_path": "learner-example/platform-lab",
  "target_type": "Member",
  "target_details": "fixture-user",
  "created_at": "2026-08-22T01:14:22Z",
  "event_type": "TRAINING_FIXTURE_EVENT",
  "details": {"custom_message": "Synthetic fixture: membership changed"}
}

The event above intentionally uses a training-only event type. The schema illustrates the fields you correlate; automation must use the current documented event types from the exact GitLab version rather than a course fixture string.

4. Logs: follow the request across service boundaries

A Linux-package Self-Managed installation can produce distinct logs for GitLab Rails, Puma, Sidekiq, Gitaly, Workhorse, NGINX, Registry, PostgreSQL, Redis and other services. Runner logs are produced by the Runner process, which may live on another host entirely. Each source sees a different part of the transaction.

For most HTTP requests, GitLab uses a correlation ID. It appears in structured logs as correlation_id and in response headers as x-request-id. That ID is a join key: the Workhorse access record, Rails request record, and Gitaly RPC record can describe the same request without relying only on timestamp proximity.

Component evidence What it can reveal Example question
Workhorse / proxy HTTP status, duration, URI, upstream behavior Did the request reach GitLab and how long did the edge path take?
Rails / Puma Controller/action, DB/Gitaly time, status, request metadata Did application processing spend time in DB or Gitaly?
Sidekiq Background job execution/retries/failures/latency Is asynchronous work delayed or failing?
Gitaly Git RPC timing, repository storage, concurrency limiting Are Git operations queued, slow, or rejected?
Runner Job polling/execution/configuration failures Why are CI jobs waiting or failing on execution infrastructure?
# OPTIONAL, read-only, Linux-package Self-Managed example.
# Do not run on a production node without your operational procedure.
sudo gitlab-ctl status
sudo gitlab-ctl tail gitlab-rails/production_json.log
sudo gitlab-ctl tail gitlab-workhorse
sudo gitlab-ctl tail gitaly

# Once you have a correlation ID, search narrowly instead of dumping all logs.
CORRELATION_ID="01JFAKECORRELATIONID000000"
find /var/log/gitlab -type f -mtime 0 -exec grep -F "$CORRELATION_ID" '{}' '+' 2>/dev/null

5. Health, comprehensive checks, readiness, and liveness answer different questions

GitLab’s Self-Managed health endpoints deliberately have different semantics. /-/health confirms that the application server is running; it does not prove that database or other dependencies work. /health_check performs broader backend checks and is intended for diagnostics/monitoring, not load-balancer eviction. /-/readiness indicates whether Rails is ready to accept traffic; with ?all=1 it checks dependencies such as database, Redis, and Gitaly. /-/liveness is meant to detect a stuck application process.

Endpoint What a success proves What it does not prove
/-/health Application server responds Dependencies, Git storage, background processing, CI runner health
/health_check Configured backend checks pass That it is safe to use as a load-balancer/autoscaling probe
/-/readiness Rails readiness check passes All dependencies unless ?all=1 is requested
/-/readiness?all=1 Readiness plus reported dependencies pass User journeys, external services, runner fleet, every node in a distributed topology
/-/liveness Application server is not judged deadlocked End-to-end service health
# Illustrative read-only requests. Replace the .invalid host only in an isolated lab.
curl -i https://gitlab.example.invalid/-/health
curl -i 'https://gitlab.example.invalid/-/readiness?all=1'
curl -i https://gitlab.example.invalid/-/liveness

# Expected shapes include plain "GitLab OK" for /-/health and JSON/status codes
# for readiness/liveness. A readiness failure returns HTTP 503.

6. Metrics and Prometheus: trends, rates, saturation, and capacity evidence

Prometheus-style metrics are time-series samples. They let you ask whether a condition is rising, falling, persistent, or correlated with load. On Self-Managed GitLab, the application exposes Prometheus metrics and Linux-package installations bundle Prometheus/exporters by default. Access to monitoring endpoints is an administrative/network decision; GitLab documents an IP allowlist for these resources. Runner has its own embedded metrics server, disabled until a listen address is configured, and its metrics endpoint must not be exposed publicly because it has no built-in authorization.

Metric interpretation starts with type. A counter only rises (until process restart), so alert on a rate or increase. A gauge represents current state such as queue depth. A histogram supports latency distributions/quantiles. Labels make metrics useful but can make them expensive when values have unbounded cardinality.

# Rate of Gitaly requests dropped by concurrency limiting.
sum(rate(gitaly_requests_dropped_total[5m])) by (reason)

# Current queued Gitaly work (metric names/labels can change; verify your version).
sum(gitaly_concurrency_limiting_queued)

# Runner API errors over five minutes.
sum(rate(gitlab_runner_errors_total[5m]))
Security boundary: do not publish Prometheus or Runner metrics endpoints to the Internet merely to make scraping easy. Restrict network access or put an authenticated proxy in front of them according to your architecture. Metrics can disclose topology, version, workload, and operational state.

7. Correlation is an evidence join, not “same minute = same cause”

Correlate with the strongest identifiers available: correlation/request ID for one HTTP request, pipeline/job ID for CI, project path/ID for resource scope, runner identity for execution, and UTC timestamps for sequence. When identifiers are absent, time correlation is weaker evidence and must be treated as such.

09:14:21.902Z  user sees 502 for repository page
09:14:21.903Z  x-request-id = 01JINCIDENTA
09:14:21.906Z  Workhorse status=502 correlation_id=01JINCIDENTA
09:14:21.914Z  Rails gitaly_duration=4.87s correlation_id=01JINCIDENTA
09:14:26.780Z  Gitaly dropped request reason=max_time correlation_id=01JINCIDENTA
09:14:30.000Z  metric: gitaly queued=27, dropped_total increases
09:14:22.000Z  audit: unrelated member role change

The audit event is close in time, but the request ID and Gitaly evidence form a much stronger causal chain for the 502. A disciplined incident report can explicitly record the audit event as investigated and rejected as causal.

8. Evidence itself is sensitive operational data

Logs and audit records can contain usernames, project paths, IP addresses, query parameters, user agents, identifiers, error messages, and occasionally data that should not leave the operational boundary. Debug logging can increase this exposure. Treat “turn on debug and upload everything” as a security-sensitive action, not a default troubleshooting step.

Preserve the original privately, create a sanitized working copy, and document every redaction. Do not replace all identifiers with the same string because you may destroy correlation. Use stable pseudonyms such as USER_A, PROJECT_X, and IP_REDACTED_1 so relationships remain analyzable.

Credential rule: if evidence reveals a real token, password, private key, webhook secret, session cookie, or cloud credential, containment starts with revoke/rotate/disable. Redacting the evidence or deleting a log does not make an exposed credential safe.

9. Read-only inspection before a configuration change

Before changing logging level, metrics exposure, runner configuration, or application settings, capture current version, service state, endpoint state, and the exact symptom. The goal is a reproducible “before” snapshot. In production this also supports rollback: you know what you changed and what evidence justified it.

# OPTIONAL read-only Self-Managed preflight. Commands are Bash/Linux-package specific.
sudo gitlab-rake gitlab:env:info
sudo gitlab-ctl status
curl -fsS https://gitlab.example.invalid/-/health
curl -fsS 'https://gitlab.example.invalid/-/readiness?all=1'

# Inspect headers for the request ID without printing cookies/tokens.
curl -sS -D - -o /dev/null https://gitlab.example.invalid/users/sign_in   | grep -i '^x-request-id:'

# Never use `env`, `set`, or blanket secret dumps as an incident shortcut.

Knowledge check

A project page is slow but /-/health returns GitLab OK. What can you conclude?

Why is a correlation ID stronger evidence than “these two log lines happened in the same second”?

Why can a Free-tier learner not rely on project audit-event API access for the required lab?

A Runner metrics endpoint works on 0.0.0.0:9252. Is that automatically a good production configuration?

What is the first operational action if a real token appears in a captured log bundle?

10. Lesson summary and bridge

  • Audit events establish governance history; logs explain individual execution; metrics show trends; health probes answer narrow readiness/liveness questions; symptoms define user impact.
  • Correlation IDs, resource IDs, pipeline/job IDs, and normalized timestamps turn disconnected evidence into a defensible incident timeline.
  • A successful simple health endpoint is not proof that dependencies, background jobs, repositories, runners, or user journeys are healthy.
  • Observability data is sensitive: preserve originals privately, work from sanitized copies, and rotate any exposed credential.

Next, you will build and analyze a disposable evidence bundle so the model becomes an executable troubleshooting workflow rather than a diagram.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. Audit-event tier/role behavior and monitoring endpoint exposure are especially version-sensitive; re-check the exact deployed version before using these controls in production.

Next lesson

Guided Hands-On Workflow and Core Operations

Generate and analyze a disposable evidence bundle, correlate a failed request, and write a structured troubleshooting runbook.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.