Chapter 32Lesson 04~360 minutes

Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics: Diagnostics, Failure Modes, Security, and Performance

Diagnose realistic GitLab failures by preserving evidence, normalizing time and scope, following correlation IDs, separating liveness from dependency readiness, and validating corrective actions.

TroubleshootingGitalySidekiqRunnerFailure Analysis

Learning objectives

  • Apply a repeatable evidence-preserving diagnostic sequence to GitLab incidents.
  • Diagnose false correlation caused by time-zone mismatch or missing request IDs.
  • Explain how healthy web/liveness checks can coexist with degraded background jobs, repositories, or runners.
  • Recognize sensitive-data leakage caused by indiscriminate debug logging and evidence sharing.
  • Diagnose noisy or blind alerts caused by thresholds chosen without baseline, rate semantics, or user impact.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. All failure drills use synthetic outputs and are free-compatible. Commands that inspect real Self-Managed services are optional, read-only examples unless explicitly marked otherwise. Any logging-level change, service restart, metrics exposure change, or configuration edit is treated as production-sensitive and is not required.

1. Evidence-first diagnostic sequence

Use the same order under pressure: preserve evidence → define offering/instance/project/ref/pipeline/job/runner scope → normalize time → inspect health and user impact → follow correlation IDs or resource IDs → quantify with metrics → review audit/change history → rank hypotheses → make the least destructive correction → verify independently.

INCIDENT WORKSHEET
Impact:       which user journey, how many users/projects, first observed UTC
Scope:        offering / instance / namespace / project / ref / pipeline / job / runner
Evidence:     originals preserved? source host? collection time? version?
Identifiers:  request/correlation ID, pipeline/job ID, runner ID, project ID
Health:       /-/health vs readiness?all=1 vs actual user journey
Metrics:      error rate, latency, queue/saturation, resource pressure, runner state
Changes:      GitLab audit events + OS/cloud/config change records
Hypotheses:   evidence for / evidence against / next discriminating test
Correction:   owner, risk, rollback, expected signal change
Verify:       user journey + dependency + causal metric + no regression

2. Failure mode: time zones/request IDs are ignored

Operator A exports audit data in UTC. Operator B reads UI timestamps in local time. A log aggregator stores a mixture of UTC and local timestamps. They see a membership change “at 09:14” and a 502 “at 09:14” and declare the membership change causal.

Repair the analysis by converting all sources to an explicit UTC timeline and then looking for a stronger join key. If the 502 has x-request-id=01JINCIDENTA, search Workhorse/Rails/Gitaly for that value. If the audit event has no request/component link, keep it as an investigated but unproven change.

Diagnostic rule: timestamp proximity ranks a hypothesis; it does not prove it. Correlation ID, pipeline/job ID, runner identity, and exact resource scope provide stronger joins.

3. Failure mode: audit availability or export behavior is assumed

A runbook says “open project audit events” but the control is absent. The operator opens a permission ticket, assuming their role is wrong. Current GitLab behavior has several other explanations: project/group audit views and API are Premium/Ultimate; instance audit is offering/admin scoped; streaming is Ultimate; the UI location can move; or administrator policy can restrict a feature.

A second failure is to treat an export as unlimited. Instance audit CSV export on Self-Managed is currently capped at 100,000 events, while API requests constrain date windows and pagination behavior. A reliable runbook documents tier/offering/role plus query/export limits and provides an alternative fixture or external-retention path.

Symptom Check first Do not conclude
Audit page absent Tier + offering + minimum role + current UI docs “GitLab is broken”
API returns 403 Token identity + role + endpoint availability “Token authentication failed” without checking authorization
Export smaller than expected Filters + documented event/export cap “Events were deleted”
External stream has duplicates Receiver dedupe by event ID “GitLab emitted two independent changes”

4. Failure mode: debug logging exposes tokens or personal data

During an incident, someone enables broad debug logging and uploads the bundle to an issue. The diagnosis may improve while the incident severity gets worse because the bundle contains session identifiers, project paths, personal data, or credentials.

Containment order matters. If an actual credential is exposed, revoke/rotate it first. Restrict the evidence, identify who could access it, and create a sanitized derivative. Then reduce logging detail once the diagnostic window ends. Avoid commands that dump all environment variables, runner configuration with secrets, cookies, or request bodies.

BAD INCIDENT HABIT
- print all environment variables
- attach full runner config.toml
- copy browser Cookie header
- upload unreviewed /var/log/gitlab archive publicly

SAFER PATTERN
- capture only required fields
- preserve original privately
- rotate exposed credential immediately
- pseudonymize stable identifiers in a working copy
- document every redaction
- time-box increased log verbosity
Never teach “hide the leaked token in Git history/logs” as remediation. Once exposed, assume compromise and revoke/rotate first.

5. Failure mode: a healthy web endpoint hides a degraded component

The status page says green because /-/health returns GitLab OK. Users still cannot browse repositories. The simple endpoint intentionally does not test dependencies; this is expected semantics, not a false response.

Use all-dependency readiness and actual user-journey checks. If readiness identifies Gitaly, follow a failed request’s correlation ID into Gitaly logs and pair it with queue/drop/storage metrics. If readiness is green but CI jobs are pending, move to pipeline/Runner evidence instead of repeatedly probing Rails.

/-/health                 -> 200 GitLab OK
/-/readiness?all=1        -> 503, gitaly_check failed
repository page request   -> 502, x-request-id=01JINCIDENTA
Gitaly same request ID    -> dropped, reason=max_time
Gitaly queued metric      -> rising
Runner metrics            -> normal

Conclusion: app process alive; repository dependency degraded.

6. Failure mode: alert threshold is tuned without a baseline

An operator sees a queue depth of 5 during one incident and sets “queue > 5 = page.” On the next business day normal burst traffic pages the team repeatedly. They raise the threshold to 100; the next sustained degradation begins at 40 and no alert fires. Both settings are arbitrary.

Build a baseline over representative workload, alert on rates/increases for counters, and connect the alert to user impact or SLO risk. Use severity tiers: a small queue may be dashboard-only; sustained queue plus rising dropped requests and 5xx can page.

Signal combination Likely response
Queue briefly rises; no drops; latency normal Observe, no page
Queue sustained; drops rising; user latency/5xx rising Actionable incident alert
Runner errors rise; jobs pending; GitLab readiness healthy Investigate runner fleet/polling/executor
Health 200 only Insufficient for platform-wide “healthy” statement

7. Failure mode: metric semantics are misread

A raw counter such as gitaly_requests_dropped_total=18000 looks alarming, but it may represent months of history. The operational question is whether it is increasing now and at what rate. Similarly, a histogram’s bucket series should not be interpreted as a single latency value.

# Better: current drop activity.
sum(rate(gitaly_requests_dropped_total[5m])) by (reason)

# Better: compare queue saturation with user impact.
sum(gitaly_concurrency_limiting_queued)
sum(rate(http_requests_total{status=~"5.."}[5m]))

# Verify the exact metric names and labels available on your GitLab version.

8. Intentionally broken diagnosis: preserve the cause, then repair it

Broken diagnosis: “At 01:14:22Z a member-change audit event occurred. At 01:14:21Z users saw 502. Therefore the role change caused the outage. Restart all GitLab services.”

There are three faults. First, the audit event does not share a causal identifier with the request. Second, the restart destroys volatile evidence and changes many components at once. Third, it does not address the evidence that Rails spent almost all request time in Gitaly and Gitaly dropped the same correlation ID after queue wait.

Repaired diagnosis: preserve logs/metrics, trace 01JINCIDENTA, verify Gitaly dependency readiness and queue/drop rate, check host/storage resource pressure and recent Gitaly/infrastructure changes, then reduce load or restore known-good capacity/configuration according to the current runbook. Verify repository journey, readiness, queue/drop trend, and absence of new errors.

Broken versus evidence-first response
flowchart TD
  BAD[502 + nearby audit event] --> GUESS[Assume audit change caused outage]
  GUESS --> RESTART[Restart everything]
  RESTART --> LOST[Evidence lost / many variables changed]
  BAD --> PRESERVE[Preserve evidence]
  PRESERVE --> CID[Follow request ID]
  CID --> GITALY[Gitaly queue/drop evidence]
  GITALY --> CAP[Check capacity + recent config/load]
  CAP --> FIX[Least-destructive correction]
  FIX --> VERIFY[User journey + readiness + metrics]

9. Security, reliability, and performance are causal—not appendix warnings

Observability configuration itself consumes resources. Excessive log volume can pressure disk and ingestion systems. High-cardinality metrics can overload Prometheus. Public unauthenticated metrics endpoints expose operational data. Over-aggressive health-based eviction can amplify failures by removing usable nodes. Debug logging can reveal sensitive data.

Therefore every instrumentation change should have an owner, purpose, expected volume, access boundary, rollback, and success criterion. Performance and security are properties of the diagnostic system, not just of the application being diagnosed.

Knowledge check

What is wrong with correlating two events only because their local-clock timestamps match?

Why can /-/health be green during a Gitaly outage?

What should happen before sharing a log bundle that contains a real token?

Why is a lifetime counter value usually a poor alert condition?

What is the least-destructive next step after evidence points to Gitaly queue saturation?

10. Lesson summary and bridge

  • Normalize time and use strong identifiers before assigning causality.
  • Feature absence can be tier/offering/role/policy/version—not an outage.
  • Simple health, all-dependency readiness, component metrics, and user journeys must be interpreted together.
  • Debugging changes can leak secrets or create resource pressure; instrumentation needs governance and rollback.
  • The best correction changes the smallest justified state and is verified by the same causal evidence that motivated it.

The checkpoint combines these failure modes into one multi-source incident that you must diagnose, correct on paper/fixture, and hand off with sanitized evidence.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. Metric names in examples are representative of current documentation; production alerting must be tested against the exact exporter output from the deployed release.

Next lesson

Checkpoint Lab

Diagnose the complete synthetic incident, preserve evidence, reject unsupported hypotheses, and verify recovery.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.