Chapter 32Lesson 02~390 minutes

Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics: Guided Hands-On Workflow and Core Operations

Practice a free-compatible diagnostic workflow with synthetic audit, log, health, GitLab, Gitaly, Sidekiq, and Runner evidence, plus optional read-only checks on an isolated Self-Managed instance.

FixturesPromQLCorrelation IDRunbookVerification

Learning objectives

  • Generate a local incident fixture that requires no GitLab instance, paid tier, runner, or cloud service.
  • Inspect synthetic audit events, Rails/Workhorse/Gitaly/Sidekiq logs, health responses, and Prometheus samples without modifying state.
  • Correlate a request across components by request ID and test hypotheses against metrics and readiness evidence.
  • Write a minimal Prometheus query/dashboard plan that measures impact, saturation, and recovery rather than collecting every metric.
  • Execute a structured runbook from symptom to evidence to correction proposal to independently verified recovery.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. The mandatory workflow is an offline fixture and works with Python 3 plus a browser/editor. Optional live checks are read-only and apply only to an isolated Self-Managed instance you are authorized to inspect. Paid audit-event views/APIs, external SIEM, and live Prometheus/Runner configuration are not required.

1. Disposable scenario and preflight

The incident has two reported symptoms: repository pages intermittently return 502/timeout, and a team says pipelines “feel slow.” Your job is to determine which symptom is supported by the evidence and which is merely concurrent noise. The fixture intentionally contains an unrelated audit event and normal Runner metrics so that “everything near the timestamp is causal” fails.

Create an empty directory such as gitlab-ch32-lab. Do not place real production logs in it. The fixture generator below writes only synthetic data with fake project/user identifiers.

Preflight: use a disposable directory. If you later substitute real operational evidence, preserve originals outside the working directory, restrict access, and sanitize before sharing.
from pathlib import Path
import json

root = Path("gitlab-ch32-lab")
root.mkdir(exist_ok=True)

def write(name, text):
    (root / name).write_text(text.strip() + "\n", encoding="utf-8")

# Paid-tier audit data is represented as a fixture; event_type is explicitly training-only.
audit = {
  "id": 4102, "author_id": 42, "entity_type": "Project",
  "entity_path": "learner-example/platform-lab",
  "target_type": "Member", "target_details": "USER_B",
  "created_at": "2026-08-22T01:14:22Z",
  "event_type": "TRAINING_FIXTURE_MEMBER_CHANGE"
}
write("audit_events.jsonl", json.dumps(audit))

cid = "01JINCIDENTA000000000000000"
write("workhorse.log", json.dumps({
  "time":"2026-08-22T01:14:21.906Z", "status":502,
  "duration_ms":4912, "uri":"/learner-example/platform-lab/-/tree/main",
  "correlation_id":cid
}))
write("rails.log", json.dumps({
  "time":"2026-08-22T01:14:21.914Z", "status":502,
  "duration":4888.2, "db":18.4, "gitaly_duration":4862.1,
  "path":"/learner-example/platform-lab/-/tree/main",
  "correlation_id":cid
}))
write("gitaly.log", json.dumps({
  "time":"2026-08-22T01:14:26.780Z", "grpc.code":"Unavailable",
  "limit.concurrency_dropped":"max_time", "limit.concurrency_queue_length":27,
  "grpc.method":"FindCommit", "correlation_id":cid
}))
write("sidekiq.log", json.dumps({
  "time":"2026-08-22T01:14:25.000Z", "job_status":"done",
  "queue_duration_s":0.18, "class":"TrainingFixtureWorker"
}))
write("health.txt", "200 GitLab OK")
write("readiness_all.json", json.dumps({
  "http_status":503, "status":"failed",
  "gitaly_check":[{"status":"failed","message":"synthetic timeout"}],
  "db_check":[{"status":"ok"}], "redis_check":[{"status":"ok"}]
}, indent=2))
write("metrics.prom", """
gitaly_concurrency_limiting_queued 27
gitaly_requests_dropped_total{reason="max_time"} 18
gitaly_requests_dropped_total{reason="max_size"} 0
""")
write("runner_metrics.prom", """
gitlab_runner_concurrent 4
gitlab_runner_errors_total 0
gitlab_runner_request_concurrency 1
""")
print(root.resolve())

2. Inventory evidence before interpreting it

Start with filenames, timestamps, and source semantics—not grep patterns. Record whether each source is original or derived, what system produced it, its time zone, and whether it can contain sensitive fields. In a real incident, also record the collection time and hash if chain-of-custody matters.

from pathlib import Path
root = Path("gitlab-ch32-lab")
for p in sorted(root.iterdir()):
    print(f"{p.name:24} {p.stat().st_size:6} bytes")

Expected evidence includes one audit fixture, four log sources, two health snapshots, GitLab/Gitaly-style metrics, and Runner metrics. Nothing has been changed yet.

3. Normalize the timeline and use the correlation ID as the primary join key

Now parse only the fields needed for the hypothesis. Do not dump every log field to the terminal. The request begins at Workhorse, passes through Rails, and ends with a Gitaly concurrency drop. The shared correlation ID makes that chain much stronger than timestamp proximity.

from pathlib import Path
import json
root = Path("gitlab-ch32-lab")
for name in ["workhorse.log", "rails.log", "gitaly.log", "sidekiq.log"]:
    record = json.loads((root / name).read_text(encoding="utf-8"))
    print(name, record.get("time"), record.get("correlation_id", "NO_REQUEST_ID"),
          record.get("status", record.get("grpc.code", record.get("job_status"))))
workhorse.log 2026-08-22T01:14:21.906Z 01JINCIDENTA000000000000000 502
rails.log     2026-08-22T01:14:21.914Z 01JINCIDENTA000000000000000 502
gitaly.log    2026-08-22T01:14:26.780Z 01JINCIDENTA000000000000000 Unavailable
sidekiq.log   2026-08-22T01:14:25.000Z NO_REQUEST_ID done

The Sidekiq record occurs during the same interval but has no shared request ID and reports normal completion. It is evidence against a platform-wide background-job failure for this specific request.

4. Reconcile apparently contradictory health evidence

health.txt is green while readiness_all.json is 503. That is not inconsistent: the simple health endpoint proves only that the application server responds. The all-dependencies readiness fixture isolates Gitaly while DB and Redis remain healthy.

from pathlib import Path
import json
root = Path("gitlab-ch32-lab")
print((root / "health.txt").read_text().strip())
ready = json.loads((root / "readiness_all.json").read_text())
print("readiness HTTP", ready["http_status"])
for key in ("gitaly_check", "db_check", "redis_check"):
    print(key, ready[key][0]["status"])
200 GitLab OK
readiness HTTP 503
gitaly_check failed
db_check ok
redis_check ok

5. Use metrics to test whether the Gitaly failure is isolated or a saturation pattern

One failed RPC could be a transient request. The metric fixture adds fleet/service context: 27 requests are queued and the drop counter has increased for max_time. That pattern is consistent with sustained concurrency pressure rather than a single malformed project request.

# On a real Prometheus server, use rates/increases for counters.
sum(gitaly_concurrency_limiting_queued)
sum(rate(gitaly_requests_dropped_total[5m])) by (reason)

# Keep runner evidence separate from GitLab application/Gitaly evidence.
sum(gitlab_runner_request_concurrency)
sum(rate(gitlab_runner_errors_total[5m]))
Panel / query Purpose Interpretation in fixture
Gitaly queued work Saturation indicator High: 27 queued
Gitaly dropped rate by reason Failure impact Drop reason points to queue wait timeout
Request 5xx rate / latency User impact Would confirm whether failures are widespread
Runner errors / request concurrency CI execution health Fixture is normal; do not blame Runner
Readiness dependency status Immediate dependency gate Gitaly failed; DB/Redis healthy

6. Investigate the nearby audit event without declaring it causal

The audit event happened less than a second after the failing request began. That makes it worth investigating, not guilty. Its target is a project member. The failing request, readiness result, and Gitaly logs instead point to repository service saturation. Record the event as a checked hypothesis and reject it unless additional evidence connects the role change to the failure.

from pathlib import Path
import json
p = Path("gitlab-ch32-lab/audit_events.jsonl")
event = json.loads(p.read_text())
print(event["created_at"], event["entity_path"], event["target_type"], event["event_type"])
print("Conclusion: temporally adjacent; no causal request/component link in the fixture")

7. Write the structured troubleshooting runbook

A useful runbook contains decision points. For this fixture, the conclusion should be: repository traffic is degraded by Gitaly concurrency saturation; Rails and Workhorse expose the symptom, all-dependency readiness identifies Gitaly, Gitaly logs show the drop reason, metrics show queue pressure, and Runner/Sidekiq evidence does not support a platform-wide CI/background failure.

Runbook phase Action Exit criterion
Preserve Copy original evidence read-only; record source/time/version Original bundle remains unchanged
Scope Project path, request ID, UTC window, user journey One reproducible affected operation
Differentiate Health vs readiness; Rails vs Gitaly; Runner vs app Competing hypotheses narrowed
Quantify Queue, drop rate, 5xx/latency, runner errors Impact and saturation measured
Correct Least-destructive capacity/load/limit action per current runbook Change has owner and rollback
Verify Repeat request/readiness and observe queue/drop trend User journey and causal metrics recover
Document Timeline, cause, rejected hypotheses, prevention Another operator can reproduce reasoning
Do not “fix” this fixture by blindly increasing concurrency. A higher limit can move saturation from Gitaly queueing into CPU, memory, disk, or downstream storage pressure. Real changes require baseline/load evidence and the current Gitaly guidance.

8. Optional read-only live preflight on an isolated Self-Managed instance

If you administer a disposable Self-Managed instance, you may compare the fixture workflow with real read-only evidence. Do not enable debug logging, alter Prometheus exposure, restart services, or change runner settings for this optional step.

# Bash/Linux-package examples; read-only.
sudo gitlab-ctl status
curl -i https://gitlab.example.invalid/-/health
curl -i 'https://gitlab.example.invalid/-/readiness?all=1'

# Capture only the response header needed for correlation.
curl -sS -D - -o /dev/null https://gitlab.example.invalid/users/sign_in   | grep -i '^x-request-id:'

# After identifying one request ID, search specifically for it.
CID="01JFAKECORRELATIONID000000"
find /var/log/gitlab -type f -mtime 0 -exec grep -F "$CID" '{}' '+' 2>/dev/null

9. Small challenge: choose the next evidence source

Suppose /-/health is 200, /-/readiness?all=1 shows all dependencies healthy, and jobs are pending with no runner assigned. Which evidence should you inspect next? The best next step is Runner/project CI state—runner availability/tags/protection and Runner metrics/logs—not Gitaly logs. The resource boundary changed, so the evidence source must change with it.

Suppose instead jobs are running but cloning the repository inside the job fails with a correlation ID and Gitaly Unavailable. Now the Runner can be healthy while the GitLab/Gitaly path is degraded. “Pipeline problem” does not automatically mean “Runner problem.”

10. Cleanup and retained evidence

The mandatory lab created only synthetic files. Review the directory, retain your incident summary if you want it for the checkpoint, then delete the generated fixture directory. In a real incident, cleanup is governed by evidence-retention policy: do not destroy originals merely because the service recovered.

from pathlib import Path
import shutil
root = Path("gitlab-ch32-lab")
# Review the path before deletion. This directory must contain only the synthetic fixture.
print(root.resolve())
# shutil.rmtree(root)  # Uncomment only after verifying the disposable path.

Knowledge check

Why does the fixture include a nearby audit event that is not causal?

The Runner metrics are normal while a pipeline-related complaint exists. What does that tell you?

Why should you alert on a rate/increase of a counter rather than its raw lifetime value?

What is the safe interpretation of a 200 from /-/health and 503 from all-dependency readiness?

Why is “increase Gitaly concurrency” not an automatic fix for a queue?

11. Lesson summary and bridge

  • The fixture workflow preserves evidence before interpretation and uses request ID, component scope, and UTC time as join keys.
  • Health, readiness, logs, metrics, and audit events can disagree without contradiction because they answer different questions.
  • Normal Runner/Sidekiq evidence is useful negative evidence; it prevents blaming every platform component at once.
  • A diagnostic dashboard should measure user impact, saturation, failures, and recovery—not maximize metric count.

Next, you will turn these mechanics into an observability design: what to retain, what to alert on, where to aggregate it, and how to balance detail against privacy and cost.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. The fixture uses synthetic event types and samples; treat official current metric/event names and endpoint behavior as the production contract.

Next lesson

Configuration, Design Choices, and Tradeoffs

Turn the workflow into an observability design for retention, cardinality, privacy, SLOs, and operational ownership.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.