Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics: Guided Hands-On Workflow and Core Operations
Practice a free-compatible diagnostic workflow with synthetic audit, log, health, GitLab, Gitaly, Sidekiq, and Runner evidence, plus optional read-only checks on an isolated Self-Managed instance.
Learning objectives
- Generate a local incident fixture that requires no GitLab instance, paid tier, runner, or cloud service.
- Inspect synthetic audit events, Rails/Workhorse/Gitaly/Sidekiq logs, health responses, and Prometheus samples without modifying state.
- Correlate a request across components by request ID and test hypotheses against metrics and readiness evidence.
- Write a minimal Prometheus query/dashboard plan that measures impact, saturation, and recovery rather than collecting every metric.
- Execute a structured runbook from symptom to evidence to correction proposal to independently verified recovery.
1. Disposable scenario and preflight
The incident has two reported symptoms: repository pages intermittently return 502/timeout, and a team says pipelines “feel slow.” Your job is to determine which symptom is supported by the evidence and which is merely concurrent noise. The fixture intentionally contains an unrelated audit event and normal Runner metrics so that “everything near the timestamp is causal” fails.
Create an empty directory such as gitlab-ch32-lab. Do
not place real production logs in it. The fixture generator below
writes only synthetic data with fake project/user identifiers.
from pathlib import Path
import json
root = Path("gitlab-ch32-lab")
root.mkdir(exist_ok=True)
def write(name, text):
(root / name).write_text(text.strip() + "\n", encoding="utf-8")
# Paid-tier audit data is represented as a fixture; event_type is explicitly training-only.
audit = {
"id": 4102, "author_id": 42, "entity_type": "Project",
"entity_path": "learner-example/platform-lab",
"target_type": "Member", "target_details": "USER_B",
"created_at": "2026-08-22T01:14:22Z",
"event_type": "TRAINING_FIXTURE_MEMBER_CHANGE"
}
write("audit_events.jsonl", json.dumps(audit))
cid = "01JINCIDENTA000000000000000"
write("workhorse.log", json.dumps({
"time":"2026-08-22T01:14:21.906Z", "status":502,
"duration_ms":4912, "uri":"/learner-example/platform-lab/-/tree/main",
"correlation_id":cid
}))
write("rails.log", json.dumps({
"time":"2026-08-22T01:14:21.914Z", "status":502,
"duration":4888.2, "db":18.4, "gitaly_duration":4862.1,
"path":"/learner-example/platform-lab/-/tree/main",
"correlation_id":cid
}))
write("gitaly.log", json.dumps({
"time":"2026-08-22T01:14:26.780Z", "grpc.code":"Unavailable",
"limit.concurrency_dropped":"max_time", "limit.concurrency_queue_length":27,
"grpc.method":"FindCommit", "correlation_id":cid
}))
write("sidekiq.log", json.dumps({
"time":"2026-08-22T01:14:25.000Z", "job_status":"done",
"queue_duration_s":0.18, "class":"TrainingFixtureWorker"
}))
write("health.txt", "200 GitLab OK")
write("readiness_all.json", json.dumps({
"http_status":503, "status":"failed",
"gitaly_check":[{"status":"failed","message":"synthetic timeout"}],
"db_check":[{"status":"ok"}], "redis_check":[{"status":"ok"}]
}, indent=2))
write("metrics.prom", """
gitaly_concurrency_limiting_queued 27
gitaly_requests_dropped_total{reason="max_time"} 18
gitaly_requests_dropped_total{reason="max_size"} 0
""")
write("runner_metrics.prom", """
gitlab_runner_concurrent 4
gitlab_runner_errors_total 0
gitlab_runner_request_concurrency 1
""")
print(root.resolve())
2. Inventory evidence before interpreting it
Start with filenames, timestamps, and source semantics—not grep patterns. Record whether each source is original or derived, what system produced it, its time zone, and whether it can contain sensitive fields. In a real incident, also record the collection time and hash if chain-of-custody matters.
from pathlib import Path
root = Path("gitlab-ch32-lab")
for p in sorted(root.iterdir()):
print(f"{p.name:24} {p.stat().st_size:6} bytes")
Expected evidence includes one audit fixture, four log sources, two health snapshots, GitLab/Gitaly-style metrics, and Runner metrics. Nothing has been changed yet.
3. Normalize the timeline and use the correlation ID as the primary join key
Now parse only the fields needed for the hypothesis. Do not dump every log field to the terminal. The request begins at Workhorse, passes through Rails, and ends with a Gitaly concurrency drop. The shared correlation ID makes that chain much stronger than timestamp proximity.
from pathlib import Path
import json
root = Path("gitlab-ch32-lab")
for name in ["workhorse.log", "rails.log", "gitaly.log", "sidekiq.log"]:
record = json.loads((root / name).read_text(encoding="utf-8"))
print(name, record.get("time"), record.get("correlation_id", "NO_REQUEST_ID"),
record.get("status", record.get("grpc.code", record.get("job_status"))))
workhorse.log 2026-08-22T01:14:21.906Z 01JINCIDENTA000000000000000 502
rails.log 2026-08-22T01:14:21.914Z 01JINCIDENTA000000000000000 502
gitaly.log 2026-08-22T01:14:26.780Z 01JINCIDENTA000000000000000 Unavailable
sidekiq.log 2026-08-22T01:14:25.000Z NO_REQUEST_ID done
The Sidekiq record occurs during the same interval but has no shared request ID and reports normal completion. It is evidence against a platform-wide background-job failure for this specific request.
4. Reconcile apparently contradictory health evidence
health.txt is green while
readiness_all.json is 503. That is not inconsistent:
the simple health endpoint proves only that the application server
responds. The all-dependencies readiness fixture isolates Gitaly
while DB and Redis remain healthy.
from pathlib import Path
import json
root = Path("gitlab-ch32-lab")
print((root / "health.txt").read_text().strip())
ready = json.loads((root / "readiness_all.json").read_text())
print("readiness HTTP", ready["http_status"])
for key in ("gitaly_check", "db_check", "redis_check"):
print(key, ready[key][0]["status"])
200 GitLab OK
readiness HTTP 503
gitaly_check failed
db_check ok
redis_check ok
5. Use metrics to test whether the Gitaly failure is isolated or a saturation pattern
One failed RPC could be a transient request. The metric fixture adds
fleet/service context: 27 requests are queued and the drop counter
has increased for max_time. That pattern is consistent
with sustained concurrency pressure rather than a single malformed
project request.
# On a real Prometheus server, use rates/increases for counters.
sum(gitaly_concurrency_limiting_queued)
sum(rate(gitaly_requests_dropped_total[5m])) by (reason)
# Keep runner evidence separate from GitLab application/Gitaly evidence.
sum(gitlab_runner_request_concurrency)
sum(rate(gitlab_runner_errors_total[5m]))
| Panel / query | Purpose | Interpretation in fixture |
|---|---|---|
| Gitaly queued work | Saturation indicator | High: 27 queued |
| Gitaly dropped rate by reason | Failure impact | Drop reason points to queue wait timeout |
| Request 5xx rate / latency | User impact | Would confirm whether failures are widespread |
| Runner errors / request concurrency | CI execution health | Fixture is normal; do not blame Runner |
| Readiness dependency status | Immediate dependency gate | Gitaly failed; DB/Redis healthy |
6. Investigate the nearby audit event without declaring it causal
The audit event happened less than a second after the failing request began. That makes it worth investigating, not guilty. Its target is a project member. The failing request, readiness result, and Gitaly logs instead point to repository service saturation. Record the event as a checked hypothesis and reject it unless additional evidence connects the role change to the failure.
from pathlib import Path
import json
p = Path("gitlab-ch32-lab/audit_events.jsonl")
event = json.loads(p.read_text())
print(event["created_at"], event["entity_path"], event["target_type"], event["event_type"])
print("Conclusion: temporally adjacent; no causal request/component link in the fixture")
7. Write the structured troubleshooting runbook
A useful runbook contains decision points. For this fixture, the conclusion should be: repository traffic is degraded by Gitaly concurrency saturation; Rails and Workhorse expose the symptom, all-dependency readiness identifies Gitaly, Gitaly logs show the drop reason, metrics show queue pressure, and Runner/Sidekiq evidence does not support a platform-wide CI/background failure.
| Runbook phase | Action | Exit criterion |
|---|---|---|
| Preserve | Copy original evidence read-only; record source/time/version | Original bundle remains unchanged |
| Scope | Project path, request ID, UTC window, user journey | One reproducible affected operation |
| Differentiate | Health vs readiness; Rails vs Gitaly; Runner vs app | Competing hypotheses narrowed |
| Quantify | Queue, drop rate, 5xx/latency, runner errors | Impact and saturation measured |
| Correct | Least-destructive capacity/load/limit action per current runbook | Change has owner and rollback |
| Verify | Repeat request/readiness and observe queue/drop trend | User journey and causal metrics recover |
| Document | Timeline, cause, rejected hypotheses, prevention | Another operator can reproduce reasoning |
8. Optional read-only live preflight on an isolated Self-Managed instance
If you administer a disposable Self-Managed instance, you may compare the fixture workflow with real read-only evidence. Do not enable debug logging, alter Prometheus exposure, restart services, or change runner settings for this optional step.
# Bash/Linux-package examples; read-only.
sudo gitlab-ctl status
curl -i https://gitlab.example.invalid/-/health
curl -i 'https://gitlab.example.invalid/-/readiness?all=1'
# Capture only the response header needed for correlation.
curl -sS -D - -o /dev/null https://gitlab.example.invalid/users/sign_in | grep -i '^x-request-id:'
# After identifying one request ID, search specifically for it.
CID="01JFAKECORRELATIONID000000"
find /var/log/gitlab -type f -mtime 0 -exec grep -F "$CID" '{}' '+' 2>/dev/null
9. Small challenge: choose the next evidence source
Suppose /-/health is 200,
/-/readiness?all=1 shows all dependencies healthy, and
jobs are pending with no runner assigned. Which evidence should you
inspect next? The best next step is Runner/project CI state—runner
availability/tags/protection and Runner metrics/logs—not Gitaly
logs. The resource boundary changed, so the evidence source must
change with it.
Suppose instead jobs are running but cloning the repository inside
the job fails with a correlation ID and Gitaly
Unavailable. Now the Runner can be healthy while the
GitLab/Gitaly path is degraded. “Pipeline problem” does not
automatically mean “Runner problem.”
10. Cleanup and retained evidence
The mandatory lab created only synthetic files. Review the directory, retain your incident summary if you want it for the checkpoint, then delete the generated fixture directory. In a real incident, cleanup is governed by evidence-retention policy: do not destroy originals merely because the service recovered.
from pathlib import Path
import shutil
root = Path("gitlab-ch32-lab")
# Review the path before deletion. This directory must contain only the synthetic fixture.
print(root.resolve())
# shutil.rmtree(root) # Uncomment only after verifying the disposable path.
Knowledge check
Why does the fixture include a nearby audit event that is not causal?
To force evidence-based reasoning. Temporal proximity is a clue, not proof. The request ID, readiness result, Gitaly log, and queue/drop metrics provide the causal chain.
The Runner metrics are normal while a pipeline-related complaint exists. What does that tell you?
It weakens the Runner-failure hypothesis but does not by itself prove the entire pipeline path is healthy. Inspect the actual pipeline/job state and dependencies.
Why should you alert on a rate/increase of a counter rather than its raw lifetime value?
Counters accumulate from process start. A large total may be old history; a rate or increase shows current failure activity.
What is the safe interpretation of a 200 from
/-/health and 503 from all-dependency
readiness?
The application server is alive, but at least one required dependency check is failing. The fixture identifies Gitaly while DB/Redis remain healthy.
Why is “increase Gitaly concurrency” not an automatic fix for a queue?
The queue can be protecting CPU, memory, disk, or downstream resources. Raising the limit without baseline/capacity evidence can move or amplify the bottleneck.
11. Lesson summary and bridge
- The fixture workflow preserves evidence before interpretation and uses request ID, component scope, and UTC time as join keys.
- Health, readiness, logs, metrics, and audit events can disagree without contradiction because they answer different questions.
- Normal Runner/Sidekiq evidence is useful negative evidence; it prevents blaming every platform component at once.
- A diagnostic dashboard should measure user impact, saturation, failures, and recovery—not maximize metric count.
Next, you will turn these mechanics into an observability design: what to retain, what to alert on, where to aggregate it, and how to balance detail against privacy and cost.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. The fixture uses synthetic event types and samples; treat official current metric/event names and endpoint behavior as the production contract.
- GitLab 19.3 release
- Audit events
- Audit events API
- Audit events administration and CSV export
- Audit event streaming for top-level groups
- Audit event streaming for instances
- GitLab log system
- Trace logs with a correlation ID
- Health checks
- Monitoring GitLab with Prometheus
- GitLab Prometheus metrics
- GitLab performance monitoring
- Monitor GitLab Runner usage
- Monitoring Gitaly
- GitLab Admin area monitoring
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.