Checkpoint Lab — Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics
Diagnose a multi-symptom synthetic GitLab incident from preserved audit, log, metric, health, and runner evidence, produce a defensible timeline/root cause, and verify recovery without leaking sensitive data.
Learning objectives
- Preserve and inventory a synthetic incident bundle before changing or deleting any evidence.
- Correlate at least three independent evidence sources and explicitly reject a plausible but unsupported hypothesis.
- Produce a UTC incident timeline, root cause, corrective action, prevention plan, and verification criteria.
- Redact sensitive fields in a working copy while preserving stable correlation relationships.
- Demonstrate that recovery is verified by user impact, dependency state, and causal metrics—not by one green status endpoint.
1. Checkpoint scenario and acceptance criteria
At 01:14 UTC, users report intermittent 502 errors when browsing a repository. A developer also reports that “CI feels slow.” A successful simple health check led the first responder to mark GitLab healthy. A project membership audit event occurred near the first failure. Your assignment is to decide what is causal, what is unrelated, and what evidence would prove recovery.
Your final handoff must contain: (1) evidence inventory, (2) normalized UTC timeline, (3) root cause with evidence for/against, (4) rejected hypotheses, (5) least-destructive corrective action, (6) prevention/alert plan, (7) verification checklist, and (8) sanitized evidence handling/cleanup notes.
2. Setup and preflight evidence
Reuse the synthetic fixture from Lesson 2 or recreate it from that generator. Record your local directory and confirm it contains no valuable data. Make a read-only original copy and a separate working copy. The originals represent evidence preservation; all annotation/redaction happens in the working copy.
from pathlib import Path
import shutil, hashlib
src = Path("gitlab-ch32-lab")
original = Path("gitlab-ch32-checkpoint/original")
working = Path("gitlab-ch32-checkpoint/working")
if not src.exists():
raise SystemExit("Create the Lesson 2 synthetic fixture first.")
if original.parent.exists():
raise SystemExit("Checkpoint directory already exists; inspect it before reuse.")
shutil.copytree(src, original)
shutil.copytree(src, working)
for p in sorted(original.iterdir()):
digest = hashlib.sha256(p.read_bytes()).hexdigest()
print(p.name, digest)
Store the hashes in your notes. You do not need a formal forensic chain of custody for a synthetic lab; the purpose is to practice proving that your original evidence was not silently edited.
3. Predict before action
Write two predictions before reading deeper evidence:
- If Gitaly saturation is the cause, the failing HTTP request should share an identifier with a Gitaly error and dependency readiness should identify Gitaly while unrelated services can remain healthy.
- If Runner failure is the main cause of the 502 repository symptom, Runner errors or unavailable execution state should correlate with the same user-visible failure path. Because repository browsing does not execute on Runner, this prediction is structurally weak and should likely fail.
Also predict what recovery would look like: repository request success, all-dependency readiness recovered, Gitaly queue/drop activity returned to baseline, while Runner metrics may be unchanged.
4. Build the normalized incident timeline
Extract only operationally relevant fields from the working copy and produce a single UTC timeline. Keep the audit event, even if it becomes a rejected hypothesis.
from pathlib import Path
import json
root = Path("gitlab-ch32-checkpoint/working")
records=[]
for name in ["workhorse.log","rails.log","gitaly.log","sidekiq.log","audit_events.jsonl"]:
obj=json.loads((root/name).read_text())
t=obj.get("time") or obj.get("created_at")
records.append((t,name,obj))
for t,name,obj in sorted(records):
print(t, name,
obj.get("correlation_id","NO_REQUEST_ID"),
obj.get("status",obj.get("grpc.code",obj.get("event_type",obj.get("job_status")))))
2026-08-22T01:14:21.906Z workhorse.log 01JINCIDENTA... 502
2026-08-22T01:14:21.914Z rails.log 01JINCIDENTA... 502
2026-08-22T01:14:22Z audit_events.jsonl NO_REQUEST_ID TRAINING_FIXTURE_MEMBER_CHANGE
2026-08-22T01:14:25.000Z sidekiq.log NO_REQUEST_ID done
2026-08-22T01:14:26.780Z gitaly.log 01JINCIDENTA... Unavailable
The order alone does not prove causality. The shared request ID links Workhorse → Rails → Gitaly. The audit and Sidekiq records need separate evidence if they are to enter the causal chain.
5. Correlate at least three independent evidence sources
Source 1 — logs: Rails spends nearly the whole failed request duration waiting on Gitaly, and Gitaly drops the same correlation ID after concurrency queue wait.
Source 2 — health/readiness: the simple health endpoint is 200, but all-dependency readiness is 503 with only the synthetic Gitaly check failed; database and Redis checks remain OK.
Source 3 — metrics: Gitaly queued work is high and
dropped requests have accumulated for max_time. Runner
error metrics are zero in the fixture.
Source 4 — audit: a member change occurred at the same project/time window but has no request/component link. It is a plausible coincidence, not the root cause.
| Evidence | Supports Gitaly saturation? | Supports Runner failure? | Supports member-change cause? |
|---|---|---|---|
| Shared Workhorse/Rails/Gitaly request ID | Strong | No | No |
| Readiness: Gitaly failed, DB/Redis OK | Strong | No | No |
| Gitaly queue/drop metrics | Strong | No | No |
| Runner errors = 0 | Neutral | Against | Neutral |
| Nearby member-change audit fixture | No direct link | No | Only temporal proximity |
6. State root cause and rejected hypotheses precisely
Root cause supported by the fixture: repository
requests are failing because Gitaly work is saturated/queued long
enough to be dropped for max_time. Rails and Workhorse
surface the 502; they are not independently identified as the
bottleneck. The fixture does not provide host/storage metrics or a
real configuration history, so you must not invent the
deeper reason for the Gitaly saturation.
Rejected hypothesis — member change caused outage: timing is close, but there is no causal identifier/component mechanism.
Rejected hypothesis — Runner caused repository 502: repository browsing does not run on Runner, and the Runner metrics fixture contains no error signal. The separate “CI feels slow” report requires pipeline/job timing evidence before it becomes a confirmed incident symptom.
7. Corrective action: choose the least destructive control
You cannot safely choose a real production fix from this synthetic bundle because it omits the underlying resource/load/configuration cause. Your correct operational action is to halt speculative changes, collect current Gitaly host/storage/load evidence and recent configuration/infrastructure change records, then follow the current Gitaly concurrency/capacity guidance.
Examples of justified corrections could include stopping an abnormal workload, restoring lost compute/storage capacity, reverting a known bad concurrency/storage configuration, or scaling according to validated capacity. “Restart everything” and “raise every limit” are not justified by the evidence.
| Candidate action | Evidence required first | Rollback / risk |
|---|---|---|
| Stop abnormal high-load operation | Identify load source and owner | Work can be retried; confirm business impact |
| Restore failed/lost capacity | Host/storage/network fault evidence | Return to known topology |
| Revert recent Gitaly limit/config change | Exact diff + timeline + known-good config | Restore previous version/config |
| Increase concurrency/scale | Baseline proving capacity headroom | Can overload CPU/disk/downstream; incremental rollout |
| Restart all services | No specific evidence in fixture | High blast radius, destroys state/evidence; reject |
8. Verify recovery independently
A correction is successful only if the user journey and the causal
signals recover. Do not accept one green
/-/health response as closure.
RECOVERY CHECK
[ ] repository browse/clone succeeds for the previously failing scope
[ ] /-/readiness?all=1 reports dependencies healthy
[ ] Gitaly queued work returns toward established baseline
[ ] drop rate/increase for max_time stops rising materially
[ ] new failed requests no longer show Gitaly queue timeout
[ ] database/Redis/Sidekiq remain healthy (no shifted bottleneck)
[ ] Runner/pipeline complaint is separately measured and either resolved or tracked
[ ] no new security exposure was created by troubleshooting
The last two checks matter. A service can recover while a second complaint remains unrelated, and a performance fix can create a security regression if it exposes a metrics endpoint or logs sensitive data.
9. Redact a working copy without destroying correlation
Practice on the synthetic data. Replace project/user/IP fields if present, but keep the same request ID across Workhorse/Rails/Gitaly so the causal chain remains visible. In real evidence, whether the request ID itself is sensitive depends on your threat model and sharing destination; the principle is stable pseudonymization, not blind deletion.
from pathlib import Path
root = Path("gitlab-ch32-checkpoint/working")
replacements = {
"learner-example/platform-lab": "PROJECT_X",
"USER_B": "USER_REDACTED_1"
}
for p in root.iterdir():
if not p.is_file():
continue
text = p.read_text(encoding="utf-8")
for old,new in replacements.items():
text = text.replace(old,new)
p.write_text(text, encoding="utf-8")
print("Sanitized working copy created; originals remain untouched.")
10. Prevention and minimal dashboard/alert plan
The incident shows two observability gaps: simple health was overinterpreted, and no alert connected Gitaly queue/drop signals to repository user impact. A prevention plan should add all-dependency readiness for diagnostics, a repository/API user-journey SLI, Gitaly queue/drop and host/storage panels, and an alert that requires sustained saturation plus impact.
# Illustrative components of a diagnostic view; adapt to exact current metrics.
sum(gitaly_concurrency_limiting_queued)
sum(rate(gitaly_requests_dropped_total[5m])) by (reason)
# Pair service saturation with your real HTTP/Git success and latency SLIs.
# Do not page on queue depth alone without a baseline.
Also update the incident runbook: “/-/health green”
must never be the only platform-health criterion, and nearby audit
changes must be treated as hypotheses until connected by
scope/mechanism/evidence.
11. Final operational handoff
Your concise handoff can use this structure:
INCIDENT: Repository 502 / Gitaly saturation fixture
Window: 2026-08-22 01:14:21Z–01:14:27Z
Impact: repository browse request failed; CI-slowness report not independently confirmed
Root cause: Gitaly concurrency queue saturation; request dropped after max wait
Evidence: shared request ID across Workhorse/Rails/Gitaly; readiness Gitaly failed; queue/drop metrics elevated
Rejected: nearby member audit event (no causal link); Runner failure (normal fixture metrics, wrong boundary)
Correction: production action requires current host/storage/load/config evidence; avoid blind restart/limit increase
Verify: user journey + readiness + queue/drop trend + no shifted bottleneck
Prevention: SLI + saturation/impact alert, readiness-aware runbook, stable correlation workflow
Evidence handling: originals hashed/restricted; working copy pseudonymized; no real credentials present
The wording distinguishes observed facts from proposed next investigation. That boundary is essential in real incident response.
12. Cleanup and retention decision
The fixture contains no real sensitive data. After reviewing your handoff, you may delete both the Lesson 2 source fixture and checkpoint copies. In a real incident, deletion follows retention/legal/security policy; recovery is not permission to destroy audit or diagnostic evidence.
from pathlib import Path
import shutil
for name in ["gitlab-ch32-lab", "gitlab-ch32-checkpoint"]:
p=Path(name)
print("REVIEW BEFORE DELETE:", p.resolve())
# shutil.rmtree(p) # Uncomment only after confirming each path is disposable.
Knowledge check
What is the strongest causal join in the checkpoint?
The same correlation/request ID across the failed Workhorse, Rails, and Gitaly records, supported by Gitaly readiness and queue/drop metrics.
Why is the nearby audit event retained in the final report if it is not causal?
It documents a plausible hypothesis that was investigated and rejected, which makes the reasoning auditable and avoids silently ignoring a relevant change.
What deeper root cause does the fixture not support?
It does not prove why Gitaly saturated—such as disk failure, bad limits, or abnormal load—because host/storage/configuration evidence is absent.
What three categories of evidence should be green before declaring recovery?
The affected user journey, dependency/readiness state, and causal metrics/log pattern. Also verify no shifted bottleneck or new security regression.
Why keep an original and sanitized working copy?
The original preserves evidentiary integrity; the working copy can be shared/analyzed with minimized sensitive data while retaining stable correlation relationships.
13. Lesson summary and bridge
- You diagnosed a multi-source incident without paid services or destructive actions.
- The root cause was bounded to what the evidence actually supports: Gitaly queue saturation causing repository request failure.
- You used audit evidence as a hypothesis source rather than assuming a nearby governance event was causal.
- Recovery requires user-journey success, dependency readiness, causal metric improvement, and no shifted failure.
- A production handoff must distinguish observed facts, rejected hypotheses, unknown deeper causes, proposed corrections, and evidence-retention handling.
Chapter 32 adds evidence-driven operations to the GitLab production model. Chapter 33 moves into GitLab Duo and Agentic DevSecOps: AI Features, Security Remediation, Governance, and Usage Controls, where AI/agent activity adds another governance and audit boundary: generated or agent-executed work must remain reviewable, scoped, and independently verified.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. The mandatory incident data is synthetic. If you adapt the runbook to real evidence, apply your organization’s privacy, security, legal, and retention requirements in addition to GitLab product guidance.
- GitLab 19.3 release
- Audit events
- Audit events API
- Audit events administration and CSV export
- Audit event streaming for top-level groups
- Audit event streaming for instances
- GitLab log system
- Trace logs with a correlation ID
- Health checks
- Monitoring GitLab with Prometheus
- GitLab Prometheus metrics
- GitLab performance monitoring
- Monitor GitLab Runner usage
- Monitoring Gitaly
- GitLab Admin area monitoring
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.