Chapter 32Lesson 05~430 minutes

Checkpoint Lab — Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics

Diagnose a multi-symptom synthetic GitLab incident from preserved audit, log, metric, health, and runner evidence, produce a defensible timeline/root cause, and verify recovery without leaking sensitive data.

CheckpointIncident TimelineRoot CauseRedactionHandoff

Learning objectives

  • Preserve and inventory a synthetic incident bundle before changing or deleting any evidence.
  • Correlate at least three independent evidence sources and explicitly reject a plausible but unsupported hypothesis.
  • Produce a UTC incident timeline, root cause, corrective action, prevention plan, and verification criteria.
  • Redact sensitive fields in a working copy while preserving stable correlation relationships.
  • Demonstrate that recovery is verified by user impact, dependency state, and causal metrics—not by one green status endpoint.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. This checkpoint is entirely offline and free-compatible. It simulates paid audit-event data and Self-Managed operational outputs with clearly synthetic fixtures. No production GitLab instance, administrator access, Premium/Ultimate subscription, Prometheus server, external SIEM, runner registration, cloud service, or destructive action is required.

1. Checkpoint scenario and acceptance criteria

At 01:14 UTC, users report intermittent 502 errors when browsing a repository. A developer also reports that “CI feels slow.” A successful simple health check led the first responder to mark GitLab healthy. A project membership audit event occurred near the first failure. Your assignment is to decide what is causal, what is unrelated, and what evidence would prove recovery.

Your final handoff must contain: (1) evidence inventory, (2) normalized UTC timeline, (3) root cause with evidence for/against, (4) rejected hypotheses, (5) least-destructive corrective action, (6) prevention/alert plan, (7) verification checklist, and (8) sanitized evidence handling/cleanup notes.

Checkpoint rule: do not “solve” the incident by deleting evidence, restarting every service, or declaring the nearby audit event causal. Predictions and evidence come before proposed state changes.

2. Setup and preflight evidence

Reuse the synthetic fixture from Lesson 2 or recreate it from that generator. Record your local directory and confirm it contains no valuable data. Make a read-only original copy and a separate working copy. The originals represent evidence preservation; all annotation/redaction happens in the working copy.

from pathlib import Path
import shutil, hashlib
src = Path("gitlab-ch32-lab")
original = Path("gitlab-ch32-checkpoint/original")
working = Path("gitlab-ch32-checkpoint/working")

if not src.exists():
    raise SystemExit("Create the Lesson 2 synthetic fixture first.")
if original.parent.exists():
    raise SystemExit("Checkpoint directory already exists; inspect it before reuse.")

shutil.copytree(src, original)
shutil.copytree(src, working)

for p in sorted(original.iterdir()):
    digest = hashlib.sha256(p.read_bytes()).hexdigest()
    print(p.name, digest)

Store the hashes in your notes. You do not need a formal forensic chain of custody for a synthetic lab; the purpose is to practice proving that your original evidence was not silently edited.

3. Predict before action

Write two predictions before reading deeper evidence:

  1. If Gitaly saturation is the cause, the failing HTTP request should share an identifier with a Gitaly error and dependency readiness should identify Gitaly while unrelated services can remain healthy.
  2. If Runner failure is the main cause of the 502 repository symptom, Runner errors or unavailable execution state should correlate with the same user-visible failure path. Because repository browsing does not execute on Runner, this prediction is structurally weak and should likely fail.

Also predict what recovery would look like: repository request success, all-dependency readiness recovered, Gitaly queue/drop activity returned to baseline, while Runner metrics may be unchanged.

4. Build the normalized incident timeline

Extract only operationally relevant fields from the working copy and produce a single UTC timeline. Keep the audit event, even if it becomes a rejected hypothesis.

from pathlib import Path
import json
root = Path("gitlab-ch32-checkpoint/working")
records=[]
for name in ["workhorse.log","rails.log","gitaly.log","sidekiq.log","audit_events.jsonl"]:
    obj=json.loads((root/name).read_text())
    t=obj.get("time") or obj.get("created_at")
    records.append((t,name,obj))
for t,name,obj in sorted(records):
    print(t, name,
          obj.get("correlation_id","NO_REQUEST_ID"),
          obj.get("status",obj.get("grpc.code",obj.get("event_type",obj.get("job_status")))))
2026-08-22T01:14:21.906Z workhorse.log 01JINCIDENTA... 502
2026-08-22T01:14:21.914Z rails.log     01JINCIDENTA... 502
2026-08-22T01:14:22Z     audit_events.jsonl NO_REQUEST_ID TRAINING_FIXTURE_MEMBER_CHANGE
2026-08-22T01:14:25.000Z sidekiq.log   NO_REQUEST_ID done
2026-08-22T01:14:26.780Z gitaly.log    01JINCIDENTA... Unavailable

The order alone does not prove causality. The shared request ID links Workhorse → Rails → Gitaly. The audit and Sidekiq records need separate evidence if they are to enter the causal chain.

5. Correlate at least three independent evidence sources

Source 1 — logs: Rails spends nearly the whole failed request duration waiting on Gitaly, and Gitaly drops the same correlation ID after concurrency queue wait.

Source 2 — health/readiness: the simple health endpoint is 200, but all-dependency readiness is 503 with only the synthetic Gitaly check failed; database and Redis checks remain OK.

Source 3 — metrics: Gitaly queued work is high and dropped requests have accumulated for max_time. Runner error metrics are zero in the fixture.

Source 4 — audit: a member change occurred at the same project/time window but has no request/component link. It is a plausible coincidence, not the root cause.

Evidence Supports Gitaly saturation? Supports Runner failure? Supports member-change cause?
Shared Workhorse/Rails/Gitaly request ID Strong No No
Readiness: Gitaly failed, DB/Redis OK Strong No No
Gitaly queue/drop metrics Strong No No
Runner errors = 0 Neutral Against Neutral
Nearby member-change audit fixture No direct link No Only temporal proximity

6. State root cause and rejected hypotheses precisely

Root cause supported by the fixture: repository requests are failing because Gitaly work is saturated/queued long enough to be dropped for max_time. Rails and Workhorse surface the 502; they are not independently identified as the bottleneck. The fixture does not provide host/storage metrics or a real configuration history, so you must not invent the deeper reason for the Gitaly saturation.

Rejected hypothesis — member change caused outage: timing is close, but there is no causal identifier/component mechanism.

Rejected hypothesis — Runner caused repository 502: repository browsing does not run on Runner, and the Runner metrics fixture contains no error signal. The separate “CI feels slow” report requires pipeline/job timing evidence before it becomes a confirmed incident symptom.

Source-supported boundary: this fixture supports “Gitaly concurrency saturation caused repository request failure.” It does not support “disk failure,” “bad concurrency configuration,” or “malicious membership change” because those facts are not present.

7. Corrective action: choose the least destructive control

You cannot safely choose a real production fix from this synthetic bundle because it omits the underlying resource/load/configuration cause. Your correct operational action is to halt speculative changes, collect current Gitaly host/storage/load evidence and recent configuration/infrastructure change records, then follow the current Gitaly concurrency/capacity guidance.

Examples of justified corrections could include stopping an abnormal workload, restoring lost compute/storage capacity, reverting a known bad concurrency/storage configuration, or scaling according to validated capacity. “Restart everything” and “raise every limit” are not justified by the evidence.

Candidate action Evidence required first Rollback / risk
Stop abnormal high-load operation Identify load source and owner Work can be retried; confirm business impact
Restore failed/lost capacity Host/storage/network fault evidence Return to known topology
Revert recent Gitaly limit/config change Exact diff + timeline + known-good config Restore previous version/config
Increase concurrency/scale Baseline proving capacity headroom Can overload CPU/disk/downstream; incremental rollout
Restart all services No specific evidence in fixture High blast radius, destroys state/evidence; reject

8. Verify recovery independently

A correction is successful only if the user journey and the causal signals recover. Do not accept one green /-/health response as closure.

RECOVERY CHECK
[ ] repository browse/clone succeeds for the previously failing scope
[ ] /-/readiness?all=1 reports dependencies healthy
[ ] Gitaly queued work returns toward established baseline
[ ] drop rate/increase for max_time stops rising materially
[ ] new failed requests no longer show Gitaly queue timeout
[ ] database/Redis/Sidekiq remain healthy (no shifted bottleneck)
[ ] Runner/pipeline complaint is separately measured and either resolved or tracked
[ ] no new security exposure was created by troubleshooting

The last two checks matter. A service can recover while a second complaint remains unrelated, and a performance fix can create a security regression if it exposes a metrics endpoint or logs sensitive data.

9. Redact a working copy without destroying correlation

Practice on the synthetic data. Replace project/user/IP fields if present, but keep the same request ID across Workhorse/Rails/Gitaly so the causal chain remains visible. In real evidence, whether the request ID itself is sensitive depends on your threat model and sharing destination; the principle is stable pseudonymization, not blind deletion.

from pathlib import Path
root = Path("gitlab-ch32-checkpoint/working")
replacements = {
    "learner-example/platform-lab": "PROJECT_X",
    "USER_B": "USER_REDACTED_1"
}
for p in root.iterdir():
    if not p.is_file():
        continue
    text = p.read_text(encoding="utf-8")
    for old,new in replacements.items():
        text = text.replace(old,new)
    p.write_text(text, encoding="utf-8")
print("Sanitized working copy created; originals remain untouched.")
Do not sanitize by overwriting the original evidence bundle. Preserve originals under access control; share only the minimum sanitized derivative required for the audience.

10. Prevention and minimal dashboard/alert plan

The incident shows two observability gaps: simple health was overinterpreted, and no alert connected Gitaly queue/drop signals to repository user impact. A prevention plan should add all-dependency readiness for diagnostics, a repository/API user-journey SLI, Gitaly queue/drop and host/storage panels, and an alert that requires sustained saturation plus impact.

# Illustrative components of a diagnostic view; adapt to exact current metrics.
sum(gitaly_concurrency_limiting_queued)
sum(rate(gitaly_requests_dropped_total[5m])) by (reason)

# Pair service saturation with your real HTTP/Git success and latency SLIs.
# Do not page on queue depth alone without a baseline.

Also update the incident runbook: “/-/health green” must never be the only platform-health criterion, and nearby audit changes must be treated as hypotheses until connected by scope/mechanism/evidence.

11. Final operational handoff

Your concise handoff can use this structure:

INCIDENT: Repository 502 / Gitaly saturation fixture
Window: 2026-08-22 01:14:21Z–01:14:27Z
Impact: repository browse request failed; CI-slowness report not independently confirmed
Root cause: Gitaly concurrency queue saturation; request dropped after max wait
Evidence: shared request ID across Workhorse/Rails/Gitaly; readiness Gitaly failed; queue/drop metrics elevated
Rejected: nearby member audit event (no causal link); Runner failure (normal fixture metrics, wrong boundary)
Correction: production action requires current host/storage/load/config evidence; avoid blind restart/limit increase
Verify: user journey + readiness + queue/drop trend + no shifted bottleneck
Prevention: SLI + saturation/impact alert, readiness-aware runbook, stable correlation workflow
Evidence handling: originals hashed/restricted; working copy pseudonymized; no real credentials present

The wording distinguishes observed facts from proposed next investigation. That boundary is essential in real incident response.

12. Cleanup and retention decision

The fixture contains no real sensitive data. After reviewing your handoff, you may delete both the Lesson 2 source fixture and checkpoint copies. In a real incident, deletion follows retention/legal/security policy; recovery is not permission to destroy audit or diagnostic evidence.

from pathlib import Path
import shutil
for name in ["gitlab-ch32-lab", "gitlab-ch32-checkpoint"]:
    p=Path(name)
    print("REVIEW BEFORE DELETE:", p.resolve())
    # shutil.rmtree(p)  # Uncomment only after confirming each path is disposable.

Knowledge check

What is the strongest causal join in the checkpoint?

Why is the nearby audit event retained in the final report if it is not causal?

What deeper root cause does the fixture not support?

What three categories of evidence should be green before declaring recovery?

Why keep an original and sanitized working copy?

13. Lesson summary and bridge

  • You diagnosed a multi-source incident without paid services or destructive actions.
  • The root cause was bounded to what the evidence actually supports: Gitaly queue saturation causing repository request failure.
  • You used audit evidence as a hypothesis source rather than assuming a nearby governance event was causal.
  • Recovery requires user-journey success, dependency readiness, causal metric improvement, and no shifted failure.
  • A production handoff must distinguish observed facts, rejected hypotheses, unknown deeper causes, proposed corrections, and evidence-retention handling.

Chapter 32 adds evidence-driven operations to the GitLab production model. Chapter 33 moves into GitLab Duo and Agentic DevSecOps: AI Features, Security Remediation, Governance, and Usage Controls, where AI/agent activity adds another governance and audit boundary: generated or agent-executed work must remain reviewable, scoped, and independently verified.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. Self-Managed operators must use the documentation for the exact version they run; metric names, event types, UI placement, availability, and operational procedures can change between releases. The mandatory incident data is synthetic. If you adapt the runbook to real evidence, apply your organization’s privacy, security, legal, and retention requirements in addition to GitLab product guidance.

Next chapter

GitLab Duo and Agentic DevSecOps: AI Features, Security Remediation, Governance, and Usage Controls

Carry evidence, auditability, least privilege, and independent verification into AI-assisted and agentic GitLab workflows.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.