Chapter 37Lesson 05~240 minutes

Checkpoint Lab — Monitoring Jenkins with Metrics, Logs, Health Checks, Audit Trails, Queue Telemetry, and Observability

Build and test a compact Jenkins observability pack: establish baseline evidence, induce bounded queue pressure and one disposable-agent failure, diagnose the causal layer, and define an alert that is actionable without leaking operational secrets.

checkpointqueue pressureagent failureSLOevidence packetcleanup

Learning objectives

  • Predict queue/executor/agent state changes before fault injection.
  • Capture metrics, API snapshots, logs and audit evidence around one incident.
  • Diagnose the causal layer from correlated identifiers and timestamps.
  • Define one meaningful alert and explain what it does not prove.
  • Produce a reviewable evidence packet and perform guarded cleanup.

1. Mission

On a disposable Jenkins controller, create one synthetic Pipeline that requires a single disposable observe-linux executor. Capture a baseline, trigger enough builds to create a small queue, disconnect the agent during a controlled run, and use the observability evidence—not guesswork—to explain the resulting queue/build behavior. Finish by writing one bounded alert rule/definition for queue wait or eligible capacity.

2. Preflight: stop if any boundary is unclear

Check Required state
Controller Disposable jenkins-observe-lab, Jenkins 2.568.3, Java 21.
Agent Disposable only; label observe-linux; Jenkins agent JVM Java 21.
Built-in node No routine build executors used for the lab workload.
Plugins Record Metrics/Prometheus/Audit Trail exact installed/resolved versions if used.
Job Exact full name observe-lab/slow-ci; synthetic repository and Jenkinsfile revision.
Credentials Only disposable read/trigger identity if needed; no real secret in scripts/evidence.
External systems None required; local files/endpoints only.
Cleanup Exact lab controller/agent/evidence resource identities recorded before failure injection.

3. Predict state changes before triggering anything

Prediction A — queue pressure
Given: 1 eligible observe-linux executor; 3 builds triggered.
Expect: 1 running; up to 2 queued; queued items retain queue IDs/reasons.
Do NOT expect: controller heap necessarily rises or Windows/other labels become eligible.

Prediction B — agent failure
Given: running synthetic sleep step on observe-linux.
Expect: agent online/channel state changes; affected build records interruption/failure behavior; queued builds remain unschedulable until eligible capacity returns.
Do NOT expect: every Pipeline step automatically resumes exactly where it stopped.

4. Capture the clean baseline

Create /tmp/jenkins-observe-checkpoint and record UTC time, plugin/runtime inventory, raw Prometheus scrape (if plugin enabled), queue API, computer API and exact job metadata. Do not store authentication headers or token values.

set -euo pipefail
OUT='/tmp/jenkins-observe-checkpoint'
JENKINS_URL='http://127.0.0.1:8080'
mkdir -p "$OUT"
date -u +baseline_utc=%Y-%m-%dT%H:%M:%SZ | tee "$OUT/timeline.txt"
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" > "$OUT/queue-before.json"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,assignedLabels[name]]" > "$OUT/nodes-before.json"
# If configured and authorized locally:
curl -fsS "$JENKINS_URL/prometheus/" > "$OUT/metrics-before.txt"

5. Trigger exactly three synthetic builds

Use the approved local lab trigger mechanism. Immediately capture queue snapshots at short intervals and record queue item IDs. Once builds start, record build numbers/URLs and the allocated node. Your evidence must show the queue-to-build transition rather than assuming trigger success means execution started.

6. Prove queue pressure causally

Evidence What to record
Queue API queue item ID, inQueueSince, why, task full name.
Node API agent online state, labels, executor count.
Build metadata job full name, build number, cause, source SHA if SCM-backed, node/workspace.
Metrics queue/executor/controller signals available in your installed plugin baseline.
UTC timeline trigger, queue, allocation timestamps.
Interpretation Which queue delay is explained by the single eligible executor?

7. Induce one agent failure

Record the affected build number, exact agent name, failure-injection UTC timestamp and first visible console/controller/agent messages. Capture another queue/node/metrics snapshot while the agent is offline.

8. Diagnose before repair

Write one short causal statement using this form:

At <UTC>, build <job>#<number> was executing on <agent>.
The agent channel became unavailable at <UTC>.
Queue items <ids> required label observe-linux and had no eligible online executor.
Controller/JVM baseline did/did not show independent resource pressure.
Audit evidence did/did not show an administrative node-disable action.
Therefore the primary causal layer is <agent connectivity/availability>, with queue delay as a consequence.

If your evidence does not support this conclusion, do not force it. The checkpoint is about method; a different causal layer is acceptable when the preserved evidence proves it.

9. Restore the agent and verify recovery

Reconnect/restart only the disposable agent. Capture the recovery timestamp, node online state, subsequent queue transition and the result of any affected build. Do not blindly replay an external side effect; this lab has no external side effect by design.

10. Define one meaningful alert

Choose an alert that follows from the drill. Example policy—not copy-paste production syntax:

SLI: eligible queue wait for workload=ci, pool=observe-linux
SLO: 95% of eligible builds begin execution within 120 seconds
Alert: page only if p95 wait > 120s for 10 minutes AND eligible pool capacity < required minimum
Context: controller, pool/label, oldest queue age, affected job count, dashboard/runbook link
Do not label metric series with build number or commit SHA.

Explain what the alert does not prove: it does not identify source-code defects, prove a specific agent failure, or guarantee controller health. It points the operator toward scheduling/capacity evidence.

11. Required evidence packet

Group Required contents
Baseline Jenkins/core Java/plugin inventory; controller identity and UTC timestamp.
Job/source Exact job full name, Jenkinsfile/source revision where applicable.
Queue At least one queue item ID, reason and wait timestamps before/during incident.
Build Affected build number/URL/cause/result and assigned node/workspace before loss.
Agent Exact agent identity, labels, online/offline transition and Java/runtime metadata.
Metrics Before/during/after raw or summarized controller/queue/executor signals, with metric-version note.
Logs First-failure build/controller/agent excerpts with secrets reviewed/redacted for sharing.
Audit Actor/action evidence for the benign admin event and incident window; limitations documented.
External latency Explicitly “none/local-only” or measured mock dependency state.
Diagnosis One causal statement and one rejected alternative hypothesis.
Alert SLI, SLO/threshold, window, dimensions, runbook context and known non-claims.
Limitations What this lab does not prove about production scale, HA, centralized logging or real dependency latency.

12. Verification checklist

  • Three builds created bounded pressure; no uncontrolled trigger loop was used.
  • Queue item identity was preserved separately from build identity.
  • The failed agent was disposable and exact.
  • First-failure evidence was captured before recovery/retry.
  • Queue pressure was correlated to eligible capacity, not raw executor total.
  • Metrics/log/audit evidence contained no real credential/token/private key.
  • The alert has a clear operational owner and does not use build/SHA as long-lived metric labels.
  • Recovery was measured with a post-restoration queue/node snapshot.

13. Cleanup

After copying the non-sensitive evidence packet to your chosen training output location, remove only the known checkpoint path and dispose of the lab controller/agent according to the same environment recreation procedure used at setup.

set -euo pipefail
OUT='/tmp/jenkins-observe-checkpoint'
case "$OUT" in
  /tmp/jenkins-observe-checkpoint) rm -rf -- "$OUT" ;;
  *) echo 'refusing unexpected cleanup path' >&2; exit 70 ;;
esac

14. What Chapter 37 adds to a production Jenkins operating model

Chapter 36 proved you can restore Jenkins state. Chapter 37 gives you the instrumentation to detect degradation, preserve the pre-failure timeline, prove the post-recovery service state and distinguish controller, queue, eligibility, agent, build and external-service failures. Observability becomes a versioned operational contract rather than an afterthought.

Chapter 38 uses this evidence to engineer performance: executors, queueing, heap/GC, disk I/O, workspaces and controller load are tuned only after a measured bottleneck is identified.

Next chapter

Chapter 38 — Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load

Carry the verified evidence and operating discipline from this chapter into the next chapter.

Knowledge check

Answer before revealing the explanation.

1. Why must queue item IDs and build numbers both appear in the evidence packet?

2. What evidence distinguishes missing eligible capacity from controller memory pressure?

3. Why should the alert avoid build-number and commit-SHA labels?

4. What should you do before reconnecting the failed agent?

5. What is the bridge from Chapter 37 to Chapter 38?

Official references and version notes

Observability interfaces, plugin versions and security behavior evolve. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.