Checkpoint Lab — Monitoring Jenkins with Metrics, Logs, Health Checks, Audit Trails, Queue Telemetry, and Observability
Build and test a compact Jenkins observability pack: establish baseline evidence, induce bounded queue pressure and one disposable-agent failure, diagnose the causal layer, and define an alert that is actionable without leaking operational secrets.
Learning objectives
- Predict queue/executor/agent state changes before fault injection.
- Capture metrics, API snapshots, logs and audit evidence around one incident.
- Diagnose the causal layer from correlated identifiers and timestamps.
- Define one meaningful alert and explain what it does not prove.
- Produce a reviewable evidence packet and perform guarded cleanup.
1. Mission
On a disposable Jenkins controller, create one synthetic Pipeline
that requires a single disposable
observe-linux executor. Capture a baseline, trigger
enough builds to create a small queue, disconnect the agent during a
controlled run, and use the observability evidence—not guesswork—to
explain the resulting queue/build behavior. Finish by writing one
bounded alert rule/definition for queue wait or eligible capacity.
2. Preflight: stop if any boundary is unclear
| Check | Required state |
|---|---|
| Controller |
Disposable jenkins-observe-lab, Jenkins
2.568.3, Java 21.
|
| Agent |
Disposable only; label observe-linux; Jenkins
agent JVM Java 21.
|
| Built-in node | No routine build executors used for the lab workload. |
| Plugins | Record Metrics/Prometheus/Audit Trail exact installed/resolved versions if used. |
| Job |
Exact full name observe-lab/slow-ci; synthetic
repository and Jenkinsfile revision.
|
| Credentials | Only disposable read/trigger identity if needed; no real secret in scripts/evidence. |
| External systems | None required; local files/endpoints only. |
| Cleanup | Exact lab controller/agent/evidence resource identities recorded before failure injection. |
3. Predict state changes before triggering anything
Prediction A — queue pressure
Given: 1 eligible observe-linux executor; 3 builds triggered.
Expect: 1 running; up to 2 queued; queued items retain queue IDs/reasons.
Do NOT expect: controller heap necessarily rises or Windows/other labels become eligible.
Prediction B — agent failure
Given: running synthetic sleep step on observe-linux.
Expect: agent online/channel state changes; affected build records interruption/failure behavior; queued builds remain unschedulable until eligible capacity returns.
Do NOT expect: every Pipeline step automatically resumes exactly where it stopped.
4. Capture the clean baseline
Create /tmp/jenkins-observe-checkpoint and record UTC
time, plugin/runtime inventory, raw Prometheus scrape (if plugin
enabled), queue API, computer API and exact job metadata. Do not
store authentication headers or token values.
set -euo pipefail
OUT='/tmp/jenkins-observe-checkpoint'
JENKINS_URL='http://127.0.0.1:8080'
mkdir -p "$OUT"
date -u +baseline_utc=%Y-%m-%dT%H:%M:%SZ | tee "$OUT/timeline.txt"
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" > "$OUT/queue-before.json"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,assignedLabels[name]]" > "$OUT/nodes-before.json"
# If configured and authorized locally:
curl -fsS "$JENKINS_URL/prometheus/" > "$OUT/metrics-before.txt"
5. Trigger exactly three synthetic builds
Use the approved local lab trigger mechanism. Immediately capture queue snapshots at short intervals and record queue item IDs. Once builds start, record build numbers/URLs and the allocated node. Your evidence must show the queue-to-build transition rather than assuming trigger success means execution started.
6. Prove queue pressure causally
| Evidence | What to record |
|---|---|
| Queue API |
queue item ID, inQueueSince, why,
task full name.
|
| Node API | agent online state, labels, executor count. |
| Build metadata | job full name, build number, cause, source SHA if SCM-backed, node/workspace. |
| Metrics | queue/executor/controller signals available in your installed plugin baseline. |
| UTC timeline | trigger, queue, allocation timestamps. |
| Interpretation | Which queue delay is explained by the single eligible executor? |
7. Induce one agent failure
Record the affected build number, exact agent name, failure-injection UTC timestamp and first visible console/controller/agent messages. Capture another queue/node/metrics snapshot while the agent is offline.
8. Diagnose before repair
Write one short causal statement using this form:
At <UTC>, build <job>#<number> was executing on <agent>.
The agent channel became unavailable at <UTC>.
Queue items <ids> required label observe-linux and had no eligible online executor.
Controller/JVM baseline did/did not show independent resource pressure.
Audit evidence did/did not show an administrative node-disable action.
Therefore the primary causal layer is <agent connectivity/availability>, with queue delay as a consequence.
If your evidence does not support this conclusion, do not force it. The checkpoint is about method; a different causal layer is acceptable when the preserved evidence proves it.
9. Restore the agent and verify recovery
Reconnect/restart only the disposable agent. Capture the recovery timestamp, node online state, subsequent queue transition and the result of any affected build. Do not blindly replay an external side effect; this lab has no external side effect by design.
10. Define one meaningful alert
Choose an alert that follows from the drill. Example policy—not copy-paste production syntax:
SLI: eligible queue wait for workload=ci, pool=observe-linux
SLO: 95% of eligible builds begin execution within 120 seconds
Alert: page only if p95 wait > 120s for 10 minutes AND eligible pool capacity < required minimum
Context: controller, pool/label, oldest queue age, affected job count, dashboard/runbook link
Do not label metric series with build number or commit SHA.
Explain what the alert does not prove: it does not identify source-code defects, prove a specific agent failure, or guarantee controller health. It points the operator toward scheduling/capacity evidence.
11. Required evidence packet
| Group | Required contents |
|---|---|
| Baseline | Jenkins/core Java/plugin inventory; controller identity and UTC timestamp. |
| Job/source | Exact job full name, Jenkinsfile/source revision where applicable. |
| Queue | At least one queue item ID, reason and wait timestamps before/during incident. |
| Build | Affected build number/URL/cause/result and assigned node/workspace before loss. |
| Agent | Exact agent identity, labels, online/offline transition and Java/runtime metadata. |
| Metrics | Before/during/after raw or summarized controller/queue/executor signals, with metric-version note. |
| Logs | First-failure build/controller/agent excerpts with secrets reviewed/redacted for sharing. |
| Audit | Actor/action evidence for the benign admin event and incident window; limitations documented. |
| External latency | Explicitly “none/local-only” or measured mock dependency state. |
| Diagnosis | One causal statement and one rejected alternative hypothesis. |
| Alert | SLI, SLO/threshold, window, dimensions, runbook context and known non-claims. |
| Limitations | What this lab does not prove about production scale, HA, centralized logging or real dependency latency. |
12. Verification checklist
- Three builds created bounded pressure; no uncontrolled trigger loop was used.
- Queue item identity was preserved separately from build identity.
- The failed agent was disposable and exact.
- First-failure evidence was captured before recovery/retry.
- Queue pressure was correlated to eligible capacity, not raw executor total.
- Metrics/log/audit evidence contained no real credential/token/private key.
- The alert has a clear operational owner and does not use build/SHA as long-lived metric labels.
- Recovery was measured with a post-restoration queue/node snapshot.
13. Cleanup
After copying the non-sensitive evidence packet to your chosen training output location, remove only the known checkpoint path and dispose of the lab controller/agent according to the same environment recreation procedure used at setup.
set -euo pipefail
OUT='/tmp/jenkins-observe-checkpoint'
case "$OUT" in
/tmp/jenkins-observe-checkpoint) rm -rf -- "$OUT" ;;
*) echo 'refusing unexpected cleanup path' >&2; exit 70 ;;
esac
14. What Chapter 37 adds to a production Jenkins operating model
Chapter 36 proved you can restore Jenkins state. Chapter 37 gives you the instrumentation to detect degradation, preserve the pre-failure timeline, prove the post-recovery service state and distinguish controller, queue, eligibility, agent, build and external-service failures. Observability becomes a versioned operational contract rather than an afterthought.
Chapter 38 uses this evidence to engineer performance: executors, queueing, heap/GC, disk I/O, workspaces and controller load are tuned only after a measured bottleneck is identified.
Knowledge check
Answer before revealing the explanation.
1. Why must queue item IDs and build numbers both appear in the evidence packet?
Queue and build are different lifecycle states; a queue item may wait or never become a build, and correlation proves the transition.
2. What evidence distinguishes missing eligible capacity from controller memory pressure?
Queue reason/requested label plus node/executor eligibility, alongside normal JVM/controller metrics.
3. Why should the alert avoid build-number and commit-SHA labels?
They are unbounded/high-cardinality identities better preserved in logs/events.
4. What should you do before reconnecting the failed agent?
Preserve exact job/build/queue/agent identities and first-failure metrics/log evidence.
5. What is the bridge from Chapter 37 to Chapter 38?
Use observed queue/JVM/disk/executor/workspace measurements to form and test one performance bottleneck hypothesis at a time.
Official references and version notes
Observability interfaces, plugin versions and security behavior evolve. Re-check current primary documentation before applying these patterns to a real controller.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.