Monitoring Jenkins with Metrics, Logs, Health Checks, Audit Trails, Queue Telemetry, and Observability: Guided Hands-On Workflow and Core Operations
Create a disposable observability pack: pin the monitoring plugins, collect controller/queue/executor signals, generate bounded load, induce one safe agent failure, and correlate metrics, logs and build IDs without exposing real infrastructure.
Learning objectives
- Preflight a disposable controller/agent before enabling monitoring plugins.
- Use a pinned Prometheus/Metrics path without exposing it publicly.
- Capture queue and executor evidence before, during and after synthetic pressure.
- Preserve controller, build and agent logs around a controlled failure.
- Build a small correlation packet keyed by UTC time, job/build and agent identity.
1. Lab scenario and safety boundary
Use a local disposable controller named conceptually
jenkins-observe-lab, a synthetic Pipeline job
observe-lab/slow-ci, and at least one disposable agent
labeled observe-linux. Do not run routine lab builds on
the built-in node. No production SCM, chat, cloud, identity provider
or deployment target is involved.
2. Preflight the exact baseline
Evidence record
---------------
controller = jenkins-observe-lab
jenkins_lts = 2.568.3
controller_java = 21
agent_java = 21
metrics = 4.2.37-494.v06f9a_939d33a_
prometheus = 860.v532442b_44e9a_
audit-trail = 456.v39d2fd1ed556
job_full_name = observe-lab/slow-ci
agent_label = observe-linux
external_targets = none (local-only lab)
Before installation, inspect the plugin pages for current minimum Jenkins versions, dependencies and security warnings. The dated versions above are compatible with this chapter baseline, but that does not make them evergreen recommendations.
3. Pin the optional observability plugins in a disposable controller image
The safest repeatable pattern is to build/recreate the lab controller with an explicit plugin list rather than clicking “latest” in Plugin Manager. Example for an image build context:
# plugins.txt
metrics:4.2.37-494.v06f9a_939d33a_
prometheus:860.v532442b_44e9a_
audit-trail:456.v39d2fd1ed556
FROM jenkins/jenkins:2.568.3-jdk21
COPY plugins.txt /usr/share/jenkins/ref/plugins.txt
RUN jenkins-plugin-cli --plugin-file /usr/share/jenkins/ref/plugins.txt
After startup, capture
/pluginManager/api/json?depth=1 or the installed-plugin
UI as evidence. Never assume the requested plugin list proves the
resolved dependency graph actually loaded.
4. Inspect the metrics endpoint locally
The Prometheus plugin default endpoint is /prometheus/.
Keep the trailing slash and keep the endpoint restricted to the lab
network/loopback or an authenticated reverse proxy. Do not publish a
real controller’s metrics endpoint to the Internet.
set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
OUT='/tmp/jenkins-observe-evidence'
mkdir -p "$OUT"
# Supply approved lab authentication if your controller requires it.
curl -fsS "$JENKINS_URL/prometheus/" > "$OUT/prometheus-before.txt"
grep -E 'jenkins|vm_|jvm_|executor|queue' "$OUT/prometheus-before.txt" | head -n 80
Metric names can change with plugin versions. Treat the scrape itself as authoritative for the installed baseline; do not hard-code an old tutorial’s metric names into production alerts without checking them.
5. Capture queue, node and build identity snapshots
set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
OUT='/tmp/jenkins-observe-evidence'
TS="$(date -u +%Y%m%dT%H%M%SZ)"
# Add the disposable read-only lab auth in your environment; do not embed it here.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" > "$OUT/queue-$TS.json"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,assignedLabels[name]]" > "$OUT/nodes-$TS.json"
printf 'snapshot_utc=%s\n' "$TS" | tee -a "$OUT/timeline.txt"
Queue item IDs and build numbers are different identities. A queue item may never become a build; preserve both when the transition matters.
6. Create a bounded synthetic workload
The synthetic Pipeline requests only observe-linux,
records its build identity, sleeps long enough to create measurable
pressure, and writes a tiny evidence file. Keep concurrency
deliberately small.
pipeline {
agent { label 'observe-linux' }
options { timestamps(); disableConcurrentBuilds(abortPrevious: false) }
stages {
stage('Identify') {
steps {
sh '''set -eu
printf 'job=%s\n' "$JOB_NAME"
printf 'build=%s\n' "$BUILD_NUMBER"
printf 'node=%s\n' "$NODE_NAME"
printf 'workspace=%s\n' "$WORKSPACE"
date -u +observed_start_utc=%Y-%m-%dT%H:%M:%SZ
'''
}
}
stage('Synthetic work') {
steps { sh 'sleep 45' }
}
stage('Evidence') {
steps {
sh 'printf "build=%s node=%s\\n" "$BUILD_NUMBER" "$NODE_NAME" > observe.txt'
archiveArtifacts artifacts: 'observe.txt', fingerprint: true
}
}
}
}
Use synthetic source and a fixed Jenkinsfile revision in your local repository so every run’s code identity is known.
7. Generate queue pressure without creating a load test against Jenkins itself
Trigger a small number of builds—enough to exceed the eligible
executor count. For example, with one
observe-linux executor, start three builds. Record the
queue snapshots every few seconds for one minute. Do not use
unbounded loops or hundreds of jobs.
set -euo pipefail
# Pseudocode/CLI shape: trigger exactly three disposable builds using your approved lab token/crumb handling.
for i in 1 2 3; do
printf 'trigger synthetic build %s at %s\n' "$i" "$(date -u +%H:%M:%SZ)"
# curl -f -X POST "$JENKINS_URL/job/observe-lab/job/slow-ci/build" ...
sleep 2
done
Expected observation: one build executes, later builds remain
queued, and queue why/wait timestamps explain that
eligible capacity is occupied.
8. Induce one controlled agent failure
Depending on the step and Pipeline resilience behavior, the running build may fail, wait, or be retried only if you explicitly designed it that way. Do not teach “agent comes back, therefore every step resumes” as a universal rule.
9. Preserve logs from the right layers
| Evidence | Where/why |
|---|---|
| Build console | Exact job/build; stage/step timing and first visible error. |
| Controller log | Scheduling, Remoting/channel and plugin-side messages. |
| Agent process log | Connection/reconnect, Java/runtime or process failure. |
| Queue snapshots | Why waiting items could/could not schedule. |
| Metrics snapshots | How queue/executor/JVM signals changed over time. |
| Audit trail | Whether an authenticated actor changed node/job/controller state. |
| External/mock latency | Proves whether a dependency, rather than Jenkins scheduling, was slow. |
Jenkins uses java.util.logging; custom log recorders
can narrow specific packages when troubleshooting. Do not leave
verbose diagnostic logging enabled indefinitely without evaluating
volume and sensitive content.
10. Add an audit event safely
On the disposable controller, configure Audit Trail to a local file or test sink. Perform one benign, reversible administrative action such as temporarily changing the description of the disposable job, then restore it. Verify that the audit record contains the actor/action/target/time information you expect. Never point the chapter lab at an enterprise SIEM or production endpoint.
11. Build the correlation timeline
UTC timeline
------------
18:00:00 build #41 triggered
18:00:01 queue item 912 created; requires observe-linux
18:00:03 build #41 allocated to agent observe-01
18:00:05 build #42 queued; eligible executor busy
18:00:12 observe-01 channel lost
18:00:13 controller log records Remoting disconnect
18:00:14 build #41 records agent/step failure
18:00:15 queue wait and executor availability diverge
18:00:22 audit: no node-disable action by administrator
18:00:30 observe-01 reconnect begins
This timeline turns independent telemetry into a causal hypothesis: agent connectivity loss reduced eligible capacity and affected the running build. It is stronger than “queue graph went up.”
12. Verification checklist
| Check | Pass criterion |
|---|---|
| Versions | Jenkins/Java/metrics/prometheus/audit versions recorded. |
| Endpoint | Metrics scrape succeeds only through intended lab access path. |
| Queue | At least one queued item has ID, in-queue timestamp and reason captured. |
| Executor | Eligible executor state is independently recorded. |
| Agent failure | Exact agent and affected build identified; first-failure logs preserved. |
| Audit | Benign administrative action produces expected actor/action/target evidence. |
| Correlation | UTC timeline links build/queue/agent/log/metric observations. |
| Secrets | No token, API key, credential value or inbound-agent secret appears in evidence. |
13. Cleanup and rollback
Archive the non-sensitive evidence packet outside the disposable controller if you want to retain it. Then stop/remove only the lab controller and agent, remove the local monitoring sink, and delete temporary evidence paths. Do not uninstall plugins from a shared controller as part of this lab.
set -euo pipefail
EVIDENCE='/tmp/jenkins-observe-evidence'
case "$EVIDENCE" in
/tmp/jenkins-observe-evidence) rm -rf -- "$EVIDENCE" ;;
*) echo 'refusing unexpected path' >&2; exit 70 ;;
esac
14. Challenge: identify the causal layer
Queue wait rises sharply. Controller heap is normal, three agents
are online, but every queued job requests
observe-gpu and none of the online agents has that
label. Should you increase controller heap or executor count?
No. The evidence identifies label/eligibility capacity, not JVM pressure. Fix the intended agent/label design or restore the required pool, then re-measure queue wait.
Knowledge check
Answer before revealing the explanation.
1. Why pin plugin versions in the disposable controller image?
It makes the lab reproducible and records the actual privileged controller dependencies that produced the telemetry.
2. What should you preserve before reconnecting a failed agent?
Affected job/build and queue IDs, node identity, first-failure console/controller/agent logs, and relevant metrics snapshots.
3. Why is a trailing slash important for the default Prometheus endpoint?
Current plugin documentation notes that omitting it produces a redirect that some scrapers do not handle correctly.
4. Why trigger only a few synthetic builds?
The goal is bounded queue evidence, not uncontrolled load or denial of service.
5. What does the audit trail help rule in or out during the agent incident?
Whether an authenticated administrative/user action changed the relevant Jenkins state around the incident time.
Official references and version notes
Observability interfaces, plugin versions and security behavior evolve. Re-check current primary documentation before applying these patterns to a real controller.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.