Chapter 37Lesson 02~220 minutes

Monitoring Jenkins with Metrics, Logs, Health Checks, Audit Trails, Queue Telemetry, and Observability: Guided Hands-On Workflow and Core Operations

Create a disposable observability pack: pin the monitoring plugins, collect controller/queue/executor signals, generate bounded load, induce one safe agent failure, and correlate metrics, logs and build IDs without exposing real infrastructure.

PrometheusMetricsAudit Trailqueue APIsynthetic loadcorrelation

Learning objectives

  • Preflight a disposable controller/agent before enabling monitoring plugins.
  • Use a pinned Prometheus/Metrics path without exposing it publicly.
  • Capture queue and executor evidence before, during and after synthetic pressure.
  • Preserve controller, build and agent logs around a controlled failure.
  • Build a small correlation packet keyed by UTC time, job/build and agent identity.

1. Lab scenario and safety boundary

Use a local disposable controller named conceptually jenkins-observe-lab, a synthetic Pipeline job observe-lab/slow-ci, and at least one disposable agent labeled observe-linux. Do not run routine lab builds on the built-in node. No production SCM, chat, cloud, identity provider or deployment target is involved.

2. Preflight the exact baseline

Evidence record
---------------
controller = jenkins-observe-lab
jenkins_lts = 2.568.3
controller_java = 21
agent_java = 21
metrics = 4.2.37-494.v06f9a_939d33a_
prometheus = 860.v532442b_44e9a_
audit-trail = 456.v39d2fd1ed556
job_full_name = observe-lab/slow-ci
agent_label = observe-linux
external_targets = none (local-only lab)

Before installation, inspect the plugin pages for current minimum Jenkins versions, dependencies and security warnings. The dated versions above are compatible with this chapter baseline, but that does not make them evergreen recommendations.

3. Pin the optional observability plugins in a disposable controller image

The safest repeatable pattern is to build/recreate the lab controller with an explicit plugin list rather than clicking “latest” in Plugin Manager. Example for an image build context:

# plugins.txt
metrics:4.2.37-494.v06f9a_939d33a_
prometheus:860.v532442b_44e9a_
audit-trail:456.v39d2fd1ed556
FROM jenkins/jenkins:2.568.3-jdk21
COPY plugins.txt /usr/share/jenkins/ref/plugins.txt
RUN jenkins-plugin-cli --plugin-file /usr/share/jenkins/ref/plugins.txt

After startup, capture /pluginManager/api/json?depth=1 or the installed-plugin UI as evidence. Never assume the requested plugin list proves the resolved dependency graph actually loaded.

4. Inspect the metrics endpoint locally

The Prometheus plugin default endpoint is /prometheus/. Keep the trailing slash and keep the endpoint restricted to the lab network/loopback or an authenticated reverse proxy. Do not publish a real controller’s metrics endpoint to the Internet.

set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
OUT='/tmp/jenkins-observe-evidence'
mkdir -p "$OUT"
# Supply approved lab authentication if your controller requires it.
curl -fsS "$JENKINS_URL/prometheus/" > "$OUT/prometheus-before.txt"
grep -E 'jenkins|vm_|jvm_|executor|queue' "$OUT/prometheus-before.txt" | head -n 80

Metric names can change with plugin versions. Treat the scrape itself as authoritative for the installed baseline; do not hard-code an old tutorial’s metric names into production alerts without checking them.

5. Capture queue, node and build identity snapshots

set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
OUT='/tmp/jenkins-observe-evidence'
TS="$(date -u +%Y%m%dT%H%M%SZ)"
# Add the disposable read-only lab auth in your environment; do not embed it here.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" > "$OUT/queue-$TS.json"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,assignedLabels[name]]" > "$OUT/nodes-$TS.json"
printf 'snapshot_utc=%s\n' "$TS" | tee -a "$OUT/timeline.txt"

Queue item IDs and build numbers are different identities. A queue item may never become a build; preserve both when the transition matters.

6. Create a bounded synthetic workload

The synthetic Pipeline requests only observe-linux, records its build identity, sleeps long enough to create measurable pressure, and writes a tiny evidence file. Keep concurrency deliberately small.

pipeline {
  agent { label 'observe-linux' }
  options { timestamps(); disableConcurrentBuilds(abortPrevious: false) }
  stages {
    stage('Identify') {
      steps {
        sh '''set -eu
          printf 'job=%s\n' "$JOB_NAME"
          printf 'build=%s\n' "$BUILD_NUMBER"
          printf 'node=%s\n' "$NODE_NAME"
          printf 'workspace=%s\n' "$WORKSPACE"
          date -u +observed_start_utc=%Y-%m-%dT%H:%M:%SZ
        '''
      }
    }
    stage('Synthetic work') {
      steps { sh 'sleep 45' }
    }
    stage('Evidence') {
      steps {
        sh 'printf "build=%s node=%s\\n" "$BUILD_NUMBER" "$NODE_NAME" > observe.txt'
        archiveArtifacts artifacts: 'observe.txt', fingerprint: true
      }
    }
  }
}

Use synthetic source and a fixed Jenkinsfile revision in your local repository so every run’s code identity is known.

7. Generate queue pressure without creating a load test against Jenkins itself

Trigger a small number of builds—enough to exceed the eligible executor count. For example, with one observe-linux executor, start three builds. Record the queue snapshots every few seconds for one minute. Do not use unbounded loops or hundreds of jobs.

set -euo pipefail
# Pseudocode/CLI shape: trigger exactly three disposable builds using your approved lab token/crumb handling.
for i in 1 2 3; do
  printf 'trigger synthetic build %s at %s\n' "$i" "$(date -u +%H:%M:%SZ)"
  # curl -f -X POST "$JENKINS_URL/job/observe-lab/job/slow-ci/build" ...
  sleep 2
done

Expected observation: one build executes, later builds remain queued, and queue why/wait timestamps explain that eligible capacity is occupied.

8. Induce one controlled agent failure

Depending on the step and Pipeline resilience behavior, the running build may fail, wait, or be retried only if you explicitly designed it that way. Do not teach “agent comes back, therefore every step resumes” as a universal rule.

9. Preserve logs from the right layers

Evidence Where/why
Build console Exact job/build; stage/step timing and first visible error.
Controller log Scheduling, Remoting/channel and plugin-side messages.
Agent process log Connection/reconnect, Java/runtime or process failure.
Queue snapshots Why waiting items could/could not schedule.
Metrics snapshots How queue/executor/JVM signals changed over time.
Audit trail Whether an authenticated actor changed node/job/controller state.
External/mock latency Proves whether a dependency, rather than Jenkins scheduling, was slow.

Jenkins uses java.util.logging; custom log recorders can narrow specific packages when troubleshooting. Do not leave verbose diagnostic logging enabled indefinitely without evaluating volume and sensitive content.

10. Add an audit event safely

On the disposable controller, configure Audit Trail to a local file or test sink. Perform one benign, reversible administrative action such as temporarily changing the description of the disposable job, then restore it. Verify that the audit record contains the actor/action/target/time information you expect. Never point the chapter lab at an enterprise SIEM or production endpoint.

11. Build the correlation timeline

UTC timeline
------------
18:00:00 build #41 triggered
18:00:01 queue item 912 created; requires observe-linux
18:00:03 build #41 allocated to agent observe-01
18:00:05 build #42 queued; eligible executor busy
18:00:12 observe-01 channel lost
18:00:13 controller log records Remoting disconnect
18:00:14 build #41 records agent/step failure
18:00:15 queue wait and executor availability diverge
18:00:22 audit: no node-disable action by administrator
18:00:30 observe-01 reconnect begins

This timeline turns independent telemetry into a causal hypothesis: agent connectivity loss reduced eligible capacity and affected the running build. It is stronger than “queue graph went up.”

12. Verification checklist

Check Pass criterion
Versions Jenkins/Java/metrics/prometheus/audit versions recorded.
Endpoint Metrics scrape succeeds only through intended lab access path.
Queue At least one queued item has ID, in-queue timestamp and reason captured.
Executor Eligible executor state is independently recorded.
Agent failure Exact agent and affected build identified; first-failure logs preserved.
Audit Benign administrative action produces expected actor/action/target evidence.
Correlation UTC timeline links build/queue/agent/log/metric observations.
Secrets No token, API key, credential value or inbound-agent secret appears in evidence.

13. Cleanup and rollback

Archive the non-sensitive evidence packet outside the disposable controller if you want to retain it. Then stop/remove only the lab controller and agent, remove the local monitoring sink, and delete temporary evidence paths. Do not uninstall plugins from a shared controller as part of this lab.

set -euo pipefail
EVIDENCE='/tmp/jenkins-observe-evidence'
case "$EVIDENCE" in
  /tmp/jenkins-observe-evidence) rm -rf -- "$EVIDENCE" ;;
  *) echo 'refusing unexpected path' >&2; exit 70 ;;
esac

14. Challenge: identify the causal layer

Queue wait rises sharply. Controller heap is normal, three agents are online, but every queued job requests observe-gpu and none of the online agents has that label. Should you increase controller heap or executor count?

No. The evidence identifies label/eligibility capacity, not JVM pressure. Fix the intended agent/label design or restore the required pool, then re-measure queue wait.

Next

Choose the observability architecture

Lesson 3 compares JMX with Prometheus/Metrics, local versus centralized logs, symptom alerts versus SLOs, per-job versus platform metrics, and safe label/cardinality design.

Knowledge check

Answer before revealing the explanation.

1. Why pin plugin versions in the disposable controller image?

2. What should you preserve before reconnecting a failed agent?

3. Why is a trailing slash important for the default Prometheus endpoint?

4. Why trigger only a few synthetic builds?

5. What does the audit trail help rule in or out during the agent incident?

Official references and version notes

Observability interfaces, plugin versions and security behavior evolve. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.