Chapter 37Lesson 01~190 minutes

Monitoring Jenkins with Metrics, Logs, Health Checks, Audit Trails, Queue Telemetry, and Observability: Concepts, Architecture, and Mental Model

Treat observability as an evidence system, not a collection of pretty graphs: connect controller/JVM health, queue pressure, agent capacity, build outcomes, logs, audit actions and external latency to explicit reliability questions.

observabilitymetricslogsqueueauditSLO

Learning objectives

  • Explain how metrics, logs, health checks and audit events answer different operational questions.
  • Model queue wait, executor utilization and agent availability as separate scheduling evidence.
  • Correlate controller/JVM, job/build, agent and external-service state with stable identifiers and timestamps.
  • Define useful SLIs/SLOs instead of alerting on raw counts.
  • Recognize observability data as potentially sensitive operational information.

1. The practical problem: a green dashboard can still hide a broken delivery system

Jenkins coordinates many independently failing layers. The controller process can answer HTTP requests while every Linux agent is offline. Agents can be online while the queue grows because the jobs require another label. Builds can fail because source code is wrong, because a credential expired, because an artifact repository is slow, or because the controller is starved for disk I/O. A single “Jenkins up” probe cannot distinguish any of those conditions.

Observability makes those layers inspectable. The goal is not to collect every possible metric. The goal is to preserve enough causal evidence to answer: what changed, which layer is constrained, who or what acted, which jobs/builds were affected, and what intervention is justified.

2. Mental model: events become evidence, then action

Mental model: events become evidence, then action
flowchart TD
  A["Controller / JVM events"] --> E["Collectors"]
  B["Queue / executor / agent events"] --> E
  C["Job / build / Pipeline events"] --> E
  D["External dependency events"] --> E
  E --> F["Metrics + logs + health + audit records"]
  F --> G["Time labels + job/build/agent IDs"]
  G --> H["Dashboards / queries / alerts"]
  H --> I{"Question answered?"}
  I -->|capacity| J["Scale or rebalance safely"]
  I -->|failure| K["Diagnose causal layer"]
  I -->|security| L["Review actor / action"]
  I -->|unclear| M["Collect narrower evidence"]
  J --> N["Measure again"]
  K --> N
  L --> N
  M --> E
  

The central arrow is correlation. A graph without job/build/agent identity may show that something changed but still fail to tell you which workload or dependency caused it.

3. Keep the observable state layers separate

Layer Representative evidence What it can prove
Controller/JVM uptime, process health, heap, GC, threads, CPU, disk I/O Whether the Jenkins service/runtime is under pressure; not why a specific test failed.
Queue queue length, age/wait time, blocked/stuck reason, requested label Whether demand is waiting and why scheduling has not advanced.
Executors busy/idle executors, node labels, executor count Whether eligible execution capacity exists; not whether an online node matches the job.
Agents/Remoting online/offline cause, channel state, reconnects, JVM/OS metadata Whether the execution endpoint is reachable and compatible.
Build/run job full name, build number, cause, source SHA, duration, result Which workload changed and what Jenkins recorded about its run.
Pipeline/steps stage/step timing, durable-task output, CPS state Where Pipeline execution spent time or failed.
Logs controller logs, agent logs, build console, plugin logs Detailed event context and first-failure messages.
Audit authenticated actor, action, item/config target, timestamp Who/what changed administratively observable state; scope depends on mechanism/plugin.
External systems SCM, registry, artifact repo, IdP, Sonar, API latency/status Whether a dependency is contributing to Jenkins-visible delay/failure.
Recovery restart/restore timestamps, post-recovery health Whether service returned to the intended operational state.

4. Metrics, logs, health checks and audit trails are not interchangeable

Metrics summarize measurements over time and are ideal for trends, saturation and alert thresholds. Logs preserve detailed events and errors. Health checks answer a bounded service-readiness question. Audit trails focus on accountable administrative/user actions. A mature incident investigation often needs all four.

5. Read-only inspection before changing anything

Start with observations that do not mutate the controller. Exact endpoints and permissions vary, so use an authenticated disposable account with only the read permissions needed.

set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
# Auth header intentionally omitted here; use a disposable read-only lab identity.
curl -fsS "$JENKINS_URL/api/json?tree=mode,nodeDescription,numExecutors,quietingDown" | python -m json.tool
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" | python -m json.tool
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,assignedLabels[name]]" | python -m json.tool

For a protected real controller, never “solve” a 403 by weakening authorization. Grant a dedicated monitoring identity only the necessary visibility or use an approved collector integration.

6. JVM evidence: heap alone is not a diagnosis

Heap usage, garbage-collection pauses and thread activity help explain controller responsiveness, but one high heap reading is not sufficient evidence to increase memory. Look for sustained pressure, GC behavior, workload timing, thread contention and disk/network symptoms together. Chapter 38 will turn these measurements into performance hypotheses.

7. Queue telemetry: measure waiting, not only queue length

A queue of ten items waiting for two seconds is different from two items waiting for twenty minutes. Track both cardinality and age/wait time, then segment by useful dimensions such as job family or required label. Read the queue reason: “waiting for next available executor”, “no nodes with label”, a lock, throttling or a quiet-down state imply different remedies.

8. From raw telemetry to an SLO

An SLI is the measured indicator; an SLO is the target. Choose an indicator tied to user-visible delivery behavior. For example, “95% of eligible CI builds start execution within 120 seconds during business hours” is more actionable than “queue length < 10”. The first can fail because of capacity, label fragmentation or agent loss and can be investigated with queue/executor evidence.

Signal Weak alert More actionable formulation
Queue Queue length > 10 95th percentile eligible queue wait > 120s for 10m, grouped by required label.
Build failures More than 20 failures Failure-rate increase above baseline after excluding intentional test/fault-injection jobs; group by causal taxonomy.
Agents Any agent offline Required trusted agent pool capacity below minimum for 5m, with offline cause attached.
Controller Heap > 70% once Sustained heap pressure + GC pause/throughput degradation + controller response latency.
External API One slow request p95 dependency latency/error budget burn correlated to affected Jenkins builds.

9. Observability has a security boundary

Metrics can reveal job names, labels, repository hints, hostnames, capacity and timing. Logs can contain parameters, URLs, stack traces and accidentally emitted secrets. Audit events identify users and actions. Thread dumps/support data may expose even more internal detail. Apply authentication, network restriction, retention, redaction and least privilege just as you would for other operational data.

10. DevOps connection: evidence closes the improvement loop

Observability connects delivery behavior to operating decisions. It lets a platform team test whether adding an agent pool reduced queue wait, whether an upgrade changed GC behavior, whether a plugin update increased controller errors, and whether a recovery actually restored expected service. The useful unit is therefore not “a dashboard”; it is a repeatable chain from exact build/source and execution context to measured system behavior and a verified action.

Next

Build a disposable observability workflow

Lesson 2 enables a maintained metrics path, captures queue/executor evidence, creates safe synthetic pressure, inspects logs/health, and correlates one controlled agent failure by timestamps and build IDs.

Knowledge check

Answer before revealing the explanation.

1. Why is controller HTTP 200 not enough to declare Jenkins healthy?

2. Why track queue wait time as well as queue length?

3. What does an audit event add that a metric usually does not?

4. Why can high heap usage be a symptom rather than the root cause?

5. What makes an SLO more useful than a decorative graph?

Official references and version notes

Observability interfaces, plugin versions and security behavior evolve. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.