Chapter 37Lesson 03~200 minutes

Monitoring Jenkins with Metrics, Logs, Health Checks, Audit Trails, Queue Telemetry, and Observability: Configuration, Design Choices, and Tradeoffs

Design observability around questions and budgets: select metrics transport, log ownership, alert semantics, aggregation level and label dimensions without adding avoidable controller load or data leakage.

JMXPrometheuscentralized logsSLO alertscardinalitytradeoffs

Learning objectives

  • Compare JMX and Jenkins metrics/Prometheus paths.
  • Choose local versus centralized logs based on failure and retention needs.
  • Distinguish symptom alerts from SLO/error-budget alerts.
  • Balance per-job visibility with platform aggregate cost.
  • Control metric-label cardinality and sensitive dimensions.

1. Start with the operational question

The wrong design question is “Which dashboard product should we use?” The right question is “What decision must an operator make, and what evidence makes that decision safe?” If the question is why builds start slowly, you need queue wait, requested labels and eligible executor capacity. If the question is whether a controller upgrade hurt responsiveness, you need before/after controller latency, heap/GC, thread and disk evidence tied to the change window.

2. JMX versus Prometheus/Metrics

Choice Strengths Costs/risks Good fit
JMX/JVM tooling Deep JVM-native data; standard Java tooling; can inspect memory/threads/GC Remote JMX security/network setup can be complex; may expose sensitive runtime internals JVM-focused diagnostics and controlled platform monitoring.
Metrics plugin Jenkins-aware metric/health interfaces with permissions/API-key concepts Adds privileged plugin dependency; metric semantics/version compatibility must be tracked Controller health/metrics where Jenkins-native interfaces are desired.
Prometheus plugin Scrape-friendly endpoint; integrates with Prometheus/Grafana ecosystem; Jenkins-specific metrics Plugin/dependency overhead; endpoint/network/auth exposure; cardinality/storage cost Time-series monitoring and alerts in organizations already using Prometheus.
Remote API snapshots No dedicated metrics stack required for basic queue/node/job state Polling cost; not a full time-series model; auth/rate/retention must be handled Small lab and targeted automation/verification.

These are not mutually exclusive. Use the smallest set that answers your reliability questions. More collectors can also mean more controller work, dependencies and data exposure.

3. Local logs versus centralized logging

Local controller/agent logs are closest to the process and remain essential during collector outages. Centralized logs improve retention, cross-host search and correlation—especially for ephemeral agents that may disappear after a failed build. The tradeoff is transport cost, access control, sensitive-data handling and the risk that a central pipeline drops or transforms the exact first-failure detail you need.

Design Use when Required control
Local only Small disposable lab, short retention Ensure logs survive long enough for the exercise and are copied before deleting ephemeral resources.
Centralized controller logs Multiple controllers/long retention/cross-service incidents TLS/auth, tenant boundaries, redaction, retention and ingestion backpressure.
Ephemeral-agent log shipping Agents/pods terminate quickly after work Flush before termination or preserve platform/container logs keyed by pod/node/build.
Build console retention Pipeline failure diagnosis and audit needs Build-retention policy, secret masking limits, access permissions and size controls.

4. Symptom alert versus SLO alert

A symptom alert says a component looks unusual; an SLO alert says user-visible reliability is burning faster than accepted. Both can be useful. A controller-process-down alert is immediate and binary. A queue-wait SLO often needs a sustained window and workload context so a planned batch burst does not page an operator unnecessarily.

Avoid making every failed build an infrastructure page. Source-code test failures are developer feedback; widespread agent-connectivity failures or queue-wait budget burn are platform signals.

5. Per-job metrics versus platform aggregates

Per-job dimensions help identify noisy or slow workloads, but thousands of branch jobs can create expensive cardinality and long-term storage. Platform aggregates are cheap but can hide one critical team. Start with low-cardinality dimensions—controller, trusted pool, label class, workload class—and add job-level drill-down only where there is a clear operational owner and retention budget.

6. High-cardinality labels: the hidden observability failure mode

A metric label creates one time series for every unique combination. Labeling metrics with build number, commit SHA, workspace path, PR number and ephemeral pod name can explode series count. Those values are often better in logs/traces or short-lived evidence records.

Candidate dimension Metric label? Reason
controller Usually yes Small bounded set and primary service boundary.
agent pool / stable label Usually yes Useful capacity dimension if vocabulary is governed.
job family Sometimes Useful but can grow; define ownership/retention.
job full name Cautiously Multibranch/organization jobs can produce many unique names.
build number Usually no Unbounded per run; preserve in logs/events.
commit SHA No Near-unique; use build metadata/log correlation.
ephemeral pod UID No for long-lived metrics High churn; retain in event/log evidence instead.
result class Yes Small finite set useful for rates.

7. Health endpoints and probe semantics

A health endpoint should answer a specific readiness/liveness question. The Jenkins Metrics APIs include health concepts, but permission and exposure rules matter. A load balancer probe that only checks “HTTP responds” should not be presented as proof that the queue can schedule or that downstream delivery dependencies are healthy.

8. Audit plugin versus platform/reverse-proxy audit

The Audit Trail plugin can record Jenkins requests/actions in a controller-aware way, while an identity provider, reverse proxy or OS may provide complementary records. Choose the source based on the question. “Who authenticated?” may be answered by IdP logs; “who changed this Jenkins item?” needs Jenkins-context evidence. Correlate rather than pretending one log is complete.

9. Worked decision: a 40-team Jenkins platform

Requirement Decision Why/evidence
Queue SLO Prometheus time series grouped by stable pool/label class Supports p95 wait and executor saturation without build-number labels.
Deep JVM incident JMX/local JVM tools plus controller logs Used on demand for heap/GC/thread diagnosis rather than as the only platform SLI.
Ephemeral agent failures Centralize platform/agent logs with build/node correlation Agent filesystem may vanish before investigation.
Job failure feedback Keep developer-facing build result/report in Jenkins; platform alerts use causal taxonomy Avoid paging platform operators for ordinary test failures.
Administrative accountability Audit Trail to restricted local/central sink plus IdP/proxy logs Separates Jenkins action from authentication context.
Cost control No build-number/SHA metric labels; bounded retention for per-job series Prevents cardinality-driven monitoring failure.

10. Configuration ownership and rollback

Version monitoring configuration, dashboard queries, alert rules and log-routing policy just like other platform code. Record plugin versions and endpoint exposure. Before adding a new collector or high-frequency scrape, establish a baseline of controller load; if the monitoring system causes measurable degradation, roll back that one change and re-measure.

Next

Diagnose observability failures

Lesson 4 intentionally breaks the monitoring design: public metrics, raw-failure alerts, missing ephemeral logs, cardinality explosion and queue/executor miscorrelation.

Knowledge check

Answer before revealing the explanation.

1. Why might a platform use both JMX and Prometheus rather than choosing exactly one?

2. Why is build number usually a poor long-lived metric label?

3. What is the problem with paging on every failed Jenkins build?

4. Why centralize ephemeral-agent logs?

5. What must accompany a new high-frequency collector?

Official references and version notes

Observability interfaces, plugin versions and security behavior evolve. Re-check current primary documentation before applying these patterns to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.