Chapter 03Lesson 04~105 minutes

JENKINS_HOME, Filesystem Layout, Controller Configuration, System Settings, Tools, and Global Properties: Diagnostics, Failure Modes, Security, and Performance

Diagnose JENKINS_HOME and controller-state failures with evidence: live XML edits, missing recovery keys, oversized controller data, workspace confusion, and bad filesystem ownership.

DiagnosticsPermissionsSecretsStorage growthRecovery

Learning objectives

  • Use an evidence-first sequence to distinguish controller/JVM/plugin failures from filesystem, permission, workspace, and external-system problems.
  • Explain why direct XML edits while Jenkins is running can be overwritten or leave disk and memory inconsistent.
  • Diagnose a deliberately broken restore volume with wrong ownership without weakening container security or running Jenkins permanently as root.
  • Detect controller disk growth caused by retained builds, archives, plugins/tools/caches, or accidental large files before tuning blindly.
  • Protect support evidence by avoiding secret dumps and by preserving first-failure logs before restarts or cleanup.

1. Evidence-first diagnostic sequence for controller state

  1. Record controller/core/Java/image/plugin baseline and exact container/service identity.
  2. Preserve controller logs and the first error before restart.
  3. Confirm active JENKINS_HOME, mount, ownership, free space, and filesystem health.
  4. Classify the symptom: startup/config load, plugin load, job/build data, permission, queue/agent, or external service.
  5. Compare safe metadata to a known-good snapshot; do not dump secret contents.
  6. Apply the smallest repair on a disposable copy when possible.
  7. Restart/reload only when the repair requires it, then verify controller health and affected state independently.
Preserve before repair: container/service logs, exact image/core/Java baseline, mount identity, file ownership, free space, relevant file timestamps, and backup identifiers. “Restart until it works” destroys chronology.

2. Failure mode: editing internal XML while Jenkins is running

Jenkins owns its serialized configuration. If an administrator edits config.xml behind a running controller, several bad outcomes are possible: the running process may ignore the change, a later save may overwrite it, malformed XML may fail on restart, or plugin/core schema expectations may be violated.

Use a safe comparison rather than a live mutation:

docker exec jenkins-ch03-controller sh -lc '
  cp "$JENKINS_HOME/config.xml" /tmp/config.xml.readonly-copy
  stat -c "%y %s %n" "$JENKINS_HOME/config.xml" /tmp/config.xml.readonly-copy
'

If a real low-level repair is ever unavoidable, stop Jenkins, take a consistent backup, work on a copy, preserve the original file, and validate the repaired controller in isolation. Routine changes belong in supported UI/JCasC/plugin interfaces.

3. Intentionally broken example: wrong ownership on a restore volume

This exercise uses only disposable volumes. First capture the UID/GID that the official image uses for the Jenkins account; do not hard-code it:

JENKINS_UID=$(docker run --rm --entrypoint /bin/sh jenkins/jenkins:2.568.3-jdk21 -c 'id -u jenkins')
JENKINS_GID=$(docker run --rm --entrypoint /bin/sh jenkins/jenkins:2.568.3-jdk21 -c 'id -g jenkins')
printf 'jenkins uid=%s gid=%s\n' "$JENKINS_UID" "$JENKINS_GID"

docker volume create jenkins-ch03-broken-home
# Make the volume deliberately inaccessible to the normal Jenkins user.
docker run --rm --user root --entrypoint /bin/sh \
  -v jenkins-ch03-broken-home:/var/jenkins_home \
  jenkins/jenkins:2.568.3-jdk21 -c '
    chown 0:0 /var/jenkins_home && chmod 700 /var/jenkins_home
  '

docker run -d --name jenkins-ch03-broken \
  -p 127.0.0.1:18085:8080 \
  -v jenkins-ch03-broken-home:/var/jenkins_home \
  jenkins/jenkins:2.568.3-jdk21
sleep 5
docker logs --tail 80 jenkins-ch03-broken

The expected evidence is a startup/write-permission failure, not a healthy wizard. Repair ownership using the image’s actual UID/GID, then retry:

docker rm -f jenkins-ch03-broken

docker run --rm --user root --entrypoint /bin/sh \
  -e JUID="$JENKINS_UID" -e JGID="$JENKINS_GID" \
  -v jenkins-ch03-broken-home:/var/jenkins_home \
  jenkins/jenkins:2.568.3-jdk21 -c '
    chown -R "$JUID:$JGID" /var/jenkins_home
  '

docker run -d --name jenkins-ch03-broken \
  -p 127.0.0.1:18085:8080 \
  -v jenkins-ch03-broken-home:/var/jenkins_home \
  jenkins/jenkins:2.568.3-jdk21
sleep 5
docker logs --tail 80 jenkins-ch03-broken
Do not “fix” production by running Jenkins as root. Correct the volume ownership/permissions for the intended service identity. Root is used here only in a short-lived repair helper container against a disposable volume.

4. Failure mode: backup without the recovery keys

Encrypted credential records are not useful if the keys required to decrypt them are lost. Conversely, backup data plus key material in the same broadly accessible archive collapses the protection boundary. Diagnose this as a recovery-design failure, not by printing key files.

Observation Likely layer Safe next action
Controller starts but encrypted values cannot be used/decrypted Secret-key recovery mismatch/loss Stop; verify you are using the matching separately protected secrets recovery material from that controller snapshot.
Backup archive contains readable secret-key files alongside routine config Backup trust boundary Do not share it; redesign storage/access and rotate exposed external credentials if compromise is possible.
Restored controller has different instance identity unexpectedly Incomplete/mismatched recovery state Compare snapshot provenance and recovery material; do not copy random keys from another controller.

5. Failure mode: controller disk growth

Large build histories, archived artifacts, logs, plugin/tool caches, and accidental files can fill controller storage. Diagnose the owner before deleting anything:

docker exec jenkins-ch03-controller sh -lc '
  df -h "$JENKINS_HOME"
  du -sh "$JENKINS_HOME"/* 2>/dev/null | sort -h | tail -25
  find "$JENKINS_HOME/jobs" -type f -size +100M \
    ! -path "*/secrets/*" -printf "%s %p\n" 2>/dev/null \
    | sort -nr | head -20
'

Do not immediately delete the largest directory. A large jobs/.../builds tree may be required audit history; a large archive may be an intentionally retained deliverable; a cache may be safely reconstructable. Define retention and external artifact storage before cleanup.

6. Failure mode: confusing agent workspace with controller home

A user reports, “The file was in the workspace yesterday, so Jenkins should restore it.” Ask where the build ran. A static agent may keep a workspace locally; an ephemeral container/pod may already be gone. Neither is automatically part of the controller backup.

The production pattern is explicit dataflow: source from SCM, dependencies from controlled repositories/caches, build outputs archived or published, and release artifacts identified by digest/version. Workspaces are disposable execution surfaces.

7. Failure mode: restoring data with the wrong core/plugin baseline

Controller data is interpreted by Jenkins core and plugins. Restoring a home directory while simultaneously changing Java, Jenkins core, or many plugin versions turns one recovery operation into several migrations at once. Keep the original baseline as evidence and restore it first where practical.

docker inspect --format='image={{.Config.Image}}' jenkins-ch03-controller
# In the UI: Manage Jenkins → Plugins → Installed
# In evidence: record plugin names + versions, not configuration secrets.

Only after the recovered controller is stable should an upgrade be treated as its own planned change with compatibility testing and rollback limits.

8. Performance diagnosis: measure the controller before tuning

Disk I/O and filesystem scale can affect controller responsiveness, especially with many builds, logs, small files, plugin activity, and backups. But adding heap or executors will not fix a saturated storage device. Correlate:

  • controller CPU/heap/GC and thread state,
  • free disk/inode capacity and I/O latency,
  • size/count of jobs/builds/artifacts/logs,
  • queue wait versus executor/agent availability,
  • backup/snapshot windows, antivirus/indexing, or network filesystems.

Change one bottleneck hypothesis at a time and re-measure. Chapter 37/38 develops observability and performance engineering in depth.

9. Recovery playbook

  1. Stop destructive automation and preserve first-failure evidence.
  2. Verify exact controller/image/core/Java/plugin baseline.
  3. Verify mount/path/owner/permissions/capacity.
  4. Identify whether the broken state is controller configuration, plugin data, job/build data, key material, workspace, or external system.
  5. Restore/repair on an isolated copy if the change is high risk.
  6. Use matching separately protected recovery keys only for the matching controller backup.
  7. Start on a non-conflicting loopback port and inspect logs before exposing it.
  8. Verify expected settings/jobs/history and external integrations independently.
  9. Document cause, repair, and prevention before cleanup.

10. Clean up the deliberately broken controller

docker rm -f jenkins-ch03-broken 2>/dev/null || true
docker volume rm jenkins-ch03-broken-home 2>/dev/null || true

11. Summary

  • Controller-state incidents are diagnosed from evidence, not by blind restart or raw XML surgery.
  • Filesystem ownership must match the intended Jenkins service identity; running permanently as root is not a repair.
  • Encrypted data and recovery keys form a matched recovery set but belong in separate protection domains.
  • Disk growth needs classification and retention decisions before deletion.
  • Agent workspaces and external systems remain outside controller-home recovery semantics.
Next lesson

Checkpoint Lab — snapshot and restore

Prove the entire operating model by rebuilding a disposable controller from separated backup volumes and verifying what returns and what intentionally does not.

Knowledge check

Why is “run Jenkins as root” the wrong fix for a volume permission failure?

What is the first thing to do before a restart when diagnosing a controller-state failure?

Why should large controller directories not be deleted solely because du ranks them first?

A restored controller starts, but credentials fail. Which layer should you investigate before the network?

Why is a workspace file poor recovery evidence?

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked on 2026-09-14. The disposable baseline continues Chapter 02 with Jenkins 2.568.3 LTS on Java 21 using the official jenkins/jenkins:2.568.3-jdk21 image. Jenkins and plugins evolve; regenerate the storage map from the actual controller, keep plugin-specific files opaque unless the plugin documents them, and re-check backup/security guidance before applying the patterns to a production controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.