JENKINS_HOME, Filesystem Layout, Controller Configuration, System Settings, Tools, and Global Properties: Diagnostics, Failure Modes, Security, and Performance
Diagnose JENKINS_HOME and controller-state failures with evidence: live XML edits, missing recovery keys, oversized controller data, workspace confusion, and bad filesystem ownership.
Learning objectives
- Use an evidence-first sequence to distinguish controller/JVM/plugin failures from filesystem, permission, workspace, and external-system problems.
- Explain why direct XML edits while Jenkins is running can be overwritten or leave disk and memory inconsistent.
- Diagnose a deliberately broken restore volume with wrong ownership without weakening container security or running Jenkins permanently as root.
- Detect controller disk growth caused by retained builds, archives, plugins/tools/caches, or accidental large files before tuning blindly.
- Protect support evidence by avoiding secret dumps and by preserving first-failure logs before restarts or cleanup.
1. Evidence-first diagnostic sequence for controller state
- Record controller/core/Java/image/plugin baseline and exact container/service identity.
- Preserve controller logs and the first error before restart.
-
Confirm active
JENKINS_HOME, mount, ownership, free space, and filesystem health. - Classify the symptom: startup/config load, plugin load, job/build data, permission, queue/agent, or external service.
- Compare safe metadata to a known-good snapshot; do not dump secret contents.
- Apply the smallest repair on a disposable copy when possible.
- Restart/reload only when the repair requires it, then verify controller health and affected state independently.
2. Failure mode: editing internal XML while Jenkins is running
Jenkins owns its serialized configuration. If an administrator edits
config.xml behind a running controller, several bad
outcomes are possible: the running process may ignore the change, a
later save may overwrite it, malformed XML may fail on restart, or
plugin/core schema expectations may be violated.
Use a safe comparison rather than a live mutation:
docker exec jenkins-ch03-controller sh -lc '
cp "$JENKINS_HOME/config.xml" /tmp/config.xml.readonly-copy
stat -c "%y %s %n" "$JENKINS_HOME/config.xml" /tmp/config.xml.readonly-copy
'
If a real low-level repair is ever unavoidable, stop Jenkins, take a consistent backup, work on a copy, preserve the original file, and validate the repaired controller in isolation. Routine changes belong in supported UI/JCasC/plugin interfaces.
3. Intentionally broken example: wrong ownership on a restore volume
This exercise uses only disposable volumes. First capture the UID/GID that the official image uses for the Jenkins account; do not hard-code it:
JENKINS_UID=$(docker run --rm --entrypoint /bin/sh jenkins/jenkins:2.568.3-jdk21 -c 'id -u jenkins')
JENKINS_GID=$(docker run --rm --entrypoint /bin/sh jenkins/jenkins:2.568.3-jdk21 -c 'id -g jenkins')
printf 'jenkins uid=%s gid=%s\n' "$JENKINS_UID" "$JENKINS_GID"
docker volume create jenkins-ch03-broken-home
# Make the volume deliberately inaccessible to the normal Jenkins user.
docker run --rm --user root --entrypoint /bin/sh \
-v jenkins-ch03-broken-home:/var/jenkins_home \
jenkins/jenkins:2.568.3-jdk21 -c '
chown 0:0 /var/jenkins_home && chmod 700 /var/jenkins_home
'
docker run -d --name jenkins-ch03-broken \
-p 127.0.0.1:18085:8080 \
-v jenkins-ch03-broken-home:/var/jenkins_home \
jenkins/jenkins:2.568.3-jdk21
sleep 5
docker logs --tail 80 jenkins-ch03-broken
The expected evidence is a startup/write-permission failure, not a healthy wizard. Repair ownership using the image’s actual UID/GID, then retry:
docker rm -f jenkins-ch03-broken
docker run --rm --user root --entrypoint /bin/sh \
-e JUID="$JENKINS_UID" -e JGID="$JENKINS_GID" \
-v jenkins-ch03-broken-home:/var/jenkins_home \
jenkins/jenkins:2.568.3-jdk21 -c '
chown -R "$JUID:$JGID" /var/jenkins_home
'
docker run -d --name jenkins-ch03-broken \
-p 127.0.0.1:18085:8080 \
-v jenkins-ch03-broken-home:/var/jenkins_home \
jenkins/jenkins:2.568.3-jdk21
sleep 5
docker logs --tail 80 jenkins-ch03-broken
4. Failure mode: backup without the recovery keys
Encrypted credential records are not useful if the keys required to decrypt them are lost. Conversely, backup data plus key material in the same broadly accessible archive collapses the protection boundary. Diagnose this as a recovery-design failure, not by printing key files.
| Observation | Likely layer | Safe next action |
|---|---|---|
| Controller starts but encrypted values cannot be used/decrypted | Secret-key recovery mismatch/loss |
Stop; verify you are using the matching separately protected
secrets recovery material from that controller
snapshot.
|
| Backup archive contains readable secret-key files alongside routine config | Backup trust boundary | Do not share it; redesign storage/access and rotate exposed external credentials if compromise is possible. |
| Restored controller has different instance identity unexpectedly | Incomplete/mismatched recovery state | Compare snapshot provenance and recovery material; do not copy random keys from another controller. |
5. Failure mode: controller disk growth
Large build histories, archived artifacts, logs, plugin/tool caches, and accidental files can fill controller storage. Diagnose the owner before deleting anything:
docker exec jenkins-ch03-controller sh -lc '
df -h "$JENKINS_HOME"
du -sh "$JENKINS_HOME"/* 2>/dev/null | sort -h | tail -25
find "$JENKINS_HOME/jobs" -type f -size +100M \
! -path "*/secrets/*" -printf "%s %p\n" 2>/dev/null \
| sort -nr | head -20
'
Do not immediately delete the largest directory. A large
jobs/.../builds tree may be required audit history; a
large archive may be an intentionally retained deliverable; a cache
may be safely reconstructable. Define retention and external
artifact storage before cleanup.
6. Failure mode: confusing agent workspace with controller home
A user reports, “The file was in the workspace yesterday, so Jenkins should restore it.” Ask where the build ran. A static agent may keep a workspace locally; an ephemeral container/pod may already be gone. Neither is automatically part of the controller backup.
The production pattern is explicit dataflow: source from SCM, dependencies from controlled repositories/caches, build outputs archived or published, and release artifacts identified by digest/version. Workspaces are disposable execution surfaces.
7. Failure mode: restoring data with the wrong core/plugin baseline
Controller data is interpreted by Jenkins core and plugins. Restoring a home directory while simultaneously changing Java, Jenkins core, or many plugin versions turns one recovery operation into several migrations at once. Keep the original baseline as evidence and restore it first where practical.
docker inspect --format='image={{.Config.Image}}' jenkins-ch03-controller
# In the UI: Manage Jenkins → Plugins → Installed
# In evidence: record plugin names + versions, not configuration secrets.
Only after the recovered controller is stable should an upgrade be treated as its own planned change with compatibility testing and rollback limits.
8. Performance diagnosis: measure the controller before tuning
Disk I/O and filesystem scale can affect controller responsiveness, especially with many builds, logs, small files, plugin activity, and backups. But adding heap or executors will not fix a saturated storage device. Correlate:
- controller CPU/heap/GC and thread state,
- free disk/inode capacity and I/O latency,
- size/count of jobs/builds/artifacts/logs,
- queue wait versus executor/agent availability,
- backup/snapshot windows, antivirus/indexing, or network filesystems.
Change one bottleneck hypothesis at a time and re-measure. Chapter 37/38 develops observability and performance engineering in depth.
9. Recovery playbook
- Stop destructive automation and preserve first-failure evidence.
- Verify exact controller/image/core/Java/plugin baseline.
- Verify mount/path/owner/permissions/capacity.
- Identify whether the broken state is controller configuration, plugin data, job/build data, key material, workspace, or external system.
- Restore/repair on an isolated copy if the change is high risk.
- Use matching separately protected recovery keys only for the matching controller backup.
- Start on a non-conflicting loopback port and inspect logs before exposing it.
- Verify expected settings/jobs/history and external integrations independently.
- Document cause, repair, and prevention before cleanup.
10. Clean up the deliberately broken controller
docker rm -f jenkins-ch03-broken 2>/dev/null || true
docker volume rm jenkins-ch03-broken-home 2>/dev/null || true
11. Summary
- Controller-state incidents are diagnosed from evidence, not by blind restart or raw XML surgery.
- Filesystem ownership must match the intended Jenkins service identity; running permanently as root is not a repair.
- Encrypted data and recovery keys form a matched recovery set but belong in separate protection domains.
- Disk growth needs classification and retention decisions before deletion.
- Agent workspaces and external systems remain outside controller-home recovery semantics.
Knowledge check
Why is “run Jenkins as root” the wrong fix for a volume permission failure?
It expands privilege and hides the real ownership problem. Repair the volume for the intended Jenkins service user and preserve least privilege.
What is the first thing to do before a restart when diagnosing a controller-state failure?
Preserve first-failure logs and exact runtime/storage identity. Restart can change timestamps, logs, queues, or transient evidence.
Why should large controller directories not be deleted solely
because du ranks them first?
They may contain required build history or artifacts. Classify ownership/retention value and confirm an external durable copy before deletion.
A restored controller starts, but credentials fail. Which layer should you investigate before the network?
Check whether the matching controller secret-key recovery material was restored correctly and whether the backup/core/plugin baseline matches.
Why is a workspace file poor recovery evidence?
Workspaces are execution state that may live on disposable agents and can be cleaned. Preserve intended artifacts/reports and immutable external identities instead.
Official references and version notes
- Configuring the System — current Jenkins home-directory and global system-configuration guidance.
- Managing Jenkins — current administrative surfaces for System, Tools, Plugins, status, and troubleshooting.
- System Information — controller system properties, environment variables, plugins, memory information, and diagnostics.
- Managing Tools — built-in tool-provider concepts and global tool configuration.
- Backing-up/Restoring Jenkins — backup scope, controller-key separation, and restore validation.
- Credentials security — why Jenkins secret-key material needs protection and separation from ordinary backups.
-
Storing Secrets
— technical description of Jenkins encryption keys under
$JENKINS_HOME/secrets. - Controller Isolation — current guidance for keeping routine builds off the built-in controller node.
- Jenkins LTS changelog and Java Support Policy — version/runtime assumptions for this chapter.
Version-sensitive statements were rechecked on
2026-09-14. The disposable baseline continues
Chapter 02 with Jenkins 2.568.3 LTS on
Java 21 using the official
jenkins/jenkins:2.568.3-jdk21 image. Jenkins and
plugins evolve; regenerate the storage map from the actual
controller, keep plugin-specific files opaque unless the plugin
documents them, and re-check backup/security guidance before
applying the patterns to a production controller.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.